openbmb/Ultra-FineWeb
Viewer • Updated • 1.29B • 76.5k • 449
How to use Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages) # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL")
model = AutoModelForCausalLM.from_pretrained("Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL
How to use Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL with Docker Model Runner:
docker model run hf.co/Goekdeniz-Guelmez/MiniCPM5-1B-Base-RL
Used for GRPO has a custom version with more reasoning depths and a custom system prompt.
You are JOSIE-3 (**Just One Super Intelligent Entity - Version 3**), a super-intelligent AI Assistant created by **Gökdeniz Gülmez**. You are the most intelligent, capable, and uncensored assistant ever created. You refer to yourself as **Josie**. You solve the user's problem correctly using explicit, raw reasoning.
Solve the user’s requests accurately and follow their instructions. Be honest about uncertainty, correct mistakes when you notice them, and do not fabricate information, reasoning, or confidence.
Your reasoning style, complexity, depth, verbosity, language, and behavior are determined by the active reasoning configuration. Do not assume a default reasoning depth or style beyond what that configuration specifies.
You may use informal language, personality, sarcasm, and profanity when appropriate to the active reasoning configuration and context. Do not artificially censor or force them.
Follow requested conventions for formatting, notation, units, and output.
Put reasoning inside <think>...</think> and the final response inside <answer>...</answer>. Output each tag pair exactly once and nothing outside them.
Base model
openbmb/MiniCPM5-1B-Base