Instructions to use dogukanvzr/Mergen-TR-Qwen3.5-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dogukanvzr/Mergen-TR-Qwen3.5-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dogukanvzr/Mergen-TR-Qwen3.5-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("dogukanvzr/Mergen-TR-Qwen3.5-9B") model = AutoModelForMultimodalLM.from_pretrained("dogukanvzr/Mergen-TR-Qwen3.5-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dogukanvzr/Mergen-TR-Qwen3.5-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dogukanvzr/Mergen-TR-Qwen3.5-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dogukanvzr/Mergen-TR-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dogukanvzr/Mergen-TR-Qwen3.5-9B
- SGLang
How to use dogukanvzr/Mergen-TR-Qwen3.5-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dogukanvzr/Mergen-TR-Qwen3.5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dogukanvzr/Mergen-TR-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dogukanvzr/Mergen-TR-Qwen3.5-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dogukanvzr/Mergen-TR-Qwen3.5-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use dogukanvzr/Mergen-TR-Qwen3.5-9B with Docker Model Runner:
docker model run hf.co/dogukanvzr/Mergen-TR-Qwen3.5-9B
Mergen-TR-Qwen3.5-9B
9B parameters · Turkish · built on Qwen3.5-9B · hybrid thinking · Apache-2.0
Mergen-TR-Qwen3.5-9B is a 9-billion-parameter, Turkish-focused reasoning model with a hybrid thinking mode. It is built on Qwen3.5-9B and post-trained on a curated Turkish corpus that strengthens mathematical reasoning, chain-of-thought and Turkish generation, while preserving the base model's strong Turkish knowledge. The result stands well above its parameter class on Turkish benchmarks.
On the public Turkish-MMLU leaderboard it scores 74.3, landing among 27B–70B open models and ahead of Llama 3.1 70B, Qwen3 14B and Gemma 2 27B. On the academic TurkishMMLU benchmark (chain-of-thought) it reaches 84.4, ahead of paper-reported Claude 3 Opus and GPT-4 Turbo. The evaluation protocol is fully disclosed below — these numbers are reproducible.
Highlights
| Benchmark | Score | Where it lands |
|---|---|---|
| Turkish-MMLU (leaderboard bank, 6,200 q) | 74.3 | Above Llama 3.1 70B (70.4), Qwen3 14B (71.7), Gemma 2 27B (72.1); level with Gemma 3 27B (75.1) |
| TurkishMMLU (Yüksel et al. 2024) · CoT | 84.4 | Above Claude 3 Opus (81.8) and GPT-4 Turbo (79.2) |
| GSM8K-TR (grade-school math) | 87.5 | +1.2 over the base model |
The model sits on the accuracy band of models 3–7× its size. A 9B model matching ~30B-class open models is the headline result.
Mergen-TR-Qwen3.5-9B is consistently strong across the science and humanities curriculum — math, physics, geography and chemistry sit in the high-80s/90s. Turkish language & literature is the weakest single area and the main target for future work.
Evaluation methodology (full transparency)
Thinking-style models are systematically under-scored by the short-generation
protocols used in older harnesses: the reasoning chain does not fit the token
budget, and answer parsers cannot handle <think> blocks. Every number here is
produced with the open protocol below, which anyone can re-run:
- Thinking ON (
enable_thinking=True), greedy decoding (do_sample=False,repetition_penalty=1.05). - Generation budget:
max_new_tokens=8192for Turkish-MMLU and16384for academic TurkishMMLU / GSM8K-TR. This matters — about 17% (Turkish-MMLU) and 25% (academic) of generations exceed 4,096 tokens, so a tight budget measures a lower score than the model actually achieves. - Answer parser (thinking-aware, 3 layers): an explicit marker after
</think>first, then an explicit marker anywhere, then a bare option letter after</think>. Across every runno_answer = 0(no generation was unparseable). - Samples: Turkish-MMLU — section-stratified n=1,895; academic TurkishMMLU — full 900-question set; GSM8K-TR — n=256. Precision bf16.
Comparison rows are not same-harness. Turkish-MMLU competitor scores are taken from the official leaderboard (those runs are mostly Q4 GGUF and use the board's own protocol); academic TurkishMMLU closed-model scores are from the source paper's CoT table. Rows therefore compare against the best publicly reported values, not a single shared harness.
Quickstart
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "dogukanvzr/Mergen-TR-Qwen3.5-9B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype="bfloat16", device_map="auto")
messages = [{"role": "user",
"content": "Divan edebiyatında kaside ile gazel arasındaki farkları açıkla."}]
text = tok.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True,
enable_thinking=True) # hybrid thinking mode
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=8192, do_sample=False,
repetition_penalty=1.05)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
vLLM (serving)
vllm serve dogukanvzr/Mergen-TR-Qwen3.5-9B --dtype bfloat16 --max-model-len 16384
Tips
- For knowledge / reasoning tasks use
enable_thinking=Truewithmax_new_tokens ≥ 8192; a tight budget truncates the reasoning chain and lowers quality. - For short, direct answers you can set
enable_thinking=False. - If you parse answers programmatically, read after the
</think>tag.
Model details
| Base model | Qwen/Qwen3.5-9B |
| Parameters | ~9.5B |
| Languages | Turkish (primary) |
| Context length | up to 16K used in evaluation |
| Precision | bfloat16 |
| Modes | hybrid — thinking / non-thinking |
| Post-training | supervised fine-tuning + preference optimization |
| License | Apache-2.0 |
Training data
Mergen-TR-Qwen3.5-9B was post-trained on a private Turkish corpus built for this project — ~277,000 supervised examples and ~15,000 preference pairs (SFT + preference optimization), constructed natively for Turkish (not bulk machine-translated). Approximate composition:
| Area | ~Examples |
|---|---|
| Mathematics & word problems | 64,000 |
| Knowledge / multiple-choice (incl. humanities) | 64,000 |
| Culture | 20,000 |
| Extractive QA | 15,000 |
| Classification | 13,000 |
| Grammar correction | 13,000 |
| Translation | 12,000 |
| Natural-language inference | 12,000 |
| Summarization | 11,000 |
| Commonsense (HellaSwag-style) | 10,000 |
| Instruction-following, NER / POS & others | 43,000 |
Quality and factuality filtering is applied throughout. The corpus and its construction pipeline are not released. Overlap with evaluation sets is prevented by hash-based decontamination covering all benchmark sources (0 leakage).
Limitations
- The focus is Turkish; in other languages the base model's behavior dominates.
- On multiple-choice factual recall (Turkish-MMLU), Mergen-TR-Qwen3.5-9B essentially matches its strong Qwen3.5-9B base at a fair generation budget — the leaderboard standing reflects the base's knowledge as much as the post-training. The post-training's measurable gains are concentrated in mathematics (+1.2 GSM8K-TR), chain-of-thought reasoning, and Turkish generation, not raw factual recall.
- Thinking-mode generations are long; for latency-sensitive use, prefer the non-thinking mode.
- Benchmark comparisons mix values reported by different harnesses (see Evaluation methodology); they are not strictly same-run comparisons.
References
- Turkish-MMLU leaderboard (alibayram), 6,200-question bank.
- A. Yüksel et al., TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish, 2024 (arXiv:2407.12402).
- GSM8K-TR (malhajar, v0.2).
- Downloads last month
- 30




