Instructions to use Alextgt/qwen3-4b-kg-extraction with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Alextgt/qwen3-4b-kg-extraction with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B") model = PeftModel.from_pretrained(base_model, "Alextgt/qwen3-4b-kg-extraction") - Notebooks
- Google Colab
- Kaggle
Qwen3-4B KG extraction (LoRA, canonical masked)
LoRA adapter that extracts {entities, relationships} from open English text in the
Microsoft GraphRAG knowledge-model format.
Part of the kg-triplet-sft project: https://github.com/Alex-tangt/kg-triplet-sft
- Base model:
Qwen/Qwen3-4B - Output types:
PERSON/ORGANIZATION/GEO/EVENT/CONCEPT; relationshipstrengthis 0–10. - Entity titles follow the ALL-CAPS GraphRAG convention.
Output format
{"entities": [{"title": "...", "type": "CONCEPT", "description": "..."}],
"relationships": [{"source": "...", "target": "...", "description": "...", "strength": 8.0}]}
source/target must equal a title in entities. The model returns
{"entities": [], "relationships": []} when nothing is extractable (hard-negative training).
Prompt
The student prompt lives in the project repo (kg_contract/student_prompt.py); it is derived
from the official GraphRAG extraction prompt and uses one field vocabulary
(title/type/description, source/target/description/strength). At inference the project injects
a schema prefix ({"entities": [) to keep small models structurally compliant.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
repo = "<this repo id>"
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B", torch_dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(repo)
model = PeftModel.from_pretrained(base, repo)
prompt = STUDENT_PROMPT + "\n\n" + passage # see the repo
msgs = [{"role": "user", "content": prompt}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Training
Canonical masked recipe (Unsloth, one RTX 4090): Qwen3 in no-think chatml,
response-only loss (train_on_responses_only), non-packed (unsloth#5230),
r=32 / alpha=32, lr 5e-5 cosine, 5 epochs, cutoff 6144, bs 2 × grad-accum 4
(effective 8), bf16, 2,575 training passages.
Evaluation
Referent-level (same real-world entity) on a fixed 200-passage sample; judge = Qwen3 thinking. Reference baseline (same recipe, 0.6B) on this sample: entity R/P/F1 = 0.534 / 0.723 / 0.614.
| metric | value |
|---|---|
| entity micro R / P / F1 (strict) | 0.595 / 0.720 / 0.651 |
| entity micro R / P / F1 (case-tolerant) | 0.664 / 0.743 / 0.701 |
| schema valid (strict / case-tolerant) | 0.941 / 0.953 |
| hallucination rate (strict / case-tolerant) | 0.099 / 0.084 |
Entity recall is monotone in size across the capacity line (0.536 → 0.590 → 0.664 case-tolerant 4B). The 0.6B point matches the reference within the sample's ±2.1 pt CI.
Limitations
- One fixed-200 sample; single dataset (English Wikipedia 60% + arXiv 40%); no statistical power for small differences.
- ~17% of 4B rows emit UPPERCASE schema keys under some decodes; a case-tolerant key-normalizing parser recovers them (the "case-tolerant" numbers above). Deployment should normalize keys.
- Description factual accuracy is not verified by an LLM judge (out of scope); descriptions are grounded lexically only.
- Research artifact, not a production extractor;
serve/(GGUF / Gradio) is not implemented.
Evidence & links
- Example node–edge extraction (teacher vs base vs this model, fixed-200): https://github.com/Alex-tangt/kg-triplet-sft/blob/master/docs/example-graph.md
- Capacity-line repair report / root-cause (loss coverage): https://github.com/Alex-tangt/kg-triplet-sft/blob/master/docs/2026-09-07-capacity-line-repair.md
- Evaluated predictions/reports and full methodology: https://github.com/Alex-tangt/kg-triplet-sft
License & attribution
MIT (adapters). Base models: Qwen3 (Apache-2.0). Teacher labels from the official Microsoft GraphRAG prompt (MIT). Training data is not redistributed.
- Downloads last month
- 15