Qwen3-4B KG extraction (LoRA, canonical masked)

LoRA adapter that extracts {entities, relationships} from open English text in the Microsoft GraphRAG knowledge-model format. Part of the kg-triplet-sft project: https://github.com/Alex-tangt/kg-triplet-sft

  • Base model: Qwen/Qwen3-4B
  • Output types: PERSON / ORGANIZATION / GEO / EVENT / CONCEPT; relationship strength is 0–10.
  • Entity titles follow the ALL-CAPS GraphRAG convention.

Output format

{"entities": [{"title": "...", "type": "CONCEPT", "description": "..."}],
 "relationships": [{"source": "...", "target": "...", "description": "...", "strength": 8.0}]}

source/target must equal a title in entities. The model returns {"entities": [], "relationships": []} when nothing is extractable (hard-negative training).

Prompt

The student prompt lives in the project repo (kg_contract/student_prompt.py); it is derived from the official GraphRAG extraction prompt and uses one field vocabulary (title/type/description, source/target/description/strength). At inference the project injects a schema prefix ({"entities": [) to keep small models structurally compliant.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

repo = "<this repo id>"
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B", torch_dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained(repo)
model = PeftModel.from_pretrained(base, repo)

prompt = STUDENT_PROMPT + "\n\n" + passage          # see the repo
msgs = [{"role": "user", "content": prompt}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Training

Canonical masked recipe (Unsloth, one RTX 4090): Qwen3 in no-think chatml, response-only loss (train_on_responses_only), non-packed (unsloth#5230), r=32 / alpha=32, lr 5e-5 cosine, 5 epochs, cutoff 6144, bs 2 × grad-accum 4 (effective 8), bf16, 2,575 training passages.

Evaluation

Referent-level (same real-world entity) on a fixed 200-passage sample; judge = Qwen3 thinking. Reference baseline (same recipe, 0.6B) on this sample: entity R/P/F1 = 0.534 / 0.723 / 0.614.

metric value
entity micro R / P / F1 (strict) 0.595 / 0.720 / 0.651
entity micro R / P / F1 (case-tolerant) 0.664 / 0.743 / 0.701
schema valid (strict / case-tolerant) 0.941 / 0.953
hallucination rate (strict / case-tolerant) 0.099 / 0.084

Entity recall is monotone in size across the capacity line (0.536 → 0.590 → 0.664 case-tolerant 4B). The 0.6B point matches the reference within the sample's ±2.1 pt CI.

Limitations

  • One fixed-200 sample; single dataset (English Wikipedia 60% + arXiv 40%); no statistical power for small differences.
  • ~17% of 4B rows emit UPPERCASE schema keys under some decodes; a case-tolerant key-normalizing parser recovers them (the "case-tolerant" numbers above). Deployment should normalize keys.
  • Description factual accuracy is not verified by an LLM judge (out of scope); descriptions are grounded lexically only.
  • Research artifact, not a production extractor; serve/ (GGUF / Gradio) is not implemented.

Evidence & links

License & attribution

MIT (adapters). Base models: Qwen3 (Apache-2.0). Teacher labels from the official Microsoft GraphRAG prompt (MIT). Training data is not redistributed.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Alextgt/qwen3-4b-kg-extraction

Finetuned
Qwen/Qwen3-4B
Adapter
(1170)
this model

Collection including Alextgt/qwen3-4b-kg-extraction