Qwen3-1.7B-Libra-MF

Qwen3-1.7B fine-tuned to read Romanian registre de mijloace fixe (fixed-asset registers) in any surface form and emit a column-mapping recipe as structured JSON. A deterministic post-processor consumes the recipe and produces, per 3-digit asset category, the six accounting totals. LoRA SFT, then merged back into a single 1.7B checkpoint for drop-in inference.

Trained on surogate/mf-dataset.

Business use case

Romanian accounting software (Soft1, Saga, Mentor, SmartBill, custom Excel exports) emits the fixed-asset register in a dozen incompatible layouts. For each asset category an accountant needs the six-field totals line:

  • Valoare intrare (entry value), Valoare modernizări (improvements), Valoare de inventar (inventory value), Valoare amortizată (accumulated depreciation), Amortizare lunară (monthly depreciation), Valoare rămasă (net book value).

That mapping is not fixed. Across registers:

  • Headers differ per software (Valoare intrare vs Valoare de intrare; Amortizare inregistrata vs Uzura; Val. ramasa (neamortizata) vs Valoare ramasa).
  • Registers come in two shapes: grouped (section headers like 212 CONSTRUCTII + a Total pe … subtotal; the category comes from the section header, there is no Cont column) and column (a per-row Cont/Categorie column; the category comes from that cell).
  • Trap columns look right but aren't: a bare Valoare, the monthly Amortizare lunară vs the cumulative Valoare amortizată, or decoy integer columns (Durata funct., Luni rămase).
  • The text arrives collapsed, OCR-mangled, headerless, multi-line-header, English-mixed, …

This model reads the raw extracted text and emits a JSON recipe naming exactly which column index (and its header text) plays each of the 8 roles. The model handles layout variance; the post-processor handles the arithmetic (per-category sums, the accounting identities, and a cross-check).

Why a dedicated SLM

General GPT-4-class models on Romanian registre repeatedly:

where big models fail what they output why it matters
Confuse monthly vs cumulative depreciation amortizare_luna points at the cumulative column every monthly total is wrong
Pick a bare Valoare trap column wrong inventory/entry value category totals don't reconcile
Fabricate headers on headerless layouts invented column names indexed lookup returns nothing
Map a value role to an integer decoy (Durata, Luni) a duration counted as money totals inflate
Emit a cont column on a grouped register category leaks between sections a 215 asset lands under 211

The disambiguating information is in the input every time; the problem is attending to it. A 1.7B model trained on ~5,750 examples across 12 surface formats does this at a fraction of the inference cost.

How the model + post-processor split work

The model emits, per role, {header, index} (or null if the column is absent). cont = null ⇒ grouped shape (category from section headers); cont present ⇒ column shape. The deterministic post-processor (mf_apply) then:

  1. splits each row into cells (by separator, or by typed-token reconstruction for collapsed text),
  2. anchors each role to a column by header match, falling back to the model's index,
  3. sums the six fields per category, derives modernizări = inventar − intrare,
  4. cross-checks against the register's own printed Total pe … / Totaluri … line and against the identity inventar = amortizată + rămasă (except terenuri/211), emitting an Observatii note.

Eval results (shipped merged checkpoint)

eval set size score
Real client registers (end-to-end 6-field totals) 7 7 / 7 (every register, every category)
Held-out synthetic, all 12 formats (end-to-end) 360 95.3 %
Validation set (model recipe vs ground-truth recipe, both applied) 600 92.6 %

Per-format accuracy on the 360 held-out set (end-to-end totals):

format acc format acc
canonical_markers 100 % mixed_language 100 %
cont_prefix 100 % multi_line_header 93.3 %
header_only 100 % ocr_mangled 96.7 %
fixed_width 100 % csv 90.0 %
markdown 100 % pdf_copy (collapsed) 86.7 %
tsv 100 % headerless 76.7 %

Worked examples

Ametech (grouped register, no cont; category from section headers):

ANTET … | Denumire imobilizare | Valoare intrare | … | Amortizare inregistrata | Amortizare lunara | Val. ramasa
212 CONSTRUCTII
1  APARTAMENT 118 MUN.BUC, STR  1 805 972.00  …  149 935.72  3 762.44  1 656 036.28
…
{ "coloane": {
  "denumire":            {"header": "Denumire imobilizare", "index": 1},
  "valoare_intrare":     {"header": "Valoare intrare", "index": 3},
  "valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
  "valoare_inventar":    {"header": "Valoare de inventar", "index": 5},
  "valoare_amortizata":  {"header": "Amortizare inregistrata", "index": 9},
  "amortizare_luna":     {"header": "Amortizare lunara", "index": 8},
  "valoare_ramasa":      {"header": "Val. ramasa", "index": 6},
  "cont": null } }

cont = null → the post-processor takes the category from each NNN … section header.

Algorithm (column register, per-row Cont; category from that cell):

A/A Cod | Denumire mijloc fix | Valoare intrare | … | Cont de mijloace fixe | …
1  ONORARIU … 01/06/2017  185,45  …  208  …
{ "coloane": {
  "denumire":            {"header": "Denumire mijloc fix", "index": 1},
  "valoare_intrare":     {"header": "Valoare intrare", "index": 2},
  "valoare_modernizari": {"header": "Valoare modernizari", "index": 4},
  "valoare_inventar":    {"header": "Valoare de inventar", "index": 5},
  "valoare_amortizata":  {"header": "Amortizare inregistrata", "index": 6},
  "amortizare_luna":     {"header": "din care amortizat in luna", "index": 7},
  "valoare_ramasa":      {"header": "Val. ramasa (neamortizata)", "index": 8},
  "cont":                {"header": "Cont de mijloace fixe", "index": 10} } }

Headerless layout (no header row, columns by position):

1 INVESTITIE IMOBILIARA SIGMA 25.06.2024 63.421,57 9.509,52 72.931,09 10.824,39 607,76 62.106,70
…
{ "coloane": {
  "denumire": {"header": "", "index": 1}, "valoare_intrare": {"header": "", "index": 4},
  "valoare_modernizari": {"header": "", "index": 6}, "valoare_inventar": {"header": "", "index": 7},
  "valoare_amortizata": {"header": "", "index": 11}, "amortizare_luna": {"header": "", "index": 10},
  "valoare_ramasa": {"header": "", "index": 8}, "cont": null } }

With no header text, the model emits "header": "" and locates columns by their numeric position.

Output schema

field type content
coloane.<role> {header, index} or null one entry per role
roles list denumire, valoare_intrare, valoare_modernizari, valoare_inventar, valoare_amortizata, amortizare_luna, valoare_ramasa, cont
header str the column's header text as it appears ("" if headerless); the robust anchor
index int 0-based column position; the fallback anchor
cont = null flag grouped register (category from section headers)

Quick start

transformers

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("surogate/Qwen3-1.7B-Libra-MF")
model = AutoModelForCausalLM.from_pretrained("surogate/Qwen3-1.7B-Libra-MF",
                                             torch_dtype=torch.bfloat16, device_map="auto")
from datasets import load_dataset
SYSTEM = load_dataset("surogate/mf-dataset", split="train[:1]")[0]["instruction"]
user_text = open("my_registru.txt").read()
prompt = tok.apply_chat_template(
    [{"role": "system", "content": SYSTEM}, {"role": "user", "content": user_text}],
    tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=768, do_sample=False)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

vLLM

vllm serve surogate/Qwen3-1.7B-Libra-MF --max-model-len 4096 --gpu-memory-utilization 0.6

Then POST {system_prompt}\n{registru text} with temperature: 0, max_tokens: 768.

Training details

field value
base model Qwen/Qwen3-1.7B
method LoRA SFT, merged into base for shipping
recipe fp8-hybrid
batch per_device 1 × grad-accum 8 (effective 8), sequence_len 2048
LR 5e-5 cosine, warmup ratio 0.05
LoRA rank / alpha / dropout 16 / 32 / 0.15
LoRA target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
dataset surogate/mf-dataset (5,750 train + 600 val)
framework surogate sft

Note on batch: peak memory scales with the per_device microbatch, not the effective batch (accumulation is sequential). per_device 1 keeps the graph small and stable; effective batch 8 is reached via accumulation.

Limitations

  • Romanian only. The mixed_language format introduces some English headers, but the model is not robust to fully English registers.
  • Categories 205 to 215 (the standard 3-digit fixed-asset accounts) are the focus.
  • Inputs are token-budgeted to ≤ 2048 (no truncation in training). Very long registers should be passed first-N-rows-windowed (a header + a sample of rows is all the column mapping needs).
  • Collapsed pdf_copy and headerless are the hardest forms (86.7 % / 76.7 %): space-collapsed text is information-lossy, and headerless requires pure positional reasoning. For PDFs, the production path supplies an x-clustered cell grid as an aid, which sidesteps the collapse.
  • The post-processor is not part of this checkpoint. Without it the model output is a recipe, not the totals.

License

Apache 2.0. Inherits from Qwen/Qwen3-1.7B. Synthetic training data plus 7 anonymized real-register layout anchors.

Citation

@misc{qwen3-1.7b-libra-mf,
  title  = {Qwen3-1.7B-Libra-MF: Romanian registru de mijloace fixe column-mapping extractor},
  author = {Surogate},
  year   = {2026},
  url    = {https://huggingface.co/surogate/Qwen3-1.7B-Libra-MF}
}
Downloads last month
34
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for surogate/Qwen3-1.7B-Libra-MF

Finetuned
Qwen/Qwen3-1.7B
Adapter
(723)
this model