Abhinav-jigsawstack commited on
Commit
b9a47f2
·
verified ·
1 Parent(s): fe90e3e

lev: step 18750 of preset 4b-instruct

Browse files
.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ assets/leaderboard.png filter=lfs diff=lfs merge=lfs -text
37
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,3 +1,217 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-4B
4
+ base_model_relation: adapter
5
+ library_name: peft
6
+ pipeline_tag: zero-shot-classification
7
+ inference: false
8
+ language:
9
+ - en
10
+ tags:
11
+ - lev
12
+ - system-one
13
+ - decision-model
14
+ - calibrated-decisions
15
+ - classification
16
+ - routing
17
+ - moderation
18
+ - lora
19
  ---
20
+
21
+ # lev
22
+
23
+ **A calibrated System One decision model.** You give it a **state** (text, a ticket, an email, or JSON) and a set of **typed questions**. It returns typed answers with calibrated probabilities, all in one forward pass (about 69 ms on an H100). It never generates text, so there is nothing to parse and no label outside your option set.
24
+
25
+ | Question | You give | You get |
26
+ |---|---|---|
27
+ | `noul` | a yes/no question | `noul` = p(yes) |
28
+ | `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
29
+ | `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
30
+
31
+ The request and response contract matches TypeSafe's `POST /v1/systemone`.
32
+ Code written for the TypeSafe SDK works against lev once you change the
33
+ base URL. It is built for routing, moderation, intent detection, triage,
34
+ grading, and verifying LLM output.
35
+
36
+ ## Installation
37
+
38
+ ```bash
39
+ pip install "lev[serve] @ git+https://github.com/Abhinavexists/lev#subdirectory=packages/lev"
40
+ ```
41
+
42
+ This needs Python 3.12 or newer and, for real-time use, a CUDA GPU. The `serve` extra installs torch, transformers, peft, and the HTTP server. The first load downloads the base model (Qwen/Qwen3.5-4B, about 8 GB) and this adapter (about 200 MB).
43
+
44
+ ## Quickstart
45
+
46
+ ```python
47
+ import lev
48
+
49
+ model = lev.load("interfaze-ai/lev")
50
+
51
+ state = "Hi, I was charged twice for my order #4471 and I want a refund."
52
+ questions = {
53
+ "intent": {
54
+ "type": "choice",
55
+ "instructions": "What does the customer want?",
56
+ "criteria": {
57
+ "refund": "wants money back",
58
+ "cancel": "wants to cancel an order",
59
+ "track": "wants to know where an order is",
60
+ "other": "anything else",
61
+ },
62
+ },
63
+ "urgent": {"type": "noul", "instructions": "Does this need a human within the hour?"},
64
+ "frustration": {
65
+ "type": "score",
66
+ "instructions": "How frustrated is the customer?",
67
+ "criteria": ["calm", "mildly annoyed", "annoyed", "angry"],
68
+ },
69
+ }
70
+
71
+ result = model.system_one(state, questions)
72
+ print(result.answers["intent"].choice) # refund
73
+ print(result.answers["intent"].probabilities)
74
+ # {'refund': 0.84, 'cancel': 0.094, 'track': 0.012, 'other': 0.054}
75
+ print(result.answers["urgent"].noul) # 0.43
76
+ print(result.answers["frustration"].score) # 1.57, between "mildly annoyed" and "annoyed"
77
+ print(result.usage.output_tokens) # 0
78
+ ```
79
+
80
+ These are real outputs from this checkpoint. The answers are objects, and `result.model_dump()` gives the same JSON the HTTP server returns.
81
+
82
+ All the questions share one forward pass, so asking three questions costs about the same as asking one. `lev.load` reads `lev_release.json` from this repository to find the base model and the prompt format the adapter was trained with. It also applies the shipped calibration and loads the matching head, so there is nothing to configure.
83
+
84
+ Because the probabilities are calibrated, you can gate on them. For example, act automatically above 0.9 and send anything lower to a person.
85
+
86
+ ## Self-hosting: Jev-compatible HTTP server
87
+
88
+ ```bash
89
+ lev serve --checkpoint interfaze-ai/lev --host 0.0.0.0 --port 8000
90
+ ```
91
+
92
+ ```bash
93
+ curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
94
+ "state": "The package arrived crushed and the screen is cracked.",
95
+ "questions": {"damaged": {"type": "noul", "instructions": "Was the item damaged?"}}
96
+ }'
97
+ ```
98
+
99
+ Existing TypeSafe clients work once you point them at the server:
100
+
101
+ ```python
102
+ from typesafe_sdk import Choice, Noul, TypeSafeClient
103
+
104
+ # The first request after startup compiles kernels; allow more than the default 10 s.
105
+ client = TypeSafeClient(base_url="http://localhost:8000", api_key="local", timeout=60)
106
+ response = client.system_one(
107
+ state={"ticket": "The app crashes every time I open settings."},
108
+ questions={
109
+ "team": Choice(
110
+ instructions="Which team owns this?",
111
+ criteria={"billing": "payments, refunds", "technical": "bugs, crashes", "other": None},
112
+ ),
113
+ "bug": Noul(instructions="Is this a bug report?"),
114
+ },
115
+ )
116
+ print(response.answers["team"].choice, response.answers["bug"].noul) # technical 0.92
117
+ ```
118
+
119
+ `GET /health` reports the loaded checkpoint, whether calibration is active, and the routing settings. The server batches every question in a request into one forward pass, accepts concurrent requests, and returns 422 with the reason for a malformed question.
120
+
121
+ ## How it works
122
+
123
+ - **Backbone:** Qwen/Qwen3.5-4B, adapted with LoRA (r=32, α=64) on the q/k/v/o attention and MLP projections.
124
+ - **Label-token readout.** Each option gets a short code, and the answer is read from the next-token logits over those codes. Codes that would split into two tokens are skipped, so every option stays one token. This handles yes/no, scores, and choice lists up to several hundred options. Choices are read in two option orders and averaged, which cancels position bias.
125
+ - **Candidate-path head.** Past that point, a small learned head matches the state against each option's text, so the number of options has no fixed ceiling.
126
+ - **Calibration.** Temperatures fitted after training are applied at load time.
127
+ There is one per question type, readout mode, and choice option-count band.
128
+ They were selected by how well they carry over to task families left out of
129
+ the fit, not only by how well they fit held-out rows.
130
+
131
+ ## Benchmarks
132
+
133
+ ### S1Bench
134
+
135
+ Six S1Bench subsets, 1,999 items. lev and TypeSafe Jev were run through the same harness on the same task files. Accuracy, best in each row in bold:
136
+
137
+ | subset | task | **lev** | Jev | Qwen3.5-4B, untuned |
138
+ |---|---|---|---|---|
139
+ | aegis2 | safety moderation | **0.864** | 0.832 | 0.776 |
140
+ | boolq | yes/no reading comprehension | 0.880 | **0.910** | 0.860 |
141
+ | massive-en-US | intent, 60 classes | 0.791 | **0.814** | 0.734 |
142
+ | vitaminc-dev | claim verification | 0.738 | **0.846** | 0.733 |
143
+ | paws | adversarial paraphrase | 0.716 | **0.820** | 0.756 |
144
+ | helpsteer2 | helpfulness rating | 0.360 | 0.304 | **0.400** |
145
+ | **macro** | | 0.725 | **0.754** | 0.710 |
146
+
147
+ ![lev and Jev accuracy on each S1Bench subset](assets/accuracy-by-subset.png)
148
+
149
+ lev is ahead of Jev on safety moderation and helpfulness rating, within about 3 points on boolq and intent, and furthest behind on the two minimal-edit tasks, paws and vitaminc.
150
+
151
+ ![S1Bench leaderboard: lev ranks sixth of 22](assets/leaderboard.png)
152
+
153
+ On the S1Bench board over these subsets, lev ranks sixth of 22. Only Jev and three open models of 26B–35B parameters score higher, and lev is ahead of every model its size or smaller.
154
+
155
+ ![Macro accuracy against parameter count](assets/accuracy-per-parameter.png)
156
+
157
+ ### Held-out split
158
+
159
+ A held-out split of the 29 training sources, with no row shared with training:
160
+
161
+ | metric | value |
162
+ |---|---|
163
+ | weighted accuracy | 0.807\* |
164
+ | expected calibration error | 0.061\* (0.180 before calibration) |
165
+ | banking77 (77 intents) | 0.980 |
166
+ | clinc_oos (151 intents) | 0.968 |
167
+ | FEVER claim verification | 0.872\* |
168
+
169
+ \* Measured on this checkpoint before the last serving update. That update routes choice sets of more than 68 options to label-token readout and re-selects the temperatures. The banking77 and clinc_oos rows come from after the update, which raised banking77 from 0.818. Smaller option sets are routed the same way as before.
170
+
171
+ ### Speed
172
+
173
+ | | |
174
+ |---|---|
175
+ | compute per call, one H100 | **69 ms** |
176
+ | from 1 question to a 60-option choice | flat: one forward pass either way |
177
+ | output tokens | 0 |
178
+ | end to end from a laptop to a hosted endpoint, median | 414–463 ms (mostly network) |
179
+
180
+ ![lev compute per call on one H100](assets/compute.png)
181
+
182
+ The served path is one batched forward pass over every question, with the depthwise-conv kernel. Compute stays flat from one question to eight.
183
+
184
+ ![Per-call latency from the same laptop, lev and Jev](assets/latency.png)
185
+
186
+ Measured end to end from the same laptop, one request at a time, Jev's hosted API answered in 344–357 ms median and lev, on one Modal H100, in 414–463 ms. Most of lev's round trip is network and Modal's ingress; its compute is 69 ms. Self-hosted next to your application, that network hop disappears.
187
+
188
+ ## Training
189
+
190
+ - **Data:** 200,000 examples from 29 sources built on 26 public Hugging Face datasets. The tasks cover topic, sentiment, and emotion classification; intent detection (banking77, clinc_oos, snips); NLI (SNLI, ANLI, FEVER); paraphrase (MRPC, QQP, PARADE, plus word-swapped hard negatives); multiple-choice QA (RACE, ARC, SciQ, OpenBookQA, CommonsenseQA, StrategyQA); toxicity and safety (ToxiGen, ToxicChat, BeaverTails); and helpfulness (UltraFeedback). Questions are paraphrased, negated, and recast between types, and option sets are shuffled and resized, so the model learns the question format and not one wording.
191
+ - **Contamination guard:** the data build refuses any source that resolves to one of the 13 S1Bench subsets.
192
+ - **Recipe:** 3 epochs, 18,750 steps, batch size 32, learning rate 5e-5, in the base model's chat format, on one H100.
193
+
194
+ ## Limitations
195
+
196
+ - **Minimal-edit pairs.** On paws and vitaminc, two inputs can differ by one swapped word or one changed number. Here the model can be confidently wrong. It scores below the untuned backbone on paws and only matches it on vitaminc.
197
+ - **Fine-grained quality ratings are weak.** Helpfulness scoring (helpsteer2) sits at 0.36. Treat such scores as a rough signal.
198
+ - **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
199
+ - **Partial benchmark coverage.** S1Bench results cover 6 of its 13 subsets.
200
+ - **English only.**
201
+ - **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
202
+
203
+ ## Files
204
+
205
+ | file | role |
206
+ |---|---|
207
+ | `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
208
+ | `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
209
+ | `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
210
+ | `calibration.json` | fitted temperatures |
211
+ | `lev_release.json` | manifest: base model, prompt format, readout, training step |
212
+
213
+ ## License
214
+
215
+ The adapter is released under Apache-2.0, the same license as the base model. Some of the training datasets have their own terms, including non-commercial licenses. Review them before commercial use.
216
+
217
+ Apache 2.0 · Interfaze
adapter_config.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "Qwen/Qwen3.5-4B",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "kasa_config": null,
16
+ "layer_replication": null,
17
+ "layers_pattern": null,
18
+ "layers_to_transform": null,
19
+ "loftq_config": {},
20
+ "lora_alpha": 64,
21
+ "lora_bias": false,
22
+ "lora_dropout": 0.05,
23
+ "lora_ga_config": null,
24
+ "megatron_config": null,
25
+ "megatron_core": "megatron.core",
26
+ "modules_to_save": null,
27
+ "monteclora_config": null,
28
+ "peft_type": "LORA",
29
+ "peft_version": "0.21.0",
30
+ "qalora_group_size": 16,
31
+ "r": 32,
32
+ "rank_pattern": {},
33
+ "revision": null,
34
+ "target_modules": [
35
+ "gate_proj",
36
+ "down_proj",
37
+ "up_proj",
38
+ "k_proj",
39
+ "v_proj",
40
+ "o_proj",
41
+ "q_proj"
42
+ ],
43
+ "target_parameters": null,
44
+ "task_type": "CAUSAL_LM",
45
+ "trainable_token_indices": null,
46
+ "use_bdlora": null,
47
+ "use_dora": false,
48
+ "use_qalora": false,
49
+ "use_rslora": false,
50
+ "velora_config": null
51
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:64c718974d7ed0c9b22a0304060669fa3289ff89b57e572d08b9e9a0dc0623c1
3
+ size 169903320
assets/accuracy-by-subset.png ADDED
assets/accuracy-per-parameter.png ADDED
assets/compute.png ADDED
assets/latency.png ADDED
assets/leaderboard.png ADDED

Git LFS Details

  • SHA256: 550c4fa229b4394b6568e4a6face54414cf53658b4944c4307e89661d21b4230
  • Pointer size: 131 Bytes
  • Size of remote file: 113 kB
calibration.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "temperatures": {
3
+ "choice:B:large": 0.6333014152276044,
4
+ "choice:B": 0.6333014152276044,
5
+ "noul:A": 2.3332930831706458,
6
+ "score:A": 2.8033236630327885,
7
+ "choice:A": 1.7666659539679779,
8
+ "choice:A:small": 1.789783450971802,
9
+ "choice:A:mid": 1.6085685001661532,
10
+ "choice:A:large": 1.6641112617330986
11
+ },
12
+ "fitted_on": "calibration (transfer-selected)",
13
+ "n_samples": {
14
+ "choice:B:large": 10227,
15
+ "choice:B": 10227,
16
+ "noul:A": 12073,
17
+ "score:A": 12258,
18
+ "choice:A": 14133,
19
+ "choice:A:small": 11480,
20
+ "choice:A:mid": 1138,
21
+ "choice:A:large": 1515
22
+ }
23
+ }
chat_template.jinja ADDED
@@ -0,0 +1,154 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- if tools and tools is iterable and tools is not mapping %}
46
+ {{- '<|im_start|>system\n' }}
47
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
48
+ {%- for tool in tools %}
49
+ {{- "\n" }}
50
+ {{- tool | tojson }}
51
+ {%- endfor %}
52
+ {{- "\n</tools>" }}
53
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
54
+ {%- if messages[0].role == 'system' %}
55
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
56
+ {%- if content %}
57
+ {{- '\n\n' + content }}
58
+ {%- endif %}
59
+ {%- endif %}
60
+ {{- '<|im_end|>\n' }}
61
+ {%- else %}
62
+ {%- if messages[0].role == 'system' %}
63
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
64
+ {{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
65
+ {%- endif %}
66
+ {%- endif %}
67
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
68
+ {%- for message in messages[::-1] %}
69
+ {%- set index = (messages|length - 1) - loop.index0 %}
70
+ {%- if ns.multi_step_tool and message.role == "user" %}
71
+ {%- set content = render_content(message.content, false)|trim %}
72
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
73
+ {%- set ns.multi_step_tool = false %}
74
+ {%- set ns.last_query_index = index %}
75
+ {%- endif %}
76
+ {%- endif %}
77
+ {%- endfor %}
78
+ {%- if ns.multi_step_tool %}
79
+ {{- raise_exception('No user query found in messages.') }}
80
+ {%- endif %}
81
+ {%- for message in messages %}
82
+ {%- set content = render_content(message.content, true)|trim %}
83
+ {%- if message.role == "system" %}
84
+ {%- if not loop.first %}
85
+ {{- raise_exception('System message must be at the beginning.') }}
86
+ {%- endif %}
87
+ {%- elif message.role == "user" %}
88
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
89
+ {%- elif message.role == "assistant" %}
90
+ {%- set reasoning_content = '' %}
91
+ {%- if message.reasoning_content is string %}
92
+ {%- set reasoning_content = message.reasoning_content %}
93
+ {%- else %}
94
+ {%- if '</think>' in content %}
95
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
96
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
97
+ {%- endif %}
98
+ {%- endif %}
99
+ {%- set reasoning_content = reasoning_content|trim %}
100
+ {%- if loop.index0 > ns.last_query_index %}
101
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
102
+ {%- else %}
103
+ {{- '<|im_start|>' + message.role + '\n' + content }}
104
+ {%- endif %}
105
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
106
+ {%- for tool_call in message.tool_calls %}
107
+ {%- if tool_call.function is defined %}
108
+ {%- set tool_call = tool_call.function %}
109
+ {%- endif %}
110
+ {%- if loop.first %}
111
+ {%- if content|trim %}
112
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
113
+ {%- else %}
114
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
115
+ {%- endif %}
116
+ {%- else %}
117
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
118
+ {%- endif %}
119
+ {%- if tool_call.arguments is defined %}
120
+ {%- for args_name, args_value in tool_call.arguments|items %}
121
+ {{- '<parameter=' + args_name + '>\n' }}
122
+ {%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
123
+ {{- args_value }}
124
+ {{- '\n</parameter>\n' }}
125
+ {%- endfor %}
126
+ {%- endif %}
127
+ {{- '</function>\n</tool_call>' }}
128
+ {%- endfor %}
129
+ {%- endif %}
130
+ {{- '<|im_end|>\n' }}
131
+ {%- elif message.role == "tool" %}
132
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
133
+ {{- '<|im_start|>user' }}
134
+ {%- endif %}
135
+ {{- '\n<tool_response>\n' }}
136
+ {{- content }}
137
+ {{- '\n</tool_response>' }}
138
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
139
+ {{- '<|im_end|>\n' }}
140
+ {%- elif loop.last %}
141
+ {{- '<|im_end|>\n' }}
142
+ {%- endif %}
143
+ {%- else %}
144
+ {{- raise_exception('Unexpected message role.') }}
145
+ {%- endif %}
146
+ {%- endfor %}
147
+ {%- if add_generation_prompt %}
148
+ {{- '<|im_start|>assistant\n' }}
149
+ {%- if enable_thinking is defined and enable_thinking is false %}
150
+ {{- '<think>\n\n</think>\n\n' }}
151
+ {%- else %}
152
+ {{- '<think>\n' }}
153
+ {%- endif %}
154
+ {%- endif %}
lev_release.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "name": "lev",
3
+ "preset": "4b-instruct",
4
+ "step": 18750,
5
+ "base_model": "Qwen/Qwen3.5-4B",
6
+ "lora_rank": 32,
7
+ "mode_b_head": true,
8
+ "calibrated": true,
9
+ "noul_readout": "rating",
10
+ "prompt_style": "chat",
11
+ "files": [
12
+ "adapter_config.json",
13
+ "adapter_model.safetensors",
14
+ "chat_template.jinja",
15
+ "mode_b_head.pt",
16
+ "tokenizer.json",
17
+ "tokenizer_config.json",
18
+ "calibration.json"
19
+ ],
20
+ "metrics": {},
21
+ "created": "2026-09-24T06:29:31+00:00"
22
+ }
mode_b_head.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:27eedf7bb9e20d66432882b91297ef6e1ce3d23d81e94ed105ffadba77c6d3ad
3
+ size 14698565
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
3
+ size 19989325
tokenizer_config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": false,
13
+ "local_files_only": false,
14
+ "model_max_length": 262144,
15
+ "model_specific_special_tokens": {
16
+ "audio_bos_token": "<|audio_start|>",
17
+ "audio_eos_token": "<|audio_end|>",
18
+ "audio_token": "<|audio_pad|>",
19
+ "image_token": "<|image_pad|>",
20
+ "video_token": "<|video_pad|>",
21
+ "vision_bos_token": "<|vision_start|>",
22
+ "vision_eos_token": "<|vision_end|>"
23
+ },
24
+ "pad_token": "<|endoftext|>",
25
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
26
+ "split_special_tokens": false,
27
+ "tokenizer_class": "Qwen2Tokenizer",
28
+ "unk_token": null,
29
+ "video_token": "<|video_pad|>",
30
+ "vision_bos_token": "<|vision_start|>",
31
+ "vision_eos_token": "<|vision_end|>"
32
+ }