Zero-Shot Classification
PEFT
Safetensors
English
lev
system-one
decision-model
calibrated-decisions
classification
routing
moderation
lora
Instructions to use interfaze-ai/lev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use interfaze-ai/lev with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "interfaze-ai/lev") - Notebooks
- Google Colab
- Kaggle
lev: step 18750 of preset 4b-instruct
Browse files- .gitattributes +2 -0
- README.md +214 -0
- adapter_config.json +51 -0
- adapter_model.safetensors +3 -0
- assets/accuracy-by-subset.png +0 -0
- assets/accuracy-per-parameter.png +0 -0
- assets/compute.png +0 -0
- assets/latency.png +0 -0
- assets/leaderboard.png +3 -0
- calibration.json +23 -0
- chat_template.jinja +154 -0
- lev_release.json +22 -0
- mode_b_head.pt +3 -0
- tokenizer.json +3 -0
- tokenizer_config.json +32 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
assets/leaderboard.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -1,3 +1,217 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen3.5-4B
|
| 4 |
+
base_model_relation: adapter
|
| 5 |
+
library_name: peft
|
| 6 |
+
pipeline_tag: zero-shot-classification
|
| 7 |
+
inference: false
|
| 8 |
+
language:
|
| 9 |
+
- en
|
| 10 |
+
tags:
|
| 11 |
+
- lev
|
| 12 |
+
- system-one
|
| 13 |
+
- decision-model
|
| 14 |
+
- calibrated-decisions
|
| 15 |
+
- classification
|
| 16 |
+
- routing
|
| 17 |
+
- moderation
|
| 18 |
+
- lora
|
| 19 |
---
|
| 20 |
+
|
| 21 |
+
# lev
|
| 22 |
+
|
| 23 |
+
**A calibrated System One decision model.** You give it a **state** (text, a ticket, an email, or JSON) and a set of **typed questions**. It returns typed answers with calibrated probabilities, all in one forward pass (about 69 ms on an H100). It never generates text, so there is nothing to parse and no label outside your option set.
|
| 24 |
+
|
| 25 |
+
| Question | You give | You get |
|
| 26 |
+
|---|---|---|
|
| 27 |
+
| `noul` | a yes/no question | `noul` = p(yes) |
|
| 28 |
+
| `choice` | instructions + options (name → description or `null`) | `choice`, `probabilities`, `confidence` |
|
| 29 |
+
| `score` | instructions + 2–10 ordered levels | `score` (expected level), `probabilities`, `confidence` |
|
| 30 |
+
|
| 31 |
+
The request and response contract matches TypeSafe's `POST /v1/systemone`.
|
| 32 |
+
Code written for the TypeSafe SDK works against lev once you change the
|
| 33 |
+
base URL. It is built for routing, moderation, intent detection, triage,
|
| 34 |
+
grading, and verifying LLM output.
|
| 35 |
+
|
| 36 |
+
## Installation
|
| 37 |
+
|
| 38 |
+
```bash
|
| 39 |
+
pip install "lev[serve] @ git+https://github.com/Abhinavexists/lev#subdirectory=packages/lev"
|
| 40 |
+
```
|
| 41 |
+
|
| 42 |
+
This needs Python 3.12 or newer and, for real-time use, a CUDA GPU. The `serve` extra installs torch, transformers, peft, and the HTTP server. The first load downloads the base model (Qwen/Qwen3.5-4B, about 8 GB) and this adapter (about 200 MB).
|
| 43 |
+
|
| 44 |
+
## Quickstart
|
| 45 |
+
|
| 46 |
+
```python
|
| 47 |
+
import lev
|
| 48 |
+
|
| 49 |
+
model = lev.load("interfaze-ai/lev")
|
| 50 |
+
|
| 51 |
+
state = "Hi, I was charged twice for my order #4471 and I want a refund."
|
| 52 |
+
questions = {
|
| 53 |
+
"intent": {
|
| 54 |
+
"type": "choice",
|
| 55 |
+
"instructions": "What does the customer want?",
|
| 56 |
+
"criteria": {
|
| 57 |
+
"refund": "wants money back",
|
| 58 |
+
"cancel": "wants to cancel an order",
|
| 59 |
+
"track": "wants to know where an order is",
|
| 60 |
+
"other": "anything else",
|
| 61 |
+
},
|
| 62 |
+
},
|
| 63 |
+
"urgent": {"type": "noul", "instructions": "Does this need a human within the hour?"},
|
| 64 |
+
"frustration": {
|
| 65 |
+
"type": "score",
|
| 66 |
+
"instructions": "How frustrated is the customer?",
|
| 67 |
+
"criteria": ["calm", "mildly annoyed", "annoyed", "angry"],
|
| 68 |
+
},
|
| 69 |
+
}
|
| 70 |
+
|
| 71 |
+
result = model.system_one(state, questions)
|
| 72 |
+
print(result.answers["intent"].choice) # refund
|
| 73 |
+
print(result.answers["intent"].probabilities)
|
| 74 |
+
# {'refund': 0.84, 'cancel': 0.094, 'track': 0.012, 'other': 0.054}
|
| 75 |
+
print(result.answers["urgent"].noul) # 0.43
|
| 76 |
+
print(result.answers["frustration"].score) # 1.57, between "mildly annoyed" and "annoyed"
|
| 77 |
+
print(result.usage.output_tokens) # 0
|
| 78 |
+
```
|
| 79 |
+
|
| 80 |
+
These are real outputs from this checkpoint. The answers are objects, and `result.model_dump()` gives the same JSON the HTTP server returns.
|
| 81 |
+
|
| 82 |
+
All the questions share one forward pass, so asking three questions costs about the same as asking one. `lev.load` reads `lev_release.json` from this repository to find the base model and the prompt format the adapter was trained with. It also applies the shipped calibration and loads the matching head, so there is nothing to configure.
|
| 83 |
+
|
| 84 |
+
Because the probabilities are calibrated, you can gate on them. For example, act automatically above 0.9 and send anything lower to a person.
|
| 85 |
+
|
| 86 |
+
## Self-hosting: Jev-compatible HTTP server
|
| 87 |
+
|
| 88 |
+
```bash
|
| 89 |
+
lev serve --checkpoint interfaze-ai/lev --host 0.0.0.0 --port 8000
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
```bash
|
| 93 |
+
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
|
| 94 |
+
"state": "The package arrived crushed and the screen is cracked.",
|
| 95 |
+
"questions": {"damaged": {"type": "noul", "instructions": "Was the item damaged?"}}
|
| 96 |
+
}'
|
| 97 |
+
```
|
| 98 |
+
|
| 99 |
+
Existing TypeSafe clients work once you point them at the server:
|
| 100 |
+
|
| 101 |
+
```python
|
| 102 |
+
from typesafe_sdk import Choice, Noul, TypeSafeClient
|
| 103 |
+
|
| 104 |
+
# The first request after startup compiles kernels; allow more than the default 10 s.
|
| 105 |
+
client = TypeSafeClient(base_url="http://localhost:8000", api_key="local", timeout=60)
|
| 106 |
+
response = client.system_one(
|
| 107 |
+
state={"ticket": "The app crashes every time I open settings."},
|
| 108 |
+
questions={
|
| 109 |
+
"team": Choice(
|
| 110 |
+
instructions="Which team owns this?",
|
| 111 |
+
criteria={"billing": "payments, refunds", "technical": "bugs, crashes", "other": None},
|
| 112 |
+
),
|
| 113 |
+
"bug": Noul(instructions="Is this a bug report?"),
|
| 114 |
+
},
|
| 115 |
+
)
|
| 116 |
+
print(response.answers["team"].choice, response.answers["bug"].noul) # technical 0.92
|
| 117 |
+
```
|
| 118 |
+
|
| 119 |
+
`GET /health` reports the loaded checkpoint, whether calibration is active, and the routing settings. The server batches every question in a request into one forward pass, accepts concurrent requests, and returns 422 with the reason for a malformed question.
|
| 120 |
+
|
| 121 |
+
## How it works
|
| 122 |
+
|
| 123 |
+
- **Backbone:** Qwen/Qwen3.5-4B, adapted with LoRA (r=32, α=64) on the q/k/v/o attention and MLP projections.
|
| 124 |
+
- **Label-token readout.** Each option gets a short code, and the answer is read from the next-token logits over those codes. Codes that would split into two tokens are skipped, so every option stays one token. This handles yes/no, scores, and choice lists up to several hundred options. Choices are read in two option orders and averaged, which cancels position bias.
|
| 125 |
+
- **Candidate-path head.** Past that point, a small learned head matches the state against each option's text, so the number of options has no fixed ceiling.
|
| 126 |
+
- **Calibration.** Temperatures fitted after training are applied at load time.
|
| 127 |
+
There is one per question type, readout mode, and choice option-count band.
|
| 128 |
+
They were selected by how well they carry over to task families left out of
|
| 129 |
+
the fit, not only by how well they fit held-out rows.
|
| 130 |
+
|
| 131 |
+
## Benchmarks
|
| 132 |
+
|
| 133 |
+
### S1Bench
|
| 134 |
+
|
| 135 |
+
Six S1Bench subsets, 1,999 items. lev and TypeSafe Jev were run through the same harness on the same task files. Accuracy, best in each row in bold:
|
| 136 |
+
|
| 137 |
+
| subset | task | **lev** | Jev | Qwen3.5-4B, untuned |
|
| 138 |
+
|---|---|---|---|---|
|
| 139 |
+
| aegis2 | safety moderation | **0.864** | 0.832 | 0.776 |
|
| 140 |
+
| boolq | yes/no reading comprehension | 0.880 | **0.910** | 0.860 |
|
| 141 |
+
| massive-en-US | intent, 60 classes | 0.791 | **0.814** | 0.734 |
|
| 142 |
+
| vitaminc-dev | claim verification | 0.738 | **0.846** | 0.733 |
|
| 143 |
+
| paws | adversarial paraphrase | 0.716 | **0.820** | 0.756 |
|
| 144 |
+
| helpsteer2 | helpfulness rating | 0.360 | 0.304 | **0.400** |
|
| 145 |
+
| **macro** | | 0.725 | **0.754** | 0.710 |
|
| 146 |
+
|
| 147 |
+

|
| 148 |
+
|
| 149 |
+
lev is ahead of Jev on safety moderation and helpfulness rating, within about 3 points on boolq and intent, and furthest behind on the two minimal-edit tasks, paws and vitaminc.
|
| 150 |
+
|
| 151 |
+

|
| 152 |
+
|
| 153 |
+
On the S1Bench board over these subsets, lev ranks sixth of 22. Only Jev and three open models of 26B–35B parameters score higher, and lev is ahead of every model its size or smaller.
|
| 154 |
+
|
| 155 |
+

|
| 156 |
+
|
| 157 |
+
### Held-out split
|
| 158 |
+
|
| 159 |
+
A held-out split of the 29 training sources, with no row shared with training:
|
| 160 |
+
|
| 161 |
+
| metric | value |
|
| 162 |
+
|---|---|
|
| 163 |
+
| weighted accuracy | 0.807\* |
|
| 164 |
+
| expected calibration error | 0.061\* (0.180 before calibration) |
|
| 165 |
+
| banking77 (77 intents) | 0.980 |
|
| 166 |
+
| clinc_oos (151 intents) | 0.968 |
|
| 167 |
+
| FEVER claim verification | 0.872\* |
|
| 168 |
+
|
| 169 |
+
\* Measured on this checkpoint before the last serving update. That update routes choice sets of more than 68 options to label-token readout and re-selects the temperatures. The banking77 and clinc_oos rows come from after the update, which raised banking77 from 0.818. Smaller option sets are routed the same way as before.
|
| 170 |
+
|
| 171 |
+
### Speed
|
| 172 |
+
|
| 173 |
+
| | |
|
| 174 |
+
|---|---|
|
| 175 |
+
| compute per call, one H100 | **69 ms** |
|
| 176 |
+
| from 1 question to a 60-option choice | flat: one forward pass either way |
|
| 177 |
+
| output tokens | 0 |
|
| 178 |
+
| end to end from a laptop to a hosted endpoint, median | 414–463 ms (mostly network) |
|
| 179 |
+
|
| 180 |
+

|
| 181 |
+
|
| 182 |
+
The served path is one batched forward pass over every question, with the depthwise-conv kernel. Compute stays flat from one question to eight.
|
| 183 |
+
|
| 184 |
+

|
| 185 |
+
|
| 186 |
+
Measured end to end from the same laptop, one request at a time, Jev's hosted API answered in 344–357 ms median and lev, on one Modal H100, in 414–463 ms. Most of lev's round trip is network and Modal's ingress; its compute is 69 ms. Self-hosted next to your application, that network hop disappears.
|
| 187 |
+
|
| 188 |
+
## Training
|
| 189 |
+
|
| 190 |
+
- **Data:** 200,000 examples from 29 sources built on 26 public Hugging Face datasets. The tasks cover topic, sentiment, and emotion classification; intent detection (banking77, clinc_oos, snips); NLI (SNLI, ANLI, FEVER); paraphrase (MRPC, QQP, PARADE, plus word-swapped hard negatives); multiple-choice QA (RACE, ARC, SciQ, OpenBookQA, CommonsenseQA, StrategyQA); toxicity and safety (ToxiGen, ToxicChat, BeaverTails); and helpfulness (UltraFeedback). Questions are paraphrased, negated, and recast between types, and option sets are shuffled and resized, so the model learns the question format and not one wording.
|
| 191 |
+
- **Contamination guard:** the data build refuses any source that resolves to one of the 13 S1Bench subsets.
|
| 192 |
+
- **Recipe:** 3 epochs, 18,750 steps, batch size 32, learning rate 5e-5, in the base model's chat format, on one H100.
|
| 193 |
+
|
| 194 |
+
## Limitations
|
| 195 |
+
|
| 196 |
+
- **Minimal-edit pairs.** On paws and vitaminc, two inputs can differ by one swapped word or one changed number. Here the model can be confidently wrong. It scores below the untuned backbone on paws and only matches it on vitaminc.
|
| 197 |
+
- **Fine-grained quality ratings are weak.** Helpfulness scoring (helpsteer2) sits at 0.36. Treat such scores as a rough signal.
|
| 198 |
+
- **Calibration is fitted on the training distribution.** Temperatures are chosen to transfer across task families. Even so, a task very unlike the training mix may be less well calibrated. Check on your own data before you gate on the probabilities.
|
| 199 |
+
- **Partial benchmark coverage.** S1Bench results cover 6 of its 13 subsets.
|
| 200 |
+
- **English only.**
|
| 201 |
+
- **Needs a GPU for real-time use.** It runs on CPU, but a 4B backbone there takes seconds per call, not milliseconds.
|
| 202 |
+
|
| 203 |
+
## Files
|
| 204 |
+
|
| 205 |
+
| file | role |
|
| 206 |
+
|---|---|
|
| 207 |
+
| `adapter_model.safetensors`, `adapter_config.json` | LoRA adapter |
|
| 208 |
+
| `mode_b_head.pt` | candidate-path head (tensor state dict, loaded with `weights_only=True`) |
|
| 209 |
+
| `tokenizer*`, `chat_template.jinja` | the tokenizer that the label codes were verified against |
|
| 210 |
+
| `calibration.json` | fitted temperatures |
|
| 211 |
+
| `lev_release.json` | manifest: base model, prompt format, readout, training step |
|
| 212 |
+
|
| 213 |
+
## License
|
| 214 |
+
|
| 215 |
+
The adapter is released under Apache-2.0, the same license as the base model. Some of the training datasets have their own terms, including non-commercial licenses. Review them before commercial use.
|
| 216 |
+
|
| 217 |
+
Apache 2.0 · Interfaze
|
adapter_config.json
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen3.5-4B",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"kasa_config": null,
|
| 16 |
+
"layer_replication": null,
|
| 17 |
+
"layers_pattern": null,
|
| 18 |
+
"layers_to_transform": null,
|
| 19 |
+
"loftq_config": {},
|
| 20 |
+
"lora_alpha": 64,
|
| 21 |
+
"lora_bias": false,
|
| 22 |
+
"lora_dropout": 0.05,
|
| 23 |
+
"lora_ga_config": null,
|
| 24 |
+
"megatron_config": null,
|
| 25 |
+
"megatron_core": "megatron.core",
|
| 26 |
+
"modules_to_save": null,
|
| 27 |
+
"monteclora_config": null,
|
| 28 |
+
"peft_type": "LORA",
|
| 29 |
+
"peft_version": "0.21.0",
|
| 30 |
+
"qalora_group_size": 16,
|
| 31 |
+
"r": 32,
|
| 32 |
+
"rank_pattern": {},
|
| 33 |
+
"revision": null,
|
| 34 |
+
"target_modules": [
|
| 35 |
+
"gate_proj",
|
| 36 |
+
"down_proj",
|
| 37 |
+
"up_proj",
|
| 38 |
+
"k_proj",
|
| 39 |
+
"v_proj",
|
| 40 |
+
"o_proj",
|
| 41 |
+
"q_proj"
|
| 42 |
+
],
|
| 43 |
+
"target_parameters": null,
|
| 44 |
+
"task_type": "CAUSAL_LM",
|
| 45 |
+
"trainable_token_indices": null,
|
| 46 |
+
"use_bdlora": null,
|
| 47 |
+
"use_dora": false,
|
| 48 |
+
"use_qalora": false,
|
| 49 |
+
"use_rslora": false,
|
| 50 |
+
"velora_config": null
|
| 51 |
+
}
|
adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:64c718974d7ed0c9b22a0304060669fa3289ff89b57e572d08b9e9a0dc0623c1
|
| 3 |
+
size 169903320
|
assets/accuracy-by-subset.png
ADDED
|
assets/accuracy-per-parameter.png
ADDED
|
assets/compute.png
ADDED
|
assets/latency.png
ADDED
|
assets/leaderboard.png
ADDED
|
Git LFS Details
|
calibration.json
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"temperatures": {
|
| 3 |
+
"choice:B:large": 0.6333014152276044,
|
| 4 |
+
"choice:B": 0.6333014152276044,
|
| 5 |
+
"noul:A": 2.3332930831706458,
|
| 6 |
+
"score:A": 2.8033236630327885,
|
| 7 |
+
"choice:A": 1.7666659539679779,
|
| 8 |
+
"choice:A:small": 1.789783450971802,
|
| 9 |
+
"choice:A:mid": 1.6085685001661532,
|
| 10 |
+
"choice:A:large": 1.6641112617330986
|
| 11 |
+
},
|
| 12 |
+
"fitted_on": "calibration (transfer-selected)",
|
| 13 |
+
"n_samples": {
|
| 14 |
+
"choice:B:large": 10227,
|
| 15 |
+
"choice:B": 10227,
|
| 16 |
+
"noul:A": 12073,
|
| 17 |
+
"score:A": 12258,
|
| 18 |
+
"choice:A": 14133,
|
| 19 |
+
"choice:A:small": 11480,
|
| 20 |
+
"choice:A:mid": 1138,
|
| 21 |
+
"choice:A:large": 1515
|
| 22 |
+
}
|
| 23 |
+
}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,154 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- set image_count = namespace(value=0) %}
|
| 2 |
+
{%- set video_count = namespace(value=0) %}
|
| 3 |
+
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
|
| 4 |
+
{%- if content is string %}
|
| 5 |
+
{{- content }}
|
| 6 |
+
{%- elif content is iterable and content is not mapping %}
|
| 7 |
+
{%- for item in content %}
|
| 8 |
+
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
|
| 9 |
+
{%- if is_system_content %}
|
| 10 |
+
{{- raise_exception('System message cannot contain images.') }}
|
| 11 |
+
{%- endif %}
|
| 12 |
+
{%- if do_vision_count %}
|
| 13 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 14 |
+
{%- endif %}
|
| 15 |
+
{%- if add_vision_id %}
|
| 16 |
+
{{- 'Picture ' ~ image_count.value ~ ': ' }}
|
| 17 |
+
{%- endif %}
|
| 18 |
+
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
|
| 19 |
+
{%- elif 'video' in item or item.type == 'video' %}
|
| 20 |
+
{%- if is_system_content %}
|
| 21 |
+
{{- raise_exception('System message cannot contain videos.') }}
|
| 22 |
+
{%- endif %}
|
| 23 |
+
{%- if do_vision_count %}
|
| 24 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 25 |
+
{%- endif %}
|
| 26 |
+
{%- if add_vision_id %}
|
| 27 |
+
{{- 'Video ' ~ video_count.value ~ ': ' }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
|
| 30 |
+
{%- elif 'text' in item %}
|
| 31 |
+
{{- item.text }}
|
| 32 |
+
{%- else %}
|
| 33 |
+
{{- raise_exception('Unexpected item type in content.') }}
|
| 34 |
+
{%- endif %}
|
| 35 |
+
{%- endfor %}
|
| 36 |
+
{%- elif content is none or content is undefined %}
|
| 37 |
+
{{- '' }}
|
| 38 |
+
{%- else %}
|
| 39 |
+
{{- raise_exception('Unexpected content type.') }}
|
| 40 |
+
{%- endif %}
|
| 41 |
+
{%- endmacro %}
|
| 42 |
+
{%- if not messages %}
|
| 43 |
+
{{- raise_exception('No messages provided.') }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- if tools and tools is iterable and tools is not mapping %}
|
| 46 |
+
{{- '<|im_start|>system\n' }}
|
| 47 |
+
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
|
| 48 |
+
{%- for tool in tools %}
|
| 49 |
+
{{- "\n" }}
|
| 50 |
+
{{- tool | tojson }}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{{- "\n</tools>" }}
|
| 53 |
+
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
|
| 54 |
+
{%- if messages[0].role == 'system' %}
|
| 55 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 56 |
+
{%- if content %}
|
| 57 |
+
{{- '\n\n' + content }}
|
| 58 |
+
{%- endif %}
|
| 59 |
+
{%- endif %}
|
| 60 |
+
{{- '<|im_end|>\n' }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{%- if messages[0].role == 'system' %}
|
| 63 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 64 |
+
{{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
|
| 65 |
+
{%- endif %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 68 |
+
{%- for message in messages[::-1] %}
|
| 69 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 70 |
+
{%- if ns.multi_step_tool and message.role == "user" %}
|
| 71 |
+
{%- set content = render_content(message.content, false)|trim %}
|
| 72 |
+
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
|
| 73 |
+
{%- set ns.multi_step_tool = false %}
|
| 74 |
+
{%- set ns.last_query_index = index %}
|
| 75 |
+
{%- endif %}
|
| 76 |
+
{%- endif %}
|
| 77 |
+
{%- endfor %}
|
| 78 |
+
{%- if ns.multi_step_tool %}
|
| 79 |
+
{{- raise_exception('No user query found in messages.') }}
|
| 80 |
+
{%- endif %}
|
| 81 |
+
{%- for message in messages %}
|
| 82 |
+
{%- set content = render_content(message.content, true)|trim %}
|
| 83 |
+
{%- if message.role == "system" %}
|
| 84 |
+
{%- if not loop.first %}
|
| 85 |
+
{{- raise_exception('System message must be at the beginning.') }}
|
| 86 |
+
{%- endif %}
|
| 87 |
+
{%- elif message.role == "user" %}
|
| 88 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 89 |
+
{%- elif message.role == "assistant" %}
|
| 90 |
+
{%- set reasoning_content = '' %}
|
| 91 |
+
{%- if message.reasoning_content is string %}
|
| 92 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 93 |
+
{%- else %}
|
| 94 |
+
{%- if '</think>' in content %}
|
| 95 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 96 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 97 |
+
{%- endif %}
|
| 98 |
+
{%- endif %}
|
| 99 |
+
{%- set reasoning_content = reasoning_content|trim %}
|
| 100 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 101 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
|
| 102 |
+
{%- else %}
|
| 103 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 104 |
+
{%- endif %}
|
| 105 |
+
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
|
| 106 |
+
{%- for tool_call in message.tool_calls %}
|
| 107 |
+
{%- if tool_call.function is defined %}
|
| 108 |
+
{%- set tool_call = tool_call.function %}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- if loop.first %}
|
| 111 |
+
{%- if content|trim %}
|
| 112 |
+
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 113 |
+
{%- else %}
|
| 114 |
+
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 115 |
+
{%- endif %}
|
| 116 |
+
{%- else %}
|
| 117 |
+
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 118 |
+
{%- endif %}
|
| 119 |
+
{%- if tool_call.arguments is defined %}
|
| 120 |
+
{%- for args_name, args_value in tool_call.arguments|items %}
|
| 121 |
+
{{- '<parameter=' + args_name + '>\n' }}
|
| 122 |
+
{%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
|
| 123 |
+
{{- args_value }}
|
| 124 |
+
{{- '\n</parameter>\n' }}
|
| 125 |
+
{%- endfor %}
|
| 126 |
+
{%- endif %}
|
| 127 |
+
{{- '</function>\n</tool_call>' }}
|
| 128 |
+
{%- endfor %}
|
| 129 |
+
{%- endif %}
|
| 130 |
+
{{- '<|im_end|>\n' }}
|
| 131 |
+
{%- elif message.role == "tool" %}
|
| 132 |
+
{%- if loop.previtem and loop.previtem.role != "tool" %}
|
| 133 |
+
{{- '<|im_start|>user' }}
|
| 134 |
+
{%- endif %}
|
| 135 |
+
{{- '\n<tool_response>\n' }}
|
| 136 |
+
{{- content }}
|
| 137 |
+
{{- '\n</tool_response>' }}
|
| 138 |
+
{%- if not loop.last and loop.nextitem.role != "tool" %}
|
| 139 |
+
{{- '<|im_end|>\n' }}
|
| 140 |
+
{%- elif loop.last %}
|
| 141 |
+
{{- '<|im_end|>\n' }}
|
| 142 |
+
{%- endif %}
|
| 143 |
+
{%- else %}
|
| 144 |
+
{{- raise_exception('Unexpected message role.') }}
|
| 145 |
+
{%- endif %}
|
| 146 |
+
{%- endfor %}
|
| 147 |
+
{%- if add_generation_prompt %}
|
| 148 |
+
{{- '<|im_start|>assistant\n' }}
|
| 149 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 150 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 151 |
+
{%- else %}
|
| 152 |
+
{{- '<think>\n' }}
|
| 153 |
+
{%- endif %}
|
| 154 |
+
{%- endif %}
|
lev_release.json
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"name": "lev",
|
| 3 |
+
"preset": "4b-instruct",
|
| 4 |
+
"step": 18750,
|
| 5 |
+
"base_model": "Qwen/Qwen3.5-4B",
|
| 6 |
+
"lora_rank": 32,
|
| 7 |
+
"mode_b_head": true,
|
| 8 |
+
"calibrated": true,
|
| 9 |
+
"noul_readout": "rating",
|
| 10 |
+
"prompt_style": "chat",
|
| 11 |
+
"files": [
|
| 12 |
+
"adapter_config.json",
|
| 13 |
+
"adapter_model.safetensors",
|
| 14 |
+
"chat_template.jinja",
|
| 15 |
+
"mode_b_head.pt",
|
| 16 |
+
"tokenizer.json",
|
| 17 |
+
"tokenizer_config.json",
|
| 18 |
+
"calibration.json"
|
| 19 |
+
],
|
| 20 |
+
"metrics": {},
|
| 21 |
+
"created": "2026-09-24T06:29:31+00:00"
|
| 22 |
+
}
|
mode_b_head.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:27eedf7bb9e20d66432882b91297ef6e1ce3d23d81e94ed105ffadba77c6d3ad
|
| 3 |
+
size 14698565
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:06b9509352d2af50381ab2247e083b80d32d5c0aba91c272ca9ff729b6a0e523
|
| 3 |
+
size 19989325
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,32 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"audio_bos_token": "<|audio_start|>",
|
| 4 |
+
"audio_eos_token": "<|audio_end|>",
|
| 5 |
+
"audio_token": "<|audio_pad|>",
|
| 6 |
+
"backend": "tokenizers",
|
| 7 |
+
"bos_token": null,
|
| 8 |
+
"clean_up_tokenization_spaces": false,
|
| 9 |
+
"eos_token": "<|im_end|>",
|
| 10 |
+
"errors": "replace",
|
| 11 |
+
"image_token": "<|image_pad|>",
|
| 12 |
+
"is_local": false,
|
| 13 |
+
"local_files_only": false,
|
| 14 |
+
"model_max_length": 262144,
|
| 15 |
+
"model_specific_special_tokens": {
|
| 16 |
+
"audio_bos_token": "<|audio_start|>",
|
| 17 |
+
"audio_eos_token": "<|audio_end|>",
|
| 18 |
+
"audio_token": "<|audio_pad|>",
|
| 19 |
+
"image_token": "<|image_pad|>",
|
| 20 |
+
"video_token": "<|video_pad|>",
|
| 21 |
+
"vision_bos_token": "<|vision_start|>",
|
| 22 |
+
"vision_eos_token": "<|vision_end|>"
|
| 23 |
+
},
|
| 24 |
+
"pad_token": "<|endoftext|>",
|
| 25 |
+
"pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
|
| 26 |
+
"split_special_tokens": false,
|
| 27 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 28 |
+
"unk_token": null,
|
| 29 |
+
"video_token": "<|video_pad|>",
|
| 30 |
+
"vision_bos_token": "<|vision_start|>",
|
| 31 |
+
"vision_eos_token": "<|vision_end|>"
|
| 32 |
+
}
|