A 1.5B Model That Writes E2E Tests — Fully Offline
Qwen2.5-Coder-1.5B-Instruct to turn a plain-English requirement into a runnable Playwright or Cypress test file — and it runs entirely on a laptop through Ollama. No API key, no page markup leaving the machine.
The interesting part isn't the fine-tune; SFT is well-trodden. It's where the data came from.
The data trained its own replacement
I already had a cloud-backed tool that generates E2E tests. So I let it seed the dataset: I fed it a categorized bank of requirements — logins, forms, dropdowns, tables, waits, e-commerce flows — run against three public demo sites (the-internet, demoqa, saucedemo).
Every generated file passed mechanical gates before it was allowed in: no invalid selectors, no hard-coded sleeps, no fabricated paths, valid JSON throughout. Anything that failed was dropped, not patched.
What survived: 202 clean pairs — 95 Playwright/TypeScript, 107 Cypress/JavaScript — stored as chat-format JSONL, ready for TRL's SFTTrainer.
The fine-tune
A LoRA SFT on a code-specialized base. Because the base already writes valid TS/JS, I only had to teach it the shape of the task — the house style, not programming. That's why 1.5B was enough.
One honest caveat: the tests aren't execution-verified yet. Expected values came from a model, not from running against live sites. An execution-pass gate is the obvious v2.
Running it locally
ollama run hf.co/aiqualitylab/ai-natural-language-tests
~1.6 GB, pulled once, then fully offline. For privacy-sensitive or air-gapped QA work, that's the whole point.
Takeaways
- A working cloud tool can bootstrap the data that trains its cheaper local replacement.
- 200 consistent examples beat a larger, messier pile.
- Start from a base that knows the domain — then you're only teaching format.
Everything's open:
- Model → aiqualitylab/ai-natural-language-tests
- Dataset → datasets/aiqualitylab/ai-natural-language-tests
- Source → github.com/aiqualitylab/ai-natural-language-tests
- Space → spaces/aiqualitylab/ai-natural-language-tests
Model → https://huggingface.co/ayan4m1/Qwen3.5-4B-E2E-Tests-GGUF , It's SFT on → https://huggingface.co/datasets/aiqualitylab/ai-natural-language-tests (202 requirement→test pairs)