Qwen3-Coder-Next, Red Lite CF2 (19.3 GB)

A 2-bit-class GGUF of Qwen/Qwen3-Coder-Next with the same size as Bartowski's IQ2_XXS file, a 3.3 % lower perplexity on code, and 12 of 12 harder coding-agent tasks passed on a 24 GB Mac where Bartowski's file passed 7.

Made for DwarfStar Red Lite, a native Metal runtime for this model family on Apple Silicon. It is a standard GGUF and loads in llama.cpp too.

file size SHA-256
Qwen3-Coder-Next-RedLite-CF2.gguf 19.32 GB 17da099e1a7b0f5caff1536ef56f5f616a7e9bccc0216bd1030d1e0855092bac

What is different

The recipe of Red Lite F2 for Qwen3-Next-80B-A3B-Instruct, applied to the coding model:

  • Dense projections at Q4_K. The DeltaNet and attention input projections and part of the shared experts are IQ2_XXS (2.06 bits) in Bartowski's file. They are only about 290 MiB, but every token reads them.
  • IQ2_XS experts on layers 40โ€“47, IQ1_M elsewhere, to keep the size.
  • Every other tensor keeps the type of Bartowski's IQ2_XXS.

Quality, measured

Quantized from Bartowski's Q8_0 with his imatrix.gguf and the same llama.cpp. Perplexity at context 512 on a frozen text corpus (111 chunks) and a frozen code corpus (90 chunks, llama.cpp's llama-vocab.cpp).

Bartowski IQ2_XXS CE3 (expert layers moved) CF2
size 19.30 GB 19.30 GB 19.32 GB
perplexity, code 2.623 2.584 2.537 (โˆ’3.3 %, t = โˆ’4.0)
perplexity, text 18.53 18.24 18.72
agent tasks, M4 Pro 24 GiB 7 / 12 11 / 12 12 / 12
agent tasks, M4 Max 48 GiB 10 / 12 12 / 12 12 / 12
slowest agent task, M4 Pro 600 s (limit) 506 s 270 s
  • Agent tasks: the pi coding agent on a copy of the Red Lite repository, four tasks ร— three runs:

    • extend a CLI with a test;
    • explain part of a 1,900-line C server;
    • fix a planted bug without touching the tests;
    • a four-prompt session.

    Every expert was resident on the GPU and the context was 32K.

  • On prose CF2 is no better than Bartowski's file; it is a file for code and agents.

  • Temperature: at 0.7 and above the 2-bit Coder sometimes repeats a tool call until the time limit; at 0.3 it did not, in 12 sessions. Red Lite's setup-pi uses 0.3.

Records and method: dev63.

With Red Lite 0.9

Measured on an Apple M4 Max 48 GiB unless noted.

  • A 64K agent window matters more than the file. A harder six-task suite, three runs each, was run on the same Mac:
    • CF2 passed 16 / 18 with a 64K window and 10 / 18 with 32K, because these sessions reach 10โ€“30K tokens per request;
    • CF2 at 64K matched Bartowski's 3-bit Coder (16 / 18), which is 12 GB larger.
    • redlite setup-pi writes a 64K window.
  • The full 262K context is usable. Three passphrases hidden in a prompt of real code were found at 62K, 127K and 256K tokens. Decoding the answer ran at 48, 33 and 22 tok/s.
  • Prompt lookup. The server drafts the next token from the conversation and verifies it exactly, so the output is unchanged.
    • Coding-agent decode is 15โ€“25 % faster (M4 Max) and 28โ€“29 % faster (M4 Pro 24 GiB, every expert resident).
    • Rewriting a file: 80 โ†’ 109 tok/s on the M4 Max.
  • Details: dev65, dev70.

With Red Lite 0.9.2โ€“0.9.5

Decode kernels rewritten after profiling a real token (0.9.2), and Apple's GEMM for the dense prompt ingestion (0.9.3). Same file, faster runtime:

Red Lite 0.9.1 Red Lite 0.9.3
Decode, M4 Max, every expert resident 85 tok/s 99 tok/s
Decode, M4 Pro 24 GiB, every expert resident 48 tok/s 59 tok/s
Decode, M4 Pro 24 GiB, 4 GiB expert cache 32 tok/s 36 tok/s
Prompt ingestion, M4 Pro 24 GiB, 4 GiB cache / every expert resident 354 / 403 tok/s 377 / 447 tok/s
Prompt ingestion, M4 Max, 1.7K tokens 929 tok/s 1,070 tok/s
Rewriting a file with prompt lookup, M4 Max 109 tok/s 125 tok/s
  • Against MLX's 3-bit Qwen3-Coder-Next (32.5 GB), same Mac, same prompts, measured the same day: CF2 decodes at 99 against 87โ€“90 tok/s with half the memory (18.2 GiB against 36.8 GB). MLX still ingests prompts faster (1,435 against 1,070 tok/s).
  • Coding agents (0.9.4): when pi sends the model's last answer back re-rendered, the server returns to the end of the previous prompt instead of recomputing the whole conversation (20โ€“40K tokens, 30โ€“75 s each time on an M4 Max). The hard agent suite went from 5 to 2 full recomputes in two runs (the rest are pi's own compactions) and from 386 to 282 s of prompt processing per run. The answers are exactly those of a fresh server.
  • Output and quality are unchanged: greedy file rewrites are token-identical to 0.9.1, and engine perplexity moved 2.5623 โ†’ 2.5637 (code) / 16.658 โ†’ 16.655 (text).
  • Records: dev74, dev75, dev77.

Use

  • Red Lite (24 GiB and larger Apple Silicon Macs):

    redlite download coder
    redlite setup-pi --port 8080
    redlite serve --native ~/.redlite/models/Qwen3-Coder-Next-RedLite-CF2.gguf --context 65536 --port 8080
    pi --provider redlite --model qwen3-next-80b-a3b-redlite
    

    On a 48 GiB Mac every expert stays on the GPU. Prompt lookup is on in both cases.

    On a 24 GiB Mac (Red Lite 0.9.5):

    • run redlite gpu-limit --boot once (it asks for your password);
    • add --kv f16 to the serve line;
    • every expert then stays on the GPU at the 64K window too.

    On the M4 Pro, the hard agent suite decoded at 35โ€“51 tok/s with this setup, against 17โ€“25 with the experts streamed from the SSD; it passed 5 of 6 tasks in both runs, against 5โ€“6 of 6. The half KV cache's outputs can differ slightly from the float one. redlite gpu-limit --off undoes the limit.

  • llama.cpp: llama-server -m Qwen3-Coder-Next-RedLite-CF2.gguf --jinja. It needs a build with Qwen3-Next support.

Reproduce

python3 scripts/dev/quant_mix.py --like Qwen_Qwen3-Coder-Next-IQ2_XXS.gguf \
    --q8 Qwen_Qwen3-Coder-Next-Q8_0-00001-of-00003.gguf \
    --imatrix Qwen_Qwen3-Coder-Next-imatrix.gguf --iq2xs-layers 40-47 --dense-type q4_K --out CF2.gguf

The script is in the Red Lite repository; it calls llama-quantize with one type per tensor.

Credits and limits

  • Model: Qwen team, Apache-2.0.
  • Q8_0 source, importance matrix and the reference IQ2_XXS layout: bartowski.
  • Quantizer: llama.cpp.
  • Limits:
    • the agent suite has four tasks, so 12/12 against 10/12 is a small sample;
    • CE3 and CF2 both pass every task on the 48 GiB Mac;
    • perplexity was measured on one text corpus and one code corpus.
Downloads last month
654
GGUF
Model size
80B params
Architecture
qwen3next
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for alfodaniello/Qwen3-Coder-Next-RedLite-GGUF

Quantized
(115)
this model