Instructions to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF # Run inference directly in the terminal: llama cli -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF # Run inference directly in the terminal: llama cli -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF # Run inference directly in the terminal: ./llama-cli -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Use Docker
docker model run hf.co/alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
- LM Studio
- Jan
- vLLM
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alfodaniello/Qwen3-Coder-Next-RedLite-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alfodaniello/Qwen3-Coder-Next-RedLite-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
- Ollama
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with Ollama:
ollama run hf.co/alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
- Unsloth Desktop
- Pi
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "alfodaniello/Qwen3-Coder-Next-RedLite-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with Docker Model Runner:
docker model run hf.co/alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
- Lemonade
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Run and chat with the model
lemonade run user.Qwen3-Coder-Next-RedLite-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use alfodaniello/Qwen3-Coder-Next-RedLite-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "alfodaniello/Qwen3-Coder-Next-RedLite-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3-Coder-Next, Red Lite CF2 (19.3 GB)
A 2-bit-class GGUF of Qwen/Qwen3-Coder-Next with the same size as Bartowski's IQ2_XXS file, a 3.3 % lower perplexity on code, and 12 of 12 harder coding-agent tasks passed on a 24 GB Mac where Bartowski's file passed 7.
Made for DwarfStar Red Lite, a native Metal runtime for this model family on Apple Silicon. It is a standard GGUF and loads in llama.cpp too.
| file | size | SHA-256 |
|---|---|---|
Qwen3-Coder-Next-RedLite-CF2.gguf |
19.32 GB | 17da099e1a7b0f5caff1536ef56f5f616a7e9bccc0216bd1030d1e0855092bac |
What is different
The recipe of Red Lite F2 for Qwen3-Next-80B-A3B-Instruct, applied to the coding model:
- Dense projections at Q4_K. The DeltaNet and attention input projections and part of the shared experts are IQ2_XXS (2.06 bits) in Bartowski's file. They are only about 290 MiB, but every token reads them.
- IQ2_XS experts on layers 40โ47, IQ1_M elsewhere, to keep the size.
- Every other tensor keeps the type of Bartowski's IQ2_XXS.
Quality, measured
Quantized from Bartowski's Q8_0 with his imatrix.gguf and the same llama.cpp. Perplexity at context 512 on a frozen
text corpus (111 chunks) and a frozen code corpus (90 chunks, llama.cpp's llama-vocab.cpp).
| Bartowski IQ2_XXS | CE3 (expert layers moved) | CF2 | |
|---|---|---|---|
| size | 19.30 GB | 19.30 GB | 19.32 GB |
| perplexity, code | 2.623 | 2.584 | 2.537 (โ3.3 %, t = โ4.0) |
| perplexity, text | 18.53 | 18.24 | 18.72 |
| agent tasks, M4 Pro 24 GiB | 7 / 12 | 11 / 12 | 12 / 12 |
| agent tasks, M4 Max 48 GiB | 10 / 12 | 12 / 12 | 12 / 12 |
| slowest agent task, M4 Pro | 600 s (limit) | 506 s | 270 s |
Agent tasks: the pi coding agent on a copy of the Red Lite repository, four tasks ร three runs:
- extend a CLI with a test;
- explain part of a 1,900-line C server;
- fix a planted bug without touching the tests;
- a four-prompt session.
Every expert was resident on the GPU and the context was 32K.
On prose CF2 is no better than Bartowski's file; it is a file for code and agents.
Temperature: at 0.7 and above the 2-bit Coder sometimes repeats a tool call until the time limit; at 0.3 it did not, in 12 sessions. Red Lite's
setup-piuses 0.3.
Records and method: dev63.
With Red Lite 0.9
Measured on an Apple M4 Max 48 GiB unless noted.
- A 64K agent window matters more than the file. A harder six-task suite, three runs each, was run on the same Mac:
- CF2 passed 16 / 18 with a 64K window and 10 / 18 with 32K, because these sessions reach 10โ30K tokens per request;
- CF2 at 64K matched Bartowski's 3-bit Coder (16 / 18), which is 12 GB larger.
redlite setup-piwrites a 64K window.
- The full 262K context is usable. Three passphrases hidden in a prompt of real code were found at 62K, 127K and 256K tokens. Decoding the answer ran at 48, 33 and 22 tok/s.
- Prompt lookup. The server drafts the next token from the conversation and verifies it exactly, so the output is
unchanged.
- Coding-agent decode is 15โ25 % faster (M4 Max) and 28โ29 % faster (M4 Pro 24 GiB, every expert resident).
- Rewriting a file: 80 โ 109 tok/s on the M4 Max.
- Details: dev65, dev70.
With Red Lite 0.9.2โ0.9.5
Decode kernels rewritten after profiling a real token (0.9.2), and Apple's GEMM for the dense prompt ingestion (0.9.3). Same file, faster runtime:
| Red Lite 0.9.1 | Red Lite 0.9.3 | |
|---|---|---|
| Decode, M4 Max, every expert resident | 85 tok/s | 99 tok/s |
| Decode, M4 Pro 24 GiB, every expert resident | 48 tok/s | 59 tok/s |
| Decode, M4 Pro 24 GiB, 4 GiB expert cache | 32 tok/s | 36 tok/s |
| Prompt ingestion, M4 Pro 24 GiB, 4 GiB cache / every expert resident | 354 / 403 tok/s | 377 / 447 tok/s |
| Prompt ingestion, M4 Max, 1.7K tokens | 929 tok/s | 1,070 tok/s |
| Rewriting a file with prompt lookup, M4 Max | 109 tok/s | 125 tok/s |
- Against MLX's 3-bit Qwen3-Coder-Next (32.5 GB), same Mac, same prompts, measured the same day: CF2 decodes at 99 against 87โ90 tok/s with half the memory (18.2 GiB against 36.8 GB). MLX still ingests prompts faster (1,435 against 1,070 tok/s).
- Coding agents (0.9.4): when pi sends the model's last answer back re-rendered, the server returns to the end of the previous prompt instead of recomputing the whole conversation (20โ40K tokens, 30โ75 s each time on an M4 Max). The hard agent suite went from 5 to 2 full recomputes in two runs (the rest are pi's own compactions) and from 386 to 282 s of prompt processing per run. The answers are exactly those of a fresh server.
- Output and quality are unchanged: greedy file rewrites are token-identical to 0.9.1, and engine perplexity moved 2.5623 โ 2.5637 (code) / 16.658 โ 16.655 (text).
- Records: dev74, dev75, dev77.
Use
Red Lite (24 GiB and larger Apple Silicon Macs):
redlite download coder redlite setup-pi --port 8080 redlite serve --native ~/.redlite/models/Qwen3-Coder-Next-RedLite-CF2.gguf --context 65536 --port 8080 pi --provider redlite --model qwen3-next-80b-a3b-redliteOn a 48 GiB Mac every expert stays on the GPU. Prompt lookup is on in both cases.
On a 24 GiB Mac (Red Lite 0.9.5):
- run
redlite gpu-limit --bootonce (it asks for your password); - add
--kv f16to theserveline; - every expert then stays on the GPU at the 64K window too.
On the M4 Pro, the hard agent suite decoded at 35โ51 tok/s with this setup, against 17โ25 with the experts streamed from the SSD; it passed 5 of 6 tasks in both runs, against 5โ6 of 6. The half KV cache's outputs can differ slightly from the float one.
redlite gpu-limit --offundoes the limit.- run
llama.cpp:
llama-server -m Qwen3-Coder-Next-RedLite-CF2.gguf --jinja. It needs a build with Qwen3-Next support.
Reproduce
python3 scripts/dev/quant_mix.py --like Qwen_Qwen3-Coder-Next-IQ2_XXS.gguf \
--q8 Qwen_Qwen3-Coder-Next-Q8_0-00001-of-00003.gguf \
--imatrix Qwen_Qwen3-Coder-Next-imatrix.gguf --iq2xs-layers 40-47 --dense-type q4_K --out CF2.gguf
The script is in the Red Lite repository; it calls llama-quantize with one type per tensor.
Credits and limits
- Model: Qwen team, Apache-2.0.
- Q8_0 source, importance matrix and the reference IQ2_XXS layout: bartowski.
- Quantizer: llama.cpp.
- Limits:
- the agent suite has four tasks, so 12/12 against 10/12 is a small sample;
- CE3 and CF2 both pass every task on the 48 GiB Mac;
- perplexity was measured on one text corpus and one code corpus.
- Downloads last month
- 654
We're not able to determine the quantization variants.
Model tree for alfodaniello/Qwen3-Coder-Next-RedLite-GGUF
Base model
Qwen/Qwen3-Coder-Next