Qwen3-TTS 0.6B Base on the Apple Neural Engine
This is a Core ML build of Qwen/Qwen3-TTS-12Hz-0.6B-Base, the voice-cloning model. It clones a voice from a few seconds of reference audio plus its transcript, and runs it on the Neural Engine of an iPhone, iPad or Mac. It needs no GPU, MLX or ONNX Runtime.
It's built for Gloam Voice Studio and its radio app by Tiny Trash Labs.
The Swift runtime that loads it is the open-source QwenANE target in
TinyTrashLabs/gloam-voice-studio
(see Sources/QwenANE/README.md).
What's in it
| Path | What it is | Runs on |
|---|---|---|
coreml/talker0.mlmodelc, coreml/talker1.mlmodelc |
The talker transformer, as two stateful chunks (KV cache inside Core ML) | Neural Engine |
coreml/cp_ane.mlmodelc |
The code predictor (15 sub-codes per frame) | Neural Engine |
coreml/upF.mlmodelc, coreml/upMall.mlmodelc |
The vocoder upsampler | Neural Engine |
coreml/QwenSpeechEncoder.mlmodelc |
The reference audio → codec codes encoder, for on-device voice prep | Neural Engine |
coreml/QwenSpeakerEncoder.mlmodelc |
The reference audio → speaker embedding encoder, for on-device voice prep | Neural Engine |
host/ |
Tokenizer, config, the text embedding and projection, the codec embeddings | CPU (Accelerate) |
vochead/ |
The vocoder head (pre-transformer and quantizer) | CPU (Accelerate) |
Requires iOS 18 / macOS 15 or later (stateful Core ML models). The set is about 1.9 GB.
What's different from the other Core ML ports
- Base, not CustomVoice. It clones any voice rather than offering preset speakers, and both encoders are included so a new voice can be prepared on the device.
- One read is one performance.
QwenTalkSession/renderBreaksplit long text and carry each part's text and codec frames into the next part's prompt. They draw the whole read from one sampler stream and redraw a derailed take, so pace and tone don't reset at every split. - Exact sampler. It reproduces numpy's
default_rng(seed), so a seeded render is reproducible and comparable to the Python reference. - Streaming. Decoded audio arrives in 0.96 s chunks, and the first chunk can be shortened for lower latency.
Speed
In our tests it renders at about real time on an iPhone 15 Pro, at roughly 5% CPU. It also renders in the background with the screen locked. Your numbers will vary by device and thermal state.
Use
import QwenANE
let engine = try QwenANEEngine(modelsDirectory: modelsURL) // the downloaded folder
let voice = try QwenVoicePrep.prepare(referenceWAV: wavData, transcript: "exactly what the clip says",
modelsDirectory: modelsURL)
let session = QwenTalkSession(engine: engine, voice: voice, language: "en", seedText: text)
let read = try session.renderBreak(text, onPart: { part in print(part.logLine) })
// read.samples: mono Float at read.sampleRate (24 kHz)
License and attribution
The original model is © the Qwen team (Alibaba Cloud) and licensed under the
Apache License 2.0. This repository is a
modified redistribution under the same license (see LICENSE). Changes from the original:
- converted to Core ML (fp16, stateful talker split into two chunks)
- the vocoder split between Core ML (upsampler) and CPU arrays (head)
- the text embedding stored quantized
The model weights were not fine-tuned.
Only clone your own voice, or a voice you have permission to use.
Model tree for tinytrashlabs/Qwen3-TTS-0.6B-Base-ANE
Base model
Qwen/Qwen3-TTS-12Hz-0.6B-Base