Qwen3-TTS 0.6B Base on the Apple Neural Engine

This is a Core ML build of Qwen/Qwen3-TTS-12Hz-0.6B-Base, the voice-cloning model. It clones a voice from a few seconds of reference audio plus its transcript, and runs it on the Neural Engine of an iPhone, iPad or Mac. It needs no GPU, MLX or ONNX Runtime.

It's built for Gloam Voice Studio and its radio app by Tiny Trash Labs. The Swift runtime that loads it is the open-source QwenANE target in TinyTrashLabs/gloam-voice-studio (see Sources/QwenANE/README.md).

What's in it

Path What it is Runs on
coreml/talker0.mlmodelc, coreml/talker1.mlmodelc The talker transformer, as two stateful chunks (KV cache inside Core ML) Neural Engine
coreml/cp_ane.mlmodelc The code predictor (15 sub-codes per frame) Neural Engine
coreml/upF.mlmodelc, coreml/upMall.mlmodelc The vocoder upsampler Neural Engine
coreml/QwenSpeechEncoder.mlmodelc The reference audio → codec codes encoder, for on-device voice prep Neural Engine
coreml/QwenSpeakerEncoder.mlmodelc The reference audio → speaker embedding encoder, for on-device voice prep Neural Engine
host/ Tokenizer, config, the text embedding and projection, the codec embeddings CPU (Accelerate)
vochead/ The vocoder head (pre-transformer and quantizer) CPU (Accelerate)

Requires iOS 18 / macOS 15 or later (stateful Core ML models). The set is about 1.9 GB.

What's different from the other Core ML ports

  • Base, not CustomVoice. It clones any voice rather than offering preset speakers, and both encoders are included so a new voice can be prepared on the device.
  • One read is one performance. QwenTalkSession / renderBreak split long text and carry each part's text and codec frames into the next part's prompt. They draw the whole read from one sampler stream and redraw a derailed take, so pace and tone don't reset at every split.
  • Exact sampler. It reproduces numpy's default_rng(seed), so a seeded render is reproducible and comparable to the Python reference.
  • Streaming. Decoded audio arrives in 0.96 s chunks, and the first chunk can be shortened for lower latency.

Speed

In our tests it renders at about real time on an iPhone 15 Pro, at roughly 5% CPU. It also renders in the background with the screen locked. Your numbers will vary by device and thermal state.

Use

import QwenANE

let engine = try QwenANEEngine(modelsDirectory: modelsURL)      // the downloaded folder
let voice  = try QwenVoicePrep.prepare(referenceWAV: wavData, transcript: "exactly what the clip says",
                                       modelsDirectory: modelsURL)
let session = QwenTalkSession(engine: engine, voice: voice, language: "en", seedText: text)
let read = try session.renderBreak(text, onPart: { part in print(part.logLine) })
// read.samples: mono Float at read.sampleRate (24 kHz)

License and attribution

The original model is © the Qwen team (Alibaba Cloud) and licensed under the Apache License 2.0. This repository is a modified redistribution under the same license (see LICENSE). Changes from the original:

  • converted to Core ML (fp16, stateful talker split into two chunks)
  • the vocoder split between Core ML (upsampler) and CPU arrays (head)
  • the text embedding stored quantized

The model weights were not fine-tuned.

Only clone your own voice, or a voice you have permission to use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinytrashlabs/Qwen3-TTS-0.6B-Base-ANE

Finetuned
(31)
this model