SupertonicTTS-3 β€” LiteRT (.tflite, Android / Qualcomm NPU)

First-party LiteRT export of Supertonic-3's four non-autoregressive flow-matching graphs. Built by our own pipeline (speech-models/stmodels): weights lifted from the Supertone/supertonic-3 ONNX initializers → PyTorch nn.Module → litert_torch.convert (torch.export → StableHLO → TFLite). This avoids the onnx2tf NCHW/ConvNeXt layout failures that block direct ONNX→TFLite for this model.

Graphs & parity (FP32, vs ONNX Runtime)

Module tflite parity max|Ξ”|
duration_predictor.tflite 3.4 MB 4.1e-05 βœ“
vector_estimator.tflite (ODE denoiser) 244 MB 5.6e-03 βœ“
vocoder.tflite 97 MB 2.6e-04 βœ“
text_encoder.tflite 34 MB 1.1e-01 (localized; mean ~2.5e-4) ⚠️
vector_estimator_L128.tflite (L=128 bucket) 244 MB 1.7e-02 (relative 2.7e-3) βœ“
vocoder_L128.tflite (L=128 bucket) 97 MB 2.7e-04 βœ“

Shapes are fixed (T=128 text tokens; latent length L per graph) β€” litert_torch cannot lower the relpos attention + ConvNeXt stack with a symbolic L, so the latent axis ships in buckets: the base pair is L=64 (β‰ˆ4.5 s of audio), and *_L128.tflite is the same two L-dependent graphs exported at L=128 (β‰ˆ9 s). Pick the smallest bucket whose window holds a piece's predicted duration β€” short sentences run on the cheap L=64 pair, a long sentence is generated in one pass on L=128 instead of being split (speech-core's LiteRTSupertonicTts does this automatically and loads the L=128 pair on first use). The host runs the flow-matching ODE loop (vector_estimator Γ—total_steps). Assets to drive them: tts.json, unicode_indexer.json (G2P-free tokenizer), voice_styles/*.json.

Graph inputs are bound by tensor name (serving_default_args_N, N = signature position): the converter permutes input slots differently per toolchain version, so do not rely on slot order.

Running on Android / Qualcomm NPU

  • CPU/GPU: LiteRT (ai_edge_litert / TFLite) interpreter with XNNPACK/GPU delegate.
  • Qualcomm HTP/NPU: the LiteRT QNN delegate at runtime, or compile to a QNN context binary via Qualcomm AI Hub (qai_hub) from these graphs (static shapes are HTP-friendly). int8/int4 PTQ via ai-edge-quantizer for full HTP residency is a follow-up.

Attribution & license

Other Supertonic-3 formats

Ecosystem

  • soniqo.audio β€” website / use-case explorer (transcription, voice cloning, live ASR, voice agents).
  • speech-core β€” C++ orchestration library; Supertonic plugs in as a TTSInterface LiteRT model.
  • speech-swift β€” Apple Silicon MLX + CoreML runtime.
  • speech-android β€” Android SDK consuming on-device LiteRT bundles.

Other LiteRT models in this collection

Downloads last month
427
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for soniqo/Supertonic-3-LiteRT

Finetuned
(12)
this model

Collection including soniqo/Supertonic-3-LiteRT

Paper for soniqo/Supertonic-3-LiteRT