Babbeldoos β€” Flemish LoRA adapters for Chatterbox

β–Ά Try it in the browser β€” Babbeldoos demo Space

Type Flemish, pick a voice or record your own, hear it back. The demo also plays pre-rendered examples that cost no GPU time, so you can hear what the adapters sound like before spending any.


Two LoRA adapters that shift Chatterbox Multilingual from Netherlandic Dutch towards Flemish (Belgian Dutch). Voice identity stays an inference-time argument β€” you supply a reference clip. These adapters change accent and register, not who is speaking.

adapter register trained on size
Conversational spontaneous Flemish 18.6 h Flemish broadcast speech 22.6 M params
Narration polite Belgian Dutch, read aloud 11.4 h Flemish broadcast speech 22.6 M params

Base model weights are not included or modified β€” the adapters are a 90 MB delta on t3.tfmr and Chatterbox is downloaded separately at load time.

Which one

Conversational reads the register off the text. Give it spontaneous Flemish and it drops word-final -n and -t the way many Flemish speakers do; give it a formal script and it tightens up on its own. Measured word-final consonant retention: 0.818 on conversational text against 0.948 on narration text β€” the same weights, a 0.13 swing.

Narration keeps those consonants in both registers (0.915 / 0.979) and is the more stable of the two across reference voices. Use it for anything read aloud.

Usage

from babbeldoos import Babbeldoos

tts = Babbeldoos.from_pretrained("Howest-AI-Lab/babbeldoos-flemish-tts")

wav = tts.generate("Goeiemiddag, waarmee kan ik u vandaag helpen?",
                   reference="your_voice.wav", model="Narration")
tts.save(wav, "out.wav")

Adapter names are case-insensitive β€” "Narration", "narration" and "NARRATION" all resolve to the same adapter. The repo directories are lowercase (conversational/, narration/); nothing depends on how you spell the name at the call site.

babbeldoos.py is vendored in this repo β€” Chatterbox's T3 has no adapter API, so the adapters attach by wrapping nn.Linear.forward on the modules they were trained for. Copy the file next to your code, or hf download the repo and add it to sys.path.

The reference clip

You must supply one, and none ships with this repo β€” see Reference voices below. 5 s of clean speech is enough. Chatterbox reads only the first 6 s for the T3 prompt and 10 s for the vocoder, so anything past ~10 s is ignored: a clean 8-second recording beats a noisy 30-second one.

Longer text

Chatterbox stops at about 1000 speech tokens (~40 s) and these adapters saw nothing over 15 s in training. read() splits at sentence boundaries, generates each piece against the same reference and joins them:

wav = tts.read(open("article.txt").read(), reference="your_voice.wav",
               model="Narration")
tts.save(wav, "article.wav")

Streaming

generate() samples every speech token before it decodes any audio, so you hear nothing until the whole utterance is done. generate_stream() yields audio while the sentence is still being generated, and read_stream() does the same for longer text, sentence by sentence, with the same pauses as read():

for piece in tts.read_stream(text, reference="your_voice.wav", model="Narration"):
    player.write(piece)          # float32 mono at tts.sr (24 kHz)

Measured on a DGX Spark (GB10), Narration adapter, warm model:

text generate(): audio after generate_stream(): first audio gap after the first piece
45 characters (~2.5 s of speech) 1.9 s 1.3 s 0.4 s
106 characters (~5 s) 3.8 s 1.05 s 0.4 s
155 characters (~8.5 s) 6.1 s 1.3 s 0.45 s

Time to first audio does not grow with the length of the text. The first piece is short (0.64 s), and the second takes slightly longer than that to arrive. From then on the stream stays ahead of playback. Buffer about 0.5 s before you start playing and the audio has no gaps, which is still well ahead of generate() for anything longer than a short sentence.

Quality. The same sampled tokens were decoded both ways: 4 voices Γ— 10 texts, short sentences and 10-second paragraphs:

decoding ASR word error (NeLF) defective clips speaker similarity
generate() 0.045 2% 0.886
generate_stream() 0.045 2% 0.888

The one defective clip is the same clip in both: it comes from the sampled tokens, not from how they were decoded. In a blind listening test the two could not be told apart.

How it works. Chatterbox's decoder was designed for chunked streaming, but the released package never streams. generate_stream() runs the token sampler as a generator. Every 40 speech tokens (1.6 s of audio; 20 for the first piece), it decodes the whole prefix again in a background thread that overlaps with sampling. Three things keep the pieces consistent with each other:

  • the flow gets the same noise on every pass;
  • the vocoder reuses the excitation it already emitted;
  • the newest 40 ms is held back and crossfaded into the next piece.

Knobs. Most of a decoding pass is spent on the reference prompt that is decoded along with it.

  • prompt_seconds (default 4.0) keeps only the last 4 s of the 10 s reference for the decoder. T3's voice conditioning is untouched. This makes each pass ~45% cheaper, and it is the setting measured above. prompt_seconds=None uses the full 10 s.
  • flow_steps (default 10, as upstream) can go down to 6 for another ~40%, but speaker similarity dropped slightly (0.876), so it is not the default.
  • first_chunk and chunk trade time to first audio against the number of decoding passes.

Things to know.

  • Loudness. save() matches a whole file to -19 LUFS, which a stream cannot do. For a consistent level, measure a fixed gain per voice once, on a calibration sentence generated with generate(), and apply it to every piece.
  • Watermark. The Perth watermark is applied to each piece.
  • Concurrency. Like generate(), one stream at a time per Babbeldoos instance.
  • Threads. Consume the generator from a single thread. In a web server, run it in a dedicated thread and pass the pieces over a queue. Don't hand a plain generator to a framework that may resume it on different worker threads: PyTorch's inference mode is per thread.
  • Chatterbox upgrades. The sampling loop is a copy of the one in chatterbox-tts 0.1.7, so re-check it before upgrading Chatterbox. In 0.1.7 the decoder's own streaming mode (finalize=False) crashes on a mask-length mismatch, and generate_stream() works around it.

Adapter strength

A LoRA is an additive low-rank delta, so scaling it walks the straight line between base and adapter β€” nothing retrained, nothing approximated:

tts.generate(text, reference="v.wav", model="Conversational", strength=0.6)

strength=0.0 reproduces the base model bit for bit; 1.0 is the adapter as trained. Lower values trade Flemish accent for crisper final consonants (retention rises monotonically: 0.874 at 1.0 β†’ 0.946 at 0.5 β†’ 0.966 at base).

We ship 1.0 because that is what listeners preferred. At 0.4 and 0.7 the Netherlandic accent starts coming back β€” the thing these adapters exist to remove β€” and the silences get less stable. The metric prefers low strength; the ear does not.

The deltas also compose, since they add:

tts.attach("Conversational", 0.7)
tts.add("Narration", 0.5)          # W + 0.7Β·d1 + 0.5Β·d2

Fitted independently, so a blend is worth listening to rather than assuming.

Reference voices

No reference clips are included in this repo, deliberately. The voices used during development were either synthesised by a commercial service, whose terms we will not extend to redistribution, or real broadcast speakers who never consented to having their voice cloned by strangers. Record five seconds of your own voice, or of someone who agreed to it.

The demo Space offers a handful of voices so you can try the adapters without recording anything first; it also takes an upload or a microphone recording.

Reference choice matters more than you would expect: across ten references on the same adapter, measured articulation spanned 0.084–0.094 β€” a bigger spread than between the two adapters on the same voice. It is worth trying several, which is quickest in the demo.

Training

LoRA r=32, Ξ±=64, dropout 0.05, on t3.tfmr q/k/v/o/gate/up/down projections (210 layers, 22.6 M parameters). 400 optimizer steps, batch 1 Γ— grad-accum 8, lr 1e-5, fp32 master weights. Audio is Flemish broadcast speech, segmented at sentence boundaries, transcribed with NeLF β€” a Flemish ASR, because general-purpose models normalise the Flemish away (Whisper rewrites goeiemiddag β†’ goedemiddag and da's β†’ dat is, which teaches the TTS model to do the same). Loudness-normalised, gender-balanced, language-filtered.

The corpus is not published.

Deploying on a Hugging Face Space (ZeroGPU)

Four things bit us getting the demo Space to run. All four produce error messages that point somewhere other than the cause, so they are written out here rather than left as folklore. A working configuration is in the demo Space.

Build the model at startup, not inside the GPU function. Constructing Chatterbox pulls ~2 GB of base weights, and the wall clock inside a @spaces.GPU(duration=...) call is budgeted β€” a first request that also has to download the model gets killed. Load on CPU while the app boots, move in per request:

tts = Babbeldoos.from_pretrained(REPO, device="cpu")   # at import

@spaces.GPU(duration=120)
def speak(...):
    tts.to("cuda")        # ~0.5 s; the adapters move with it
    return tts.generate(...)

to() re-attaches the adapter branches afterwards so their buffers follow the model. Moving back to CPU works too.

Install chatterbox-tts in a separate pip pass, and let a newer torch win. It pins torch==2.6.0, whose wheels carry kernels only up to sm_90. ZeroGPU now runs Blackwell (sm_120), so a generation dies with "no kernel image is available for execution on the device". sm_120 needs CUDA 12.8, first in torch 2.8's default wheel. Put chatterbox-tts alone in pre-requirements.txt and torch==2.8.0 / torchaudio==2.8.0 in requirements.txt: chatterbox is not part of the second resolve, so its pin becomes a warning rather than a constraint. pip's "requires torch==2.6.0, but you have 2.8.0" line is expected β€” these adapters are developed against torch 2.13.

Keep sdk_version equal to your gradio pin. Spaces installs gradio[oauth,mcp]==<sdk_version> from the Space README, and every chatterbox-tts release pins gradio with ==. If the two disagree, pip backtracks through older chatterbox releases and lands on 0.1.4, the last one depending on the unmaintained pkuseg β€” which ships no wheels and fails to build with "ModuleNotFoundError: No module named 'numpy'". The numpy error is a symptom of the gradio mismatch.

Pin setuptools<81. resemble-perth, which Chatterbox uses for its watermark, imports pkg_resources at module load. Newer setuptools removed it; perth catches the ImportError and sets its watermarker to None, and Chatterbox then calls None() β€” "TypeError: 'NoneType' object is not callable" while building the model.

Limitations

  • English proper nouns are unreliable. Trained on Flemish speech, these adapters mispronounce embedded English names β€” Piper β†’ "Peper", Whisper Large β†’ "wisper lage". Retrying with a different seed does not help: it is systematic, not stochastic.
  • Acronyms are not spelled out. TTS comes out as "bijt", T T S as "thees". Expand them in the text before synthesis (TTS β†’ text-to-speech).
  • Quality depends on the reference clip, sometimes a lot. Try several.
  • Dutch only. A LoRA can steer an accent; it cannot install a language.
  • Chatterbox embeds a Perth audio watermark in its output. That is base-model behaviour, not ours.
  • Not evaluated for, and not suitable for, speaker impersonation. Use references you have the right to use.
  • Long text has to be split at sentence boundaries: Chatterbox stops at 1000 speech tokens (40 s) and these adapters saw nothing over 15 s in training. read() and read_stream() do the splitting and joining.

Licence and status

cc-by-nc-4.0. The base model (Chatterbox, Resemble AI) is MIT, but the transcription model used to build the training corpus is CC-BY-NC, so non-commercial is the consistent choice downstream.

Built for the PWO Physical AI project at Howest (Kortrijk, Belgium). Version 1.1 β€” the same adapters as version 1, plus streaming (generate_stream(), read_stream()). A first step, not an endpoint.

Citation

@misc{babbeldoos2026,
  title  = {Babbeldoos: Flemish LoRA adapters for Chatterbox},
  author = {Howest AI Lab},
  year   = {2026},
  note   = {https://huggingface.co/Howest-AI-Lab/babbeldoos-flemish-tts}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Howest-AI-Lab/babbeldoos-flemish-tts

Adapter
(19)
this model

Space using Howest-AI-Lab/babbeldoos-flemish-tts 1