We use Fishaudio already and are curious about this

#6
by YellowjacketGames - opened

We make 3d games and use fish for our voiceovers by using voice cloning. what does this directly improve atop the existing fishaudio workflow?

happy to try this on some higher end GPUs if i can understand what i'm comparing against

Thanks for the question! We are actually building on top of Fish’s DualAR architecture, which we think is a very strong design for voice cloning.
The main improvement is not a completely different workflow, but rather scaling efficiency: we focus on achieving comparable voice quality and cloning capability with a much smaller 0.6B model, while preserving the advantages of the DualAR architecture.
Compared with the existing FishAudio workflow using the 4.5B model, our goal is to provide a much lighter alternative with significantly lower memory usage, easier deployment, and faster inference, while keeping similar voice cloning quality.

Sign up or log in to comment