Follow-up to last month's post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that. Reduced roughness and break-up on some of the voices that struggled before Same 120M model, same speed (~300 ms to first audio on a laptop CPU) Apache-2.0 English, European Portuguese, French, German More languages are planned More control over the generated voice is also planned Still struggles with very high-pitched or cartoon-like voices, noisy reference audio, and some unusual OOD voices. We're continuing to improve those cases. If you want to contribute and help, PM me with the samples that failed. If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you. Run it locally: uvx --from sopro soprotts serve Repo: https://github.com/samuel-vitorino/sopro Weights: https://huggingface.co/samuel-vitorino/sopro-v2-turbo In-browser demo (desktop): https://samuel-vitorino.github.io/sopro/ Space: https://huggingface.co/spaces/samuel-vitorino/sopro-v2-turbo-tts Blog (evals, samples, how it works): https://research.haloneuro.ai/posts/sopro-v2 Video: six voices, ~5 seconds of reference audio each, followed by a generated line. https://reddit.com/link/1wwrw0v/video/yb63ar836ath1/player   submitted by   /u/SammyDaBeast [link]   [comments]