Full disclosure: I work on ViiTorVoice. We recently evaluated our open-source NAR TTS model on the Seed-TTS test sets. The current results were: English WER: 1.30 Chinese WER: 0.99 English speaker similarity: 0.734 https://preview.redd.it/hssoah04fgrh1.png?width=2904&format=png&auto=webp&s=31b8b0ed55d58075d43bc4d91c92dd67b2012ecd I’ve attached the comparison table for context. The numbers look encouraging, but a benchmark only covers a narrow and relatively clean slice of real-world usage. It doesn’t tell us enough about mixed-language text, unusual names, numbers, abbreviations, long-form stability or generally cursed inputs. We’d like to build a more practical test set around the cases people actually struggle with. If you use local TTS, what would you test next? Some examples: mixed Chinese and English names and obscure place names numbers, acronyms and CLI commands tongue twisters long-form narration strange punctuation or internet slang If you have a sentence that regularly breaks TTS systems, please share it. Failed cases are more useful to us than polite benchmark wins. Open-source model and inference code: https://github.com/viitor-ai/viitor-voice-nar   submitted by   /u/Important_Drag_6890 [link]   [comments]