I want to share a TTS server with an OpenAI-compatible API that generates speech really fast (about 0.3 seconds for a sentence on an RTX 3080) and can clone a voice from a short reference clip. I’ve optimized the server so generation starts quickly, and it processes long text in sequence, paragraph by paragraph. I use it to turn school books into audiobooks in my own voice, so I can listen to them while driving. Out of the box, it’s already tuned for the best settings, but you can change them however you like, for example, the CFG (guidance) scale. Here are the links to the repo and to a page that showcases it, where you can listen to all the voices. As always, it’s open source and free for anyone to use and modify. Repo: https://github.com/hypersniper05/open-omnivoice-tts Page: https://hypersniper05.github.io/open-omnivoice-tts/   submitted by   /u/Delicious-Farmer-234 [link]   [comments]