[audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2× faster, PocketTTS 2.2× faster on CPU, and WebUI generation history feature

Wait 5 sec.

Hi all, a bunch of performance improvements have been landed in audio.cpp. The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM, a 48% reduction in peak memory usage compared to the previous implementation. Thanks to https://github.com/mirek190 We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU. No compromises in parity and correctness. Here's a summary of the improvements: Model Peak memory reduction Speedup Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA ACE-Step family 6–7% VRAM 1.06–1.08× CUDA, 1.16–1.20× Vulkan MOSS-TTS v1.5 cloning 21% VRAM 1.05× CUDA MOSS-TTSD Q8 cloning 11% VRAM 1.04× CUDA Echo-TTS (Memory Saver) 20% VRAM — Qwen3-TTS 16–20% VRAM — IndexTTS2 / 2.5 12% VRAM — HTDemucs — 2.21× CUDA, 1.95× Vulkan HTDemucs six-stem — 1.99× CUDA PocketTTS 9% RAM 2.23× CPU They're runtime-level optimizations that make existing models more practical to run locally. The WebUI now includes an experimental generation history feature that lets you revisit previous outputs and restore their settings. audio.cpp now supports 110+ audio model families and 190+ variants (and counting)! We're continuing to improve memory efficiency and inference speed across CUDA, Vulkan, Metal, AMD/HIP, and CPU. The next release will bring even more optimizations! We're also looking for contributors to help improve the audio.cpp WebUI. With so many models and features now supported, we'd love some help making the UI more polished, intuitive, and enjoyable to use. If you're interested in frontend development or UI/UX design, contributions are very welcome! Thanks to everyone contributing improvements, testing builds, and reporting issues. Curious how these changes work on your setup! &#32; submitted by &#32; /u/Acceptable-Cycle4645 [link] &#32; [comments]