Repository navigation
Pocket-TTS - options to speed up generation? #799
|
I'm doing some testing to see how I can speed up generation with Pocket-TTS as it seems rather slow. Depending on what voice I choose I'm seeing around 104ms (alba) to 85ms (vera) getting logged for pocket_tts.mimi.streaming_decoder_step_ms. Pocket-TTS' documentation highly suggests CPU only and that it only makes use of two threads. I am using the CPU build for testing and I have not seen any difference in generation speed by exposing 4 threads to audio.cpp. Per their documentation using the --quantize option: "Use int8 quantization for the model (default: False). This can reduce memory usage and increase speed, with minimal impact on audio quality." Is there a way to pass this option on via a server.json config? I've attempted a few settings but have not seen any changes so I likely don't have the right settings or it's not an option in audiocpp. Anyone have any suggestions? I do want to try using the vulkan build and see if there's a noticeable boost from running on the igpu and just try to run Pocket-TTS natively to see if there's any change. The system I am running this on has a Ryzen 7 8745HS with 5600mhz modules. Not blazing fast but not a slouch, especially considering their claim of fast generation on pure CPU. Some logging traces:
Startup:
|
Replies: 2 comments 4 replies
|
@rustysixslinger Thanks for reporting the issue. The PocketTTS CPU path is less optimized, as the primary focus has been on CUDA. I made some changes to improve CPU efficiency, and it's now 2–3x faster in my setup. Please try the latest commit or next release. Also please check audio.cpp doc for PocketTTS for support options. https://github.com/0xShug0/audio.cpp/blob/main/docs/tts.md#pockettts We tried to normalize option names instead of using those from the Python implementations, as their naming conventions can vary widely.
|
|
Came across an issue with this last build but it could be a problem on my side or possibly something with the change made. I'm having a hard time pinning it down. In case it may be of help, I am running audiocpp in a Debian LXC container in Incus on an AMD system. The portable build (audio-v0.9.1-bin-ubuntu-x64-cpu-portable.tar.gz) works flawlessly and the speed increase is insane. Zero pauses in audio now when streaming is enabled. Thank you! Went from mid 80ms to 15-16ms with just two threads enabled ([TIMING ts=20261008-133931] pocket_tts.mimi.streaming_decoder_step_ms 16.110693) I have historically been using the non-portable version without issue until the 091 update. Getting a crash now on the non-portable so I decided to test the portable to verify if there was something else was going on. The non-portable version (audio-v0.9.1-bin-ubuntu-x64-cpu.tar.gz) launches and starts the web server and is accessible but crashes once a message is sent to. The crash happens in both the Web-UI when running tests there and with trying to access the API from another system. I did test Supertonic and it runs just fine. I can log this as an official issue if you like. Below is the output of the console log (running, loading, message and the crash). The main thing that stands out to me is "AMX is not ready to be used!" No issue at all with the portable version and I'll likely just use it going forward. Mainly just wanted to pass this on just in case there's some other underlying issue that could crop up elsewhere.
|
Thank you for the quick fix on, especially since you're focused on cuda. That's a good bump in performance. I'll wait for the latest release to test.
I had been poking around all of the documentation, looking for a specific pocket-tts doc and completely overlooked the general tts.md doc. Thank you for pointing it out and I'll use it for reference in the future.
One thing I'd suggest to add to the Pocket-TTS cloning section is the ability to output a .safetensors file from a wav. I was about ready to ask here if there was a way to do it and the answer was in audiocpp's --help. From testing it does shave a bit of time off of generation and getting dumped into embeddings allows the voice to …