Kyutai Text-to-Speech
[
](https://huggingface.co/collections/kyutai/text-to-speech-6866192e7e004ed04fd39e29) [
More details can be found on the project page.
We provide different implementations of Kyutai TTS for different use cases. Here is how to choose which one to use:
- PyTorch: for research and tinkering. If you want to call the model from Python for research or experimentation, use our PyTorch implementation.
- Rust: for production. If you want to serve Kyutai TTS in a production setting, use our Rust server. Our robust Rust server provides streaming access to the model over websockets. We use this server to run Unmute.
- MLX: for on-device inference on iPhone and Mac. MLX is Apple's ML framework that allows you to use hardware acceleration on Apple silicon. If you want to run the model on a Mac or an iPhone, choose the MLX implementation.
PyTorch implementation
[
Check out our Colab notebook or use the script:
# From stdin, plays audio immediately
echo "Hey, how are you?" | python scripts/tts_pytorch.py - -
# From text file to audio file
python scripts/tts_pytorch.py text_to_say.txt audio_output.wavThe tts_pytorch.py script waits for all the text to be available before
starting the audio generation. A fully streaming implementation is available in
the tts_pytorch_streaming.py script, which can be used as follows:
echo "Hey, how are you?" | python scripts/tts_pytorch_streaming.py audio_output.wavThis requires the moshi package, which can be installed via pip.
If you have uv installed, you can skip the installation step
and just prefix the command above with uvx --with moshi.
Rust server
The Rust implementation provides a server that can process multiple streaming queries in parallel.
Installing the Rust server is a bit tricky because it uses our Python implementation under the hood,
which also requires installing the Python dependencies.
Use the start_tts.sh script to properly install the Rust server.
If you already installed the moshi-server crate before and it's not working, you might need to force a reinstall by running cargo uninstall moshi-server first.
Feel free to open an issue if the installation is still broken.
Once installed, the server can be started via the following command using the config file from this repository.
moshi-server worker --config configs/config-tts.tomlOnce the server has started you can connect to it using our script as follows:
# From stdin, plays audio immediately
echo "Hey, how are you?" | python scripts/tts_rust_server.py - -
# From text file to audio file
python scripts/tts_rust_server.py text_to_say.txt audio_output.wavYou can configure the server by modifying configs/config-tts.toml. See comments in that file to see what options are available.
MLX implementation
MLX is Apple's ML framework that allows you to use hardware acceleration on Apple silicon.
Use our example script to run Kyutai TTS on MLX.
The script takes text from stdin or a file and can output to a file or stream the resulting audio.
When streaming the output, if the model is not fast enough to keep with
real-time, you can use the --quantize 8 or --quantize 4 flags to quantize
the model resulting in faster inference.
# From stdin, plays audio immediately
echo "Hey, how are you?" | python scripts/tts_mlx.py - - --quantize 8
# From text file to audio file
python scripts/tts_mlx.py text_to_say.txt audio_output.wavThis requires the moshi-mlx package, which can be installed via pip.
If you have uv installed, you can skip the installation step
and just prefix the command above with uvx --with moshi-mlx.
Source captured: 2026-10-11