You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

This is a repackaging of kyutai/pocket-tts, which Kyutai distributes under these terms. Prohibited use: Use of our model must comply with all applicable laws and regulations and must not result in, involve, or facilitate any illegal, harmful, deceptive, fraudulent, or unauthorized activity. Prohibited uses include, without limitation, voice impersonation or cloning without explicit and lawful consent; misinformation, disinformation, or deception (including fake news, fraudulent calls, or presenting generated content as genuine recordings of real people or events); and the generation of unlawful, harmful, libelous, abusive, harassing, discriminatory, hateful, or privacy-invasive content. We disclaim all liability for any non-compliant use.

Log in or Sign Up to review the conditions and access this model content.

Pocket TTS (French, English) โ€” GGUF for pocket-tts-rs

Single-file GGUF conversions of kyutai/pocket-tts (~100M params) for pocket-tts-rs, a Rust/Candle port. Each file holds the weights (linear layers quantized), config, tokenizer, the voice-cloning encoder and the 27 predefined voices.

Access terms. Kyutai gates the original model behind the prohibited-use terms above (notably: no voice cloning without explicit consent). These files include the voice-cloning encoder, so the same terms apply here.

Not for llama.cpp. pocket-tts-rs layout.

file size note
pocket-tts-french-q8_0.gguf 233 MB
pocket-tts-french-q4k.gguf 184 MB same intelligibility
pocket-tts-english-q8_0.gguf 236 MB
pocket-tts-english-q4k.gguf 187 MB same intelligibility

Embedded voices: cosette, marius, javert, alba, jean, anna, vera, fantine, charles, paul, eponine, azelma, george, mary, jane, michael, eve, bill_boerst, peter_yearsley, stuart_bell, caro_davy, giovanni, lola, juergen, rafael, daan, estelle. They are Kyutai's precomputed voice states from kyutai/pocket-tts-without-voice-cloning (CC-BY-4.0); the source recordings and their licenses are listed in kyutai/tts-voices.

Quality and speed

Output at temperature 0 matches the Python reference (correlation 1.00000). Word error rate (Parakeet TDT 0.6B v3, 10 sentences ร— 2 generations per language):

system French WER English WER
Python reference 2.5% 3.2%
GGUF q8_0 4.7% 3.2%
GGUF q4k 4.2% 2.8%

Speed: ~7ร— real time on CPU (q8_0), ~28ร— on an RTX 4090 (680 MiB VRAM).

Usage

git clone https://github.com/rzafiamy/pocket-tts-rs && cd pocket-tts-rs   # build: see the README
pocket-tts generate -m pocket-tts-french-q8_0.gguf -v estelle -t "Bonjour tout le monde." -o out.wav
pocket-tts serve -m pocket-tts-french-q8_0.gguf --port 8000      # POST /v1/audio/speech

Also served by zallama (backend: pocket-tts-server).

Reproduce: pocket-tts convert --variant french --dtype q8_0 --voices all -o pocket-tts-french-q8_0.gguf.

License

CC-BY-4.0, as the original model by Kyutai, with Kyutai's prohibited-use terms above.

Downloads last month
-
GGUF
Model size
0.2B params
Architecture
pocket-tts
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rleo/pocket-tts-GGUF

Quantized
(51)
this model