Kokoro-82M for Glade
Model assets and runtime configuration for Kokoro-82M v1.0
in Glade. Accepts ordinary text or explicit phonemes and produces mono 24 kHz PCM.
The bundle contains all 54 original voices across American/British English,
Spanish, French, Hindi, Italian, Brazilian Portuguese, Japanese and Mandarin.
Voice/language inventory and runtime conventions are in metadata.json.
Text and execution
Language downloads are optional. languages.json declares exact file groups,
shared resources and sizes. The default US English download is approximately
210 MB; all languages and voices together are approximately 767 MB.
Adding a language reuses unchanged shared neural and frontend files. Japanese's
approximately 446 MB packed dictionary is downloaded only when Japanese is selected.
The frontend is part of each selected language. English ports Misaki's context processing, trained spaCy tagger and original American English BART fallback. Japanese uses Cutlet and BSD-licensed MeCab/UniDic; Mandarin uses native normalization, Jieba and pinyin transcription. The other language paths and British out-of-dictionary words use permissive pronunciation data and Charsiu inference. Those substitutions have their own pronunciation behavior.
English planning retains upstream's up-to-510-phoneme context, with 512 text positions and 1,280 acoustic frames. Utterance normalization excludes padding. The ordinary-text interface preserves the complete input and splits capacity failures before delivering audio. Explicit phonemes are also available for controls.
The text graph runs on ANE. Prosody, acoustic and waveform graphs use GPU;
bidirectional recurrence, harmonic synthesis and FFT/overlap-add run on CPU.
Four shared-weight .aimodel assets, CPU parameters, voice tables and frontend
data are included. Japanese dictionaries expand losslessly into a disposable
memory-mapped cache on first use. No compiled specialization cache is supplied.
Measured performance
Complete reference text of JFK's “We choose to go to the Moon” speech: 2,201 words, Heart voice, speed 1, seed 42, one synthesis lane, resident execution. These are warmed Release text-to-audio medians of three measured passes. Phonemization, planning and all synthesis work are included; initial preparation, warmup and WAV writing are separate.
| Device | Generated audio | Text-to-audio | RTFx |
|---|---|---|---|
| M3 MacBook Air, 16 GB, macOS 27.0.1 | 12 min 13 s | 38.24 s | 19.2× |
| iPhone 15 Pro Max, A17 Pro, 8 GB, iOS 27.0.1 | 12 min 13 s | 96.34 s | 7.6× |
RTFx is generated audio seconds divided by generation seconds. The phone's first pass took 66.46 s (11.0×); later passes reached fair thermal state. Its sampled client peak was 1,159 MiB, including benchmark PCM. Compiler/service memory is excluded. Repeated phone outputs were byte-identical. Separate traces confirm ANE text and GPU prosody/acoustic/waveform execution.
Observed cached preparation was 0.62 s on each device, including native frontend initialization. Compiler caches were not purged; this is not a cold specialization measurement. Source assets require device specialization on first use.
Original frontend reference checks cover English context/numbers/fallback and chunk planning, Japanese and Mandarin transcription, and Charsiu numerical output. Phone functional checks cover all nine language profiles. These checks and sample recognition controls do not establish general perceptual or voice-similarity scores. Watch qualification is not claimed.
Attribution
Upstream weights/source: f3ff3571791e39611d31c381e3a41a3af07b4987, Apache-2.0.
Kokoro acknowledges StyleTTS2 and iSTFTNet. Misaki, trained frontend components,
pronunciation data and Japanese dictionaries retain their bundled license notices.
No additional training is performed. Download weights and configuration from one
repository snapshot.