Azemari β Ethiopian Multilingual Voice (TTS)
146 language tokens Β· 4 languages Β· 8 emotions Β· voice cloning Β· built in Addis Ababa
lab.et Β· voice.et Β· yakal.et
What is Azemari?
Named after the Azemari β Ethiopia's traditional improvising singer who can take any voice and any mood β Azemari is a fine-tuned IndexTTS-2.5 model that speaks Amharic, Afaan Oromoo, and Tigrinya with controlled emotion and voice cloning.
Built by lab.et, the AI lab of Ethiopia.
Capabilities
| Control | Token |
|---|---|
| Amharic | <|am|> |
| Afaan Oromoo | <|om|> |
| Tigrinya | <|ti|> |
| Shewa dialect | <|am|><|DIALECT_SHEWA|> |
| Emotion | 8-dim vector (happy, angry, sad, afraid, disgusted, melancholic, surprised, calm) |
| Voice | CAMPPlus speaker embedding from reference audio |
Usage
Everything needed for inference is in this repository β download it and run; no external model downloads required.
pip install -U "huggingface_hub[cli]"
hf download Lab-et/azemari --local-dir azemari
- Set up the IndexTTS-2.5 codebase
- Point it at the downloaded
azemari/directory as your checkpoint dir (it replaces the officialcheckpoints/) - Generate. Control language with
<|am|>,<|om|>,<|ti|>prefixes and dialects with<|DIALECT_SHEWA|>/<|DIALECT_GOJJAM|>
Repository contents
| File | What it is |
|---|---|
gpt.pth |
Azemari β the fine-tuned GPT (this is the model) |
config.yaml |
adapted config (60,513 text tokens, +2 language rows) |
model_v2.py, tokenizer_ethiopic.py |
adapted inference code + Ethiopian tokenizer |
codec.pth, s2mel.pth, feat1.pt, feat2.pt, wav2vec2bert_stats.pt |
official IndexTTS-2.5 inference components |
qwen0.6bemo4-merge/ |
official emotion model |
multilingual_zh_ja_yue_char_del.tiktoken |
base tokenizer (Ethiopic special tokens extend it) |
bigvgan_generator.pth |
BigVGAN vocoder (official IndexTTS component) |
The only runtime fetch is the CAMPPlus speaker encoder, which the official code path loads from ModelScope (iic/speech_campplus_sv_zh-cn_16k-common).
Samples
Listen to generated samples at Lab-et/azemari-samples β emotions, voice cloning, multi-language tests.
Training
Fine-tuned from the IndexTTS-2.5 base checkpoint on 41,371 clips of human ground-truth Ethiopian speech (605 hours β Amharic 61.6k, Afaan Oromoo 40.3k, Tigrinya 40.0k), with language-specific embedding rows added (109 total, +2 for OM/TI) and 5 dialect/emotion tokens extended (60,514 total, +4).
- Training: 58,000 steps, batch duration 1600s/update, learning rate 3.2e-3 β warmup β cosine
- Post-training: language/token embedding-only refinement pass (backbone frozen)
- Validation loss: 5.018 (step 3,000) β ~3.0 (step 58,000)
License
Inherited from IndexTeam/IndexTTS-2.5 β bilibili Model Use License Agreement.
About lab.et
lab.et is the AI lab of Ethiopia β building speech AI, models and infrastructure for Ethiopian languages in Addis Ababa. Our products: voice.et (speech APIs) and yakal.et (the AI console, prepaid in Birr).
Open data: amharic-corpus β 408.5 hours of Amharic speech.
Citation
@model{labet_azemari,
title = {Azemari β Ethiopian Multilingual Voice (TTS)},
author = {lab.et},
year = {2026},
url = {https://huggingface.co/Lab-et/azemari},
note = {Fine-tuned from IndexTTS-2.5}
}
- Downloads last month
- 24