Azemari β€” Ethiopian Multilingual Voice (TTS)

146 language tokens Β· 4 languages Β· 8 emotions Β· voice cloning Β· built in Addis Ababa

lab.et Β· voice.et Β· yakal.et


What is Azemari?

Named after the Azemari β€” Ethiopia's traditional improvising singer who can take any voice and any mood β€” Azemari is a fine-tuned IndexTTS-2.5 model that speaks Amharic, Afaan Oromoo, and Tigrinya with controlled emotion and voice cloning.

Built by lab.et, the AI lab of Ethiopia.

Capabilities

Control Token
Amharic <|am|>
Afaan Oromoo <|om|>
Tigrinya <|ti|>
Shewa dialect <|am|><|DIALECT_SHEWA|>
Emotion 8-dim vector (happy, angry, sad, afraid, disgusted, melancholic, surprised, calm)
Voice CAMPPlus speaker embedding from reference audio

Usage

Everything needed for inference is in this repository β€” download it and run; no external model downloads required.

pip install -U "huggingface_hub[cli]"
hf download Lab-et/azemari --local-dir azemari
  1. Set up the IndexTTS-2.5 codebase
  2. Point it at the downloaded azemari/ directory as your checkpoint dir (it replaces the official checkpoints/)
  3. Generate. Control language with <|am|>, <|om|>, <|ti|> prefixes and dialects with <|DIALECT_SHEWA|> / <|DIALECT_GOJJAM|>

Repository contents

File What it is
gpt.pth Azemari β€” the fine-tuned GPT (this is the model)
config.yaml adapted config (60,513 text tokens, +2 language rows)
model_v2.py, tokenizer_ethiopic.py adapted inference code + Ethiopian tokenizer
codec.pth, s2mel.pth, feat1.pt, feat2.pt, wav2vec2bert_stats.pt official IndexTTS-2.5 inference components
qwen0.6bemo4-merge/ official emotion model
multilingual_zh_ja_yue_char_del.tiktoken base tokenizer (Ethiopic special tokens extend it)
bigvgan_generator.pth BigVGAN vocoder (official IndexTTS component)

The only runtime fetch is the CAMPPlus speaker encoder, which the official code path loads from ModelScope (iic/speech_campplus_sv_zh-cn_16k-common).

Samples

Listen to generated samples at Lab-et/azemari-samples β€” emotions, voice cloning, multi-language tests.

Training

Fine-tuned from the IndexTTS-2.5 base checkpoint on 41,371 clips of human ground-truth Ethiopian speech (605 hours β€” Amharic 61.6k, Afaan Oromoo 40.3k, Tigrinya 40.0k), with language-specific embedding rows added (109 total, +2 for OM/TI) and 5 dialect/emotion tokens extended (60,514 total, +4).

  • Training: 58,000 steps, batch duration 1600s/update, learning rate 3.2e-3 β†’ warmup β†’ cosine
  • Post-training: language/token embedding-only refinement pass (backbone frozen)
  • Validation loss: 5.018 (step 3,000) β†’ ~3.0 (step 58,000)

License

Inherited from IndexTeam/IndexTTS-2.5 β€” bilibili Model Use License Agreement.

About lab.et

lab.et is the AI lab of Ethiopia β€” building speech AI, models and infrastructure for Ethiopian languages in Addis Ababa. Our products: voice.et (speech APIs) and yakal.et (the AI console, prepaid in Birr).

Open data: amharic-corpus β€” 408.5 hours of Amharic speech.

Citation

@model{labet_azemari,
  title  = {Azemari β€” Ethiopian Multilingual Voice (TTS)},
  author = {lab.et},
  year   = {2026},
  url    = {https://huggingface.co/Lab-et/azemari},
  note   = {Fine-tuned from IndexTTS-2.5}
}
Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support