SIBA Azerbaijani tokenizer

A byte-level BPE tokenizer with 32,768 pieces, trained by SIBA in Baku on 471 million words of Azerbaijani text. Apache 2.0. It is the first component of Meristem Native, SIBA's from-scratch model line, and it is free for anyone to use.

It encodes Azerbaijani in about 2 pieces per word, a third fewer than Qwen3.5's tokenizer and a fifth fewer than OpenAI's gpt-oss tokenizer, with a vocabulary 6 to 7 times smaller than either. Fewer pieces per word means cheaper training, faster generation and more text in the same context window.

Pieces per Azerbaijani word (lower is better)

Tokenizer Held-out Azerbaijani text azbench prompts Vocabulary
SIBA 1.98 2.00 32,768
gpt-oss (OpenAI) 2.52 2.68 200,019
Qwen3.5 2.97 3.02 248,070
Qwen2.5 3.53 3.49 151,665

Held-out: 4,000 documents of our corpus (every 50th document) that were not used for training. azbench: the 2,148 prompts of azbench (Belebele, INCLUDE and TUMLU in Azerbaijani), which were not used for training either. Words are whitespace-separated tokens. The full numbers are in report.json.

Example: "Azərbaycan Respublikasının Konstitusiyası 1995-ci il noyabrın 12-də ümumxalq səsverməsi ilə qəbul edilmişdir."

  • SIBA, 24 pieces: Azərbaycan · Respublikasının · Konstitusiyası · 1 · 9 · 9 · 5 · -ci · il · noyabrın · 1 · 2 · -də · ümumxalq · səsver · məsi · ilə · qəbul · edilmişdir · .
  • Qwen3.5, 47 pieces: Az · ərbay · can · Res · pub · lik · asının · Kon · stit · us · iy · ası · … · ü · m · um · x · al · q · … · q · ə · bul · edil · miş · dir · .

Use

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("aghasalim/siba-az-tokenizer")
ids = tok("Azərbaycan Respublikasının Konstitusiyası")["input_ids"]
print(len(ids), tok.convert_ids_to_tokens(ids))

Or with the tokenizers library alone: Tokenizer.from_pretrained("aghasalim/siba-az-tokenizer").

Details

  • Model: byte-level BPE (every byte is in the base alphabet, so any text, Azerbaijani or not, round-trips losslessly), 32,768 entries, minimum merge frequency 3.
  • Normalisation: Unicode NFC. Digits are split one per piece.
  • Special tokens: <|endoftext|> (0), <|im_start|> (1), <|im_end|> (2), <|pad|> (3). A ChatML-style chat template is included.
  • Training text: 471 million words, from SIBA's cleaned Azerbaijani corpus: Azerbaijani Wikipedia (CC BY-SA 4.0), Azerbaijani Wikisource (CC BY-SA 3.0 / GFDL), laws, decrees and orders from president.az and meclis.gov.az (official documents, not subject to copyright under Azerbaijan's Law on Copyright and Related Rights, art. 7), a 25 % sample of FinePDFs Azerbaijani and a 7 % sample of FineWeb-2 Azerbaijani (both ODC-By 1.0). The corpus was cleaned (NFC, Azerbaijani-only filter, exact deduplication, emails and phone numbers masked) before training.
  • Trained with Hugging Face tokenizers 0.23 on a single CPU machine in Baku, in 19 minutes.

Limits

  • It is built for Azerbaijani. Other languages still work (byte-level fallback) but take more pieces than in a tokenizer trained for them.
  • It is a tokenizer, not a model: it has no knowledge of its own. No model uses it yet; Meristem Native will.

Licence

Apache License 2.0. The training data's sources are credited above.

Contact

Aghasalim Mustafazada, SIBA ("Süni İntellekt Biznes Avtomatlaşdırma" MMC) · siba.az · info@siba.az

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support