Instructions to use aghasalim/siba-az-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aghasalim/siba-az-tokenizer with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("aghasalim/siba-az-tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SIBA Azerbaijani tokenizer
A byte-level BPE tokenizer with 32,768 pieces, trained by SIBA in Baku on 471 million words of Azerbaijani text. Apache 2.0. It is the first component of Meristem Native, SIBA's from-scratch model line, and it is free for anyone to use.
It encodes Azerbaijani in about 2 pieces per word, a third fewer than Qwen3.5's tokenizer and a fifth fewer than OpenAI's gpt-oss tokenizer, with a vocabulary 6 to 7 times smaller than either. Fewer pieces per word means cheaper training, faster generation and more text in the same context window.
Pieces per Azerbaijani word (lower is better)
| Tokenizer | Held-out Azerbaijani text | azbench prompts | Vocabulary |
|---|---|---|---|
| SIBA | 1.98 | 2.00 | 32,768 |
| gpt-oss (OpenAI) | 2.52 | 2.68 | 200,019 |
| Qwen3.5 | 2.97 | 3.02 | 248,070 |
| Qwen2.5 | 3.53 | 3.49 | 151,665 |
Held-out: 4,000 documents of our corpus (every 50th document) that were not used for training. azbench: the 2,148 prompts of azbench (Belebele, INCLUDE and TUMLU in Azerbaijani), which were not used for training either. Words are whitespace-separated tokens. The full numbers are in report.json.
Example: "Azərbaycan Respublikasının Konstitusiyası 1995-ci il noyabrın 12-də ümumxalq səsverməsi ilə qəbul edilmişdir."
- SIBA, 24 pieces:
Azərbaycan · Respublikasının · Konstitusiyası · 1 · 9 · 9 · 5 · -ci · il · noyabrın · 1 · 2 · -də · ümumxalq · səsver · məsi · ilə · qəbul · edilmişdir · . - Qwen3.5, 47 pieces:
Az · ərbay · can · Res · pub · lik · asının · Kon · stit · us · iy · ası · … · ü · m · um · x · al · q · … · q · ə · bul · edil · miş · dir · .
Use
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("aghasalim/siba-az-tokenizer")
ids = tok("Azərbaycan Respublikasının Konstitusiyası")["input_ids"]
print(len(ids), tok.convert_ids_to_tokens(ids))
Or with the tokenizers library alone: Tokenizer.from_pretrained("aghasalim/siba-az-tokenizer").
Details
- Model: byte-level BPE (every byte is in the base alphabet, so any text, Azerbaijani or not, round-trips losslessly), 32,768 entries, minimum merge frequency 3.
- Normalisation: Unicode NFC. Digits are split one per piece.
- Special tokens:
<|endoftext|>(0),<|im_start|>(1),<|im_end|>(2),<|pad|>(3). A ChatML-style chat template is included. - Training text: 471 million words, from SIBA's cleaned Azerbaijani corpus: Azerbaijani Wikipedia (CC BY-SA 4.0), Azerbaijani Wikisource (CC BY-SA 3.0 / GFDL), laws, decrees and orders from president.az and meclis.gov.az (official documents, not subject to copyright under Azerbaijan's Law on Copyright and Related Rights, art. 7), a 25 % sample of FinePDFs Azerbaijani and a 7 % sample of FineWeb-2 Azerbaijani (both ODC-By 1.0). The corpus was cleaned (NFC, Azerbaijani-only filter, exact deduplication, emails and phone numbers masked) before training.
- Trained with Hugging Face
tokenizers0.23 on a single CPU machine in Baku, in 19 minutes.
Limits
- It is built for Azerbaijani. Other languages still work (byte-level fallback) but take more pieces than in a tokenizer trained for them.
- It is a tokenizer, not a model: it has no knowledge of its own. No model uses it yet; Meristem Native will.
Licence
Apache License 2.0. The training data's sources are credited above.
Contact
Aghasalim Mustafazada, SIBA ("Süni İntellekt Biznes Avtomatlaşdırma" MMC) · siba.az · info@siba.az