Iris-300M

Iris-300M is a 290.6M-parameter encoder-decoder Transformer for machine translation between 11 European languages, trained from scratch on a single RTX 3070. It translates directly between any pair of supported languages (no English pivot).

Part of the Iris model family by BranchingNLP (Michelangelo Di Nicola).

Supported languages

  • en English
  • it Italian
  • fr French
  • es Spanish
  • pt Portuguese
  • de German
  • nl Dutch
  • ro Romanian
  • pl Polish
  • cs Czech
  • sv Swedish

Architecture

Parameters 290,641,920
Encoder / decoder layers 10 / 10
d_model / heads / FFN 1024 / 16 / 2816
Vocabulary 48,000 (SentencePiece BPE, shared)
Max sequence length 512 tokens
Positional encoding sinusoidal
Language conditioning learned language embeddings added to encoder and decoder inputs
Tied embeddings input embedding = output projection
Training steps 5,930,000

Evaluation β€” FLORES-200 devtest

1012 sentences per direction, beam size 5, sacreBLEU.

Direction BLEU spBLEU chrF
it β†’ en 24.79 28.80 56.39
en β†’ it 20.95 26.26 52.49
fr β†’ en 29.18 32.09 57.26
en β†’ fr 26.33 29.58 54.70
es β†’ en 18.88 22.09 50.75
en β†’ es 17.00 19.89 47.05
pt β†’ en 33.34 36.42 60.55
en β†’ pt 27.84 31.45 56.72
de β†’ en 24.34 26.80 52.58
en β†’ de 15.88 19.10 46.68
nl β†’ en 18.60 21.36 48.09
en β†’ nl 14.01 17.66 45.00
ro β†’ en 26.95 29.92 56.88
en β†’ ro 18.55 22.73 49.15
pl β†’ en 15.46 18.28 45.94
en β†’ pl 8.69 14.45 38.95
cs β†’ en 22.03 24.61 51.64
en β†’ cs 13.50 18.33 42.38
sv β†’ en 29.98 31.76 56.83
en β†’ sv 23.74 26.52 52.72
Average X β†’ en 24.36 27.21 53.69
Average en β†’ X 18.65 22.60 48.58
Average (all 20) 21.50 24.91 51.14

Usage

pip install torch sentencepiece safetensors
python inference.py --src it --tgt en "Domani mattina devo passare in farmacia."
python inference.py --src en --tgt de --beam 5 "Please back up your data before updating."
from inference import load, translate
model, sp, cfg, device = load()
print(translate("Il treno Γ¨ in ritardo.", "it", "fr", model, sp, cfg, device))

Training data

  • FineTranslations (HuggingFaceFW/finetranslations) β€” web documents paired with English translations, all 10 languages
  • Europarl, Tatoeba
  • Italian YouTube subtitles translated into 10 languages with NLLB-600M (distilled)
  • Small synthetic sets targeting specific weaknesses (proper names, colloquial phrases)

Every pair is used in both directions during training.

Limitations

  • Idioms and colloquial expressions are often translated literally ("stanco morto" β†’ "tired of death").
  • Domain jargon (gaming/streaming, some IT terms) is inconsistent.
  • Polish and Czech are the weakest languages, especially en β†’ pl / en β†’ cs; basic vocabulary errors still occur.
  • Inputs longer than ~512 tokens are truncated: split long texts into sentences.
  • Not suitable for high-stakes use (medical, legal) without human review.

Citation

@misc{iris300m,
  author = {Michelangelo Di Nicola, Roberto D'Arcangelo},
  title  = {Iris-300M: a compact multilingual European translation model},
  year   = {2026},
  publisher = {Hugging Face}
}
Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support