YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Catalan evaluation β€” MiniCPM5-2B base vs v1 vs v2

Evaluated 2026-09-10, one a10g-small job per model (~7 min each), loglikelihood MCQ scoring (options as continuations; chat-template context for v1/v2, raw prompt for base), fixed seeds. Raw JSONs: results-base.json, results-v1.json, results-v2.json.

Benchmark base (openbmb/MiniCPM5-2B) v1 (MiniCPM5-2B-catalan-chat) v2 (MiniCPM5-2B-catalan-chat-v2) Chance
Belebele ca β€” reading comprehension, 900 items 27.6% 28.3% 27.4% 25%
CaBBQ β€” social-context QA, 500 items 30.8% 34.0% 33.4% 33%
multi_lmentry ca MCQ β€” elementary skills, 500 items 47.8% 49.8% 48.2% ~mixed (2/5-choice)
PPL β€” InstruCAT validation responses (15,081 tokens) 971.2 451.0 873.7 β€”

Honest reading

  • MCQ accuracy is at chance for all three models. Belebele β‰ˆ 25% chance, CaBBQ β‰ˆ 33% chance, multi_lmentry weighted β‰ˆ mixed chance β€” all three models sit on those floors. At 2.5B parameters, with option-continuation loglikelihood scoring, these benchmarks do not discriminate between the checkpoints. The v1/v2 MCQ deltas (+0.8–3.2 points) are within sampling noise (n=500/900, Β±~2-3 points).
  • PPL differentiates, but with a domain caveat: InstruCAT validation responses are the round-1 training domain (train split; validation held out), so v1's lower PPL (451 vs 874) partly reflects domain familiarity, and the corpus itself is dominated by short task-style responses (15K tokens over 2,000 docs β€” mean ~7.5 tokens/doc), making the metric noisy. Treat PPL as directional only.
  • Caveat on method: scoring chat models by option-loglikelihood under a chat template is a non-standard format that can depress accuracy; a generative multiple-choice protocol (model writes the answer, parsed) would likely produce different absolute numbers.
  • What these numbers do not capture: conversational quality, multi-turn coherence, system-prompt adherence β€” the actual round-2 objective. Those require human/generative evaluation.

Models

Benchmarks: facebook/belebele (cat_Latn), BSC-LT/CaBBQ, BSC-LT/multi_lmentry (MCQ tasks only: 2-choice, 5-choice, sΓ­/no), projecte-aina/InstruCAT (validation responses for PPL).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support