YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Catalan evaluation β MiniCPM5-2B base vs v1 vs v2
Evaluated 2026-09-10, one a10g-small job per model (~7 min each), loglikelihood MCQ scoring (options as continuations; chat-template context for v1/v2, raw prompt for base), fixed seeds. Raw JSONs: results-base.json, results-v1.json, results-v2.json.
| Benchmark | base (openbmb/MiniCPM5-2B) |
v1 (MiniCPM5-2B-catalan-chat) |
v2 (MiniCPM5-2B-catalan-chat-v2) |
Chance |
|---|---|---|---|---|
| Belebele ca β reading comprehension, 900 items | 27.6% | 28.3% | 27.4% | 25% |
| CaBBQ β social-context QA, 500 items | 30.8% | 34.0% | 33.4% | 33% |
| multi_lmentry ca MCQ β elementary skills, 500 items | 47.8% | 49.8% | 48.2% | ~mixed (2/5-choice) |
| PPL β InstruCAT validation responses (15,081 tokens) | 971.2 | 451.0 | 873.7 | β |
Honest reading
- MCQ accuracy is at chance for all three models. Belebele β 25% chance, CaBBQ β 33% chance, multi_lmentry weighted β mixed chance β all three models sit on those floors. At 2.5B parameters, with option-continuation loglikelihood scoring, these benchmarks do not discriminate between the checkpoints. The v1/v2 MCQ deltas (+0.8β3.2 points) are within sampling noise (n=500/900, Β±~2-3 points).
- PPL differentiates, but with a domain caveat: InstruCAT validation responses are the round-1 training domain (train split; validation held out), so v1's lower PPL (451 vs 874) partly reflects domain familiarity, and the corpus itself is dominated by short task-style responses (15K tokens over 2,000 docs β mean ~7.5 tokens/doc), making the metric noisy. Treat PPL as directional only.
- Caveat on method: scoring chat models by option-loglikelihood under a chat template is a non-standard format that can depress accuracy; a generative multiple-choice protocol (model writes the answer, parsed) would likely produce different absolute numbers.
- What these numbers do not capture: conversational quality, multi-turn coherence, system-prompt adherence β the actual round-2 objective. Those require human/generative evaluation.
Models
- base: openbmb/MiniCPM5-2B
- v1: luispoveda93/MiniCPM5-2B-catalan-chat
- v2: luispoveda93/MiniCPM5-2B-catalan-chat-v2
Benchmarks: facebook/belebele (cat_Latn), BSC-LT/CaBBQ, BSC-LT/multi_lmentry (MCQ tasks only: 2-choice, 5-choice, sΓ/no), projecte-aina/InstruCAT (validation responses for PPL).
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support