MiniCPM5 SFT mix Run1 v3

Final checkpoint from a fresh one-epoch full SFT of MiniCPM5-2B-Midtrain on the v3 filtered Run1 data mix. This is checkpoint-5913, not a continuation of the old Run1 checkpoints.

  • Base revision: 0a45344e; base weight SHA-256: 38a28680f6208242a0de7c84627343d44cee517b49c2b39d1afd706be6beabad.
  • LLaMA-Factory frontend, Megatron/MCore adapter, 12 nodes / 192 Ascend 910C cards; TP4, DP48, micro batch 1, accumulation4, global batch192.
  • One epoch, 5913 updates; LR5e-5 cosine to1e-5, warmup3%, weight decay0.01, seed20260916, BF16.
  • Maximum training sequence131072; packed and balanced scheduling. Assistant header masking repaired; source labels remain at the same positions and the MCore collator shifts once.
  • Packed input tokens:36,161,705,387; supervised tokens:15,591,318,474. These describe the tokenized dataset, not an independent runtime token counter.
  • Final held-out validation loss:0.2753264904022217. SWE evaluation of this checkpoint is pending; training loss alone is not a benchmark result.

HF BF16 weights were converted once on node0 without modifying the original MCore checkpoint. Structural config, tokenizer, RoPE semantics and safetensors shards were checked. conversion-receipt.json records hashes and source provenance. Model uses XML-style MiniCPM5 tool calling; use the MiniCPM5 parser and the provided chat template.

Related fixed-200 evaluation task set: eigentom/minicpm5-sft-swe-validation-200. Evaluation results will be reported separately using the same frozen protocol as the baseline runs.

Downloads last month
184
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eigentom/nanocode-sft-mix-run1-v3

Finetuned
(6)
this model