LFM2.5-350M · IFStruct

LiquidAI/LFM2.5-350M fine-tuned with reinforcement learning (GRPO) to follow structured-output instructions: emit valid JSON/YAML that matches a requested schema, wrapper key, item count, code-block and no-commentary constraints. This is the LoRA adapter ichetandhembre/lfm2.5-350m-ifstruct-lora-adaptor merged into the base weights (bf16). Distributed under the base model's LFM Open License v1.0 (see LICENSE).

IFStruct v1.0 result

Full 2,000-row test set of LiquidAI/ifstruct-v1.0, scored with the benchmark's own validator (byte-identical to Liquid4All/ifstruct).

Model pass@1 JSON YAML
this model (merged) 52.25 52.6 51.9
same LoRA served un-merged 52.55 53.0 52.1
LFM2.5-350M (base, same eval setup on vLLM 0.24) 22.85 18.7 27.0

Eval setup: greedy (temperature 0), 1 sample per prompt, max_tokens 16000, no system prompt, the model's chat template, no constrained decoding, vLLM 0.25. No response exceeded ~2,900 tokens. Per-row outputs are in eval/.

Largest remaining failure kinds: missing required fields and extraneous fields — the model often does not produce exactly the requested set of keys. Enum, code-block, value-range and type errors dropped 3–4× vs base.

Training

  • GRPO (verl 0.9.0), CISPO loss, KL loss (low_var_kl, 0.01) to the base model, zero-variance group filtering, 16 prompts × 16 rollouts per step, rollout temperature 1.0, lr 1e-4, max response 4096 tokens. LoRA rank 8 / alpha 16 on attention, short-conv in_proj/out_proj and MLP w1/w2/w3. Hybrid conv layers: trained without sequence packing.
  • Reward: binary pass/fail from the IFStruct validator. No judge model.
  • Data: 4,293 synthetic prompts, 3 epochs; step 270 of ~276.
  • None of the 2,000 benchmark prompts or their entity types appear in the training data.

Disclosure: how the benchmark was used

This is not a blind held-out score. (1) The synthetic training set's failure-mode mix was steered toward failure types observed on this benchmark (for a different model); no benchmark rows were copied. (2) The checkpoint was chosen using validation slices drawn from this benchmark (≈288 rows). Expect a lower number on genuinely unseen structured-output distributions.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
name = "ichetandhembre/LFM2.5-350M-ifstruct"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype="auto", device_map="auto")

Or vllm serve ichetandhembre/LFM2.5-350M-ifstruct.

Downloads last month
10
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ChetanDhembreAI/LFM2.5-350M-ifstruct

Finetuned
(78)
this model
Quantizations
1 model

Evaluation results