Instructions to use ChetanDhembreAI/LFM2.5-350M-ifstruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ChetanDhembreAI/LFM2.5-350M-ifstruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ChetanDhembreAI/LFM2.5-350M-ifstruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ChetanDhembreAI/LFM2.5-350M-ifstruct") model = AutoModelForCausalLM.from_pretrained("ChetanDhembreAI/LFM2.5-350M-ifstruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ChetanDhembreAI/LFM2.5-350M-ifstruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ChetanDhembreAI/LFM2.5-350M-ifstruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChetanDhembreAI/LFM2.5-350M-ifstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ChetanDhembreAI/LFM2.5-350M-ifstruct
- SGLang
How to use ChetanDhembreAI/LFM2.5-350M-ifstruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ChetanDhembreAI/LFM2.5-350M-ifstruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChetanDhembreAI/LFM2.5-350M-ifstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ChetanDhembreAI/LFM2.5-350M-ifstruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChetanDhembreAI/LFM2.5-350M-ifstruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ChetanDhembreAI/LFM2.5-350M-ifstruct with Docker Model Runner:
docker model run hf.co/ChetanDhembreAI/LFM2.5-350M-ifstruct
LFM2.5-350M · IFStruct
LiquidAI/LFM2.5-350M fine-tuned with reinforcement learning (GRPO)
to follow structured-output instructions: emit valid JSON/YAML that matches a requested schema, wrapper key, item
count, code-block and no-commentary constraints. This is the LoRA adapter
ichetandhembre/lfm2.5-350m-ifstruct-lora-adaptor
merged into the base weights (bf16). Distributed under the base model's LFM Open License v1.0 (see LICENSE).
IFStruct v1.0 result
Full 2,000-row test set of LiquidAI/ifstruct-v1.0,
scored with the benchmark's own validator (byte-identical to Liquid4All/ifstruct).
| Model | pass@1 | JSON | YAML |
|---|---|---|---|
| this model (merged) | 52.25 | 52.6 | 51.9 |
| same LoRA served un-merged | 52.55 | 53.0 | 52.1 |
| LFM2.5-350M (base, same eval setup on vLLM 0.24) | 22.85 | 18.7 | 27.0 |
Eval setup: greedy (temperature 0), 1 sample per prompt, max_tokens 16000, no system prompt, the model's chat
template, no constrained decoding, vLLM 0.25. No response exceeded ~2,900 tokens. Per-row outputs are in
eval/.
Largest remaining failure kinds: missing required fields and extraneous fields — the model often does not produce exactly the requested set of keys. Enum, code-block, value-range and type errors dropped 3–4× vs base.
Training
- GRPO (verl 0.9.0), CISPO loss, KL loss (low_var_kl, 0.01) to the base model, zero-variance group filtering,
16 prompts × 16 rollouts per step, rollout temperature 1.0, lr 1e-4, max response 4096 tokens.
LoRA rank 8 / alpha 16 on attention, short-conv
in_proj/out_projand MLPw1/w2/w3. Hybrid conv layers: trained without sequence packing. - Reward: binary pass/fail from the IFStruct validator. No judge model.
- Data: 4,293 synthetic prompts, 3 epochs; step 270 of ~276.
- None of the 2,000 benchmark prompts or their entity types appear in the training data.
Disclosure: how the benchmark was used
This is not a blind held-out score. (1) The synthetic training set's failure-mode mix was steered toward failure types observed on this benchmark (for a different model); no benchmark rows were copied. (2) The checkpoint was chosen using validation slices drawn from this benchmark (≈288 rows). Expect a lower number on genuinely unseen structured-output distributions.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "ichetandhembre/LFM2.5-350M-ifstruct"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype="auto", device_map="auto")
Or vllm serve ichetandhembre/LFM2.5-350M-ifstruct.
- Downloads last month
- 10
Model tree for ChetanDhembreAI/LFM2.5-350M-ifstruct
Evaluation results
- LiquidAI/ifstruct-v1.0 · Ifstruct V1 View evaluation results source leaderboard52.25 *