LFM2.5-350M · IFStruct LoRA adapter

A LoRA adapter (rank 8, alpha 16; attention q/k/v_proj, out_proj, short-conv in_proj/out_proj, MLP w1/w2/w3) for LiquidAI/LFM2.5-350M, trained with reinforcement learning (GRPO) to follow structured-output instructions: emit valid JSON/YAML that matches a requested schema, wrapper key, item count, code-block and no-commentary constraints.

The base model's license applies (see LiquidAI/LFM2.5-350M).

IFStruct v1.0 result

Scored on the full 2,000-row test set of LiquidAI/ifstruct-v1.0 with the benchmark's own validator (byte-identical to Liquid4All/ifstruct). Both rows below were run on the same machine and vLLM version.

Model pass@1 JSON YAML
this adapter 52.55 53.0 52.1
LFM2.5-350M (base) 22.85 18.7 27.0

Eval setup: greedy decoding (temperature 0), 1 sample per prompt, max_tokens 16000, no system prompt, the model's chat template, no constrained decoding, vLLM 0.24. No response exceeded 2,859 tokens. Full per-row outputs (prompt, response, validator errors) for both runs are in eval/.

Largest remaining failure kinds: missing required fields and extraneous fields — the model often does not produce exactly the requested set of keys. Enum, code-block, value-range and type errors dropped 3-4x vs base.

Training

  • Method: GRPO (verl 0.9.0), CISPO policy loss, KL loss (low_var_kl, 0.01) to the base model, zero-variance group filtering, 16 prompts × 16 rollouts per step, rollout temperature 1.0, lr 1e-4, max response 4096 tokens. Hybrid conv layers: trained without sequence packing (use_remove_padding: false).
  • Reward: binary pass/fail from the IFStruct validator. No judge model.
  • Data: 4,293 synthetic prompts, 3 epochs; this is the step-270 checkpoint (of ~276).
  • None of the 2,000 benchmark prompts or their entity types appear in the training data (checked by exact prompt and prompt-prefix match).

Disclosure: how the benchmark was used

This is not a blind held-out score:

  1. Data targeting. The synthetic training set was generated with its failure-mode mix steered toward failure types observed on this benchmark (for a different model). No benchmark rows were copied.
  2. Checkpoint selection. The checkpoint was chosen using validation slices drawn from this benchmark (≈288 of the 2,000 rows).

Expect a lower number on genuinely unseen structured-output distributions.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "LiquidAI/LFM2.5-350M"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "ichetandhembre/lfm2.5-350m-ifstruct-lora-adaptor")

With vLLM: vllm serve LiquidAI/LFM2.5-350M --enable-lora --max-lora-rank 8 --lora-modules ifstruct=ichetandhembre/lfm2.5-350m-ifstruct-lora-adaptor.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ichetandhembre/lfm2.5-350m-ifstruct-lora-adaptor

Adapter
(42)
this model