Instructions to use Nurymanau/LFM2.5-350M-Halo-Structured-Output-Research with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Nurymanau/LFM2.5-350M-Halo-Structured-Output-Research with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-350M") model = PeftModel.from_pretrained(base_model, "Nurymanau/LFM2.5-350M-Halo-Structured-Output-Research") - Notebooks
- Google Colab
- Kaggle
LFM2.5-350M + Halo: structured-output research adapters
An independent, small-budget experiment with White Circle Halo online GRPO and a frozen LiquidAI LFM2.5-350M model. The first pilot improved IFStruct from 440/2000 to 555/2000. A mixed JSON/YAML/extraction continuation was not reliably better: two training seeds produced 434/2000 and 543/2000 against a matched 444/2000 baseline.
This is a research artifact with positive and negative results. It is not an official Liquid AI or White Circle release, a state-of-the-art claim, or a validated tool-calling model. No Halo-versus-TRL speed comparison was performed.
Included adapters
| Folder | Role | Training |
|---|---|---|
| Repository root | Fixed first pilot checkpoint; the default artifact, not selected across final-test results | Seed 42, original recipe, 100 steps |
variants/phase2-selected |
Mixed recipe selected on separate dev before final evaluation | Seed 42, pilot 100 + mixed 50 |
variants/phase2-repeat |
Independent repeat of the selected procedure | Seed 43, fresh original 100 + mixed 50 |
variants/phase2-json-control |
Dev-only continuation comparator; no final IFStruct evaluation | Seed 42, pilot 100 + original JSON 50 |
Each adapter contains 5,996,544 FP32 parameters (184 tensors), about 24 MB. LoRA rank 16, alpha 32. Original weight bytes are unchanged. Only exported adapter metadata was changed to replace the training-machine path with the public base-model identifier and exact revision. See MANIFEST.json and NOTICE.md.
Base: LiquidAI/LFM2.5-350M, revision 9e6c6ccf47cd318696e137d381a7ded8fe4df09f. This is the instruction-tuned checkpoint, not LFM2.5-350M-Base.
Measured results
IFStruct strict success requires all evaluator conditions to pass. It is not a general intelligence or semantic correctness score.
| Experiment / arm | JSON /1000 | YAML /1000 | Overall /2000 |
|---|---|---|---|
| Pilot baseline | 177 | 263 | 440 (22.00%) |
| Pilot step100 | 302 | 253 | 555 (27.75%) |
| Phase2 baseline | 180 | 264 | 444 (22.20%) |
| Phase2 selected, seed42 | 281 | 153 | 434 (21.70%) |
| Phase2 repeat, seed43 | 273 | 270 | 543 (27.15%) |
Pilot: +5.75 percentage points, 236 new successes and 121 regressions. Paired task-bootstrap 95% interval [+3.95, +7.60] pp. Phase2 selected: −0.50 pp [−2.60, +1.65]; repeat: +4.95 pp [+3.00, +6.95]. These intervals resample test tasks; they do not measure uncertainty across training seeds. Pilot and phase2 used different GPU models; compare each adapter against its own matched baseline.
On the phase2 dev set, strict successes were 13/186 (base), 57/186 (pilot step100), 87/186 (JSON continuation), 94/186 (mixed selected), and 38/186 (repeat). Selection used a predeclared four-category macro gain of at least 2 pp and a maximum 5 pp loss per category. The independent repeat was report-only.
A narrow held-out synthetic extraction test contained 32 English orders rendered in two formats:
| Model | All requirements /64 | Exact object, including shape and types /64 |
|---|---|---|
| Base | 5 | 26 |
| Pilot step100 | 33 | 64 |
| Phase2 selected | 64 | 64 |
| Phase2 repeat | 22 | 64 |
The trained variants already returned the exact objects. Their strict-score differences mainly reflect serialization and code-fence instructions. This does not establish general reasoning, production readiness, or tool-use ability.
Method and limitations
- Halo commit:
425f04103eedbeb3ea1a6f7fa6d142ed02c9b7a8. Native training used one trainer GPU and a separate vLLM GPU. No one-GPU implementation was evaluated. - First pilot: 476 training / 61 dev source examples, 100 steps. Phase2: 1,208 train / 186 dev, with JSON/YAML and synthetic extraction tasks. Continuations load an adapter exactly but restart the optimizer and learning-rate schedule.
- IFStruct is evaluation-only. It was not training data, reward input, or the phase2 recipe-selection signal. It is the same reporting benchmark already used in the pilot, not a newly unseen holdout.
- Phase2 combined data and reward changes; it does not isolate their causal effects. Twelve of 61 structural dev prompts retained an explicit user-level JSON-array request conflicting with the added YAML system request. The frozen evaluation was not rewritten after discovering this.
- Exact/normalized partition checks were performed; semantic near-duplicates were not independently audited.
- Training/evaluation used pinned BF16 GPU runtimes. The included CPU scripts use FP32 and are not numerical parity reproductions of those measurements.
- Legacy recipe files contain inert
checkpoint=step100/runtime=unverifiedlabels inherited from preparation. Actual phase2 adapter lineage is 100+50; these labels did not control model loading. Recipe files are preserved as historical metadata. - Original saved-response audits re-scored all 4,000 pilot and 7,186 phase2 dev/test responses. Public ledgers contain decision metadata, not raw prompts/responses. Ledger verification checks arithmetic and pairing; it cannot independently establish that unpublished response text deserved its verdict.
Total measured GPU rental estimate for this isolated project, including probes and a failed host: $9.88, of which phase2 was $6.24. Estimates use elapsed rental time and quoted rates, exclude storage, and are not provider invoices or future price guarantees. Repeated full backups and startup/evaluation overhead contributed to rental time. All rented resources were removed after verified rescue.
Verification without rented GPUs
Clone/download this repository, then from its root:
python3 code/verify_release.py
python3 code/verify_results.py
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python code/smoke_cpu.py --output cpu-smoke.json
Release checks on an Apple Silicon Mac passed for all four adapters. The public re-scoring code also reproduced all 10,000 saved IFStruct verdicts from the five measured arms; receipts are in validation/. A two-task, 64-token CPU pipeline smoke test is labelled partial and is not a quality result.
The first two commands need only Python's standard library. The smoke test downloads the pinned public base (about 714 MB), loads all four adapters on CPU, checks finite logits, and generates short responses. It does not claim benchmark quality. A local pinned base can be supplied with --base /path/to/model.
Generate a new CPU evaluation (10 tasks by default; use --limit 2000 for the full set):
.venv/bin/python code/evaluate_cpu.py --adapter base --output eval-base
.venv/bin/python code/evaluate_cpu.py --adapter . --output eval-pilot
.venv/bin/python code/check_saved.py --outcomes eval-pilot/outcomes.jsonl --output rechecked.json
The scripts download the original evaluation dataset and validator at pinned revisions and verify their hashes. They do not download training data. A full CPU run can be slow. Compare only runs using the same device/runtime/settings; do not substitute partial CPU scores into the GPU table above. check_saved.py can also re-score a supplied original-format outcomes file.
results/summary.json and 14 decision ledgers reproduce reported counts and paired transitions. recipes/ contains historical hyperparameters, not a turnkey training launcher. Private orchestration, credentials, raw datasets, optimizer states and internal infrastructure are excluded from this distribution.
Load one adapter
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_id = "LiquidAI/LFM2.5-350M"
base_revision = "9e6c6ccf47cd318696e137d381a7ded8fe4df09f"
repo = "Nurymanau/LFM2.5-350M-Halo-Structured-Output-Research"
tokenizer = AutoTokenizer.from_pretrained(base_id, revision=base_revision)
base = AutoModelForCausalLM.from_pretrained(
base_id, revision=base_revision, dtype=torch.float32,
attn_implementation="eager",
)
model = PeftModel.from_pretrained(base, repo).eval()
# Other variants: add subfolder="variants/phase2-selected", etc.
Pin the release commit as well when reproducing a result.
Attribution and license
Model/adapter distribution carries the LFM Open License v1.0, included as LICENSE; its commercial-use conditions apply. Independent verification scripts carry the MIT license in CODE_LICENSE. Evaluation and training-source licenses remain with their authors; no raw dataset is redistributed here.
- Liquid AI base model
- Liquid AI / Hugging Face GRPO tutorial: methodological starting point, not a claim that its published numbers are ours.
- White Circle Halo
- IFStruct evaluator and evaluation dataset
- NVIDIA structured-output training source, CC BY 4.0. Source prompts were split, deduplicated and augmented as described above.
No author endorsement is implied.
- Downloads last month
- 9