# Sounio LoRA Adapter Candidate — Qwen2.5-Coder-1.5B Status: candidate for Hugging Face upload, pending operator authorisation. This adapter is the first measured Sounio fine-tuning artefact in this lane that beats the zero-training base-model reference on the full 200-row held-out eval split. ## Identity | Field | Value | |---|---| | Base model | `Qwen/Qwen2.5-Coder-1.5B` | | Training objective | Sounio complete-file SFT with surgical repair pairs | | Branch | `feature/dataset-expansion` | | Dataset commit | `12c471d1a` | | Run ID | `sounio-qwen15-lora-full200-surgical-repair-20260522T021535` | | Slurm job | `1603` | | GPU | 1 GPU on `gpu-orangefs` | | Training rows requested | 4,800 | | Training rows used | 4,769 | | Optimiser steps | 120 | | Prompt style | `sounio-compact-completion` | | Repair pair ratio | 0.20 | | Run-expected ratio | 0.10 | | Separate runtime repair ratio | 0 | | Targeted runtime ratio | 0 | ## Artifact Location ```text /orangefs/training/sounio/hf-examples/lora-smoke/sounio-qwen15-lora-full200-surgical-repair-20260522T021535/results/adapter ``` Adapter files present: | File | Size | |---|---:| | `adapter_model.safetensors` | 73,911,112 bytes | | `adapter_config.json` | 797 bytes | | `README.md` | 5,097 bytes | | `tokenizer.json` | 11,421,896 bytes | | `tokenizer_config.json` | 7,338 bytes | | `vocab.json` | 2,776,833 bytes | | `merges.txt` | 1,671,853 bytes | | `special_tokens_map.json` | 616 bytes | | `added_tokens.json` | 605 bytes | Checksums: ```text 9d8540e218439bb28026115fbe8d4be5fe18162a0a1c07474d9850579ee36c4a adapter_model.safetensors 0ce16cbdfa4be1b4fb97868ee0b5f2fed33dfb3ba65c7961530f7fe7cd544057 train_summary.json 6eca4fdaaf6e3bb48f302f8090e681cddb3fc6e08b7a99dcde0f04962c09d794 lora_smoke_eval_report.json ``` ## Evaluation Result Held-out split: `datasets/hf_examples/instruction_pairs_eval.jsonl`, 200 rows. | Model/objective | Compile pass | Compile + contract-clean | Exact runtime pass | |---|---:|---:|---:| | Base model, zero training | 111/200 (55.5%) | not recorded | 25/65 (38.5%) | | Qwen2.5-Coder-1.5B LoRA candidate | 192/200 (96.0%) | 192/200 (96.0%) | 58/65 (89.2%) | | Candidate + structural repair layer | 199/200 (99.5%) | 198/200 (99.0%) | 59/65 (90.8%) | The adapter improves the two acceptance metrics used for this lane: - `souc check` pass rate on the full held-out eval split - exact stdout match on all runnable held-out rows This is a utility result, not just a training-loss result. The generated outputs extract as complete Sounio files, mostly satisfy the no-prose/no-Rust contract, compile with `souc check`, and run correctly on most rows with expected stdout. The 98%+ publication-confidence run is adapter-only: it reuses the saved adapter and widens the deterministic structural repair layer; it does not retrain model weights. Result path: ```text /orangefs/training/sounio/hf-examples/lora-smoke/sounio-qwen15-adapter-eval-98plus-20260522T083913/results ``` The structural repair layer catches outputs still outside the complete-file objective, including Rust-shaped `String`/`Vec` surfaces, invalid synthetic declarations, algebra blocks used as structs, and one observed scalar typo. ## Remaining Misses Residual failures after the 98%+ adapter-only evaluation: - compile or contract failures: 2/200 - runtime misses: 6/65 runnable rows - residual Rust-shaped surface: - `String`: 1 row - runtime families still missing: - one HumanEval runnable - two FFI logical extent rows printing `16` - two GPU tile rows printing `4` - one reserve advanced row printing `64` These are the next repair-pair targets if the goal is to improve this adapter before publication. ## Upload Policy Do not upload automatically. Upload requires explicit operator authorisation and an operator-owned Hugging Face token. Suggested repository name: ```text chiuratto-AIgourakis/sounio-qwen25-coder-15b-lora ``` Suggested upload payload: ```text adapter/ train_summary.json lora_smoke_eval_report.json lora_smoke_predictions.jsonl ADAPTER_CANDIDATE_QWEN15.md ``` The adapter should be described as an experimental Sounio LoRA adapter, not as a general-purpose production code model.