OneReason-8B LoRA: R3 Replay Step20/25 Interp075 KVO

This is an experimental rank-32/alpha-32 LoRA adapter prepared for formal OneReason platform evaluation. It must be loaded on the exact local epoch-2 full-SFT base represented by /data/sft_yaml/onereason_sft_epoch2_bf16; it is not an adapter for the untouched public pretraining checkpoint.

The project ledger associates that epoch-2 base with the platform result 1.2501, but the historical upload hash is unavailable. That score-to-artifact mapping is therefore provenance information, not a cryptographic identity claim. This candidate itself has not yet received an official platform score.

Construction

The adapter was built offline from two checkpoints on one continuous training trajectory:

  • step 20: the R0 raw-SID route-distillation + balanced R3 replay run;
  • step 25: an exact optimizer/scheduler/RNG continuation for five more updates;
  • interpolation coefficient: 0.75 from step 20 toward step 25;
  • interpolated LoRA factors: every layer's k_proj, v_proj, and o_proj;
  • unchanged at step 20: q_proj, gate_proj, up_proj, and down_proj.

Both endpoints use LoRA rank/alpha 32/32. The interpolation acts directly on the matching LoRA A/B factors. Since effective LoRA weights are products of those factors, this is not algebraically identical to dense-weight linear interpolation. No extra training was run to create this candidate.

The source trajectory used 1,192 rows: 952 raw-SID route-distillation examples and 240 balanced, held-out-safe no-thinking R3 replay examples. Training used cutoff_len=1024, global batch size 8, peak LR 5e-5, AdamW, cosine scheduling, and a schedule horizon of 122 updates.

Adapter SHA-256: efc2152249a3958a7aa854105743c68c33428f706c54b5ba4b298df262f77e5f

Local selection evidence

All values below are paired changes versus step 20. These are deterministic local proxies, not official platform scores and not LLM-as-Judge results.

Task / metric Change 95% CI
R0 raw SID char-F1, 400 rows +0.000173 [-0.003221, +0.003524]
R0 generic char-F1, 400 rows -0.002056 [-0.005840, +0.001640]
R2 overall proxy, 128 rows -0.000171 [-0.004844, +0.004240]
R3 teacher domain probability, 256 rows +0.003112 [+0.001217, +0.005004]
R3 free domain accuracy, 256 rows +0.019531 [-0.003906, +0.046875]
World relaxed parse, 2,000 rows -0.005000 [-0.012000, +0.002000]
World relaxed correct, 2,000 rows -0.004500 [-0.011000, +0.002000]

R3 teacher s_b probability was statistically unchanged; s_c changed by -0.002627 with CI [-0.004860, -0.000438]. The candidate was selected because it retained most of the R3 domain gain of the all-module interpolation while substantially reducing its fine-grained hierarchy regressions.

Loading

Use the exact epoch-2 full-SFT base, then attach this adapter with PEFT:

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_path = "/data/sft_yaml/onereason_sft_epoch2_bf16"
adapter_id = "yujackein/onereason-8b-lora-r3replay-step20-step25-interp075-kvo-r32a32"

tokenizer = AutoTokenizer.from_pretrained(adapter_id, trust_remote_code=True)
base_model = AutoModelForCausalLM.from_pretrained(
    base_path,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)
model = PeftModel.from_pretrained(base_model, adapter_id)

Evaluation status

Official platform evaluation: pending.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support