nanojev-hard-lr1e5-repro

Personal reproducibility run of NanoJev hard CE SFT (hard_lr1e5, seed 17, 600 steps), starting from the public initialization in C-Tianyu/NanoJev (training_initialization).

Layout mirrors the root checkpoint directory of C-Tianyu/NanoJev (weights + tokenizer + eval artifacts), so you can download this folder and continue training or serve with NanoJev tooling.

This is not the official unified-games-v1 release. Checkpoint selection used min dev selection_ce β†’ best_step = 300.

Files (same style as official unified root)

Path Role
best.safetensors Selected weights (step 300)
config.json Config from this run (may contain original machine paths)
training_config.json Same hyperparams with local absolute paths redacted
tokenizer/ Tokenizer
backbone_config/ Backbone config
summary.json Metrics + weights_sha256
train_log.json Step log
predictions_{dev,test,ood,calibration}.jsonl Split predictions
initial_dev.jsonl / initial_dev_metrics.json Init-time eval dump
target_audit.json Target audit

Recipe

  • Backbone: Qwen/Qwen3-0.6B + set_head=attention
  • Init SHA256: 38116340795de1c82369b7fe15819d92d79600a7b4dc7a3cd0d4390cb6782639
  • Data: NanoJev unified hard mix
  • Loss: hard CE; backbone lr 1e-5, head lr 1e-4; seed 17; 600 steps; eval every 100
  • Selected weights SHA256: f53eadbd34eb5a1040d9d579c64bebde45a7180fa3b012d187e64c118c09022f
  • best_dev_selection_ce β‰ˆ 0.636 Β· test β‰ˆ 0.678 Β· ood β‰ˆ 0.759

Download β†’ continue train / serve

huggingface-cli download maxonxie/nanojev-hard-lr1e5-repro \
  --local-dir ./nanojev-hard-lr1e5-repro

# serve
python -m research.toy.serve_decisions \
  --checkpoint-dir ./nanojev-hard-lr1e5-repro \
  --host 127.0.0.1 --port 8765

# continue SFT: point --init-checkpoint (or your NanoJev flag) at this directory
# and set --input to your hard data path on the new machine.

Exact CLI flags follow your NanoJev checkout; the checkpoint directory shape matches what serve_decisions / train scripts expect for a finished run.

License / attribution

Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for maxonxie/nanojev-hard-lr1e5-repro

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1325)
this model