Canopy-258M-R3 v7

Experimental BF16 checkpoint, released September 6, 2026. This is the answer-averaged SFT candidate initialized from the v4/v5 CMA weights. Native PyTorch implementation is bundled to preserve the evaluated architecture and decoding. This package uses run.py, not Transformers AutoModel remote-code loading.

Run

hf download psikosen/canopy-258m-r3-v7 --local-dir canopy-v7
cd canopy-v7
pip install -r requirements.txt
python run.py 'Which number is larger: 17 or 38? Answer with only the number.'

Measured quality

A small procedural suite, with unseen test phrasings and disjoint problems within the experiment:

Task v5 CMA Matched token-average SFT v7 answer-average SFT
Arithmetic 0/27 2/27 2/27
Comparison 0/25 0/25 15/25
Exact copying 0/32 0/32 1/32
JSON integer fields 0/22 15/22 20/22
Total 0/106 17/106 38/106

Selected using validation before testing. Both SFT arms used 240 steps, identical starting weights and sampling, FP32 master parameters with BF16 autocast, and chat/code replay. Per-answer averaging prevents long replay answers from dominating short task answers. See v7_quality_findings.md and evaluation/report.json.

Limitations

This is not a general benchmark score, HumanEval result, or browser speed claim. Historical training contamination is not exhaustively excluded. Arithmetic and exact copying remain weak. Conversational and code smoke tests still fail. Retention NLL is a proxy, not evidence of executable code quality. A separate 24-pair diagnostic found comparison wording/order sensitivity: 29/48 versus 43/48 with a shorter alternative prompt, not a weight improvement. No claim of general reasoning or production readiness is made.

The weights are BF16, not packed 1.58-bit. Prefix sliding and stochastic recurrence are disabled in the evaluated decoding. Config, raw evaluation outputs and checkpoint provenance are included. Earlier releases are unchanged.

Weight SHA256: e731148fba3f634c82502eb2a6d7c35a3ae06c79d44222cd45031baa5f1ff8d2

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support