Ternary-Bonsai-2-27B-cmf / evaluation /MEASUREMENTS_0.6.9.md
infosave's picture
Document Cortiq 0.6.9 and validated Vulkan/Metal results; preserve model weights
475c37e verified
|
Raw
History Blame Contribute Delete
2.4 kB

Measurement provenance for the 0.6.9 model card

The model is unchanged: SHA-256 fb0f1a9cdb0434bc9ce50eb6473bc2ca4a71d6994ad61314bbd7e2f11c4e29e8.

The three accompanying Metal JSON/JSONL files are raw development measurements retained without rewriting their contents. Their development model paths and binary version labels are historical, not installation instructions. The release integrates the accepted kernels plus scope/admission repairs and is compiled/tested independently.

  • metal-matched-pre-release-0.6.9.json: Apple M4 / 10 GPU cores / 24 GiB, fixed 1,122-ID prompt, 128 greedy generated tokens, cold + three warm rows. Read actual generated counts and 127 callback intervals, not total request time as decode time.
  • metal-serial-nll-pre-release-0.6.9.jsonl: two independently reset 512-target serial scores, 1,024 total.
  • metal-batch-nll-pre-release-0.6.9.jsonl: corresponding ordinary batched-prefill scores, 34 total chunks and 1,026 processed rows. Two overlapping source/reset rows are not extra scored targets.

Raw stage labels inherited an instrumentation issue: some GDN projection counters are classified as FFN after the first layer of a run. Do not infer per-stage arithmetic or physical memory bandwidth from those bucket labels. Aggregate dispatch counts and output IDs are the route checks. GPU allocations are not equal to process RSS or model file size.

The model card's approximately42 tok/s Vulkan figure is the measured resident current-split development path on RTX4090. The final endpoint discriminator used exactly381 warm decode spans (3×127), excluding prompt frames; endpoint instrumentation was present. It is not a guarantee for all release binaries or application sampling policies. The approximately87 tok/s Prism CUDA reference uses a different internal timing boundary.

The historical5.23tok/s Metal baseline was not paired in the same thermal/timing session with the final7.50–8.74tok/s samples. Do not turn that comparison into a controlled percentage-gain claim. No temperature/clock trace established the cause of warm-run variation.

The preserved original evaluation/ files document the earlier public-development patch and reference corpus. The current model card explicitly distinguishes those historical values from later resident execution measurements. This small correspondence corpus is not a broad benchmark suite.