jcbtc's picture
Release Ciru v4 (4.0.0): integrated runtime and patch notes
d824ec9 verified
|
Raw
History Blame Contribute Delete
4.17 kB

Final unified native performance comparison

Exact unified BUILD: c970d175624455ca30dd3efbc6f52225f6ad32595764232050342a23ced99d94.

Final model on Ciru versus the accepted retained HF publication on Sozo. These are product comparisons with host/runtime differences, not isolated kernel estimates. One final block, unchanged native token fixtures and sampler; no new HF inference. Existing C1 HE speed and BF16 fidelity evidence retain their measured v3 identity and explicit composition link.

Integrity and scope

  • 97 measured requests and 21 excluded warmup requests; all natural stops, zero truncations/errors.
  • 90/90 first coding outputs pass both native base and extended tests; seven seed acknowledgements receive no quality credit.
  • Observed eight active and two queued requests at offered C10. Zero preemptions in every measured batch.
  • Expected prefix reuse on every cached return: 62,720 and 253,120 tokens.
  • AOT artifacts unchanged across measurement; no logged compilation event in measured intervals. Silent compilation was not independently traced.
  • Original service/unit/state restored healthy, PID3120633; no owned processes remain.
  • These non-thinking timing/health fixtures use their original output windows. They are not a full reasoning benchmark. Hermes20 remains pending.

Short concurrency

Offered C HF mean decode Final mean decode Change HF aggregate Final aggregate Change
2 111.59 110.21 -1.24% 187.07 175.43 -6.22%
4 73.54 71.35 -2.98% 222.74 203.04 -8.84%
6 60.81 61.50 +1.13% 264.84 270.40 +2.10%
8 55.84 54.86 -1.74% 294.89 304.80 +3.36%
10 55.38 55.04 -0.62% 249.20 300.00 +20.38%

Units tok/s. C10 queues against max 8. C2/C4 batch throughput is lower, and C4 mean decode is 2.98% lower. Output lengths and scheduling differ; keep these measured losses rather than claiming a uniform speedup. Short C1 is the separately completed current-HF/final block: pooled +17.97%, both mirrors positive, formal order-sensitive/noisy qualification retained.

Cold native prefill

Prompt tokens HF prefill tok/s Final prefill tok/s Change Final request wall
1,024 1868.32 2321.16 +24.24% 0.63s
8,192 1714.93 2141.95 +24.90% 3.97s
32,768 1509.21 1938.86 +28.47% 17.05s
65,536 1286.66 1624.19 +26.23% 40.51s
131,072 982.67 1280.20 +30.28% 102.57s
253,952 667.61 870.11 +30.33% 292.19s

Every cold seed naturally emits 17 tokens. Those acknowledgements do not establish sustained decode speed. Near-256K request wall 292.19s versus retained 380.83s (23.28% shorter).

Cached full ten-task return sequences

History / C HF mean decode Final mean decode Decode change HF mean TTFT Final mean TTFT Final batch wall
253952-C1 94.00 118.81 +26.40% 4.619s 4.217s 55.95s
253952-C8 11.10 19.86 +78.93% 18.033s 25.291s 49.20s
63000-C1 123.34 147.69 +19.74% 0.848s 0.754s 18.29s
63000-C8 32.08 32.98 +2.79% 3.968s 3.856s 12.61s

Near-256K C8 mean first-token latency increased from 18.03s to 25.29s, even though the whole batch completed faster. This tradeoff remains part of the result.

Decode is measured after first token; total batch time includes scheduling and prefill. Cache-seed preparation is separate. Raw native counters, per-request timings and output lengths remain in the sealed record. Concurrent prefill-counter ratios are retained in the machine-readable decision and are not described as cold ingestion speed.

Disposition

Complete the requested final performance rows without another routine repeat. Preserve all mixed short and long-latency outcomes. Remaining Ornith work is full-budget native Hermes20 C1/C8 three seeds, then public Ciru v4 packaging/card updates. Apodex receives the parallel backport plus sanity checks only.

491 files imported and hash-verified. Archive SHA256 a8cbcffd9c60e25e4457d5699453b4f041155ba8685964f58c9d16011e33105e.