Decision-1.0-Sol-2B / evaluation /QUESTION-SCALING.md
Xunzhuo's picture
Publish clean Decision model repository
53ecba4
|
Raw History Blame Contribute Delete
1.43 kB

Request latency

Sol-2B · distinct Choice questions with 499 input tokens per question. The state and individual question length stay fixed as the question count grows.

Question scaling

Questions p50 ms ↓ p95 ms ↓ Peak allocated GiB
1 24.606 25.463 3.626
2 24.896 25.796 3.665
4 30.624 31.215 3.742
8 50.828 51.260 3.896
16 101.696 102.555 3.896
32 203.602 205.315 3.896

Six independently loaded process blocks supply 30 measured requests per point after warmup on an otherwise idle AMD gfx942 GPU. End-to-end Python request latency includes rendering, tokenization, inference, output assembly and final synchronization. Model loading, JSON serialization and network are excluded. Questions run in their original order in batches of eight. Requests are sequential, with no neural prefix cache.

These measurements used the original API implementation and inference bundle. That runtime passed an offline Hub-download proof and the full 3,160-answer regression; this historical result does not validate a different serving runtime. Measurements and immutable runtime identity.

This fixed short-input Choice workload does not establish concurrent HTTP throughput, long-context scaling, other question-type performance or a cross-hardware speed ranking.