|
Download evaluation/PERFORMANCE.md from llm-semantic-router/Decision-1.0-Lex-0.6B: direct link, hf CLI and curl.
- Browser
- Download file 1.4 kB
-
https://huggingface.co/llm-semantic-router/Decision-1.0-Lex-0.6B/resolve/main/evaluation/PERFORMANCE.md
- Command line
-
hf download hf://llm-semantic-router/Decision-1.0-Lex-0.6B/evaluation/PERFORMANCE.md
-
curl -L -o PERFORMANCE.md https://huggingface.co/llm-semantic-router/Decision-1.0-Lex-0.6B/resolve/main/evaluation/PERFORMANCE.md
1.4 kB
| # Lex: measured Decision runtime latency | |
| On the evaluated AMD ROCm runtime, 128 mixed Choice, Noul and Score questions took **154.41 ms median**, 56.6% less than the previous default runtime on the same fixed workload. | |
|  | |
| | Questions | Previous p50 / p95 (ms) | Typed scheduling p50 / p95 (ms) | | |
| |---:|---:|---:| | |
| | 1 | 13.22 / 13.32 | 13.20 / 13.40 | | |
| | 8 | 28.07 / 28.32 | 28.02 / 28.31 | | |
| | 32 | 93.76 / 94.06 | 53.43 / 53.77 | | |
| | 64 | 181.78 / 184.53 | 87.39 / 88.09 | | |
| | 128 | 355.97 / 360.25 | 154.41 / 155.93 | | |
| The request repeats three fixed questions over one context; only question count and bookkeeping IDs change. Each question still receives its own contextual computation. Measurements used the exact released weights, FP32 inference, physical batch size eight, 10 warmup pairs and 30 alternating AB/BA pairs per point, with GPU synchronization. They include local request conversion, tokenization, model execution and answer assembly; transport and Studio are excluded. | |
| These results describe the evaluated runtime and workload, not a latency guarantee for a separately distributed runtime. The model-only repository contains no serving code. [Samples and validation details](MIXED_QUESTION_SCALING.json) 路 [SVG](../assets/mixed-question-scaling.svg) 路 [PDF](../assets/mixed-question-scaling.pdf) | |