# Lex: AMD SystemOne latency ## Mixed questions, faster decisions **128 mixed questions: 154.41 ms median, 56.6% lower than the previous default runtime.** ![Lex: latency as mixed question count increases](assets/mixed-question-scaling.png) | Questions | Previous p50 / p95 (ms) | Typed scheduling p50 / p95 (ms) | |---:|---:|---:| | 1 | 13.22 / 13.32 | 13.20 / 13.40 | | 8 | 28.07 / 28.32 | 28.02 / 28.31 | | 32 | 93.76 / 94.06 | 53.43 / 53.77 | | 64 | 181.78 / 184.53 | 87.39 / 88.09 | | 128 | 355.97 / 360.25 | 154.41 / 155.93 | The request repeats three fixed questions—Choice, Noul and Score—over one unchanged context. Only question count and bookkeeping IDs change. The default SystemOne path groups admitted rows by type so each encoder path processes fuller batches, then restores the original answer order. Every question still has its own contextual computation. Weights, precision, complete-input limit and request schema are unchanged. Measured using the exact release package on AMD ROCm, FP32, physical batch size eight: 10 warmup pairs and 30 alternating AB/BA measurement pairs per point, with GPU synchronization. Measurements include local request conversion, tokenization, model execution and answer assembly; they exclude transport and Studio. Models ran serially after our training and data jobs completed. Results describe this fixed workload, not a service latency guarantee. Both models passed seven source panels: 4,160 admitted decisions without an argmax change and 57 identical whole-request refusals before any forward. Exact-package checks additionally cover 512 decisions, six current Studio examples, caller IDs/order, optional auto batching and input limits. Small floating-point probability differences remain possible. [All samples, workload, checks and earlier concurrent measurements](MIXED_QUESTION_SCALING.json) · [SVG](assets/mixed-question-scaling.svg) · [PDF](assets/mixed-question-scaling.pdf) ## Earlier measurements ## Lex: AMD batch capacity ### Optional larger batches ```python from decision_inference import predict_1k, predict_auto_1k ## native and records use the same objects as the Python usage examples. outputs = predict_1k(native, records) # unchanged default: up to 8 outputs = predict_auto_1k(native, records) # opt in: up to 32 with the padding guard ``` The optional entry admits every complete input before the first forward. It uses consecutive batches up to 32 only when the request has one decision type and total padded tokens do not increase relative to B8. Mixed types or increased padding fall back to B8; requests of eight or fewer take the unchanged default path. Input order, candidates, full 1,024-token limit and FP32 weights remain unchanged. It does not share contextual activations across questions. Larger physical batches can change floating-point rounding and peak memory. #### Paired B8/B32 study These are synchronized resident **Python API measurements**, not Studio or network latency. This earlier study compared default B8 with a homogeneous cap32 policy before the final padding guard was added. All eight measured points satisfy that guard, but **the final `predict_auto_1k` entry was validated separately and was not timed**. Existing default-B8 measurements above, if present, remain their original separate run. | Workload | B8 p50 / p95 (ms) | Cap32 p50 / p95 (ms) | p50 ratio · paired 95% CI | |---|---:|---:|---:| | 242 tokens × 1 question | 13.92 / 14.01 | 13.92 / 14.05 | 1.0001 · [0.9975, 1.0020] | | 242 tokens × 8 questions | 22.65 / 22.75 | 22.60 / 22.77 | 0.9981 · [0.9967, 0.9994] | | 242 tokens × 16 questions | 38.86 / 39.42 | 33.39 / 33.55 | 0.8592 · [0.8578, 0.8603] | | 242 tokens × 32 questions | 71.03 / 71.46 | 54.83 / 60.15 | 0.7718 · [0.7706, 0.7724] | | 1024 tokens × 1 question | 18.09 / 18.24 | 18.08 / 18.21 | 0.9992 · [0.9978, 1.0003] | | 1024 tokens × 8 questions | 65.74 / 65.89 | 65.72 / 66.52 | 0.9997 · [0.9990, 1.0004] | | 1024 tokens × 32 questions | 242.37 / 243.72 | 217.63 / 218.37 | 0.8979 · [0.8961, 0.8986] | | Multi-context · 32 decisions | 44.08 / 44.25 | 28.76 / 28.87 | 0.6523 · [0.6519, 0.6537] | AMD ROCm gfx942, FP32 SDPA, one resident model and one request at a time. Each point/mode had 10 warmups and 50 timed requests; the same validation, transfer and auditing hooks were present in both modes. Ten paired blocks (five AB, five BA; five observations per mode/block) were retained. The reported p50 ratio is median(cap32)/median(B8); 2,000 whole-block bootstrap resamples within order strata used fixed seed 20260922. p95 has only 50 observations/mode and correspondingly limited precision. Model loading is excluded. Choice count curves repeat the same synthetic four-candidate question with opaque ID changes. The multi-context fixture repeats eight existing contexts with two related questions twice (32 decisions). These are throughput shapes, not quality examples. The fixed 15% utility gate remained **FAIL** for both models: Q16 improved approximately 14%, despite a confidence interval below 1. Q32 and multi-context improvements do not retroactively change that outcome. The option is an explicit engineering choice, not a universal speed guarantee. At 1024×32, peak allocated memory increased from approximately 2.652 to 3.824 GiB. Shared allocator reserved-memory observations are not independent per-mode estimates. Equal total padding does not imply equal peak memory. #### Public-entry validation The actual imported package passed 46 forwards per model (92 total across Kai and Lex) over 225 technical row-occurrences/model, including all three types, dynamic candidate counts, enabled32, tail33, unequal-length fallback and mixed64. Late invalid and empty requests added zero forwards. The original strict bounds were 1e-4 logits, 2e-5 probability and normalized Score error, with zero hard-decision flips. Both models passed; fallback outputs were exact. All 489 parameters and 66 buffers retained identical contents and versions. No training, quality evaluation or timing was added in this integration check. Fixed repetition is not independent quality support. [Exact paired measurements, numerical results and retained failure gates](BATCH_CAPACITY.json). No weights or default entry behavior changed, and no Studio speedup is claimed.