Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Download runtime-results/v44/benchmarks/fp8-dcp2-mtp3-decode.json from brandonmusic/GLM-5.3-Flash-tr3-4bpw: direct link, hf CLI and curl.
- Browser
- Download file 19.8 kB
-
https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v44/benchmarks/fp8-dcp2-mtp3-decode.json
- Command line
-
hf download hf://brandonmusic/GLM-5.3-Flash-tr3-4bpw/runtime-results/v44/benchmarks/fp8-dcp2-mtp3-decode.json
-
curl -L -o fp8-dcp2-mtp3-decode.json https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v44/benchmarks/fp8-dcp2-mtp3-decode.json
19.8 kB
| { | |
| "metadata": { | |
| "version": "0.4.29", | |
| "engine": "vllm", | |
| "model": "GLM-5.3-Flash-EXL3-4bpw", | |
| "server": "127.0.0.1:8013", | |
| "timestamp": "2026-08-27T14:17:58.509027", | |
| "decode_mode": "duration", | |
| "primary_decode_layer": "sustained_decode", | |
| "duration_per_test": 10.0, | |
| "request_count": 0, | |
| "warmup_request_count": 0, | |
| "run_burst": false, | |
| "prefill_mode": "skipped", | |
| "standalone_prefill": false, | |
| "prefill_only": false, | |
| "skip_prefill": true, | |
| "burst_e2e_status": "not_run_use_--run-burst", | |
| "burst_request_count": 0, | |
| "burst_warmup_request_count": 0, | |
| "burst_requests_per_concurrency": 5, | |
| "decode_warmup_seconds": 3.0, | |
| "decode_warmup_context": 32768, | |
| "decode_warmup_concurrency": 1, | |
| "cell_warmup_timeout_seconds": 0.0, | |
| "cell_warmup_timeout_policy": "<=32k:60s,64k:120s,>=128k:180s when override is 0", | |
| "show_capacity_limited_values": false, | |
| "max_tokens": 4096, | |
| "temperature": null, | |
| "ignore_eos": true, | |
| "max_total_tokens": 601344, | |
| "dcp_size": 2, | |
| "metrics_available": true, | |
| "metrics_warning": "", | |
| "concurrency_levels": [ | |
| 1 | |
| ], | |
| "context_lengths": [ | |
| 0, | |
| 16384, | |
| 32768 | |
| ], | |
| "startup_diagnostics_available": true, | |
| "nvidia_p2p_override_effective": true, | |
| "p2pmark_status": "not_run", | |
| "amd_fabric_status": "not_run" | |
| }, | |
| "startup_diagnostics": { | |
| "version": "0.4.29", | |
| "server_url": "http://127.0.0.1:8013", | |
| "hostname": "pop-os", | |
| "uname": "Linux pop-os 6.18.7-76061807-generic #202601231045~1769703228~24.04~cb87b5b SMP PREEMPT_DYNAMIC Thu J x86_64 x86_64 x86_64 GNU/Linux", | |
| "env": {}, | |
| "args": { | |
| "concurrency": "1", | |
| "contexts": "0,16k,32k", | |
| "max_tokens": 4096, | |
| "duration": 10.0, | |
| "request_count": 0, | |
| "run_burst": false, | |
| "standalone_prefill": false, | |
| "prefill_only": false, | |
| "skip_prefill": true, | |
| "prefill_contexts": "8k,64k,128k", | |
| "prefill_metric": "client", | |
| "dcp_size": 2, | |
| "kv_budget": 0 | |
| }, | |
| "nvidia_p2p_override": { | |
| "effective": true, | |
| "configured": true, | |
| "params_path": "/proc/driver/nvidia/params", | |
| "params_available": true, | |
| "modprobe_path": "/etc/modprobe.d/nvidia-p2p-override.conf", | |
| "modprobe_available": true, | |
| "runtime": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1", | |
| "DmaRemapPeerMmio": "1" | |
| }, | |
| "expected": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1" | |
| }, | |
| "missing": [], | |
| "mismatched": {}, | |
| "registry_dwords": "ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1", | |
| "suggested_modprobe_line": "options nvidia NVreg_RegistryDwords=\"ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1\"", | |
| "suggested_reload": "stop GPU workloads, then reload NVIDIA modules or reboot; the modprobe file alone is not enough until the nvidia module is reloaded" | |
| }, | |
| "p2pmark": { | |
| "status": "not_run" | |
| }, | |
| "amd_fabric": { | |
| "status": "not_run" | |
| }, | |
| "nvidia_smi_query": { | |
| "cmd": [ | |
| "nvidia-smi", | |
| "--query-gpu=index,name,driver_version,pci.bus_id,pcie.link.gen.current,pcie.link.width.current,power.limit", | |
| "--format=csv,noheader,nounits" | |
| ], | |
| "returncode": 0, | |
| "stdout": "0, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 610.57.04, 00000000:01:00.0, 5, 16, 300.00\n1, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 610.57.04, 00000000:21:00.0, 5, 16, 300.00\n2, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 610.57.04, 00000000:81:00.0, 1, 16, 300.00\n3, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 610.57.04, 00000000:C1:00.0, 2, 16, 275.00", | |
| "stderr": "" | |
| }, | |
| "nvidia_smi_topo": { | |
| "cmd": [ | |
| "nvidia-smi", | |
| "topo", | |
| "-m" | |
| ], | |
| "returncode": 0, | |
| "stdout": "\u001b[4mGPU0\tGPU1\tGPU2\tGPU3\tCPU Affinity\tNUMA Affinity\tGPU NUMA ID\u001b[0m\nGPU0\t X \tNODE\tNODE\tNODE\t0-47\t0\t\tN/A\nGPU1\tNODE\t X \tNODE\tNODE\t0-47\t0\t\tN/A\nGPU2\tNODE\tNODE\t X \tNODE\t0-47\t0\t\tN/A\nGPU3\tNODE\tNODE\tNODE\t X \t0-47\t0\t\tN/A\n\nLegend:\n\n X = Self\n SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)\n NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node\n PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)\n PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)\n PIX = Connection traversing at most a single PCIe bridge\n NV# = Connection traversing a bonded set of # NVLinks", | |
| "stderr": "" | |
| } | |
| }, | |
| "nvidia_p2p_override": { | |
| "effective": true, | |
| "configured": true, | |
| "params_path": "/proc/driver/nvidia/params", | |
| "params_available": true, | |
| "modprobe_path": "/etc/modprobe.d/nvidia-p2p-override.conf", | |
| "modprobe_available": true, | |
| "runtime": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1", | |
| "DmaRemapPeerMmio": "1" | |
| }, | |
| "expected": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1" | |
| }, | |
| "missing": [], | |
| "mismatched": {}, | |
| "registry_dwords": "ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1", | |
| "suggested_modprobe_line": "options nvidia NVreg_RegistryDwords=\"ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1\"", | |
| "suggested_reload": "stop GPU workloads, then reload NVIDIA modules or reboot; the modprobe file alone is not enough until the nvidia module is reloaded" | |
| }, | |
| "p2pmark": { | |
| "status": "not_run" | |
| }, | |
| "amd_fabric": { | |
| "status": "not_run" | |
| }, | |
| "hardware_run_summary": {}, | |
| "event_log": [ | |
| "14:15:58 benchmark start engine=vllm", | |
| "14:15:58 startup server=http://127.0.0.1:8013 model=GLM-5.3-Flash-EXL3-4bpw", | |
| "14:15:58 startup decode concurrency=1 contexts=0,16k,32k", | |
| "14:15:58 startup NVIDIA P2P override: enabled: runtime NVIDIA P2P override matches expected RegistryDwords", | |
| "14:15:58 startup engine vLLM 0.1.dev20111+g7f1e92bec.d20260827 models=['GLM-5.3-Flash-EXL3-4bpw']", | |
| "14:15:58 startup KV cache budget from vLLM metrics: 601,344 tokens (87 blocks x 3456; local 300,672 \u00d7 CP 2; CP source: argument)", | |
| "14:15:58 startup model context length: 131,072 tokens", | |
| "14:15:58 startup prefill tests: skipped", | |
| "14:15:58 startup calibrating padding text run=jndwypvdxuxn up_to=32k", | |
| "14:15:58 startup token targeting: estimate from 8k", | |
| "14:15:58 startup calibrated: 6.18 chars/token (cached, source=8k)", | |
| "14:15:58 startup context 16k: 101,202 chars (~16,383 tokens)", | |
| "14:15:58 startup context 32k: 202,404 chars (~32,767 tokens)", | |
| "14:15:58 startup startup preparation done", | |
| "14:15:58 startup hardware monitor disabled", | |
| "14:15:58 decode warmup start", | |
| "14:15:59 decode warmup start C=1 ctx=32k 3s", | |
| "14:15:59 cell start C=1 ctx=32k", | |
| "14:16:30 ready C=1 ctx=32k running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "14:16:33 cell done C=1 ctx=32k 138.8 tok/s", | |
| "14:16:33 decode warmup done C=1 ctx=32k", | |
| "14:16:35 cell start C=1 ctx=0", | |
| "14:16:40 ready C=1 ctx=0 running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "14:16:50 cell done C=1 ctx=0 121.4 tok/s", | |
| "14:16:52 cell start C=1 ctx=16k", | |
| "14:17:12 ready C=1 ctx=16k running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "14:17:22 cell done C=1 ctx=16k 120.8 tok/s", | |
| "14:17:24 cell start C=1 ctx=32k", | |
| "14:17:46 ready C=1 ctx=32k running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "14:17:56 cell done C=1 ctx=32k 109.5 tok/s" | |
| ], | |
| "prefill": {}, | |
| "results": [ | |
| { | |
| "concurrency": 1, | |
| "context_tokens": 0, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 9.980063, | |
| "measurement_wall_seconds": 10.000236, | |
| "client_output_tokens": 1212, | |
| "server_output_tokens": 1214, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 121.44212273062575, | |
| "per_request_avg_tps": 121.44212273062575, | |
| "ttft_avg": 0.18520536500727758, | |
| "ttft_p50": 0.18520536500727758, | |
| "ttft_p90": 0.18520536500727758, | |
| "ttft_p99": 0.18520536500727758, | |
| "time_to_second_token_avg": 0.019689216976985335, | |
| "time_to_second_token_p50": 0.019689216976985335, | |
| "time_to_second_token_p90": 0.019689216976985335, | |
| "time_to_second_token_p99": 0.019689216976985335, | |
| "request_latency_avg": 0.0, | |
| "request_latency_p50": 0.0, | |
| "request_latency_p90": 0.0, | |
| "request_latency_p99": 0.0, | |
| "inter_token_latency_avg": 0.008039274356577263, | |
| "inter_token_latency_p50": 0.008039274356577263, | |
| "inter_token_latency_p90": 0.008039274356577263, | |
| "inter_token_latency_p99": 0.008039274356577263, | |
| "output_tps_per_user_avg": 124.38933610741356, | |
| "output_tps_per_user_p50": 124.38933610741356, | |
| "output_tps_per_user_p90": 124.38933610741356, | |
| "output_tps_per_user_p99": 124.38933610741356, | |
| "e2e_output_tps_per_user_avg": 0.0, | |
| "e2e_output_tps_per_user_p50": 0.0, | |
| "e2e_output_tps_per_user_p90": 0.0, | |
| "e2e_output_tps_per_user_p99": 0.0, | |
| "chunk_inter_token_latency_avg": 0.021001227668483346, | |
| "chunk_inter_token_latency_p50": 0.021001227668483346, | |
| "chunk_inter_token_latency_p90": 0.021001227668483346, | |
| "chunk_inter_token_latency_p99": 0.021001227668483346, | |
| "input_seq_len_avg": 78.0, | |
| "output_seq_len_avg": 1908.0, | |
| "output_seq_len_p50": 1908.0, | |
| "output_seq_len_p90": 1908.0, | |
| "output_seq_len_p99": 1908.0, | |
| "request_count": 1, | |
| "completed_request_count": 0, | |
| "request_samples": [ | |
| { | |
| "ttft": 0.18520536500727758, | |
| "time_to_second_token": 0.019689216976985335, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.008039274356577263, | |
| "chunk_inter_token_latency_avg": 0.021001227668483346, | |
| "input_tokens": 78, | |
| "output_tokens": 1908, | |
| "output_tps_per_user": 124.38933610741356, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| } | |
| ], | |
| "total_tokens": 1212, | |
| "wall_time": 15.54082437895704, | |
| "num_completed": 1, | |
| "num_errors": 0, | |
| "server_gen_throughput": 121.34292275248903, | |
| "server_utilization": 0.16279069767441856, | |
| "server_spec_accept_rate": 0.4513888888888889, | |
| "server_spec_accept_length": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 0, | |
| "max_queue_reqs": 0, | |
| "queue_fraction": 0.0, | |
| "underfilled": false, | |
| "warmup_timed_out": false, | |
| "warmup_duration": 5.536, | |
| "ready_reason": "running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "timeout_reason": "", | |
| "capacity_limited": false, | |
| "hardware_summary": {} | |
| }, | |
| { | |
| "concurrency": 1, | |
| "context_tokens": 16384, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 10.00012, | |
| "measurement_wall_seconds": 10.00015, | |
| "client_output_tokens": 1208, | |
| "server_output_tokens": 1208, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 120.79854757836178, | |
| "per_request_avg_tps": 120.79854757836178, | |
| "ttft_avg": 4.503566837054677, | |
| "ttft_p50": 4.503566837054677, | |
| "ttft_p90": 4.503566837054677, | |
| "ttft_p99": 4.503566837054677, | |
| "time_to_second_token_avg": 0.02603697299491614, | |
| "time_to_second_token_p50": 0.02603697299491614, | |
| "time_to_second_token_p90": 0.02603697299491614, | |
| "time_to_second_token_p99": 0.02603697299491614, | |
| "request_latency_avg": 0.0, | |
| "request_latency_p50": 0.0, | |
| "request_latency_p90": 0.0, | |
| "request_latency_p99": 0.0, | |
| "inter_token_latency_avg": 0.008216178789141408, | |
| "inter_token_latency_p50": 0.008216178789141408, | |
| "inter_token_latency_p90": 0.008216178789141408, | |
| "inter_token_latency_p99": 0.008216178789141408, | |
| "output_tps_per_user_avg": 121.71108074249929, | |
| "output_tps_per_user_p50": 121.71108074249929, | |
| "output_tps_per_user_p90": 121.71108074249929, | |
| "output_tps_per_user_p99": 121.71108074249929, | |
| "e2e_output_tps_per_user_avg": 0.0, | |
| "e2e_output_tps_per_user_p50": 0.0, | |
| "e2e_output_tps_per_user_p90": 0.0, | |
| "e2e_output_tps_per_user_p99": 0.0, | |
| "chunk_inter_token_latency_avg": 0.02122846844869219, | |
| "chunk_inter_token_latency_p50": 0.02122846844869219, | |
| "chunk_inter_token_latency_p90": 0.02122846844869219, | |
| "chunk_inter_token_latency_p99": 0.02122846844869219, | |
| "input_seq_len_avg": 16229.0, | |
| "output_seq_len_avg": 1590.0, | |
| "output_seq_len_p50": 1590.0, | |
| "output_seq_len_p90": 1590.0, | |
| "output_seq_len_p99": 1590.0, | |
| "request_count": 1, | |
| "completed_request_count": 0, | |
| "request_samples": [ | |
| { | |
| "ttft": 4.503566837054677, | |
| "time_to_second_token": 0.02603697299491614, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.008216178789141408, | |
| "chunk_inter_token_latency_avg": 0.02122846844869219, | |
| "input_tokens": 16229, | |
| "output_tokens": 1590, | |
| "output_tps_per_user": 121.71108074249929, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| } | |
| ], | |
| "total_tokens": 1208, | |
| "wall_time": 29.74803116399562, | |
| "num_completed": 1, | |
| "num_errors": 0, | |
| "server_gen_throughput": 120.74706160198144, | |
| "server_utilization": 0.18604651162790697, | |
| "server_spec_accept_rate": 0.54421768707483, | |
| "server_spec_accept_length": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 0, | |
| "max_queue_reqs": 0, | |
| "queue_fraction": 0.0, | |
| "underfilled": false, | |
| "warmup_timed_out": false, | |
| "warmup_duration": 19.727, | |
| "ready_reason": "running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "timeout_reason": "", | |
| "capacity_limited": false, | |
| "hardware_summary": {} | |
| }, | |
| { | |
| "concurrency": 1, | |
| "context_tokens": 32768, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 10.000126, | |
| "measurement_wall_seconds": 10.000173, | |
| "client_output_tokens": 1095, | |
| "server_output_tokens": 1097, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 109.49862520114951, | |
| "per_request_avg_tps": 109.49862520114951, | |
| "ttft_avg": 8.967420397035312, | |
| "ttft_p50": 8.967420397035312, | |
| "ttft_p90": 8.967420397035312, | |
| "ttft_p99": 8.967420397035312, | |
| "time_to_second_token_avg": 0.022118687978945673, | |
| "time_to_second_token_p50": 0.022118687978945673, | |
| "time_to_second_token_p90": 0.022118687978945673, | |
| "time_to_second_token_p99": 0.022118687978945673, | |
| "request_latency_avg": 0.0, | |
| "request_latency_p50": 0.0, | |
| "request_latency_p90": 0.0, | |
| "request_latency_p99": 0.0, | |
| "inter_token_latency_avg": 0.008766973567183258, | |
| "inter_token_latency_p50": 0.008766973567183258, | |
| "inter_token_latency_p90": 0.008766973567183258, | |
| "inter_token_latency_p99": 0.008766973567183258, | |
| "output_tps_per_user_avg": 114.06444793482937, | |
| "output_tps_per_user_p50": 114.06444793482937, | |
| "output_tps_per_user_p90": 114.06444793482937, | |
| "output_tps_per_user_p99": 114.06444793482937, | |
| "e2e_output_tps_per_user_avg": 0.0, | |
| "e2e_output_tps_per_user_p50": 0.0, | |
| "e2e_output_tps_per_user_p90": 0.0, | |
| "e2e_output_tps_per_user_p99": 0.0, | |
| "chunk_inter_token_latency_avg": 0.021708031683073194, | |
| "chunk_inter_token_latency_p50": 0.021708031683073194, | |
| "chunk_inter_token_latency_p90": 0.021708031683073194, | |
| "chunk_inter_token_latency_p99": 0.021708031683073194, | |
| "input_seq_len_avg": 32322.0, | |
| "output_seq_len_avg": 1556.0, | |
| "output_seq_len_p50": 1556.0, | |
| "output_seq_len_p90": 1556.0, | |
| "output_seq_len_p99": 1556.0, | |
| "request_count": 1, | |
| "completed_request_count": 0, | |
| "request_samples": [ | |
| { | |
| "ttft": 8.967420397035312, | |
| "time_to_second_token": 0.022118687978945673, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.008766973567183258, | |
| "chunk_inter_token_latency_avg": 0.021708031683073194, | |
| "input_tokens": 32322, | |
| "output_tokens": 1556, | |
| "output_tps_per_user": 114.06444793482937, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| } | |
| ], | |
| "total_tokens": 1095, | |
| "wall_time": 31.743410580966156, | |
| "num_completed": 1, | |
| "num_errors": 0, | |
| "server_gen_throughput": 109.5323841476318, | |
| "server_utilization": 0.2093023255813954, | |
| "server_spec_accept_rate": 0.427536231884058, | |
| "server_spec_accept_length": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 0, | |
| "max_queue_reqs": 0, | |
| "queue_fraction": 0.0, | |
| "underfilled": false, | |
| "warmup_timed_out": false, | |
| "warmup_duration": 21.723, | |
| "ready_reason": "running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "timeout_reason": "", | |
| "capacity_limited": false, | |
| "hardware_summary": {} | |
| } | |
| ], | |
| "summary_table": { | |
| "0": { | |
| "1": 121.44212273062575 | |
| }, | |
| "16384": { | |
| "1": 120.79854757836178 | |
| }, | |
| "32768": { | |
| "1": 109.49862520114951 | |
| } | |
| }, | |
| "burst_results": [], | |
| "burst_summary_table": {}, | |
| "methodology": { | |
| "prefill": { | |
| "name": "Prefill", | |
| "present": false, | |
| "mode": "skipped", | |
| "formula": "prompt_tokens / TTFT", | |
| "notes": "Default mode records the required decode scout request for each non-zero decode context, so normal runs do not pay for a separate prefill phase. Standalone mode repeats cold-prefill samples. Prometheus prefill counters, when available and uncontaminated, are stored as validation." | |
| }, | |
| "sustained_decode": { | |
| "name": "Sustained Decode", | |
| "present": true, | |
| "formula": "OpenAI stream usage completion_tokens per measured window; client chunk fallback only when continuous usage is unavailable", | |
| "notes": "Duration-based steady-state cell after warmup. This is the main tuning/regression signal for kernels, NCCL, DCP, MTP, and scheduling. Prometheus metrics are stored as validation and scheduler state, not the default headline." | |
| }, | |
| "burst_e2e_decode": { | |
| "name": "Burst / E2E Decode", | |
| "present": false, | |
| "status": "not run; use --run-burst", | |
| "formula": "sum(completion_tokens) / profiling_wall_time", | |
| "notes": "Finite client-facing request burst using OpenAI stream usage. It includes request admission, scheduling, prefill/cache behavior, and completion." | |
| } | |
| } | |
| } |