vikp commited on
Commit
f18ef82
·
verified ·
1 Parent(s): 0aea19c

Update throughput numbers (5090 conc-sweep + Mac single row)

Browse files
Files changed (1) hide show
  1. README.md +13 -13
README.md CHANGED
@@ -109,27 +109,27 @@ table_predictions = table.predict_full([Image.open("page.png")]) # full <table>
109
 
110
  ## Throughput
111
 
112
- Full-page OCR, 96 DPI input (~3,000 output tokens/page average), measured client-side against a running inference server.
113
 
114
  ### RTX 5090 (vllm)
115
 
116
- `vllm/vllm-openai:v0.20.1`, single RTX 5090 (32 GB), client-side concurrency 128, prefix caching off.
117
 
118
- | Concurrency | Pages/s | Tokens/s | Peak power |
119
- |---:|---:|---:|---:|
120
- | 128 | **5.64** | **13,829** | ~478 W (80 % of 600 W TDP) |
 
 
121
 
122
- ### Apple Silicon (llama.cpp / Metal)
123
 
124
- `llama-server` with Metal backend, one process per `--parallel` level.
125
 
126
- | `--parallel` | Pages/s | Tokens/s | p50 (ms) | p95 (ms) |
127
- |---:|---:|---:|---:|---:|
128
- | 4 | 0.217 | 222 | 18,122 | 19,676 |
129
- | **8** | **0.269** | **270** | 27,549 | 42,730 |
130
- | 16 | 0.264 | 286 | 53,166 | 70,378 |
131
 
132
- Knees at `--parallel=8` — Metal is decode-saturated past that point.
 
 
133
 
134
  ## Commercial Usage
135
 
 
109
 
110
  ## Throughput
111
 
112
+ Full-page OCR, 96 DPI input (~2,400 output tokens/page average), measured client-side against a running inference server.
113
 
114
  ### RTX 5090 (vllm)
115
 
116
+ `vllm/vllm-openai:v0.20.1`, single RTX 5090 (32 GB). Sustained power ~478 W (80% of 600 W TDP) across all concurrencies.
117
 
118
+ | Concurrency | Pages/s | Tokens/s | p50 (ms) | p95 (ms) | avg tok/page |
119
+ |---:|---:|---:|---:|---:|---:|
120
+ | 32 | 3.67 | 8,870 | 6,744 | 21,741 | 2,420 |
121
+ | 64 | 4.67 | 11,280 | 10,741 | 34,639 | 2,414 |
122
+ | **128** | **5.35** | **12,884** | 18,915 | 42,538 | 2,410 |
123
 
124
+ Throughput climbs from conc=32 → 128 but latency grows faster than capacity. Pick conc=64 for the latency/throughput knee, conc=128 for max throughput.
125
 
126
+ ### Apple Silicon (llama.cpp / Metal)
127
 
128
+ `llama-server` with Metal backend.
 
 
 
 
129
 
130
+ | `--parallel` | Pages/s | Tokens/s | p50 (ms) | p95 (ms) | avg tok/page | Power |
131
+ |---:|---:|---:|---:|---:|---:|---:|
132
+ | **8** | **0.108** | **254** | 59,313 | 129,173 | 2,360 | ~30 W |
133
 
134
  ## Commercial Usage
135