Correct GPU speedup: 14-16x (baseline was measured at the wrong thread count)
Browse files
README.md
CHANGED
|
@@ -120,6 +120,20 @@ Sampling is `--temp` / `--top-k` / `--top-p` / `--seed`; the default is greedy
|
|
| 120 |
and seed-reproducible. Generation stops at `<|im_end|>`, so `max_new` is a
|
| 121 |
ceiling.
|
| 122 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 123 |
### Python
|
| 124 |
|
| 125 |
```python
|
|
|
|
| 120 |
and seed-reproducible. Generation stops at `<|im_end|>`, so `max_new` is a
|
| 121 |
ceiling.
|
| 122 |
|
| 123 |
+
### On an NVIDIA GPU
|
| 124 |
+
|
| 125 |
+
The same package runs entirely on the GPU, with byte-identical output and
|
| 126 |
+
roughly 14-16x the throughput. Opt-in, so the default build is unchanged:
|
| 127 |
+
|
| 128 |
+
```bash
|
| 129 |
+
cmake -S src/c -B build -DCASCADIA_CUDA=ON -DCMAKE_BUILD_TYPE=Release \
|
| 130 |
+
&& cmake --build build -j
|
| 131 |
+
./build/cascadia_generate_cuda ./pkg 512 --chat "Explain gradient descent."
|
| 132 |
+
```
|
| 133 |
+
|
| 134 |
+
Same arguments, same tokens. Needs CUDA and sm_75 or newer.
|
| 135 |
+
|
| 136 |
+
|
| 137 |
### Python
|
| 138 |
|
| 139 |
```python
|