ZTFlynn commited on
Commit
53aabf9
·
verified ·
1 Parent(s): f7dc786

Correct GPU speedup: 14-16x (baseline was measured at the wrong thread count)

Browse files
Files changed (1) hide show
  1. README.md +14 -0
README.md CHANGED
@@ -120,6 +120,20 @@ Sampling is `--temp` / `--top-k` / `--top-p` / `--seed`; the default is greedy
120
  and seed-reproducible. Generation stops at `<|im_end|>`, so `max_new` is a
121
  ceiling.
122
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
  ### Python
124
 
125
  ```python
 
120
  and seed-reproducible. Generation stops at `<|im_end|>`, so `max_new` is a
121
  ceiling.
122
 
123
+ ### On an NVIDIA GPU
124
+
125
+ The same package runs entirely on the GPU, with byte-identical output and
126
+ roughly 14-16x the throughput. Opt-in, so the default build is unchanged:
127
+
128
+ ```bash
129
+ cmake -S src/c -B build -DCASCADIA_CUDA=ON -DCMAKE_BUILD_TYPE=Release \
130
+ && cmake --build build -j
131
+ ./build/cascadia_generate_cuda ./pkg 512 --chat "Explain gradient descent."
132
+ ```
133
+
134
+ Same arguments, same tokens. Needs CUDA and sm_75 or newer.
135
+
136
+
137
  ### Python
138
 
139
  ```python