qianlihuang commited on
Commit
1d7426f
·
verified ·
1 Parent(s): c8ee49b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +6 -0
README.md CHANGED
@@ -53,6 +53,12 @@ Accepted length is calculated from the raw server counters:
53
  accepted_length = 1 + accepted_tokens / draft_calls
54
  ```
55
 
 
 
 
 
 
 
56
  ## Warm start
57
 
58
  The inherited five-layer DFlash body is already trained, while the Markov and confidence heads are newly initialized. Applying the same `6e-4` peak learning rate to every parameter caused the warm-started body to lose some early-token accuracy during the high-LR phase. The released run uses:
 
53
  accepted_length = 1 + accepted_tokens / draft_calls
54
  ```
55
 
56
+ ### Per-position acceptance
57
+
58
+ ![Per-position acceptance rate on all 13 subsets: Official DFlash vs DSpark](per_position_acceptance.png)
59
+
60
+ Per-position acceptance curves for the same runs as the tables above. DSpark shows slightly lower position-0 acceptance but substantially stronger acceptance deeper into the proposal, with the largest gains toward the tail.
61
+
62
  ## Warm start
63
 
64
  The inherited five-layer DFlash body is already trained, while the Markov and confidence heads are newly initialized. Applying the same `6e-4` peak learning rate to every parameter caused the warm-started body to lose some early-token accuracy during the high-LR phase. The released run uses: