Transformers
Safetensors
speculative-decoding
dspark
dflash
speculators
vllm
muse-glimmer
custom_code
Instructions to use DaoCloud/Muse-Glimmer-30B-DSpark with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DaoCloud/Muse-Glimmer-30B-DSpark with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("DaoCloud/Muse-Glimmer-30B-DSpark", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -53,6 +53,12 @@ Accepted length is calculated from the raw server counters:
|
|
| 53 |
accepted_length = 1 + accepted_tokens / draft_calls
|
| 54 |
```
|
| 55 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
## Warm start
|
| 57 |
|
| 58 |
The inherited five-layer DFlash body is already trained, while the Markov and confidence heads are newly initialized. Applying the same `6e-4` peak learning rate to every parameter caused the warm-started body to lose some early-token accuracy during the high-LR phase. The released run uses:
|
|
|
|
| 53 |
accepted_length = 1 + accepted_tokens / draft_calls
|
| 54 |
```
|
| 55 |
|
| 56 |
+
### Per-position acceptance
|
| 57 |
+
|
| 58 |
+

|
| 59 |
+
|
| 60 |
+
Per-position acceptance curves for the same runs as the tables above. DSpark shows slightly lower position-0 acceptance but substantially stronger acceptance deeper into the proposal, with the largest gains toward the tail.
|
| 61 |
+
|
| 62 |
## Warm start
|
| 63 |
|
| 64 |
The inherited five-layer DFlash body is already trained, while the Markov and confidence heads are newly initialized. Applying the same `6e-4` peak learning rate to every parameter caused the warm-started body to lose some early-token accuracy during the high-LR phase. The released run uses:
|