Text Generation
memra
Safetensors
qwen3
speculative-decoding
dspark
dflash
draft-model
specforge
sglang
agentic
conversational
Eval Results (legacy)
Instructions to use tiyuvta/Qwen3.8-27B-DSpark-Agentic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- memra
How to use tiyuvta/Qwen3.8-27B-DSpark-Agentic with memra:
# memra serves NVIDIA Blackwell workstation and consumer cards (sm_120a), with a # compile-gated Hopper lane. Prebuilt binaries need Linux x86_64 and driver 580+, # and no CUDA toolkit. curl -fsSL https://raw.githubusercontent.com/avifenesh/memra/main/tools/install.sh | sh
# One chat-templated generation. In a repo with several GGUF files, append # :<substring> to choose one, for example hf:tiyuvta/Qwen3.8-27B-DSpark-Agentic:Q4_K_M MEMRA_CHAT=1 run-gen hf:tiyuvta/Qwen3.8-27B-DSpark-Agentic --prompt "Explain KV caches in one sentence."
# Or an OpenAI-compatible server on 127.0.0.1:8080. MEMRA_MODELS="model=hf:tiyuvta/Qwen3.8-27B-DSpark-Agentic" memra-server
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| base_model: Qwen/Qwen3.8-27B | |
| inference: false | |
| tags: | |
| - speculative-decoding | |
| - dspark | |
| - dflash | |
| - draft-model | |
| - specforge | |
| - sglang | |
| - memra | |
| - qwen3 | |
| - agentic | |
| - conversational | |
| model-index: | |
| - name: Qwen3.8-27B-DSpark-Agentic | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Speculative decoding acceptance, DSpark block size 7 | |
| dataset: | |
| name: held-out real agentic session prompts (128 turns, two length buckets) | |
| type: agentic-sessions-heldout | |
| metrics: | |
| - name: mean acceptance length, short chat turns (n=64) | |
| type: acceptance-length | |
| value: 2.88 | |
| - name: mean acceptance length, long agentic turns (n=64) | |
| type: acceptance-length | |
| value: 2.92 | |
| - task: | |
| type: text-generation | |
| name: Speculative decoding acceptance, DSpark block size 7 | |
| dataset: | |
| name: GSM8K | |
| type: openai/gsm8k | |
| metrics: | |
| - name: mean acceptance length (128 prompts, concurrency 8) | |
| type: acceptance-length | |
| value: 4.61 | |
| - name: aggregate decode throughput, 1x RTX PRO 6000 Blackwell 96GB | |
| type: throughput | |
| value: 693 | |
| # Qwen3.8-27B-DSpark-Agentic | |
| A **DSpark speculative-decoding drafter** for | |
| [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B), trained on | |
| **real agentic coding/assistant sessions**. DSpark drafts a block of tokens | |
| conditioned on auxiliary features tapped from the target model's layers, with a | |
| Markov head for intra-block token dependency and a confidence head that predicts | |
| per-position acceptance. Trained with | |
| [SpecForge](https://github.com/sgl-project/SpecForge); serves with | |
| [SGLang](https://github.com/sgl-project/sglang). | |
| A drafter never changes model output: every drafted token is verified by the | |
| target model. It only changes decode speed, through the acceptance rate on your | |
| traffic. This one is aimed at agentic/assistant workloads. | |
| ## Model specifications | |
| | | | | |
| |---|---| | |
| | Target model | [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) (evals below ran FP8 and NVFP4 builds of this trunk) | | |
| | Draft parameters | 1,359,284,737 (1.36B), BF16, 62 tensors in one safetensors file | | |
| | Architecture | 5 full-attention transformer layers, hidden 5,120, GQA 40 query / 8 KV heads, head_dim 128, MLP 10,240 | | |
| | Target feature taps | layers 4, 16, 28, 40, 52 (of 64) | | |
| | Heads | vanilla Markov head, rank 256; confidence head with Markov features | | |
| | DSpark block size | 7 draft tokens (verify width 8 with the target bonus token) | | |
| | Positional encoding | YaRN (factor 32, original 8,192), max position 262,144 | | |
| | Vocabulary | 248,320 (uses the target model's tokenizer; no separate tokenizer shipped) | | |
| `model.safetensors` sha256: | |
| `3c4e7d51f6a7492fd3c289e392424ab0fba31dc7b4ab4896da10a5eeb21fcd6f` | |
| ## Training | |
| Trained **from scratch** (random initialization β no third-party model weights) | |
| with the SpecForge harness. Labels are the target model's own temperature-0 | |
| regenerations over a prompt mix of real agentic coding/assistant sessions and a | |
| public chat prompt blend, covering both thinking and non-thinking modes, | |
| training on the last turn of each multi-turn session. Roughly 1,000 cumulative | |
| optimizer steps on B200 / RTX PRO 6000-class hardware; this repo is the final | |
| checkpoint of that run. | |
| ## Evaluation results | |
| All numbers below are absolute measurements of **this checkpoint**. | |
| ### Acceptance length β SGLang, FP8 target, sampled thinking-mode decoding | |
| Setup: SGLang; target `Qwen/Qwen3.8-27B-FP8` on a single RTX PRO 6000 Blackwell | |
| (96 GB); this drafter unquantized BF16; DSpark block size 7; temperature 0.6, | |
| top-k 20, top-p 0.95; thinking enabled; `max_new_tokens=2048`. Acceptance | |
| length = mean accepted tokens per verification step including the target bonus | |
| token (per-request `completion_tokens / spec_verify_ct`, unweighted mean). | |
| | Workload | Prompts | Acceptance length | | |
| |---|---:|---:| | |
| | Real agentic sessions β short chat turns | 64 | **2.88** (per-request 2.09β5.94) | | |
| | Real agentic sessions β long agentic turns | 64 | **2.92** (per-request 2.03β5.45) | | |
| | GSM8K | 128, concurrency 8 | **4.61** | | |
| The GSM8K cell decoded 60,365 output tokens in 87.1 s β **693 tok/s aggregate** | |
| at concurrency 8 on the single card. The two session buckets produced 81,199 | |
| and 115,263 output tokens respectively. The session prompts are held-out real | |
| agentic/assistant turns, not used in training. | |
| ### Acceptance β greedy argmax, NVFP4 GGUF trunk (memra engine) | |
| Setup: [memra](https://github.com/avifenesh/memra) engine, greedy decoding, | |
| single stream, RTX PRO 6000 Blackwell (96 GB); trunk = | |
| [NVFP4 GGUF build of the same target](https://huggingface.co/tiyuvta/Qwen3.8-27B-NVFP4-MTP-GGUF); | |
| DSpark block size 7, with target features harvested under the shifted-label | |
| (DSPARK-strategy) convention this drafter was trained on. The speculative | |
| stream is gated **byte-identical** to plain greedy decode on every prompt | |
| (verified across 5 repetitions). Tokens per round = 1 + accepted/rounds, read | |
| from the server's per-request speculation counters; 256-token completions. | |
| | Workload | Prompts | Tokens per round β adaptive verify | Tokens per round β full-block verify | | |
| |---|---:|---:|---:| | |
| | Agentic-session prompts | 16 | **2.412** | 2.784 | | |
| | Short chat prompts | 16 | **2.228** | 2.415 | | |
| | GSM8K | 12 | **3.664** | 4.620 | | |
| The adaptive column uses the engine's default adaptive verification-window | |
| ladder; the full-block column verifies the whole 7-token draft block plus the | |
| bonus token every round β the same verification protocol as the SGLang table | |
| above. | |
| Single-stream greedy throughput on held-out agentic-session prompts, 128 | |
| tokens per prompt: **96.3 tok/s** with this drafter vs **65.0 tok/s** plain | |
| decode on the same engine, card, and prompts (**1.48Γ**; per-repetition means | |
| 95.4β96.3 vs 63.1β65.1 across 5 runs), with speculative output byte-identical | |
| to plain decode on every prompt. | |
| These memra numbers require the shifted-label (DSPARK-strategy) feature | |
| harvest in the engine. That harvest mode is landing in an upcoming memra | |
| release; it is not part of any tagged memra release yet. | |
| Acceptance is protocol-dependent: sampled thinking-mode decoding on an FP8 | |
| trunk (first table) and greedy argmax on an NVFP4 trunk (second table) are | |
| different observables. Both are reported as measured; expect your own numbers | |
| to depend on trunk quantization, sampling settings, verification width, and | |
| traffic shape. | |
| ## Serving | |
| SGLang (use a version with DSPARK support): | |
| ```bash | |
| sglang serve \ | |
| --trust-remote-code \ | |
| --model-path Qwen/Qwen3.8-27B-FP8 \ | |
| --tp-size 1 \ | |
| --speculative-algorithm DSPARK \ | |
| --speculative-draft-model-path tiyuvta/Qwen3.8-27B-DSpark-Agentic \ | |
| --speculative-dspark-block-size 7 \ | |
| --speculative-draft-model-quantization unquant \ | |
| --mamba-scheduler-strategy extra_buffer | |
| ``` | |
| [memra](https://github.com/avifenesh/memra) (Rust + CUDA engine for RTX | |
| Blackwell) has the DSpark route wired and exactness-gated β spec output | |
| byte-identical to plain decode β landing in an upcoming release behind an | |
| env-gated flag. | |
| The team behind this drafter serves Qwen3.8-27B in production at | |
| [inference.tiyuvta.ai](https://inference.tiyuvta.ai) β native 262,144-token | |
| context, OpenAI Chat Completions / Responses and Anthropic Messages on one | |
| endpoint, tool calling included. Measured through the public endpoint on the | |
| serving build (2026-08-22, single stream, greedy, streamed, medians per | |
| output length): 136 tok/s at 128 output tokens, 259 at 512, 166 at 2048; | |
| turn-8 first-token time of an 8-turn agentic conversation 1.07 s at a | |
| 38k-token prompt (95% prefix-cache hit). Dated measurements of the live | |
| service, not commitments. | |
| ## License | |
| Apache-2.0. The drafter was trained from scratch; no third-party model weights | |
| were used for initialization. Training labels were generated by | |
| [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) (Apache-2.0); | |
| training used the open-source SpecForge harness. | |
| ## Citations | |
| - DSpark (speculative decoding method): [arXiv:2607.05147](https://arxiv.org/abs/2607.05147) | |
| - DFlash (block drafting lineage): [z-lab/dflash](https://github.com/z-lab/dflash), [arXiv:2602.06036](https://arxiv.org/abs/2602.06036) | |
| - SpecForge (training harness): [sgl-project/SpecForge](https://github.com/sgl-project/SpecForge) | |