MTP Viability Test Plan: Quant x Temperature x Task Type
Goal
Map the complete MTP speculative decoding viability landscape for Qwen3.6-27B on M2 Max 96GB. Produce a recommendation table for each quant/temperature/task combination showing whether MTP helps, hurts, or is neutral, with measured speedups.
Test Parameters
- Token budget:
-n 2500per test (expect 1500-2500 actual, model may stop early) - Context:
-c 4096(sufficient for prompt + 2500 generation) - MTP config:
--spec-type mtp --spec-draft-n-max 3(N=3) - Chat template:
--jinja --chat-template-file <path> -sys '<system prompt>' - Exit:
printf '/exit\n'piped to stdin - Verbosity:
-vto capture acceptance stats - GPU layers:
-ngl 99
Variables
Quants (4 levels)
Q4_K_M (16G), Q5_K_M (18G), Q6_K (21G), Q8_0 (27G)
Rationale: Skip iQ quants (IQ2_M, IQ3_M, IQ4_XS) β they showed inconsistent results and imatrix has no measurable effect on acceptance. Focus on Q quants for clean data.
Temperatures (3 levels)
0.0 (greedy/deterministic), 0.3 (low randomness), 0.7 (standard creative)
Task Types (5 categories, determinism-ordered)
Tasks are ordered from most deterministic (highest expected acceptance) to least.
T1: Code Generation (very high determinism)
Syntax is rigid, most tokens are forced by grammar. Greedy draft β verification.
Write a complete Python implementation of an LRU cache class with get(key), put(key, value),
delete(key), size(), and clear(). Use a doubly-linked list combined with a hash map for O(1)
operations. Include type hints, docstrings, and 10 unit tests using pytest with good coverage
of edge cases.
Expected: ~1500-2000 tokens
T2: Factual Explanation (high determinism)
Technical content with specific terminology and logical structure. One dominant token per position.
Explain in detail how the Transformer architecture works. Cover: token embedding, positional
encoding (sinusoidal and RoPE), multi-head self-attention including Q/K/V projections and
scaled dot-product attention, feed-forward networks, layer normalization (pre-norm vs post-norm),
residual connections, and the difference between encoder-decoder and decoder-only variants.
Include the key mathematical equations for attention and positional encoding.
Expected: ~1500-2500 tokens
T3: Technical Analysis (medium determinism)
Constrained by technical facts but allows structural freedom in organization and phrasing.
Analyze the tradeoffs of quantization in large language models. Cover: memory reduction,
inference speed gains on bandwidth-limited hardware, quality degradation patterns, different
quantization schemes (uniform quantization, k-quant mixtures, importance matrix guided),
perplexity impact at different bit widths, and the practical recommendations for selecting
a quantization level for deployment. Be thorough and technical.
Expected: ~1500-2500 tokens
T4: Expository Essay (low determinism)
Structured but high freedom in content organization, word choice, and emphasis.
Write a detailed 1500-word essay on the history and cultivation of lychee. Cover its origins
in southern China, spread to Southeast Asia and beyond, ideal growing conditions, major
commercial varieties, harvesting and post-harvest handling challenges, and economic importance
in global trade.
Expected: ~1500-2000 tokens
T5: Creative Writing (very low determinism)
Maximum freedom β many equally valid continuations at every position. Worst case for MTP.
Write a short story (about 1500 words) about a lone astronaut aboard a generation ship who
discovers that the ship's AI has been secretly altering the crew's memories over centuries.
The story should build tension gradually and have a surprising but logically consistent ending.
Expected: ~1500-2000 tokens
Already Completed
| Quant | Temp | Task | Type | Speed | Acceptance | Source |
|---|---|---|---|---|---|---|
| Q4_K_M | 0.0 | Essay | MTP | 15.56 | 59.4% | Previous |
| Q4_K_M | 0.0 | Essay | Base | 14.96 | β | Previous |
| Q4_K_M | 0.3 | Essay | MTP | 15.18 | 57.2% | Previous |
| Q4_K_M | 0.3 | Essay | Base | 15.09 | β | Previous |
| Q4_K_M | 0.7 | Essay | MTP | 14.95 | 55.8% | Previous |
| Q4_K_M | 0.7 | Essay | Base | 15.07 | β | Previous |
| Q4_K_M+imat | 0.7 | Essay | MTP | 14.86 | 55.0% | Previous |
| Q5_K_M | 0.7 | Essay | MTP | 13.83 | 55.1% | Previous |
| Q6_K | 0.7 | Essay | MTP | 14.97 | 54.8% | Previous |
| Q8_0 | 0.7 | Essay | MTP | 18.33 | 55.7% | Previous |
| Q8_0 | 0.7 | Essay | Base | ? | β | Running |
| IQ3_M | 0.7 | Essay | MTP | 12.48 | 54.1% | Previous |
Test Plan
Phase 1: Task Type Effect (16 tests, ~45 min)
Test all 4 quants Γ 4 new task types (T1-T3, T5) at temp 0.7 with MTP. Essay (T4) already done for all quants.
| # | Quant | Task | Temp | Type | Notes |
|---|---|---|---|---|---|
| 1 | Q4_K_M | Code | 0.7 | MTP | |
| 2 | Q4_K_M | Factual | 0.7 | MTP | |
| 3 | Q4_K_M | Analysis | 0.7 | MTP | |
| 4 | Q4_K_M | Creative | 0.7 | MTP | |
| 5 | Q5_K_M | Code | 0.7 | MTP | |
| 6 | Q5_K_M | Factual | 0.7 | MTP | |
| 7 | Q5_K_M | Analysis | 0.7 | MTP | |
| 8 | Q5_K_M | Creative | 0.7 | MTP | |
| 9 | Q6_K | Code | 0.7 | MTP | |
| 10 | Q6_K | Factual | 0.7 | MTP | |
| 11 | Q6_K | Analysis | 0.7 | MTP | |
| 12 | Q6_K | Creative | 0.7 | MTP | |
| 13 | Q8_0 | Code | 0.7 | MTP | |
| 14 | Q8_0 | Factual | 0.7 | MTP | |
| 15 | Q8_0 | Analysis | 0.7 | MTP | |
| 16 | Q8_0 | Creative | 0.7 | MTP |
Phase 2: Temperature x Task Interaction (8 tests, ~25 min)
Test code (T1, most deterministic) and creative (T5, least deterministic) at temp 0.0 and 0.3 for Q4_K_M and Q8_0 (representing low and high quality bounds).
| # | Quant | Task | Temp | Type | Notes |
|---|---|---|---|---|---|
| 17 | Q4_K_M | Code | 0.0 | MTP | Best case: low quant + deterministic task |
| 18 | Q4_K_M | Code | 0.3 | MTP | |
| 19 | Q4_K_M | Creative | 0.0 | MTP | Low quant + creative at greedy |
| 20 | Q4_K_M | Creative | 0.3 | MTP | |
| 21 | Q8_0 | Code | 0.0 | MTP | Best case: high quant + deterministic task |
| 22 | Q8_0 | Code | 0.3 | MTP | |
| 23 | Q8_0 | Creative | 0.0 | MTP | High quant + creative at greedy |
| 24 | Q8_0 | Creative | 0.3 | MTP |
Phase 3: Missing Baselines (3 tests, ~10 min)
One baseline per quant at temp 0.7. Task type doesn't significantly affect baseline speed.
| # | Quant | Temp | Type | Notes |
|---|---|---|---|---|
| 25 | Q5_K_M | 0.7 | Base | |
| 26 | Q6_K | 0.7 | Base | |
| 27 | IQ3_M | 0.7 | Base | Optional, lower priority |
Q4_K_M baselines already done. Q8_0 baseline running.
Phase 4: Fill-in Tests (conditional, ~15 min)
Based on Phase 1-3 results, test additional combinations to clarify transitions:
- If code at Q4_K_M/temp 0.7 is above breakeven: test Q4_K_M/factual at temp 0.3
- If creative at Q8_0/temp 0.7 is above breakeven: verify with analysis task
- Test any quant/task/temperature combinations near the breakeven threshold
Up to 6 additional tests, TBD based on results.
Total: 27 core tests + ~6 conditional = ~33 tests, ~90 min
Test Execution
Command Template
CLI="/Volumes/ssd/ai/llm-dev/llama.cpp-mtp/build/bin/llama-cli"
MODEL="<quant-file>"
TEMPLATE="/Volumes/ssd/ai/llm-models/froggeric/Qwen-Fixed-Chat-Templates/qwen3.6/chat_template-v9.jinja"
SYS='You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|think_off|>'
PROMPT="<task-prompt>"
# MTP test
printf '/exit\n' | $CLI -m "$MODEL" \
--spec-type mtp --spec-draft-n-max 3 \
-c 4096 -n 2500 --temp <temp> -ngl 99 \
--jinja --chat-template-file "$TEMPLATE" -sys "$SYS" -p "$PROMPT" \
-v 2>&1 | tee <output-file> | grep -E "accept|draft|statistics|tokens per second|eval time|prompt eval" | tail -10
# Baseline test
printf '/exit\n' | $CLI -m "$MODEL" \
-c 4096 -n 2500 --temp <temp> -ngl 99 \
--jinja --chat-template-file "$TEMPLATE" -sys "$SYS" -p "$PROMPT" \
-v 2>&1 | tee <output-file> | grep -E "tokens per second|eval time|prompt eval" | tail -5
Execution Strategy
Run tests sequentially (single GPU). Chain commands in groups:
- Phase 1: 4 quants Γ 4 tasks = 16 sequential tests, ~45 min
- Phase 2: 8 sequential tests, ~25 min
- Phase 3: 3 sequential tests, ~10 min
- Phase 4: conditional
Use tee to capture full output for post-hoc analysis if grep misses something.
Metrics to Extract
From each test output:
- eval time and tokens generated β decode speed (tok/s)
- draft acceptance rate (MTP tests only)
- #gen tokens and #acc tokens from statistics line
Analysis Plan
1. Acceptance Rate Heatmap
Code Factual Analysis Essay Creative
Q4_K_M 0.0 ? ? ? 59.4% ?
Q4_K_M 0.3 ? ? ? 57.2% ?
Q4_K_M 0.7 ? ? ? 55.8% ?
Q5_K_M 0.7 ? ? ? 55.1% ?
Q6_K 0.7 ? ? ? 54.8% ?
Q8_0 0.0 ? ? ? ? ?
Q8_0 0.3 ? ? ? ? ?
Q8_0 0.7 ? ? ? 55.7% ?
2. Speedup Table
Code Factual Analysis Essay Creative
Q4_K_M ?% ?% ?% -0.8% ?%
Q5_K_M ?% ?% ?% ?% ?%
Q6_K ?% ?% ?% ?% ?%
Q8_0 ?% ?% ?% ?% ?%
Speedup = (MTP_speed / Base_speed - 1) Γ 100% Breakeven acceptance for N=3 β 59% (but depends on MTP overhead ratio)
3. Final Recommendation Matrix
For each quant, recommend:
- Green: MTP recommended (measurable speedup)
- Yellow: MTP neutral (marginal, task-dependent)
- Red: MTP not recommended (net slowdown)
Organized by use case:
- Coding assistant (temp 0.0-0.3)
- General chatbot (temp 0.7)
- Technical writing (temp 0.3-0.5)
- Creative writing (temp 0.7-1.0)