|
Download TECHNICAL-DETAILS.md from Accio-Lab/occamy-1.0-MTP: direct link, hf CLI and curl.
- Browser
- Download file 5.84 kB
-
https://huggingface.co/Accio-Lab/occamy-1.0-MTP/resolve/main/TECHNICAL-DETAILS.md
- Command line
-
hf download hf://Accio-Lab/occamy-1.0-MTP/TECHNICAL-DETAILS.md
-
curl -L -o TECHNICAL-DETAILS.md https://huggingface.co/Accio-Lab/occamy-1.0-MTP/resolve/main/TECHNICAL-DETAILS.md
5.84 kB
| license: apache-2.0 | |
| base_model: Accio-Lab/occamy-1.0 | |
| library_name: sglang | |
| tags: | |
| - mtp | |
| - speculative-decoding | |
| - draft-head | |
| - qwen3_5_moe | |
| datasets: | |
| - Accio-Lab/occamy-data-1.0 | |
| # Occamy 1.0 MTP | |
| A BF16 MTP draft head adapted on [Occamy data](https://huggingface.co/datasets/Accio-Lab/occamy-data-1.0) for **512 steps**, using frozen [Occamy 1.0](https://huggingface.co/Accio-Lab/occamy-1.0) predictions as targets. This replaces the earlier 32-step pilot. Download a base checkpoint separately; this is a custom MTP head, not a standalone model or PEFT adapter. | |
| ## Training | |
| Initialized from `Qwen/Qwen3.6-35B-A3B`. Only 8,392,704 fusion and normalization parameters were trained; the rest of the 844,640,768-parameter head and the Occamy base stayed frozen. Training used 256 text examples over two epochs, with 32 heldout examples and assistant-only loss. Windows contain up to 128 context and 382 assistant tokens. Images and unnormalized tool-call roles were excluded; this is bounded text adaptation, not full-head or full-conversation training. | |
| Exported-head target top-1 agreement rose from **73.57% to 74.66%**. Data-answer NLL worsened from **2.2850 to 2.6034**: matching the target model's greedy choices differs from predicting the recorded answers. No external teacher was used. Exact revisions and settings are in [TRAINING.json](TRAINING.json). | |
| ## Verified combinations | |
| The **same head** was tested on BF16 and [Occamy NVFP4](https://huggingface.co/Accio-Lab/occamy-1.0-NVFP4). Each variant used 10 authored prompts and 8 Occamy heldout prefixes, repeated twice. Comparisons use the matching precision's no-MTP baseline. | |
| | Combination | Occamy heldout accepted / proposed | Heldout acceptance | Exact token sequences | Exact token logprobs | | |
| |---|---:|---:|---:|---:| | |
| | BF16 + original donor head | 1128 / 1292 | 87.31% | 36/36 | 36/36 | | |
| | BF16 + released head | 1130 / 1290 | **87.60%** | **36/36** | **36/36** | | |
| | NVFP4 + released head | 1178 / 1372 | **85.86%** | **36/36** | **36/36** | | |
| Across authored and heldout prompts combined, acceptance was 1470/1668 (88.13%) for BF16 and 1506/1756 (85.76%) for NVFP4. The small BF16 acceptance gain does **not establish a meaningful speed improvement**. Timings had concurrent activity and are diagnostic only. NVFP4 has no same-run original-donor-head comparison. Some heldout generations hit the 160-token limit; parity applies to the measured outputs, not unbounded continuations. These checks do not establish broad task accuracy. | |
| This release validates **single-request greedy MTP1**, with two verification nodes, context length 2048, deterministic inference, and the supplied hooks. CUDA graphs, overlap and prefix caching were disabled. FP8, GGUF, sampling, concurrency, MTP2, multimodal MTP and other engines are **not validated with this head**. Results for older heads do not apply to it. | |
| ## Required runtime | |
| Tested with SGLang **0.5.13.post1**, source revision `85fd90072d1a9f2432842b03588f63b745e524e4`, and PyTorch `2.11.0+cu130`. Use `canonical_attention:install` from the included `runtime/` directory. The hooks prevent stale prefill-state initialization during verification and reuse decode kernels for recurrent and full attention. Unsupported batch/tree shapes fail closed. Exact parity is bounded evidence for this configuration, not a universal guarantee or a claim about unmodified SGLang. | |
| ## Loading | |
| Download this repository and the desired base checkpoint separately. Run the following from this repository's directory. The helper creates a separate assembled directory with links to existing files; it does not modify the base. | |
| ```bash | |
| MTP_BASE_DIR=/path/to/occamy-bf16 | |
| python assemble_head.py \ | |
| --base "$MTP_BASE_DIR" \ | |
| --head ./mtp-trained.safetensors \ | |
| --out ./occamy-with-mtp | |
| PYTHONPATH="$PWD/runtime${PYTHONPATH:+:$PYTHONPATH}" python -m sglang.launch_server \ | |
| --model-path ./occamy-with-mtp --tokenizer-path "$MTP_BASE_DIR" \ | |
| --host 127.0.0.1 --port 30000 --dtype bfloat16 \ | |
| --context-length 2048 --max-running-requests 1 --max-total-tokens 2048 \ | |
| --max-mamba-cache-size 4 --mem-fraction-static 0.50 \ | |
| --attention-backend triton --mamba-backend triton --mm-attention-backend sdpa \ | |
| --disable-cuda-graph --disable-overlap-schedule --disable-radix-cache \ | |
| --random-seed 42 --enable-deterministic-inference --moe-runner-backend triton \ | |
| --speculative-algorithm NEXTN --speculative-num-steps 1 \ | |
| --speculative-eagle-topk 1 --speculative-num-draft-tokens 2 \ | |
| --forward-hooks '[{"name":"mtp-strict","target_modules":[""],"hook_factory":"canonical_attention:install","config":{"output":"./mtp-strict-audit.jsonl"}}]' | |
| ``` | |
| For [Occamy NVFP4](https://huggingface.co/Accio-Lab/occamy-1.0-NVFP4), assemble a new directory using that base instead. Replace `--moe-runner-backend triton` with: | |
| ```text | |
| --quantization modelopt_fp4 --moe-runner-backend flashinfer_cutlass --speculative-moe-runner-backend triton | |
| ``` | |
| The base needs sufficient device memory for the selected memory fraction. The runtime hooks use SGLang internals; different builds or settings require revalidation. | |
| ## Files and evidence | |
| - `mtp-trained.safetensors`: 19 BF16 MTP tensors; no base weights. | |
| - `mtp_head_config.json`, `assemble_head.py`, `runtime/`: assembly configuration and required runtime. | |
| - [VALIDATION.json](VALIDATION.json): aggregate results, test scope and timing limitations. | |
| - [per-request-results.json](per-request-results.json): per-request counts, parity and SHA256 fingerprints of inputs, outputs and logprobs; conversation content is not republished. | |
| - `evaluation_cases.json`, `TRAINING.json`, `SHA256SUMS`: test cases, provenance and file integrity. | |
| The base and donor models are Apache-2.0 licensed. Original base weights are unchanged. Training checkpoints and the replaced release are retained separately for recovery. | |