Qwen3.8-2B-Distill — MLC / WebGPU (q4f16_1)
An MLC compilation of empero-ai/Qwen3.8-2B-Distill for WebGPU in the browser, quantized to q4f16_1.
Text path only: the vision tower and the MTP head are stripped, since neither is used on the WebGPU path. 632 tensors → 320, 1,881,825,088 parameters, 1.06 GB at 4.503 bits/param.
🚀 Quick Start with everything-webgpu
This model runs directly in the browser via everything-webgpu with a priority scheduler:
npm install everything-webgpu
import { CreateScheduledEngine } from "everything-webgpu";
const engine = await CreateScheduledEngine({
model: "https://huggingface.co/nyaaorick/Qwen3.8-2B-q4f16_1-MLC/resolve/main/",
modelLib: "https://huggingface.co/nyaaorick/Qwen3.8-2B-q4f16_1-MLC/resolve/main/Qwen3.8-2B-q4f16_1_cs1k-webgpu.wasm"
});
const reply = await engine.ask("Explain quantum computing in simple terms.");
console.log(reply);
⚠️ Read this before using the included .wasm on Firefox
Firefox's Metal backend caps maxStorageBuffersPerShaderStage at 9. This library contains four kernels that bind 10, which is normal for MLC builds — the stock Qwen3.5-0.8B-q4f16_1-MLC has exactly the same four. MLC registers them unconditionally as a fixed positional tuple handed to the C++ PagedKVCache, so they cannot be omitted without breaking the runtime's indexing.
Three are unreachable by configuration (sliding_window_size is -1, speculative decoding is never invoked). The fourth, batch_prefill_paged_kv_kernel, is reachable — it runs whenever the engine reuses a KV cache across turns.
A WebGPU pipeline that fails to create is silent. Its dispatches become no-ops, so the model emits garbage rather than raising an error. Measured on Firefox 154: with cross-turn KV reuse the same history and temperature: 0 produced a different, corrupted continuation each run.
If you run this on Firefox, force a full re-prefill instead of reusing the KV cache — for example call resetChat() before every prefill when adapter.limits.maxStorageBuffersPerShaderStage < 10. Chrome is unaffected and keeps KV reuse.
The cost of that workaround is measured at ~5 ms per history token re-prefilled.
Measured Performance
M4 MacBook Air (16 GB), Firefox 154, macOS:
| Metric | Measured Value |
|---|---|
| Decode | 16.6 - 18.1 tok/s |
| Prefill | 48 tok/s short prompt, 100-200 tok/s at length |
| Load from Cache | ~51 s |
| Resident VRAM | ~2.4 GB |
Configuration baked into mlc-chat-config.json
- Thinking is off by default:
conv_templateis set toqwen3_5_nothinkwith an empty closed<think></think>block, avoiding seconds of TTFT latency on edge devices. - Context window:
context_window_size: 4096. - Stop tokens:
[248046, 248044](<|im_end|>,<|endoftext|>).
Contents
The model library .wasm ships directly in this repository rather than an external binary repo:
mlc-chat-config.jsontensor-cache.json,tensor-cache-b16.jsonparams_shard_0.bin…params_shard_25.bin(26 shards, 1.06 GB)tokenizer.json,tokenizer_config.json,vocab.json,merges.txtQwen3.8-2B-q4f16_1_cs1k-webgpu.wasm(6.9 MB)
License
Apache-2.0 (inherited from empero-ai/Qwen3.8-2B-Distill).
- Downloads last month
- 10