Qwen3.8-2B-Distill — MLC / WebGPU (q4f16_1)

Run with everything-webgpu WebGPU Model Size License: Apache 2.0

An MLC compilation of empero-ai/Qwen3.8-2B-Distill for WebGPU in the browser, quantized to q4f16_1.

Text path only: the vision tower and the MTP head are stripped, since neither is used on the WebGPU path. 632 tensors → 320, 1,881,825,088 parameters, 1.06 GB at 4.503 bits/param.


🚀 Quick Start with everything-webgpu

This model runs directly in the browser via everything-webgpu with a priority scheduler:

npm install everything-webgpu
import { CreateScheduledEngine } from "everything-webgpu";

const engine = await CreateScheduledEngine({
  model: "https://huggingface.co/nyaaorick/Qwen3.8-2B-q4f16_1-MLC/resolve/main/",
  modelLib: "https://huggingface.co/nyaaorick/Qwen3.8-2B-q4f16_1-MLC/resolve/main/Qwen3.8-2B-q4f16_1_cs1k-webgpu.wasm"
});

const reply = await engine.ask("Explain quantum computing in simple terms.");
console.log(reply);

⚠️ Read this before using the included .wasm on Firefox

Firefox's Metal backend caps maxStorageBuffersPerShaderStage at 9. This library contains four kernels that bind 10, which is normal for MLC builds — the stock Qwen3.5-0.8B-q4f16_1-MLC has exactly the same four. MLC registers them unconditionally as a fixed positional tuple handed to the C++ PagedKVCache, so they cannot be omitted without breaking the runtime's indexing.

Three are unreachable by configuration (sliding_window_size is -1, speculative decoding is never invoked). The fourth, batch_prefill_paged_kv_kernel, is reachable — it runs whenever the engine reuses a KV cache across turns.

A WebGPU pipeline that fails to create is silent. Its dispatches become no-ops, so the model emits garbage rather than raising an error. Measured on Firefox 154: with cross-turn KV reuse the same history and temperature: 0 produced a different, corrupted continuation each run.

If you run this on Firefox, force a full re-prefill instead of reusing the KV cache — for example call resetChat() before every prefill when adapter.limits.maxStorageBuffersPerShaderStage < 10. Chrome is unaffected and keeps KV reuse.

The cost of that workaround is measured at ~5 ms per history token re-prefilled.

Measured Performance

M4 MacBook Air (16 GB), Firefox 154, macOS:

Metric Measured Value
Decode 16.6 - 18.1 tok/s
Prefill 48 tok/s short prompt, 100-200 tok/s at length
Load from Cache ~51 s
Resident VRAM ~2.4 GB

Configuration baked into mlc-chat-config.json

  • Thinking is off by default: conv_template is set to qwen3_5_nothink with an empty closed <think></think> block, avoiding seconds of TTFT latency on edge devices.
  • Context window: context_window_size: 4096.
  • Stop tokens: [248046, 248044] (<|im_end|>, <|endoftext|>).

Contents

The model library .wasm ships directly in this repository rather than an external binary repo:

  • mlc-chat-config.json
  • tensor-cache.json, tensor-cache-b16.json
  • params_shard_0.bin … params_shard_25.bin (26 shards, 1.06 GB)
  • tokenizer.json, tokenizer_config.json, vocab.json, merges.txt
  • Qwen3.8-2B-q4f16_1_cs1k-webgpu.wasm (6.9 MB)

License

Apache-2.0 (inherited from empero-ai/Qwen3.8-2B-Distill).

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nyaaorick/Qwen3.8-2B-q4f16_1-MLC

Finetuned
Qwen/Qwen3.5-2B
Quantized
(24)
this model