Swift-Qwen3.8-Flash-Next V3 - Slipstream/Splash package (V3 build)

The package this Mac actually serves: Swift (KV-sparse) Qwen3.8-Flash-Next, Q4_0 experts with Q8_0 resident tensors, 48 target layers + MTP draft head + ngram table. Loads directly with no conversion.

Contents

  • manifest.json - schema v5, splash-packed-q4-qwen4exp, 48 layers, 512 experts
  • target/ - 48 layer-*.bin (MDFN0031) + embedding + head + mtp-layer
    • mtp-combiner + ngram.bin (32 GB) + draft-vocab
  • draft/ - 5 DFlash2 draft layers + model.bin (MDFD0004)
  • tokenizer/ - tokenizer.json, vocab.json, config, chat template

Usage

splash serve --model <dir-of-parent> --port 8090      # sees manifest.json, serves directly
splash serve-native <dir>/target <dir>/draft --tokenizer <dir>/tokenizer

Relationship to other repos

This is the V3 successor of nitinpanj/Swift-Qwen3.8-Flash-Next-Splash, which is the Sept-24 build with only 20 target layers and no ngram/MTP/tokenizer. manifest.json and draft/* are byte-identical between the two (SHA-256); target/embedding, target/head and target/layer-0..19 differ, and 44 files (87.5 GB) exist only here.

Source GGUF: nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF. Built by dev/tools/build_swift_v3_gguf.py in github.com/npanj/slipstream.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nitinpanj/Swift-Qwen3.8-Flash-Next-Splash-V3