ashxhart commited on
Commit
de59776
·
verified ·
1 Parent(s): 932f42f

Add M3 Studio performance and MTP compatibility note

Browse files
Files changed (1) hide show
  1. README.md +16 -1
README.md CHANGED
@@ -58,6 +58,9 @@ The upstream tokenizer, chat template, vision processor, and generation configur
58
  > [!IMPORTANT]
59
  > Qwen3.8 Flash Next uses the new `qwen4_exp` architecture. Use an oMLX or MLX-VLM build that explicitly lists `qwen4_exp` support. Older MLX-VLM releases cannot load this checkpoint.
60
 
 
 
 
61
  ## Quick start
62
 
63
  ```bash
@@ -74,6 +77,18 @@ python -m mlx_vlm.generate \
74
  --max-tokens 512
75
  ```
76
 
 
 
 
 
 
 
 
 
 
 
 
 
77
  ## Architecture
78
 
79
  Qwen3.8 Flash Next is an experimental vision-language architecture combining Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, and hashed bigram/trigram embeddings.
@@ -95,7 +110,7 @@ For upstream evaluations, intended use, limitations, safety guidance, and the co
95
  - Source: official BF16 checkpoint.
96
  - All 3,671 converted tensors and 22 indexed shards were checked locally.
97
  - The release payload was scanned for credentials, personal contact details, private paths, private network information, logs, caches, and private organisation data.
98
- - Generation performance will be added after the repeatable Apple-silicon benchmark finishes.
99
  - Quantisation can reduce output quality relative to BF16. Test the model on representative workloads before production use.
100
 
101
  This is a community conversion, not an official Qwen release.
 
58
  > [!IMPORTANT]
59
  > Qwen3.8 Flash Next uses the new `qwen4_exp` architecture. Use an oMLX or MLX-VLM build that explicitly lists `qwen4_exp` support. Older MLX-VLM releases cannot load this checkpoint.
60
 
61
+ > [!CAUTION]
62
+ > Do not attach a Qwen3.8 27B MTP drafter to this model. The hidden sizes differ and the drafter is incompatible with Flash Next.
63
+
64
  ## Quick start
65
 
66
  ```bash
 
77
  --max-tokens 512
78
  ```
79
 
80
+ ## Measured performance
81
+
82
+ Validated on an Apple M3 Studio with text-only generation after model load:
83
+
84
+ | Test path | Result |
85
+ | --- | ---: |
86
+ | oMLX server, warmed 543–566-token responses | 24.1–24.2 tokens/s |
87
+ | oMLX server, warmed shorter responses | 24.6–26.1 tokens/s |
88
+ | Standalone MLX exact-copy smoke test | 31.0 tokens/s |
89
+
90
+ The standalone result is a short smoke test; the longer oMLX figures better represent sustained chat generation. Results vary with prompt length, cache state, sampling settings, runtime version, and memory pressure.
91
+
92
  ## Architecture
93
 
94
  Qwen3.8 Flash Next is an experimental vision-language architecture combining Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, and hashed bigram/trigram embeddings.
 
110
  - Source: official BF16 checkpoint.
111
  - All 3,671 converted tensors and 22 indexed shards were checked locally.
112
  - The release payload was scanned for credentials, personal contact details, private paths, private network information, logs, caches, and private organisation data.
113
+ - Deterministic standalone and warmed oMLX server generation tests passed on Apple silicon.
114
  - Quantisation can reduce output quality relative to BF16. Test the model on representative workloads before production use.
115
 
116
  This is a community conversion, not an official Qwen release.