Swift-1.5-4bit-MLX / compatibility /architecture-compatibility-report.md
ukisai's picture
initial release
9fd3d5f
|
Raw History Blame
2.47 kB

Swift 1.5 MLX compatibility: validated complete checkpoint

Source: ukisai/Swift-1.5-Qwen3.8-27b at 00ccd14e006897d28cb0ed5bf26390e60d274251. All 18 BF16 shards and original runtime assets passed full SHA256 verification. Source: 55563006776 weight bytes, 1199 BF16 tensors.

The official Qwen loader dropped 333 vision and 15 MTP tensors. The isolated patch adds mlx_lm/models/qwen3_5_full.py with a real vision encoder and explicit MTP module, plus strict dispatch/index checks and config/asset preservation in mlx_lm/utils.py. tests/test_qwen3_5_full.py exercises mapping failures, Transformers numerical agreement, cache behavior, native nonquantized conversion, and exact-target quantized saving/reloading.

All 16 nonquantized architecture tests and the additional fixed affine test passed on Linux CPU. The real source loaded strictly with 851 text, 333 vision and 15 MTP tensors. There are zero ignored or unexplained tensors. The native quant has 2379 saved tensors because quantized weights have scales and biases. All 609 BF16 remainder tensors equal the source values.

Text generation passed. Vision encoder execution and an MTP step using real text hidden states passed. Image/video text integration and speculative generation are not implemented. The original tokenizer, chat template, context, processor, untied output head/shared MTP embeddings, norms and gating configuration are retained.

CPU runtime checks promote only in-memory floating values to FP32. The original Linux BF16 QMM kernel produced 256 when summing 8192 exact ones; FP32 returned 8192. This reproducible backend issue and the runtime workaround are recorded in cpu-quantized-matmul-diagnostic.json and ../USAGE.md. Stored weights were not changed.

Apple Silicon Metal checks passed for all 2379 native parameter headers and real packed Q4 samples from text, vision and MTP. Both native BF16 Metal and FP32 Metal executions matched their references. Only small actual samples were evaluated on the 16 GiB Mac; no full 27B Mac generation is claimed.

The original source was read without modification. The source tensor layout transposes are recorded individually in quant-tensor-mapping-manifest.json. No custom quantization algorithm or alternate quantization configuration was used.

Patch SHA256: f6f1d0bdafa45863bfbf93dac0398c481c993ea04fdf38b9bae98c643f89eaec. Apply to official MLX-LM commit c69d1288440a0dc4e6401fc417098b07598dccd5; see ../USAGE.md.