--- license: cc-by-nc-4.0 base_model: - MuScriptor/muscriptor-small - MuScriptor/muscriptor-medium - MuScriptor/muscriptor-large tags: - audio-to-midi - music-transcription - core-ai - aimodel - apple-silicon - macos library_name: swift-music-transcriber extra_gated_prompt: "These files are a format conversion of the MuScriptor weights released by Kyutai and Mirelo under CC BY-NC 4.0. They carry the same licence: non-commercial use only, attribution required. Requesting access confirms you accept those terms." --- # MuScriptor for Core AI (`.aimodel`) Apple Core AI conversions of **MuScriptor**, the open-weights audio-to-MIDI transcription model by [Kyutai](https://kyutai.org) and [Mirelo](https://mirelo.ai) — the same weights, in the format macOS 27 runs on-device. Nothing was retrained, pruned or fine-tuned. These files exist so that a Mac can transcribe without Python or a GPU server. | file | upstream | layers / width | size | note | |---|---|---|---|---| | `scribe-small-float32.aimodel` | [MuScriptor/muscriptor-small](https://huggingface.co/MuScriptor/muscriptor-small) | 14 / 768 | 398 MB | | | `scribe-medium-float32.aimodel` | [MuScriptor/muscriptor-medium](https://huggingface.co/MuScriptor/muscriptor-medium) | 24 / 1024 | 1.2 GB | the default | | `scribe-large-float32.aimodel` | [MuScriptor/muscriptor-large](https://huggingface.co/MuScriptor/muscriptor-large) | 48 / 1536 | 5.1 GB | | | `scribe-large-float16.aimodel` | same | 48 / 1536 | 2.6 GB | weights cast to half; twice as fast, exact on the reference clip | Each `.aimodel` is a directory (`main.mlirb`, `main.hash`, `metadata.json`) with three entry points: `main` (one decoder step with the attention cache as Core AI state), `prefix` (mel and instrument conditioning) and `embed`. Everything around the model — the log-mel front end, resampling, the token decoder, tie prologues, note cleanup, MIDI writing — lives in the Swift library. ## Faithfulness The conversion was held to upstream's Python code, not to a description of it: - **Token for token.** On 7 clips × 3 sizes (19 runs, 25,000+ tokens, including a two-minute track), the Swift pipeline with these files reproduces upstream's greedy CPU fp32 decode exactly. - **Logits.** Teacher-forced through the recorded token sequences, the decoder's logits agree with upstream's at ≥ 116 dB PSNR at every step, with no argmax flips. - **One thing had to be taken from the checkpoint rather than recomputed:** the STFT window, which the checkpoints store as a half-precision rounding of `torch.hann_window(2048)`. A float32 Hann window flips one near-tie token in 8 of 19 runs. The library carries the stored window. - **Small fp16 is not shipped**: it diverges after ~50 tokens, exactly where upstream's own fp16 run diverges. Large fp16 is exact on the reference clip and is included. Measured against 712 sample-pack loops that ship their MIDI (medium, `mir_eval`, onset 50 ms and pitch 50 cents, a per-file octave allowed because bass patches play below the written note): note F1 0.53 with recall 0.69. The same MIDI rendered through a General MIDI piano and transcribed: F1 0.95. Details and caveats in the library README. ## Use ```swift // swift-music-transcriber (MIT) — https://github.com/arraypress/swift-music-transcriber let transcriber = try await MusicTranscriber(model: URL(fileURLWithPath: "scribe-medium-float32.aimodel")) let result = try await transcriber.transcribe(audioURL) try result.midi.write(to: outputURL) ``` ```sh hf download arraypress/scribe-muscriptor --include "scribe-medium-float32.aimodel/*" --local-dir models scribe model install models/scribe-medium-float32.aimodel # the CLI built on the library ``` Requirements: macOS 27, Apple silicon. The model runs on the GPU (the ANE compiler pass rejects these fp32 graphs); `expectFrequentReshapes` is set so the changing sequence length does not respecialise per step. ## Reproduce the conversion The export is open source: `Tools/export.py` in the library repo, a `uv` script. ```sh uv run Tools/export.py --size medium # coreai-torch 0.4.2, coreai-core 1.0.0b2, torch 2.13, Python 3.12 uv run Tools/export.py --size large --dtype float16 ``` It downloads the gated upstream checkpoint with your own Hugging Face login, exports the decoder with its KV cache as Core AI state, and writes the `.aimodel`. The Swift tests (`ParityTests`, `LogitParityTests`) are the acceptance gate. ## Licence and attribution MuScriptor's weights are released under **CC BY-NC 4.0** by Kyutai and Mirelo ([github.com/muscriptor/muscriptor](https://github.com/muscriptor/muscriptor), code MIT). These files are a derivative of those weights and carry the same licence: **non-commercial use only, with attribution to the original authors.** The conversion tooling and the Swift library are MIT. If you use this, credit MuScriptor.