MOSS-Audio-Tokenizer v2 (decoder)

OpenMOSS's 48 kHz stereo audio tokenizer, decode half, exported for loom.cpp. Family 11: codec tokens in, interleaved stereo out -- and the family's first leaf with no convolution at all.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from OpenMOSS-Team/MOSS-Audio-Tokenizer-v2. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

apache-2.0, inherited from the base model above.

Language(s)

(none tagged upstream)

a codec, not a language model: it carries no vocabulary and no language.

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/moss-audio-tokenizer-v2-loom")

# The geometry a caller needs, declared by the file rather than looked up in a paper:
n_codebooks = model.hparam("codec.n_codebooks")       # code streams per frame
codebook_size = model.hparam("codec.codebook_size")   # valid id range per stream
frame_rate = model.hparam("codec.frame_rate", "f32")  # codes per second
print(n_codebooks, codebook_size, frame_rate, model.contract["sample_rate"])

# Codes are FRAME-MAJOR: all `n_codebooks` codes for frame 0, then frame 1, and so on. This file
# is the DECODE half -- real codes come from the matching encoder, or an AR model that emits them.
frames = round(frame_rate)                            # one second of audio
codes = [[0] * n_codebooks for _ in range(frames)]
audio = model.codes2speech.infer(codes)
print(len(audio), "samples at", audio.sample_rate, "Hz =", round(audio.duration, 3), "s")
audio.save("out.wav")

# A flat list works too, and is what a driver that emitted the codes hands over. One that is not a
# whole number of frames is refused rather than reinterpreted at a different width.
audio = model.codes2speech.infer([0] * (frames * n_codebooks))

The audio is interleaved stereo: audio.channels is 2, audio.samples runs L R L R ..., audio.duration accounts for it, audio.save() writes a two-channel WAV and numpy.asarray(audio) is [frames, 2].

Rows narrower than 32 decode as a prefix of the codebooks: the rest of each row is filled with the id the file declares as absent. That is how moss-tts-local-transformer-v1.5-loom's 12 codebooks decode:

print(model.hparam("codec.absent_code"), audio.channels)
audio = model.codes2speech.infer([[0] * 12 for _ in range(frames)])   # 12 of 32

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

This is the DECODE half only, like every codec in this collection: encode is audio-in/codes-out, a different contract, and no model that decodes through this codec calls it. It is also what voice cloning with MOSS-TTS would need, so cloning is not available through these two repos.

A clip is decoded in ONE call, never in chunks, because nothing else is exact here. The decoder is six causal transformer stacks (12.5 Hz up to 400 Hz) with windowed attention, and 92 layers of windows reach further back than any chunk could carry: a chunked decode with 8 s of overlap is still 48% away from the model's own answer. The attention is computed in blocks, so memory grows linearly with the clip rather than quadratically. The ceiling is 4096 frames, 5.5 minutes, which is upstream's own generation budget. Verified against upstream's decode at 1.2e-06 relative RMS on 30 s of speech.

Feed it frame-major rows of up to 32 codes at 12.5 frames per second; one frame is 3840 samples per channel at 48 kHz. On a 2-core CPU a 30 s clip takes about two minutes.

Files

  • moss-audio-tokenizer-v2.gguf -- the model, exported with loom-exporter.
Downloads last month
25
GGUF
Model size
1B params
Architecture
loom-moss-audio-tokenizer
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/moss-audio-tokenizer-v2-loom

Quantized
(1)
this model

Collection including loom-ai-org/moss-audio-tokenizer-v2-loom