jashansinghTT's picture
Update model card to the reviewed card layout
507c33d verified
|
Raw History Blame Contribute Delete
4.53 kB
metadata
tags:
  - blackhole
  - p150
  - tt-dit-server
  - tt-model-cache
  - tt-model-catalog
  - tt-model-container
license: other
license_name: minimax-music3-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE
pipeline_tag: text-to-audio
base_model:
  - MiniMaxAI/MiniMax-Music3

minimax-music3 on Tenstorrent p150

Derived from MiniMaxAI/MiniMax-Music3 by MiniMax, unmodified. See the original model card for training and evaluation details.

MiniMax Music 3, a lyrics-and-caption conditioned music generation model (8B Qwen3-based global LLM, 0.6B depth decoder for the residual RVQ codebooks, 2.4B flow-matching DiT and a DAC-style vocoder), with an SGLang-Omni compatible /v1/audio/speech API plus an async jobs API for multi-minute songs.

Runs on Tenstorrent p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Hardware p150
License minimax-music3-community-license

Intended use

Generating songs from lyrics and a text description of the music.

Prerequisites

  • The Tenstorrent CLI, tt: uv tool install tenstorrent (or use tt-model alone, see the Quickstart)
  • Docker
  • Tenstorrent hardware: p150

Quickstart

tt serve jashansinghTT/minimax-music3-blackhole

tt serve downloads the Docker image and the MiniMaxAI/MiniMax-Music3 weights at fbdf52fbaaca799592917417eb05f1899f1255ec (into your HF cache; they are not in the image), then starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy). The first start compiles kernels for your device, which takes several minutes; the server is ready when it logs Application startup complete.

Without tt-cli: tt-model serve jashansinghTT/minimax-music3-blackhole.

Try it

curl -s localhost:20000/v1/audio/speech -H 'Content-Type: application/json' -d '{
  "model": "MiniMaxAI/MiniMax-Music3",
  "input": "[Verse]\nMorning light filtering through the pine\n[Chorus]\nSoftly the world begins to breathe",
  "instructions": "A warm acoustic pop song with intimate female vocals, fingerpicked guitar and soft piano.",
  "response_format": "wav", "seed": 7, "max_new_tokens": 750}' -o song.wav

Capabilities

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API — its request and response shapes are the model's own. See the author's notes for the payload it expects.

max_new_tokens counts audio frames (25 per second). Long songs: POST /v1/music/jobs, poll GET /v1/music/jobs/{id}, then fetch GET /v1/music/jobs/{id}/audio.

Expected performance

profile request audio length concurrency N wall time
p150 POST /v1/audio/speech, max_new_tokens 250, WAV 10 s 1 1 44.7 s

Measured end to end through this package's server in the tt-model community sweep, September 2026. No quality evaluation against the reference implementation has been run yet.

Limitations

  • One song is generated at a time; further requests wait.
  • Slower than real time: about 4.5 s of compute per second of audio on one chip.
  • Only p150 is validated.
  • No quality evaluation against the reference implementation has been run yet.
  • The MiniMax-Music3 Community License governs the weights; products built with it must display "MiniMax-Music3".

Licensing

Model weights are under the MiniMax-Music3 Community License; products built with it must display "MiniMax-Music3".

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/jashansinghTT/minimax-music3-blackhole/discussions — that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout — commit not published
code/ digest 9a9c37114f86ca58 (sha256, first 16 hex digits)
built 2026-09-12T09:31:21+00:00 by tt-model 0.1.0