tomkay's picture
Upload README.md with huggingface_hub
949365d verified
|
Raw History Blame
5.61 kB
metadata
library_name: mlx
license: other
base_model: Lightricks/LTX-2.3
tags:
  - mlx
  - apple-silicon
  - safetensors
  - video-generation
  - quantized
  - ram-quantization
  - mixed-precision
pipeline_tag: text-to-video

baa-ai/LTX-2.3-MLX-RAM-12GB

Mixed-precision MLX quantization of Lightricks/LTX-2.3, targeting ~12 GB RAM using the RAM Pipeline β€” a sensitivity-driven knapsack allocator that assigns each tensor its optimal bit width independently.

Runs on Apple Silicon (M1 and later). Uses ltx-2-mlx for inference.

Model Details

Property Value
Base model Lightricks/LTX-2.3 (22B joint audio-video DiT)
Architecture LTX-2.3-22B-AudioVideoTransformer
Transformer blocks 48
Quantization method RAM mixed-precision MCKP knapsack
Sensitivity metric Kernel CKA (kernel-csk)
Group size 64
Quantized layers 1,415 of 1,760 transformer tensors
Transformer file size 9.3 GB

Quantization

The transformer uses per-layer mixed precision. Bit widths are assigned per-tensor by solving a Multiple-Choice Knapsack Problem (MCKP) against a budget constraint, using Kernel Centered Kernel Alignment (CKA) as the sensitivity signal.

Bit distribution (transformer layers):

Bits Layers %
2 29 2.1%
3 225 15.9%
4 348 24.6%
5 631 44.6%
6 125 8.8%
8 57 4.0%

Average: 4.6 bpw across quantized layers. Highest-sensitivity tensors (first/last layers, K/V projections, vision bridge) are protected at higher bit widths; lower-sensitivity layers are pushed to 2–3 bit.

The connector, VAE, audio VAE, and vocoder are kept at full BF16 precision.

Requirements

  • macOS 13+ on Apple Silicon (M1 / M2 / M3 / M4)
  • Python 3.10+
  • ~14 GB unified memory recommended (12 GB transformer + overhead)
pip install mlx mlx-lm flask ltx-pipelines-mlx ltx-core-mlx

Quick Start β€” Web App

The included webapp.py provides a browser UI with live generation log, video playback, and job history. To compare against the 24 GB version, download it alongside this model.

# Download this model
git clone https://huggingface.co/baa-ai/LTX-2.3-MLX-RAM-12GB

# Optional: also download the 24 GB version for side-by-side comparison
git clone https://huggingface.co/baa-ai/LTX-2.3-MLX-RAM-24GB

# Run the web app (single model)
cd LTX-2.3-MLX-RAM-12GB
python webapp.py

# Or with comparison enabled
python webapp.py --compare-dir ../LTX-2.3-MLX-RAM-24GB

# Open http://localhost:7860

The web app lets you:

  • Enter a prompt and set resolution, duration, FPS, and seed
  • Watch the generation log stream live
  • Play the generated video directly in the browser
  • Switch between models for comparison (if --compare-dir is set)
  • Browse job history and reload past settings

Quick Start β€” CLI

python generate.py \
    --model-dir /path/to/LTX-2.3-MLX-RAM-12GB \
    --prompt "A serene mountain lake at sunrise, mist over calm water, pine trees reflected" \
    --height 480 --width 704 --num-frames 65 \
    --output output.mp4

Common resolution + frame count combinations:

Label Height Width Frames Duration @ 24fps
Tiny (test) 256 256 33 1.4s
480p short 480 704 65 2.7s
480p medium 480 704 97 4.0s
720p short 720 1280 65 2.7s

Frame count must follow the formula 32k + 1 (33, 65, 97, 129, …).

Files in This Repo

File Size Description
transformer-distilled.safetensors 9.3 GB RAM mixed-precision quantized DiT transformer
connector.safetensors 5.9 GB Audio/video embedding connectors (BF16)
vae_decoder.safetensors 777 MB Video VAE decoder (BF16)
vae_encoder.safetensors 608 MB Video VAE encoder (BF16)
audio_vae.safetensors 102 MB Audio VAE (BF16)
vocoder.safetensors 246 MB Vocoder + bandwidth extension (BF16)
spatial_upscaler_x2_v1_1.safetensors 950 MB Optional 2Γ— spatial upscaler
split_model.json β€” Pipeline model structure descriptor
generate.py β€” Single-file generation script (CLI)
webapp.py β€” Flask web UI for interactive generation

Technical Notes

The RAM Pipeline probes each tensor's sensitivity to quantization at 2/3/4/5/6/8 bits using Kernel CKA on calibration activations, builds a per-tensor sensitivity manifest, then solves a MCKP knapsack to find the Pareto-optimal bit assignment for the target memory budget.

Key invariants applied during allocation:

  • 2-bit veto: tensors with cosine divergence > 0.0001 at 2-bit are never assigned 2 bits
  • Saturation gate: tensors whose quality plateaus early receive no further bit upgrades
  • Architectural priors: first/last layer (3Γ—), global-attention K/V (3Γ—), vision bridge (2Γ—) β€” these tensors resist low-bit allocation

Inference uses a custom apply_mixed_precision_quantization that detects each layer's bit width from its packed weight shape and calls mlx.nn.quantize once per unique bit width, enabling variable-precision loading into the ltx-2-mlx pipeline without any format modifications.

Attribution