--- library_name: mlx tags: - mlx - quantized - mixed-precision - ltx-video - video-generation - apple-silicon license: other base_model: Lightricks/LTX-2.3 base_model_relation: quantized --- # LTX-2.3-22B — 12GB (MLX) Mixed-precision quantized version of [Lightricks/LTX-2.3](https://huggingface.co/Lightricks/LTX-2.3) optimised by [baa.ai](https://baa.ai) using a proprietary Black Sheep AI method. Per-tensor bit-width allocation via advanced sensitivity analysis. Runs natively on Apple Silicon via [ltx-2-mlx](https://github.com/dgrauet/ltx-2-mlx). ## Metrics | Metric | Value | |--------|-------| | **Transformer size** | **9.3 GB** | | **Total model size** | **~18 GB** | | Average bits (transformer) | 4.6 | | Quantized layers | 1,415 / 1,760 | **Bit distribution:** | Bits | Layers | |------|--------| | 2 | 29 | | 3 | 225 | | 4 | 348 | | 5 | 631 | | 6 | 125 | | 8 | 57 | ## Requirements - Apple Silicon (M1 or later), macOS 13+ - ~14 GB unified memory - Python 3.10+ ```bash pip install mlx flask ltx-pipelines-mlx ltx-core-mlx ``` ## Usage — Web App ```bash git clone https://huggingface.co/baa-ai/LTX-2.3-22B-RAM-12GB-MLX cd LTX-2.3-22B-RAM-12GB-MLX python webapp.py # Side-by-side comparison with the 24GB version: git clone https://huggingface.co/baa-ai/LTX-2.3-22B-RAM-24GB-MLX ../LTX-2.3-22B-RAM-24GB-MLX python webapp.py --compare-dir ../LTX-2.3-22B-RAM-24GB-MLX ``` Open **http://localhost:7860** — enter a prompt, set resolution and duration, and hit Generate. The live log streams generation progress; the video plays in the browser when done. ## Usage — CLI ```bash python generate.py \ --model-dir /path/to/LTX-2.3-22B-RAM-12GB-MLX \ --prompt "A serene mountain lake at sunrise, mist over calm water, pine trees reflected" \ --height 480 --width 704 --num-frames 65 \ --output output.mp4 ``` Frame count must follow `32k + 1` (33, 65, 97, 129 …). Duration in seconds = frames ÷ fps. --- *Quantized by [baa.ai](https://baa.ai)* --- ## Black Sheep AI Products **[Shepherd](https://baa.ai/shepherd.html)** — Private AI deployment platform that shrinks frontier models by 50-60% through RAM compression, enabling enterprises to run sophisticated AI on single GPU instances or Apple Silicon hardware. Deploy in your VPC with zero data leaving your infrastructure. Includes CI/CD pipeline integration, fleet deployment across Apple Silicon clusters, air-gapped and sovereign deployment support, and multi-format export (MLX, GGUF). Annual cloud costs from ~$2,700 — or run on a Mac Studio for electricity only. **[Watchman](https://baa.ai/watchman.html)** — Capability audit and governance platform for compressed AI models. Know exactly what your quantized model can do before it goes live. Watchman predicts which capabilities survive compression in minutes — replacing weeks of benchmarking. Includes compliance-ready reporting for regulated industries, quality valley warnings for counterproductive memory allocations, instant regression diagnosis tracing issues to specific tensors, and 22 adversarial security probes scanning for injection, leakage, hallucination, and code vulnerabilities. Learn more at **[baa.ai](https://baa.ai)** — Sovereign AI.