File size: 2,496 Bytes
7de0b94
 
 
6bb81f1
7de0b94
7e65db7
 
 
 
6bb81f1
7e65db7
7de0b94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
---
base_model:
- SulphurAI/Sulphur-2-base
base_model_relation: quantized
pipeline_tag: image-text-to-video
tags:
- quantized
- mlx
- 4bit
- q4
library_name: mlx
---

This repository hosts custom implementations of the LTX2.3 video AI model, refined specifically for high-fidelity generation using the Sulphur 2 architecture. It has been converted to Apples MLX architecture and quantized down to Q4, to maximize memory efficiency. It has been tested on a 32GB M5 Silicon Mac, and that is the lowest recommended RAM for this model.

If you are not on a Mac, this madel variant is not for you.

## Model Overview

This implementation represents a highly optimized workflow built around the LTX2.3 core.

- **Base Model:** LTX2.3
- **Refinement Applied:** Sulphur 2
- **Fusion Detail:** The Sulphur 2 refinements have been successfully fused into the **`transformer-distilled.safetensors`** checkpoint, providing a unified generation experience.
- **Implementation:** MLX Conversion
- **Quantization:** FP4 (Optimized for performance and memory footprint)
- **Target Pipeline:** 8/3 Pipeline (Optimized for generation workflow)

## Usage Guide

### Core Workflow (Recommended)

For the best results and fastest generation times, users should rely on the integrated 8/3 pipeline.

- **Primary Generation:** Use the fused `transformer-distilled.safetensors` checkpoint to access the Sulphur 2 quality enhancements baked into the LTX2.3 base.
- **LoRAs:** No external LoRAs are required when using the fused model for Sulphur 2 quality, but have been included in this repo for convenience.

### Hardware & Compute Notes

- **Primary Platform:** Optimized for macOS compute environments on Apple silicon M-series SOC's.
- **AI Engine:** Built around the MLX framework integration.

## Prompting Guidelines (LTX Specific)

To achieve optimal generation quality with this model, adhere strictly to the following prompting conventions:

1.  **Structure:** Aim for a single, flowing paragraph.
2.  **Tense:** Use present tense verbs for all actions and movements.
3.  **Detail Level:** Match the level of descriptive detail to the intended shot scale (e.g., high detail for close-ups, broader strokes for wide shots).
4.  **Flow:** Describe the camera movement relative to the subject matter.
5.  **Length Target:** Aim for 4–8 descriptive sentences to maintain focus and coherence.

**Note:** Model coherence (and body horror) has a swift uptake in clips going past ~17 seconds. Test with shorter clips.