--- license: other license_name: ltx-2.x-community-license license_link: https://huggingface.co/Lightricks/LTX-2.5/blob/main/LICENSE base_model: - Lightricks/LTX-2.5 - jdopensource/JoyAI-Echo base_model_relation: merge tags: - comfy-native - comfy-quant - quantized - comfyui - ltx-2 - ltx-2.5 - joyai-echo - video-generation - text-to-video - audio-video - lip-sync - merge pipeline_tag: text-to-video --- # JoyAI-Echo x LTX-2.5 (echoVid) - comfy-native (int8 / w4a8 / w4a4 / nvfp4 / mixed) **Try it in the browser:** [ZeroGPU demo Space](https://huggingface.co/spaces/joeygambino/joyai-echo-ltx25-echovid-comfy-native) - these exact files, the two-pass ladder, no install. **LTX-2.5's engine with JoyAI-Echo's performance.** LTX-2.5 renders picture and sound in one pass, at any length, in one generation. JoyAI-Echo (a fine-tune of LTX-2.3) has the better *actor*: natural lip-sync, expressive faces, a voice that stays put. The two transformers are shape-identical, so JoyAI-Echo's video attention/feed-forward delta was transplanted onto the official **LTX-2.5 dev** transformer, and the official LTX-2.5 distilled LoRA (`ltx-2.5-22b-distilled-lora-450`) is baked in at 0.5 - so these are few-step files with the same speed, VRAM and nodes as LTX-2.5 distilled. Nothing was retrained. (**v2**: the first build put the delta on the distilled transformer and came out over-saturated with hard contrast; those files are gone. The plain dev merges, for people who want to apply their own distill LoRA at their own strength, are here: https://huggingface.co/joeygambino/joyai-echo-ltx25-echoVid-dev.) What you get over stock LTX-2.5 distilled is the acting JoyAI-Echo was trained for - lip-sync, expression, a voice that stays put - at the same speed, VRAM and nodes. > **Workflow + nodes:** https://github.com/jlucasmcrell/ComfyUI-JoyLTX25 (the *Joy-LTX 2.5* > canvases: one-prompt take with a VRAM planner, and multishot with AV-extend joins; the release zip bundles the writer). > **GGUF files (Q3_K_M .. Q8_0):** > https://huggingface.co/joeygambino/joyai-echo-ltx25-echoVid-gguf > **All models:** https://huggingface.co/joeygambino > **Try it live:** https://huggingface.co/spaces/joeygambino/joy-ltx-25 (one take, ZeroGPU) > **Civitai:** *Joy-LTX 2.5* (models being uploaded now). ## What it looks like Rendered with the files on this page (070T30, distilled LoRA baked at 0.5), the [ComfyUI-JoyLTX25](https://github.com/jlucasmcrell/ComfyUI-JoyLTX25) canvases, 8 steps at cfg 1. Sound is generated with the picture, in the same pass - turn it on. **Three shots joined into one take** Multishot, 3 x 8 s at 1280x736, AV-extend joins - the speech and the room carry across both joins with no reference photo attached. **Beach, hard sun** 10 s, single generation, picture and sound together. **Wet neon street** 10 s, single generation. Reflections and rain with a voice over them. **Snow, flat overcast** 10 s, single generation. The grade holds in high key - the failure mode of the first build. ## Two doses | dose | what it is | pick it when | |---|---|---| | **070T30** *(default)* | 0.7 x Echo delta on video attention/FF, 0.3 x on the modulation tables, distill LoRA 0.5 | the default - cleaner skin, natural grade | | **100T50** *(strong)* | 1.0 x / 0.5 x, distill LoRA 0.5 | loud, comic, animated performances - the livelier read, a touch hotter on contrast | Both were reviewed blind on 20+ paired renders: scores tie; 070T30 reads a touch less rubbery on still faces, 100T50 lands laughter and big expressions better. Start with 070T30. ## Which file (stock ComfyUI 0.32+, no custom loader - the fast family on RTX 50) These use ComfyUI's own quantisation (`comfy_quant` + comfy-kitchen kernels), the same machinery as Lightricks' official `int8-convrot` build. Load them with the plain **Load Diffusion Model** node. Sizes are decimal GB. Timings: 960x544, 8 s, two-pass x2 to 1920x1088. | file | GB | fits | RTX 5090 | RTX 3090 | |---|---|---|---|---| | `LTX25dist-echoVid--v2-DiT-comfy-w4a4.safetensors` | 11.2 | 12 GB (tight) / 16 GB | 87 s | 3121 s (avoid on Ampere) | | `LTX25dist-echoVid--v2-DiT-comfy-w4a8.safetensors` | 12.5 | 16 GB | ~90 s | ~580 s | | `LTX25dist-echoVid--v2-DiT-comfy-nvfp4.safetensors` | 12.5 | 16 GB (RTX 50 only) | ~100 s | n/a | | `LTX25dist-echoVid--v2-DiT-comfy-mix4x8-13.8GB.safetensors` | 13.8 | 16 GB | 110 s | 1685 s | | `LTX25dist-echoVid--v2-DiT-comfy-mix4x8-17.0GB.safetensors` | 17.0 | 24 GB | 111 s | 3093 s (offloads) | | `LTX25dist-echoVid--v2-DiT-comfy-int8.safetensors` | 21.5 | 32 GB (24 GB tight) | 120 s **(32 GB default)** | - | Rule of thumb: **RTX 50 -> this repo. RTX 30/40 -> the GGUF repo** (Q5_K_M / Q6_K are 4-8x faster there than any 4-bit comfy-native arm). `--enable-triton-backend` on the ComfyUI launch line roughly halves w4a8/int8 step time where triton is installed. ## fp8 and the bf16 master Two more cuts, straight from the v2 master (same bake: dev + Echo delta + distill LoRA 0.5): | file | GB | note | |---|---|---| | `LTX25dist-echoVid--v2-DiT-comfy-fp8.safetensors` | 21.5 | comfy fp8_e4m3fn scaled; stock **Load Diffusion Model** | | `LTX25dist-echoVid--v2-DiT-bf16.safetensors` | 42.0 | the master; needs a card that streams 42 GB (or offload); the file to quantise from | ## Install (ComfyUI) 1. ComfyUI 0.32 or newer (the comfy-kitchen kernels ship with it). 2. Put the `.safetensors` in `models/diffusion_models/`. 3. From [Lightricks/LTX-2.5](https://huggingface.co/Lightricks/LTX-2.5): `vae/ltx-2.5-video-vae-bf16.safetensors` and `vae/ltx-2.5-audio-vae-bf16.safetensors` -> `models/vae/`; `latent_upscale_models/ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors` -> `models/latent_upscale_models/`; text encoder `text_encoders/gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors` -> `models/text_encoders/` (16 GB cards: the 10.6 GB `gemma4-12b-ltx25-comfy-w4a8.safetensors` from [LTX-2.5-Quantized](https://huggingface.co/joeygambino/LTX-2.5-Quantized)). 4. Load the workflow from the node pack above (or any LTX-2.5 workflow: pick this file in the stock **Load Diffusion Model** loader). Distilled schedule: 8 steps pass 1, 3 steps pass 2, `euler_ancestral`, CFG 1. ## Credits JoyAI-Echo by JD ([jdopensource/JoyAI-Echo](https://huggingface.co/jdopensource/JoyAI-Echo)); LTX-2.5 by Lightricks. Merge, quantisation and workflows by joeygambino. Licensed under the LTX-2.x Community License (inherited from both parents).