Instructions to use Lightricks/LTX-2.5-22b-IC-LoRA-Restore with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LTX-2
How to use Lightricks/LTX-2.5-22b-IC-LoRA-Restore with LTX-2:
# Install the LTX-2 pipelines git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync --extra natten
# Download the adapter weights from this repo # (base components come from Lightricks/LTX-2.5 — see Files and versions) hf download Lightricks/LTX-2.5-22b-IC-LoRA-Restore --local-dir models/LTX-2.5-22b-IC-LoRA-Restore
# Video-to-video with the IC-LoRA (runs on the distilled LTX-2.5 base) uv run python -m ltx_pipelines.ic_lora \ --transformer-path path/to/distilled-transformer.safetensors \ --text-encoder-path path/to/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path path/to/video-vae.safetensors \ --audio-vae-path path/to/audio-vae.safetensors \ --spatial-upsampler-path path/to/spatial-upsampler.safetensors \ --lora models/LTX-2.5-22b-IC-LoRA-Restore/<weights>.safetensors 1.0 \ --video-conditioning reference.mp4 1.0 \ --prompt "your prompt here" \ --output-path output.mp4 - Notebooks
- Google Colab
- Kaggle
You need to agree to share your contact information to access this model
By clicking "Agree and Access" you acknowledge the Privacy Policy and consent to receive offers and updates including targeted and personalized advertisements. You can unsubscribe at any time.
Log in or Sign Up to review the conditions and access this model content.
LTX Restore IC-LoRA
An In-Context LoRA that restores old and damaged footage: it takes archive video (low-resolution web rips, low-bitrate broadcast transfers, tape, sepia or black-and-white film scans) as its reference and renders the same shot as a clean, colour, high-resolution capture, removing compression damage, tape and sepia casts, flicker, dirt and scratches.
It was trained on the LTX-2.3 foundation model and has been tested on LTX-2.5, where it works well unchanged: every example and every number on this card was produced on the LTX-2.5 distilled transformer through the LTX-2.5 tiled workflow.
- Prompt
- a 1950s girl with blonde hair at a table drinking a glass of milk, printed dress, living room, soft window light, natural colour
- Prompt
- a steam locomotive arriving at a small rural railway station in 1896, passengers in period coats and hats waiting on the platform, wooden station buildings, natural colour
- Prompt
- the Hindenburg airship close up at its mooring mast with ground crew on a grass field, ribbed silver fabric hull with painted lettering, overcast daylight, natural colour
Three public-domain archive clips restored with this LoRA and then refined (the two-stage ladder in Pipeline Details), from 640x480 or 720x544 scans to 2880x2176: a 1954 family meal (Prelinger Archives), the Lumière brothers' train arriving at La Ciotat (1896) and the Hindenburg at its mooring mast (1937). Each gallery clip is a 1:1 crop comparison (the scan lanczos-upscaled on the left, the result on the right, 800x800 result pixels); full-frame wipes: family meal, train, Hindenburg. The prompts shown are the restoration-stage prompts as run. Negative for all three: blocky compression, banding, mosquito noise, ringing, blurry, smeared, ghosting, low resolution; the train additionally carried the prohibitions shipping containers, logos, signage, lettering, ladders, modern vehicles, graffiti, modern buildings. On that clip the station side of the frame holds almost no information in the scan, so the model rebuilds it: the buildings, the hillside and the porter's cart on the left are period-plausible reconstructions, not recovered detail. The locomotive and the crowd on the right are.
Model Files
ltx-2.5-22b-ic-lora-restore-1.0.safetensors
Final checkpoint (step 4000) of the training run described below.
Model Details
- Base Model: LTX-2.3-22B Video (trained on the 2.3 dev base with the Gemma-3 12B text encoder). Tested on LTX-2.5: runs on the LTX-2.5 distilled transformer with the LTX-2.5 video VAE and the Gemma-4 text encoder at strength 1.0, with no change to the weights.
- Training Type: IC-LoRA (video-to-video)
- Control Type: reference video (the archive clip, frame-aligned); optional reference image at the -1 token; optional prefix frames for chaining windows
- Reference Downscale Factor: 1
- Pipeline details: restoration pass at 1440 wide, then the Detail Refine IC-LoRA (see Pipeline Details)
- Audio: Not trained for audio generation
Intended Use & Out-of-Scope
Intended use: restoring and colourising archive footage for delivery at HD to 4K class: silent-era and newsreel film scans, mid-century broadcast, tape-era video, low-resolution low-bitrate web uploads; recovering the mid-band texture that heavily compressed broadcast video has lost.
Out of scope: footage that is already clean (the pass is a switch, not a gentle enhancer); faithful recovery of a creative colour grade from a monochrome source (colour is inferred from content, or from a reference image); deinterlacing (deinterlace before restoring).
Control Signal Requirements
- Control signal type: the archive clip, resized (lanczos) to the working canvas and fed as the IC-LoRA reference at downscale factor 1.
- Expected input: any resolution, progressive frames, 8n+1 frames per pass (49 or 97 recommended). Working canvas 1440 wide for 4:3 sources (1440x1056 or 1440x1088), 1920x1088 for 16:9, both multiples of 32. Do not stretch a 4:3 scan to 16:9.
- Preprocessing: deinterlace telecined or interlaced sources first (
fieldmatch,yadif,decimate); otherwise none. Do not denoise or sharpen the input. - Alignment: 1:1, same frames, same frame rate.
- Reference image (optional): native-resolution crops of what the scan cannot tell the model (a face, a sign, a logo, the true colours of a place) on a grey image the size of the output canvas, attached with a second
LTXAddVideoICLoRAGuideatframe_idx-1; it is per tile (see Recommended Settings). - Mask support: none. Prefix conditioning: for clips longer than one window, the previous window's last 17 restored frames can be given as the first frames of the next window (
LTXVAddGuideat frame 0); the model was trained to continue from them, which keeps colour consistent across a long clip.
How It Works
The model learned from synthetic pairs built to match real archive uploads: a clean 1080p clip was turned monochrome or tinted (neutral broadcast, cool tape, warm sepia), blurred as a period lens would, downscaled to 240p to 360p, encoded at a low bitrate one to three times over, sometimes denoised, scaled back up, and given film damage (brightness flicker, dirt and specks, hairs, scratches, blotches and gate weave). Given such a clip as its reference, the model renders the shot as the clean original: it rebuilds resolution and texture, removes the cast and the damage, and colourises from content, so the prompt's description of the scene matters. Its window is a 960x544 tile and 97 frames; larger frames are handled as overlapping tiles fused every step, and long clips as chained windows that start from the previous window's restored frames.
Usage
ComfyUI
- Copy the LoRA weights into
models/loras. - Open LTX-2.5_V2V_TiledFusion_Upscale.json from the ComfyUI-LTXVideo repository and swap the LoRA in the Load Models group to
ltx-2.5-22b-ic-lora-restore-1.0.safetensors(strength 1.0). Setoutput_sizeto FullHD and the sampler tile to 960x544. - Drop the archive clip on
LoadVideo, describe the scene in the prompt, put what must not be invented in the negative, run. The workflow's Preprocess group first upscales the clip to the working canvas with lanczos and feeds that as the guide; the model never sees the small scan. This is the restoration stage; long clips are handled by 97-frame windows (use_streamingis on in the workflow). - For the full ladder, run the same workflow a second time on that output with the Detail Refine IC-LoRA (
ltx-2.5-22b-ic-lora-refine-details-1.0.safetensors), tile 1024x576,output_size4K: the detail stage (see Pipeline Details). - Use the tiled workflow at every output size, including full HD. The LoRA was trained on 960x544 tiles and works best when each tile sees content at that scale; a 1440x1088 canvas is a 2x3 grid of those tiles. The single-tile LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json is right only when the output itself is 960x544. Building your own graph, in this order: lanczos-resize the clip to the working canvas (a multiple of 32) →
LTXAddVideoICLoRAGuide(latent_downscale_factor1) →LTXVTiledFusionSampler(tile 960x544, overlap 0.5, blend 0.05,cfg1.0) on the 8-step distilled sigmas withLTXICLoRALoaderModelOnly(1.0) on the distilled model →VAEDecodeTiled.
Pipeline Details
Restoration to 4K class runs in two stages, both as per-step tiled latent fusion (one latent canvas, overlapping windows re-fused every denoising step, one shared noise field):
- Stage 1, restoration. Lanczos the source to 1440 wide (1440x1088 for 4:3), tile 960x544 with 50% overlap, this LoRA at strength 1.0. This stage does the restoration and colourisation and is a legitimate result on its own.
- Stage 2, detail. Lanczos the stage-1 result to 2x (2880x2176) and run the Detail Refine IC-LoRA, tile 1024x576. It adds fine texture; it must come second, because refining first sharpens the damage and desaturates the colour.
- Long clips. One temporal extent past 97 frames loses detail and colour drifts. Run 97-frame windows: either windowed guides in the fusion sampler (
use_streamingwithtile_frames97 on guide and sampler), or chained windows where each starts from the previous window's last 17 restored frames (prefix conditioning) and the overlaps are blended. Measured on a 545-frame clip, the joins add nothing to the frame-to-frame change beyond the source's own. - Resolution ceiling. Gain over a plain upscale peaks at 9x the linear size of a 480-line scan (5760x4352) and reverses beyond it; a 240p web source tops out at 4K class.
Recommended Settings
- LoRA strength / weight: 1.0. Strength behaves as a switch: at 0.7 the restoration turns off and the clip stays monochrome; 1.4 over-tints.
- Inference steps: 8 (distilled sigmas
1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0, euler) - Guidance scale: 1.0
- Input preparation: lanczos-resize the scan to the working canvas before it becomes the guide (the Upscale workflow does this); do not denoise, sharpen or deinterlace-by-blending first.
- Resolution & frames: tile 960x544 (the trained tile; 1440x1088 tiles cost a third of the gain), 49 or 97 frames per window
- Tiled at every output size: the closer each tile is to the training bucket (960x544), the better the LoRA works, so run tiled even for a full-HD canvas rather than one untiled pass.
- Prompting: describe the period, place, light, materials and what people wear: colour is a semantic decision and the prompt steers it. Keep it generic across the frame; every tile receives the whole prompt, so an object named in it can be painted into a tile that does not contain it. Name a specific object only when you prompt per tile (run that tile on its own). Put everything the model must not invent in the negative prompt (modern objects, logos, signage, lettering, vehicles); with prohibitions the model rebuilds empty regions of a scan as period-plausible background instead of painting modern objects into them.
- Reference image at the -1 token (faces, signs, logos, true colours): when real crops of the subject exist (a photograph of the person, the sign, the building), give them to the model: native-resolution crops on a grey image the size of the output canvas, each placed over its own subject, attached with a second
LTXAddVideoICLoRAGuide(frame_idx-1, strength 1.0) chained after the clip guide, on the whole-clip path (use_streamingoff). It is per tile: the fusion sampler crops the reference to each tile exactly as it crops the clip, so a crop placed elsewhere reaches the wrong tiles. On a street clip restored from monochrome, references of the signs, the shop logo and the musicians raised the sign region by 4.6 dB and the faces by 4.0 dB against the truth, brought back the real lettering and the real colours, and left the rest of the frame unchanged. - Order with the Detail Refine IC-LoRA: restoration first, refine second. Both can also be loaded into one pass (refine at 0.2 to 0.3) when throughput matters; colour fidelity is best sequential.
Example positive prompt:
a steam locomotive arriving at a railway station platform in 1896, passengers in period coats and hats, wooden station building, daylight, natural colour, sharp photographic detail, crisp faces and clothing texture, natural grain, high resolution footage
Example negative prompt:
shipping containers, logos, signage, lettering, ladders, modern vehicles, graffiti, modern buildings, blurry, soft, plastic, smeared detail, oversharpened halos, warped faces
References
- Code: GitHub Repository
- ComfyUI: ComfyUI-LTXVideo
Tips & Troubleshooting
- Stays black and white: strength is below 1.0. Set it to 1.0.
- Modern objects appear where the scan is empty (containers, ladders, painted signs): name them in the negative prompt and describe the period in the positive prompt.
- Colour drifts along a long clip: the clip ran as one extent past 97 frames, or the windows were not chained. Use 97-frame windows with the prefix hand-off; measured on 209 frames the drift then tracks the source's own tint.
- Colour appears on things that had none in a colour source: the model colourises from content; on footage that already has colour, keep the prompt neutral about colour and the strength at 1.0.
- Result is sharper but emptier: the refine pass ran first or alone. Restore first.
- Sepia or tinted sources: do not judge by HSV saturation, a strong uniform tint reads as high saturation; judge the colour by eye or by hue spread.
- Interlaced newsreels: deinterlace before restoring, the model treats field combing as texture.
- Faces are restored clean and can read slightly beautified on very degraded sources; when a likeness matters, give the model a reference image of the real face at the -1 token, placed over the face's own tile.
- Cartoons and flat art: the model under-cleans noise-type degradation on animation; expect residual grain.
- Which frames to check: the first 17 frames of a chained window are the hand-off; compare frame-to-frame changes at the joins against the source's own flicker, early film flickers more than the output.
Dataset
Paired clips built from stock footage: clean 1920x1080 targets and synthetic archive-style references (monochrome and tinted transfers, low resolution, multi-generation low-bitrate encoding, film damage) across severity tiers, same framing (the model never outpaints). 151 training pairs, 8 held out, 97 frames at 24 fps, one shared caption.
Training
- Technique: LTX-2 trainer flexible strategy:
referencecondition (the degraded video, downscale factor 1, always present),first_frame(probability 0.25) andprefix(the first 17 frames given as context, probability 0.5). - Hyperparameters: rank 128, alpha 128, targets to_q/to_k/to_v/to_out, Prodigy lr 1.0 (d_coef 1.0, bias correction, safeguard warmup), cosine schedule, batch 1, grad clip 1.0, bf16, no quantization, gradient checkpointing, seed 42
- Steps: 4000 (checkpoint every 500; step 4000 shipped)
- Resolution buckets:
960x544x97; 960x544x49 - Infrastructure: LTX-2 Community Trainer.
Validation against clean ground truth (8 pairs whose sources also appear in training under other degradations, 1920x1056, 97 frames, stage 1 only, on LTX-2.5): PSNR +2.0 dB over the degraded input, Lab colour error 24 to 16, high-frequency energy from 61% to 85% of the target, versus +0.8 dB and 75% for the previous version. On clips whose sources were never in training the model restores structure and colour at the same strength but, like any generative restorer of monochrome input, does not beat the aligned input on pixel PSNR; judge those by eye.
License
See the LTX-2-community-license for full terms.
Acknowledgments
- Base model by Lightricks
- Training infrastructure: LTX-2 Community Trainer
- Downloads last month
- 802
Model tree for Lightricks/LTX-2.5-22b-IC-LoRA-Restore
Base model
Lightricks/LTX-2.3