LTX ยท Diffusion transformer

LTX 2.5

LTX 2.5 is a pipeline of 2 networks totalling 34.14 billion parameters, of which 21.01 billion do the actual generating. At its native 1280x704 and 121 frames the denoiser works on 14,080 latent tokens at once, and attention over that sequence is what sets the render time.

Published as a bundle of variants rather than a pipeline, so its two figures are declared from the bf16 files rather than measured component by component.

Whole pipeline34.14B
Denoiser21.01B
Download63.6 GB
Default steps30
Native length121 frames
Frame rate24 fps

What the pipeline is made of

This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.

ComponentParametersPublished asOn diskWhat it does
transformer 21.01B BF16 39.13 GB Does the generating. Runs once per step, so it dominates both memory and time.
text_encoder 13.13B BF16 24.46 GB Turns your prompt into conditioning. Runs once, then can leave the GPU entirely.

These figures were stated by us, because the weights ship in a format with no readable header.

Memory by precision

Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.

PrecisionPeak VRAMQualityWhat it costs you
BF16 66.21 GB 100% How the weights are published. No loss, and the largest footprint.
FP8 34.41 GB 97% Halves the denoiser with a small, usually invisible cost. Needs Ada or newer.
GGUF Q8_0 36.4 GB 98% Works on any card, unlike FP8. Slightly slower than native precision.
GGUF Q5_K_M 25.23 GB 95% A middle step when Q8 will not fit.
GGUF Q4_K_M 21.82 GB 91% The usual way a 12B image model gets onto an 8 GB card. Detail softens.
NF4 20.11 GB 89% Aggressive 4-bit. Fast to load, noticeably looser on fine detail.

Memory by arrangement

The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 62% of the pipeline and stays resident whatever you rearrange.

ArrangementPeak VRAMTimeHow it works
Everything resident 66.21 GB 5 min 35 s All components stay on the GPU. Fastest, and needs the most memory.
Text encoder on CPU 41.75 GB 5 min 52 s The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers.
Component offload 41.75 GB 6 min 15 s One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation.
Sequential offload 10.25 GB 33 min 4 s Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower.

What a bigger output costs

Both memory and time climb faster than the frame count, because attention is quadratic in the length of the latent sequence.

OutputLatent tokensPeak VRAMTime
480p, 3 seconds 2,730 65.36 GB 37 s
480p, 5 seconds 4,290 65.36 GB 1 min 5 s
720p, 3 seconds 6,440 65.36 GB 1 min 49 s
720p, 5 seconds 10,120 65.73 GB 3 min 25 s
720p, 10 seconds 19,320 66.85 GB 9 min 10 s

Measured at BF16 with everything resident on an RTX 4090, so the columns compare with each other rather than with your machine.

Which hardware runs LTX 2.5

117 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.

DeviceMemoryVerdictArrangementPeakTime
RTX PRO 6000 Blackwell
NVIDIA
96 GB Workable BF16
everything resident
66.21 GB 3 min 42 s
GeForce RTX 5090
NVIDIA
32 GB Workable GGUF Q5_K_M
everything resident
25.23 GB 4 min 25 s
RTX 6000 Ada Generation
NVIDIA
48 GB Slow GGUF Q8_0
everything resident
36.4 GB 5 min 3 s
GeForce RTX 4090
NVIDIA
24 GB Slow GGUF Q4_K_M
everything resident
21.82 GB 5 min 35 s
GeForce RTX 5090 Laptop
NVIDIA
24 GB Slow GGUF Q4_K_M
everything resident
21.82 GB 7 min 4 s
NVIDIA DGX Spark 128GB
NVIDIA
128 GB Slow BF16
everything resident
66.21 GB 7 min 21 s
GeForce RTX 5080
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 8 min 33 s
GeForce RTX 4080 SUPER
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 9 min 15 s
GeForce RTX 4090 Laptop
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 9 min 37 s
GeForce RTX 4080
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 9 min 51 s
GeForce RTX 5080 Laptop
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 10 min 7 s
GeForce RTX 5070 Ti
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 10 min 51 s
GeForce RTX 4070 Ti SUPER
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 10 min 55 s
GeForce RTX 3090 Ti
NVIDIA
24 GB Slow GGUF Q4_K_M
everything resident
21.82 GB 11 min 26 s
RTX A6000
NVIDIA
48 GB Slow GGUF Q8_0
everything resident
36.4 GB 11 min 48 s
GeForce RTX 3090
NVIDIA
24 GB Slow GGUF Q4_K_M
everything resident
21.82 GB 12 min 52 s
RTX A5000
NVIDIA
24 GB Slow GGUF Q4_K_M
everything resident
21.82 GB 16 min 26 s
GeForce RTX 5060 Ti 16GB
NVIDIA
16 GB Slow GGUF Q4_K_M
text encoder on cpu
14.44 GB 19 min 56 s
Jetson AGX Orin 32GB
NVIDIA
32 GB Barely usable GGUF Q4_K_M
everything resident
21.82 GB 21 min 26 s
Jetson AGX Orin 64GB
NVIDIA
64 GB Barely usable GGUF Q8_0
everything resident
36.4 GB 21 min 26 s
Radeon RX 7900 XTX
AMD
24 GB Barely usable GGUF Q4_K_M
everything resident
21.82 GB 21 min 36 s
GeForce RTX 4060 Ti 16GB
NVIDIA
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 21 min 44 s
Radeon PRO W7900
AMD
48 GB Barely usable GGUF Q8_0
everything resident
36.4 GB 21 min 46 s
RTX A4000
NVIDIA
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 24 min 49 s
Radeon RX 7900 XT
AMD
20 GB Barely usable GGUF Q5_K_M
text encoder on cpu
16.53 GB 27 min 3 s
Radeon RX 9070 XT
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 28 min 43 s
Radeon RX 7900 GRE
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 30 min 17 s
Radeon RX 9070
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 35 min 42 s
Radeon RX 7800 XT
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 37 min 7 s
Radeon RX 6900 XT
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 1 h 0 min
Radeon RX 7600 XT
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 1 h 1 min
Ryzen AI Max+ 395 32GB
AMD
32 GB Barely usable GGUF Q4_K_M
everything resident
21.82 GB 1 h 7 min
Ryzen AI Max+ 395 64GB
AMD
64 GB Barely usable GGUF Q8_0
everything resident
36.4 GB 1 h 7 min
Ryzen AI Max+ 395 96GB
AMD
96 GB Barely usable BF16
everything resident
66.21 GB 1 h 7 min
Ryzen AI Max+ 395 128GB
AMD
128 GB Barely usable BF16
everything resident
66.21 GB 1 h 7 min
GeForce RTX 4070 Ti
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 8 min
GeForce RTX 4080 Laptop
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 13 min
GeForce RTX 4070 SUPER
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 17 min
GeForce RTX 3080 Ti
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 20 min
Radeon RX 6800
AMD
16 GB Barely usable GGUF Q4_K_M
text encoder on cpu
14.44 GB 1 h 26 min
GeForce RTX 5070
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 28 min
GeForce RTX 3080 12GB
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 31 min
GeForce RTX 3080 10GB
NVIDIA
10 GB Barely usable GGUF Q8_0
sequential offload
6.67 GB 1 h 31 min
GeForce RTX 4070
NVIDIA
12 GB Barely usable BF16
sequential offload
10.25 GB 1 h 33 min
GeForce RTX 2080 Ti
NVIDIA
11 GB Barely usable GGUF Q8_0
sequential offload
6.67 GB 1 h 40 min

Source: Lightricks/LTX-2.5 . Downloaded 1.2 million times in the last month. See how we calculate, or browse every video model.