LTX ยท Diffusion transformer
LTX 2.5
LTX 2.5 is a pipeline of 2 networks totalling 34.14 billion parameters, of which 21.01 billion do the actual generating. At its native 1280x704 and 121 frames the denoiser works on 14,080 latent tokens at once, and attention over that sequence is what sets the render time.
Published as a bundle of variants rather than a pipeline, so its two figures are declared from the bf16 files rather than measured component by component.
What the pipeline is made of
This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.
| Component | Parameters | Published as | On disk | What it does |
|---|---|---|---|---|
| transformer | 21.01B | BF16 | 39.13 GB | Does the generating. Runs once per step, so it dominates both memory and time. |
| text_encoder | 13.13B | BF16 | 24.46 GB | Turns your prompt into conditioning. Runs once, then can leave the GPU entirely. |
These figures were stated by us, because the weights ship in a format with no readable header.
Memory by precision
Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.
| Precision | Peak VRAM | Quality | What it costs you |
|---|---|---|---|
| BF16 | 66.21 GB | 100% | How the weights are published. No loss, and the largest footprint. |
| FP8 | 34.41 GB | 97% | Halves the denoiser with a small, usually invisible cost. Needs Ada or newer. |
| GGUF Q8_0 | 36.4 GB | 98% | Works on any card, unlike FP8. Slightly slower than native precision. |
| GGUF Q5_K_M | 25.23 GB | 95% | A middle step when Q8 will not fit. |
| GGUF Q4_K_M | 21.82 GB | 91% | The usual way a 12B image model gets onto an 8 GB card. Detail softens. |
| NF4 | 20.11 GB | 89% | Aggressive 4-bit. Fast to load, noticeably looser on fine detail. |
Memory by arrangement
The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 62% of the pipeline and stays resident whatever you rearrange.
| Arrangement | Peak VRAM | Time | How it works |
|---|---|---|---|
| Everything resident | 66.21 GB | 5 min 35 s | All components stay on the GPU. Fastest, and needs the most memory. |
| Text encoder on CPU | 41.75 GB | 5 min 52 s | The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers. |
| Component offload | 41.75 GB | 6 min 15 s | One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation. |
| Sequential offload | 10.25 GB | 33 min 4 s | Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower. |
What a bigger output costs
Both memory and time climb faster than the frame count, because attention is quadratic in the length of the latent sequence.
| Output | Latent tokens | Peak VRAM | Time |
|---|---|---|---|
| 480p, 3 seconds | 2,730 | 65.36 GB | 37 s |
| 480p, 5 seconds | 4,290 | 65.36 GB | 1 min 5 s |
| 720p, 3 seconds | 6,440 | 65.36 GB | 1 min 49 s |
| 720p, 5 seconds | 10,120 | 65.73 GB | 3 min 25 s |
| 720p, 10 seconds | 19,320 | 66.85 GB | 9 min 10 s |
Measured at BF16 with everything resident on an RTX 4090, so the columns compare with each other rather than with your machine.
Which hardware runs LTX 2.5
117 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.
| Device | Memory | Verdict | Arrangement | Peak | Time |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell NVIDIA | 96 GB | Workable | BF16 everything resident | 66.21 GB | 3 min 42 s |
| GeForce RTX 5090 NVIDIA | 32 GB | Workable | GGUF Q5_K_M everything resident | 25.23 GB | 4 min 25 s |
| RTX 6000 Ada Generation NVIDIA | 48 GB | Slow | GGUF Q8_0 everything resident | 36.4 GB | 5 min 3 s |
| GeForce RTX 4090 NVIDIA | 24 GB | Slow | GGUF Q4_K_M everything resident | 21.82 GB | 5 min 35 s |
| GeForce RTX 5090 Laptop NVIDIA | 24 GB | Slow | GGUF Q4_K_M everything resident | 21.82 GB | 7 min 4 s |
| NVIDIA DGX Spark 128GB NVIDIA | 128 GB | Slow | BF16 everything resident | 66.21 GB | 7 min 21 s |
| GeForce RTX 5080 NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 8 min 33 s |
| GeForce RTX 4080 SUPER NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 9 min 15 s |
| GeForce RTX 4090 Laptop NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 9 min 37 s |
| GeForce RTX 4080 NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 9 min 51 s |
| GeForce RTX 5080 Laptop NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 10 min 7 s |
| GeForce RTX 5070 Ti NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 10 min 51 s |
| GeForce RTX 4070 Ti SUPER NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 10 min 55 s |
| GeForce RTX 3090 Ti NVIDIA | 24 GB | Slow | GGUF Q4_K_M everything resident | 21.82 GB | 11 min 26 s |
| RTX A6000 NVIDIA | 48 GB | Slow | GGUF Q8_0 everything resident | 36.4 GB | 11 min 48 s |
| GeForce RTX 3090 NVIDIA | 24 GB | Slow | GGUF Q4_K_M everything resident | 21.82 GB | 12 min 52 s |
| RTX A5000 NVIDIA | 24 GB | Slow | GGUF Q4_K_M everything resident | 21.82 GB | 16 min 26 s |
| GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | Slow | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 19 min 56 s |
| Jetson AGX Orin 32GB NVIDIA | 32 GB | Barely usable | GGUF Q4_K_M everything resident | 21.82 GB | 21 min 26 s |
| Jetson AGX Orin 64GB NVIDIA | 64 GB | Barely usable | GGUF Q8_0 everything resident | 36.4 GB | 21 min 26 s |
| Radeon RX 7900 XTX AMD | 24 GB | Barely usable | GGUF Q4_K_M everything resident | 21.82 GB | 21 min 36 s |
| GeForce RTX 4060 Ti 16GB NVIDIA | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 21 min 44 s |
| Radeon PRO W7900 AMD | 48 GB | Barely usable | GGUF Q8_0 everything resident | 36.4 GB | 21 min 46 s |
| RTX A4000 NVIDIA | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 24 min 49 s |
| Radeon RX 7900 XT AMD | 20 GB | Barely usable | GGUF Q5_K_M text encoder on cpu | 16.53 GB | 27 min 3 s |
| Radeon RX 9070 XT AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 28 min 43 s |
| Radeon RX 7900 GRE AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 30 min 17 s |
| Radeon RX 9070 AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 35 min 42 s |
| Radeon RX 7800 XT AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 37 min 7 s |
| Radeon RX 6900 XT AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 1 h 0 min |
| Radeon RX 7600 XT AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 1 h 1 min |
| Ryzen AI Max+ 395 32GB AMD | 32 GB | Barely usable | GGUF Q4_K_M everything resident | 21.82 GB | 1 h 7 min |
| Ryzen AI Max+ 395 64GB AMD | 64 GB | Barely usable | GGUF Q8_0 everything resident | 36.4 GB | 1 h 7 min |
| Ryzen AI Max+ 395 96GB AMD | 96 GB | Barely usable | BF16 everything resident | 66.21 GB | 1 h 7 min |
| Ryzen AI Max+ 395 128GB AMD | 128 GB | Barely usable | BF16 everything resident | 66.21 GB | 1 h 7 min |
| GeForce RTX 4070 Ti NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 8 min |
| GeForce RTX 4080 Laptop NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 13 min |
| GeForce RTX 4070 SUPER NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 17 min |
| GeForce RTX 3080 Ti NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 20 min |
| Radeon RX 6800 AMD | 16 GB | Barely usable | GGUF Q4_K_M text encoder on cpu | 14.44 GB | 1 h 26 min |
| GeForce RTX 5070 NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 28 min |
| GeForce RTX 3080 12GB NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 31 min |
| GeForce RTX 3080 10GB NVIDIA | 10 GB | Barely usable | GGUF Q8_0 sequential offload | 6.67 GB | 1 h 31 min |
| GeForce RTX 4070 NVIDIA | 12 GB | Barely usable | BF16 sequential offload | 10.25 GB | 1 h 33 min |
| GeForce RTX 2080 Ti NVIDIA | 11 GB | Barely usable | GGUF Q8_0 sequential offload | 6.67 GB | 1 h 40 min |
Source: Lightricks/LTX-2.5 . Downloaded 1.2 million times in the last month. See how we calculate, or browse every video model.