Wan ยท Diffusion transformer

Wan 2.1 T2V 1.3B

Wan 2.1 T2V 1.3B is a pipeline of 3 networks totalling 7.227 billion parameters, of which 1.419 billion do the actual generating. At its native 832x480 and 81 frames the denoiser works on 32,760 latent tokens at once, and attention over that sequence is what sets the render time.

Whole pipeline7.227B
Denoiser1.419B
Download26.9 GB
Default steps50
Native length81 frames
Frame rate16 fps

What the pipeline is made of

This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.

ComponentParametersPublished asOn diskWhat it does
transformer 1.419B F32 5.29 GB Does the generating. Runs once per step, so it dominates both memory and time.
text_encoder 5.681B F32 21.16 GB Turns your prompt into conditioning. Runs once, then can leave the GPU entirely.
vae 0.127B F32 0.47 GB Converts between pixels and the compressed latent space. Small, but its decode pass is a memory spike.

These figures were measured from the weight file headers.

Memory by precision

Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.

PrecisionPeak VRAMQualityWhat it costs you
BF16 15.86 GB 100% How the weights are published. No loss, and the largest footprint.
FP8 9.25 GB 97% Halves the denoiser with a small, usually invisible cost. Needs Ada or newer.
GGUF Q8_0 9.66 GB 98% Works on any card, unlike FP8. Slightly slower than native precision.
GGUF Q5_K_M 7.34 GB 95% A middle step when Q8 will not fit.
GGUF Q4_K_M 6.63 GB 91% The usual way a 12B image model gets onto an 8 GB card. Detail softens.
NF4 6.27 GB 89% Aggressive 4-bit. Fast to load, noticeably looser on fine detail.

Memory by arrangement

The other lever: the same weights at the same precision, moved around differently. On this model the strongest lever is neither: run the text encoder on the processor. It is 79% of the pipeline, and once it is off the GPU the 20% denoiser is all that remains.

ArrangementPeak VRAMTimeHow it works
Everything resident 15.86 GB 5 min 28 s All components stay on the GPU. Fastest, and needs the most memory.
Text encoder on CPU 5.28 GB 5 min 44 s The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers.
Component offload 12.98 GB 6 min 7 s One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation.
Sequential offload 3.99 GB 32 min 40 s Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower.

What a bigger output costs

Both memory and time climb faster than the frame count, because attention is quadratic in the length of the latent sequence.

OutputLatent tokensPeak VRAMTime
480p, 3 seconds 20,280 15.29 GB 2 min 16 s
480p, 5 seconds 32,760 15.86 GB 5 min 28 s
720p, 3 seconds 46,800 16.5 GB 10 min 44 s
720p, 5 seconds 75,600 17.82 GB 27 min 0 s
720p, 10 seconds 147,600 21.12 GB 1 h 39 min

Measured at BF16 with everything resident on an RTX 4090, so the columns compare with each other rather than with your machine.

Which hardware runs Wan 2.1 T2V 1.3B

118 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.

DeviceMemoryVerdictArrangementPeakTime
RTX PRO 6000 Blackwell
NVIDIA
96 GB Workable BF16
everything resident
15.86 GB 3 min 36 s
GeForce RTX 5090
NVIDIA
32 GB Workable BF16
everything resident
15.86 GB 4 min 19 s
RTX 6000 Ada Generation
NVIDIA
48 GB Workable BF16
everything resident
15.86 GB 4 min 56 s
GeForce RTX 4090
NVIDIA
24 GB Slow BF16
everything resident
15.86 GB 5 min 28 s
GeForce RTX 5090 Laptop
NVIDIA
24 GB Slow BF16
everything resident
15.86 GB 6 min 56 s
NVIDIA DGX Spark 128GB
NVIDIA
128 GB Slow BF16
everything resident
15.86 GB 7 min 12 s
GeForce RTX 5080
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 8 min 0 s
GeForce RTX 4080 SUPER
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 8 min 39 s
GeForce RTX 4090 Laptop
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 9 min 0 s
GeForce RTX 4080
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 9 min 14 s
GeForce RTX 5080 Laptop
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 9 min 29 s
GeForce RTX 5070 Ti
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 10 min 10 s
GeForce RTX 4070 Ti SUPER
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 10 min 14 s
GeForce RTX 4070 Ti
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 11 min 15 s
GeForce RTX 3090 Ti
NVIDIA
24 GB Slow BF16
everything resident
15.86 GB 11 min 15 s
RTX A6000
NVIDIA
48 GB Slow BF16
everything resident
15.86 GB 11 min 37 s
GeForce RTX 4080 Laptop
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 12 min 9 s
GeForce RTX 3090
NVIDIA
24 GB Slow BF16
everything resident
15.86 GB 12 min 40 s
GeForce RTX 4070 SUPER
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 12 min 46 s
GeForce RTX 3080 Ti
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 13 min 14 s
GeForce RTX 5070
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 14 min 37 s
GeForce RTX 3080 12GB
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 15 min 7 s
GeForce RTX 3080 10GB
NVIDIA
10 GB Slow GGUF Q5_K_M
everything resident
7.34 GB 15 min 7 s
GeForce RTX 4070
NVIDIA
12 GB Slow GGUF Q8_0
everything resident
9.66 GB 15 min 30 s
RTX A5000
NVIDIA
24 GB Slow BF16
everything resident
15.86 GB 16 min 12 s
GeForce RTX 2080 Ti
NVIDIA
11 GB Slow GGUF Q8_0
everything resident
9.66 GB 16 min 39 s
GeForce RTX 5060 Ti 16GB
NVIDIA
16 GB Slow GGUF Q8_0
everything resident
9.66 GB 18 min 44 s
GeForce RTX 5060 Ti 8GB
NVIDIA
8 GB Slow GGUF Q4_K_M
everything resident
6.63 GB 18 min 44 s
GeForce RTX 4070 Laptop
NVIDIA
8 GB Slow GGUF Q4_K_M
everything resident
6.63 GB 19 min 20 s
GeForce RTX 4060 Ti 16GB
NVIDIA
16 GB Barely usable GGUF Q8_0
everything resident
9.66 GB 20 min 26 s
GeForce RTX 4060 Ti 8GB
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 20 min 26 s
GeForce RTX 3070 Ti
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 20 min 40 s
Jetson AGX Orin 32GB
NVIDIA
32 GB Barely usable BF16
everything resident
15.86 GB 21 min 9 s
Jetson AGX Orin 64GB
NVIDIA
64 GB Barely usable BF16
everything resident
15.86 GB 21 min 9 s
Radeon RX 7900 XTX
AMD
24 GB Barely usable BF16
everything resident
15.86 GB 21 min 19 s
Radeon PRO W7900
AMD
48 GB Barely usable BF16
everything resident
15.86 GB 21 min 29 s
GeForce RTX 3070
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 22 min 11 s
GeForce RTX 5060
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 23 min 20 s
RTX A4000
NVIDIA
16 GB Barely usable GGUF Q8_0
everything resident
9.66 GB 23 min 20 s
Radeon RX 7900 XT
AMD
20 GB Barely usable BF16
everything resident
15.86 GB 25 min 27 s
Radeon RX 9070 XT
AMD
16 GB Barely usable GGUF Q8_0
everything resident
9.66 GB 27 min 1 s
GeForce RTX 4060 Laptop
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 27 min 14 s
GeForce RTX 3060 Ti
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 27 min 39 s
Radeon RX 7900 GRE
AMD
16 GB Barely usable GGUF Q8_0
everything resident
9.66 GB 28 min 29 s
GeForce RTX 4060
NVIDIA
8 GB Barely usable GGUF Q4_K_M
everything resident
6.63 GB 29 min 27 s

Source: Wan-AI/Wan2.1-T2V-1.3B-Diffusers . Downloaded 0.2 million times in the last month. See how we calculate, or browse every video model.