Wan ยท Diffusion transformer

Wan 2.2 T2V A14B

Wan 2.2 T2V A14B is a pipeline of 4 networks totalling 34.384 billion parameters, of which 14.288 billion do the actual generating. At its native 1280x720 and 81 frames the denoiser works on 75,600 latent tokens at once, and attention over that sequence is what sets the render time.

Whole pipeline34.384B
Denoiser14.288B
Download117.5 GB
Default steps40
Native length81 frames
Frame rate16 fps

What the pipeline is made of

This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.

ComponentParametersPublished asOn diskWhat it does
transformer 14.288B F32 53.23 GB Does the generating. Runs once per step, so it dominates both memory and time.
transformer_2 14.288B F32 53.23 GB Does the generating. Runs once per step, so it dominates both memory and time.
text_encoder 5.681B BF16 10.58 GB Turns your prompt into conditioning. Runs once, then can leave the GPU entirely.
vae 0.127B F32 0.47 GB Converts between pixels and the compressed latent space. Small, but its decode pass is a memory spike.

These figures were measured from the weight file headers.

Memory by precision

Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.

PrecisionPeak VRAMQualityWhat it costs you
BF16 76.48 GB 100% How the weights are published. No loss, and the largest footprint.
FP8 44.58 GB 97% Halves the denoiser with a small, usually invisible cost. Needs Ada or newer.
GGUF Q8_0 46.57 GB 98% Works on any card, unlike FP8. Slightly slower than native precision.
GGUF Q5_K_M 35.36 GB 95% A middle step when Q8 will not fit.
GGUF Q4_K_M 31.94 GB 91% The usual way a 12B image model gets onto an 8 GB card. Detail softens.
NF4 30.22 GB 89% Aggressive 4-bit. Fast to load, noticeably looser on fine detail.

Memory by arrangement

The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 42% of the pipeline and stays resident whatever you rearrange.

ArrangementPeak VRAMTimeHow it works
Everything resident 76.48 GB 1 h 13 min All components stay on the GPU. Fastest, and needs the most memory.
Text encoder on CPU 65.9 GB 1 h 17 min The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers.
Component offload 65.66 GB 1 h 22 min One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation.
Sequential offload 20.09 GB 7 h 23 min Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower.

What a bigger output costs

Both memory and time climb faster than the frame count, because attention is quadratic in the length of the latent sequence.

OutputLatent tokensPeak VRAMTime
480p, 3 seconds 20,280 68.04 GB 7 min 47 s
480p, 5 seconds 32,760 69.94 GB 16 min 57 s
720p, 3 seconds 46,800 72.09 GB 31 min 17 s
720p, 5 seconds 75,600 76.48 GB 1 h 13 min
720p, 10 seconds 147,600 87.47 GB 4 h 18 min

Measured at BF16 with everything resident on an RTX 4090, so the columns compare with each other rather than with your machine.

Which hardware runs Wan 2.2 T2V A14B

74 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.

DeviceMemoryVerdictArrangementPeakTime
RTX PRO 6000 Blackwell
NVIDIA
96 GB Barely usable BF16
everything resident
76.48 GB 48 min 31 s
GeForce RTX 5090
NVIDIA
32 GB Barely usable NF4
everything resident
30.22 GB 58 min 13 s
RTX 6000 Ada Generation
NVIDIA
48 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 1 h 6 min
NVIDIA DGX Spark 128GB
NVIDIA
128 GB Barely usable BF16
everything resident
76.48 GB 1 h 37 min
RTX A6000
NVIDIA
48 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 2 h 37 min
Jetson AGX Orin 64GB
NVIDIA
64 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 4 h 46 min
Radeon PRO W7900
AMD
48 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 4 h 51 min
GeForce RTX 4090
NVIDIA
24 GB Barely usable BF16
sequential offload
20.09 GB 7 h 23 min
GeForce RTX 5090 Laptop
NVIDIA
24 GB Barely usable BF16
sequential offload
20.09 GB 9 h 22 min
GeForce RTX 5080
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 10 h 49 min
GeForce RTX 4080 SUPER
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 11 h 42 min
GeForce RTX 4090 Laptop
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 12 h 11 min
GeForce RTX 4080
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 12 h 29 min
GeForce RTX 5080 Laptop
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 12 h 49 min
GeForce RTX 5070 Ti
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 13 h 46 min
GeForce RTX 4070 Ti SUPER
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 13 h 50 min
Ryzen AI Max+ 395 64GB
AMD
64 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 15 h 3 min
Ryzen AI Max+ 395 96GB
AMD
96 GB Barely usable GGUF Q8_0
everything resident
46.57 GB 15 h 3 min
Ryzen AI Max+ 395 128GB
AMD
128 GB Barely usable BF16
everything resident
76.48 GB 15 h 3 min
GeForce RTX 3090 Ti
NVIDIA
24 GB Barely usable BF16
sequential offload
20.09 GB 15 h 13 min
GeForce RTX 3090
NVIDIA
24 GB Barely usable BF16
sequential offload
20.09 GB 17 h 9 min
RTX A5000
NVIDIA
24 GB Barely usable BF16
sequential offload
20.09 GB 21 h 57 min
Apple M3 Ultra 96GB
Apple
96 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 23 h 4 min
Apple M3 Ultra 256GB
Apple
256 GB Barely usable BF16
everything resident
75.98 GB 23 h 4 min
Apple M3 Ultra 512GB
Apple
512 GB Barely usable BF16
everything resident
75.98 GB 23 h 4 min
Apple M2 Ultra 64GB
Apple
64 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 23 h 45 min
Apple M2 Ultra 128GB
Apple
128 GB Barely usable BF16
everything resident
75.98 GB 23 h 45 min
Apple M2 Ultra 192GB
Apple
192 GB Barely usable BF16
everything resident
75.98 GB 23 h 45 min
GeForce RTX 5060 Ti 16GB
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 25 h 23 min
GeForce RTX 4060 Ti 16GB
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 27 h 41 min
Jetson AGX Orin 32GB
NVIDIA
32 GB Barely usable BF16
sequential offload
20.09 GB 28 h 40 min
Radeon RX 7900 XTX
AMD
24 GB Barely usable BF16
sequential offload
20.09 GB 28 h 53 min
Apple M1 Ultra 64GB
Apple
64 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 30 h 46 min
Apple M1 Ultra 128GB
Apple
128 GB Barely usable BF16
everything resident
75.98 GB 30 h 46 min
RTX A4000
NVIDIA
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 31 h 38 min
Radeon RX 7900 XT
AMD
20 GB Barely usable GGUF Q8_0
sequential offload
16.5 GB 34 h 30 min
Apple M4 Max 48GB
Apple
48 GB Barely usable GGUF Q5_K_M
everything resident
34.86 GB 35 h 6 min
Apple M4 Max 64GB
Apple
64 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 35 h 6 min
Apple M4 Max 128GB
Apple
128 GB Barely usable BF16
everything resident
75.98 GB 35 h 6 min
Radeon RX 9070 XT
AMD
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 36 h 38 min
Radeon RX 7900 GRE
AMD
16 GB Barely usable GGUF Q5_K_M
sequential offload
15.16 GB 38 h 37 min
Apple M3 Max 48GB
Apple
48 GB Barely usable GGUF Q5_K_M
everything resident
34.86 GB 45 h 30 min
Apple M3 Max 64GB
Apple
64 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 45 h 30 min
Apple M3 Max 96GB
Apple
96 GB Barely usable GGUF Q8_0
everything resident
46.07 GB 45 h 30 min
Apple M3 Max 128GB
Apple
128 GB Barely usable BF16
everything resident
75.98 GB 45 h 30 min

Source: Wan-AI/Wan2.2-T2V-A14B-Diffusers . Downloaded 0.1 million times in the last month. See how we calculate, or browse every video model.