XTTS · Autoregressive · Text to speech

XTTS v2

XTTS v2 is a single network of 0.467 billion parameters, of which 0.467 billion do the actual generating. It generates one audio token at a time, which makes it bandwidth-bound like a language model rather than compute-bound like a diffusion model. What sets its speed is not its size but its token rate: 21.5 tokens for every second of output.

Whole pipeline0.467B
Denoiser0.467B
Download1.7 GB

What the pipeline is made of

This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.

ComponentParametersPublished asOn diskWhat it does
model 0.467B F32 1.74 GB Auxiliary encoder.

These figures were stated by us, because the weights ship in a format with no readable header.

Memory by precision

Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.

PrecisionPeak VRAMQualityWhat it costs you
BF16 2.27 GB 100% How the weights are published. No loss, and the largest footprint.
FP8 1.83 GB 97% Halves the denoiser with a small, usually invisible cost. Needs Ada or newer.
GGUF Q8_0 1.86 GB 98% Works on any card, unlike FP8. Slightly slower than native precision.
GGUF Q5_K_M 1.71 GB 95% A middle step when Q8 will not fit.
GGUF Q4_K_M 1.66 GB 91% The usual way a 12B image model gets onto an 8 GB card. Detail softens.
NF4 1.64 GB 89% Aggressive 4-bit. Fast to load, noticeably looser on fine detail.

Memory by arrangement

The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 100% of the pipeline and stays resident whatever you rearrange.

ArrangementPeak VRAMTimeHow it works
Everything resident 2.27 GB 18 s All components stay on the GPU. Fastest, and needs the most memory.
Text encoder on CPU 2.27 GB 18 s The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers.
Component offload 2.27 GB 18 s One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation.
Sequential offload 1.5 GB 18 s Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower.

Which hardware runs XTTS v2

118 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.

DeviceMemoryVerdictArrangementPeakTime
GeForce RTX 5090
NVIDIA
32 GB Comfortable BF16
everything resident
2.27 GB 29.3x real time
RTX PRO 6000 Blackwell
NVIDIA
96 GB Comfortable BF16
everything resident
2.27 GB 29.3x real time
GeForce RTX 4090
NVIDIA
24 GB Comfortable BF16
everything resident
2.27 GB 16.5x real time
GeForce RTX 3090 Ti
NVIDIA
24 GB Comfortable BF16
everything resident
2.27 GB 16.5x real time
GeForce RTX 5080
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 15.7x real time
RTX 6000 Ada Generation
NVIDIA
48 GB Comfortable BF16
everything resident
2.27 GB 15.7x real time
GeForce RTX 3090
NVIDIA
24 GB Comfortable BF16
everything resident
2.27 GB 15.3x real time
GeForce RTX 3080 Ti
NVIDIA
12 GB Comfortable BF16
everything resident
2.27 GB 14.9x real time
GeForce RTX 3080 12GB
NVIDIA
12 GB Comfortable BF16
everything resident
2.27 GB 14.9x real time
GeForce RTX 5070 Ti
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 14.6x real time
GeForce RTX 5090 Laptop
NVIDIA
24 GB Comfortable BF16
everything resident
2.27 GB 14.6x real time
Radeon RX 7900 XTX
AMD
24 GB Comfortable BF16
everything resident
2.27 GB 13.8x real time
Apple M3 Ultra 96GB
Apple
96 GB Comfortable BF16
everything resident
1.77 GB 12.7x real time
Apple M3 Ultra 256GB
Apple
256 GB Comfortable BF16
everything resident
1.77 GB 12.7x real time
Apple M3 Ultra 512GB
Apple
512 GB Comfortable BF16
everything resident
1.77 GB 12.7x real time
GeForce RTX 3080 10GB
NVIDIA
10 GB Comfortable BF16
everything resident
2.27 GB 12.4x real time
GeForce RTX 5080 Laptop
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 12.5x real time
RTX A6000
NVIDIA
48 GB Comfortable BF16
everything resident
2.27 GB 12.5x real time
RTX A5000
NVIDIA
24 GB Comfortable BF16
everything resident
2.27 GB 12.5x real time
Radeon PRO W7900
AMD
48 GB Comfortable BF16
everything resident
2.27 GB 12.4x real time
Apple M1 Ultra 64GB
Apple
64 GB Comfortable BF16
everything resident
1.77 GB 12.4x real time
Apple M1 Ultra 128GB
Apple
128 GB Comfortable BF16
everything resident
1.77 GB 12.4x real time
Apple M2 Ultra 64GB
Apple
64 GB Comfortable BF16
everything resident
1.77 GB 12.4x real time
Apple M2 Ultra 128GB
Apple
128 GB Comfortable BF16
everything resident
1.77 GB 12.4x real time
Apple M2 Ultra 192GB
Apple
192 GB Comfortable BF16
everything resident
1.77 GB 12.4x real time
GeForce RTX 4080 SUPER
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 12x real time
GeForce RTX 4080
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 11.7x real time
Radeon RX 7900 XT
AMD
20 GB Comfortable BF16
everything resident
2.27 GB 11.5x real time
GeForce RTX 5070
NVIDIA
12 GB Comfortable BF16
everything resident
2.27 GB 11x real time
GeForce RTX 4070 Ti SUPER
NVIDIA
16 GB Comfortable BF16
everything resident
2.27 GB 11x real time
GeForce RTX 3070 Ti
NVIDIA
8 GB Workable BF16
everything resident
2.27 GB 9.9x real time
GeForce RTX 2080 Ti
NVIDIA
11 GB Comfortable BF16
everything resident
2.27 GB 10.1x real time
GeForce RTX 4090 Laptop
NVIDIA
16 GB Workable BF16
everything resident
2.27 GB 9.4x real time
Radeon RX 9070 XT
AMD
16 GB Workable BF16
everything resident
2.27 GB 9.3x real time
Radeon RX 9070
AMD
16 GB Workable BF16
everything resident
2.27 GB 9.3x real time
Arc A770 16GB
Intel
16 GB Workable BF16
everything resident
2.27 GB 9.1x real time
Radeon RX 7800 XT
AMD
16 GB Workable BF16
everything resident
2.27 GB 8.9x real time
Apple M4 Max 36GB
Apple
36 GB Workable BF16
everything resident
1.77 GB 8.5x real time
Apple M4 Max 48GB
Apple
48 GB Workable BF16
everything resident
1.77 GB 8.5x real time
Apple M4 Max 64GB
Apple
64 GB Workable BF16
everything resident
1.77 GB 8.5x real time
Apple M4 Max 128GB
Apple
128 GB Workable BF16
everything resident
1.77 GB 8.5x real time
Arc A750
Intel
8 GB Workable BF16
everything resident
2.27 GB 8.4x real time
GeForce RTX 4070 Ti
NVIDIA
12 GB Workable BF16
everything resident
2.27 GB 8.2x real time
GeForce RTX 4070 SUPER
NVIDIA
12 GB Workable BF16
everything resident
2.27 GB 8.2x real time
GeForce RTX 4070
NVIDIA
12 GB Workable BF16
everything resident
2.27 GB 8.2x real time

Source: coqui/XTTS-v2 . Downloaded 7.8 million times in the last month. See how we calculate, or browse every audio model.