F5 · Diffusion transformer · Text to speech
F5-TTS
F5-TTS is a single network of 0.337 billion parameters, of which 0.337 billion do the actual generating. It produces the whole utterance in parallel rather than token by token, which is why models this small can run many times faster than real time.
What the pipeline is made of
This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.
| Component | Parameters | Published as | On disk | What it does |
|---|---|---|---|---|
| F5TTS_v1_Base | 0.337B | F32 | 1.26 GB | Auxiliary encoder. |
These figures were measured from the weight file headers.
Memory by precision
Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.
| Precision | Peak VRAM | Quality | What it costs you |
|---|---|---|---|
| BF16 | 2.03 GB | 100% | How the weights are published. No loss, and the largest footprint. |
| FP8 | 1.71 GB | 97% | Halves the denoiser with a small, usually invisible cost. Needs Ada or newer. |
| GGUF Q8_0 | 1.73 GB | 98% | Works on any card, unlike FP8. Slightly slower than native precision. |
| GGUF Q5_K_M | 1.62 GB | 95% | A middle step when Q8 will not fit. |
| GGUF Q4_K_M | 1.59 GB | 91% | The usual way a 12B image model gets onto an 8 GB card. Detail softens. |
| NF4 | 1.57 GB | 89% | Aggressive 4-bit. Fast to load, noticeably looser on fine detail. |
Memory by arrangement
The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 100% of the pipeline and stays resident whatever you rearrange.
| Arrangement | Peak VRAM | Time | How it works |
|---|---|---|---|
| Everything resident | 2.03 GB | 2.7 s | All components stay on the GPU. Fastest, and needs the most memory. |
| Text encoder on CPU | 2.03 GB | 2.7 s | The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers. |
| Component offload | 2.03 GB | 2.7 s | One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation. |
| Sequential offload | 1.55 GB | 2.7 s | Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower. |
Which hardware runs F5-TTS
118 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.
| Device | Memory | Verdict | Arrangement | Peak | Time |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell NVIDIA | 96 GB | Comfortable | BF16 everything resident | 2.03 GB | 33.4x real time |
| GeForce RTX 5090 NVIDIA | 32 GB | Comfortable | BF16 everything resident | 2.03 GB | 27.9x real time |
| RTX 6000 Ada Generation NVIDIA | 48 GB | Comfortable | BF16 everything resident | 2.03 GB | 24.3x real time |
| GeForce RTX 4090 NVIDIA | 24 GB | Comfortable | BF16 everything resident | 2.03 GB | 21.9x real time |
| GeForce RTX 5090 Laptop NVIDIA | 24 GB | Comfortable | BF16 everything resident | 2.03 GB | 17.3x real time |
| NVIDIA DGX Spark 128GB NVIDIA | 128 GB | Comfortable | BF16 everything resident | 2.03 GB | 16.6x real time |
| GeForce RTX 5080 NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 15x real time |
| GeForce RTX 4080 SUPER NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 13.8x real time |
| GeForce RTX 4090 Laptop NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 13.3x real time |
| GeForce RTX 4080 NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 13x real time |
| GeForce RTX 5080 Laptop NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 12.6x real time |
| GeForce RTX 5070 Ti NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 11.8x real time |
| GeForce RTX 4070 Ti SUPER NVIDIA | 16 GB | Comfortable | BF16 everything resident | 2.03 GB | 11.7x real time |
| GeForce RTX 4070 Ti NVIDIA | 12 GB | Comfortable | BF16 everything resident | 2.03 GB | 10.6x real time |
| GeForce RTX 3090 Ti NVIDIA | 24 GB | Comfortable | BF16 everything resident | 2.03 GB | 10.6x real time |
| RTX A6000 NVIDIA | 48 GB | Comfortable | BF16 everything resident | 2.03 GB | 10.3x real time |
| GeForce RTX 4080 Laptop NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 9.8x real time |
| GeForce RTX 4070 SUPER NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 9.4x real time |
| GeForce RTX 3090 NVIDIA | 24 GB | Workable | BF16 everything resident | 2.03 GB | 9.4x real time |
| GeForce RTX 3080 Ti NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 9x real time |
| GeForce RTX 5070 NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 8.2x real time |
| Radeon RX 7900 XTX AMD | 24 GB | Workable | BF16 everything resident | 2.03 GB | 8.2x real time |
| Radeon PRO W7900 AMD | 48 GB | Workable | BF16 everything resident | 2.03 GB | 8.1x real time |
| GeForce RTX 3080 12GB NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 7.9x real time |
| GeForce RTX 3080 10GB NVIDIA | 10 GB | Workable | BF16 everything resident | 2.03 GB | 7.9x real time |
| GeForce RTX 4070 NVIDIA | 12 GB | Workable | BF16 everything resident | 2.03 GB | 7.7x real time |
| RTX A5000 NVIDIA | 24 GB | Workable | BF16 everything resident | 2.03 GB | 7.4x real time |
| GeForce RTX 2080 Ti NVIDIA | 11 GB | Workable | BF16 everything resident | 2.03 GB | 7.2x real time |
| Radeon RX 7900 XT AMD | 20 GB | Workable | BF16 everything resident | 2.03 GB | 6.8x real time |
| Radeon RX 9070 XT AMD | 16 GB | Workable | BF16 everything resident | 2.03 GB | 6.4x real time |
| GeForce RTX 5060 Ti 16GB NVIDIA | 16 GB | Workable | BF16 everything resident | 2.03 GB | 6.4x real time |
| GeForce RTX 5060 Ti 8GB NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 6.4x real time |
| GeForce RTX 4070 Laptop NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 6.2x real time |
| Radeon RX 7900 GRE AMD | 16 GB | Workable | BF16 everything resident | 2.03 GB | 6.1x real time |
| GeForce RTX 4060 Ti 16GB NVIDIA | 16 GB | Workable | BF16 everything resident | 2.03 GB | 5.8x real time |
| GeForce RTX 4060 Ti 8GB NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 5.8x real time |
| GeForce RTX 3070 Ti NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 5.8x real time |
| Jetson AGX Orin 32GB NVIDIA | 32 GB | Workable | BF16 everything resident | 2.03 GB | 5.6x real time |
| Jetson AGX Orin 64GB NVIDIA | 64 GB | Workable | BF16 everything resident | 2.03 GB | 5.6x real time |
| GeForce RTX 3070 NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 5.4x real time |
| Radeon RX 9070 AMD | 16 GB | Workable | BF16 everything resident | 2.03 GB | 5.2x real time |
| GeForce RTX 5060 NVIDIA | 8 GB | Workable | BF16 everything resident | 2.03 GB | 5.1x real time |
| RTX A4000 NVIDIA | 16 GB | Workable | BF16 everything resident | 2.03 GB | 5.1x real time |
| Radeon RX 7800 XT AMD | 16 GB | Workable | BF16 everything resident | 2.03 GB | 5x real time |
| Radeon RX 7700 XT AMD | 12 GB | Workable | BF16 everything resident | 2.03 GB | 4.7x real time |
Source: SWivid/F5-TTS . Downloaded 0.8 million times in the last month. See how we calculate, or browse every audio model.