MusicGen · Autoregressive · Music and sound
MusicGen Large
MusicGen Large is a single network of 3.3 billion parameters, of which 3.3 billion do the actual generating. It generates one audio token at a time, which makes it bandwidth-bound like a language model rather than compute-bound like a diffusion model. What sets its speed is not its size but its token rate: 50 tokens for every second of output.
What the pipeline is made of
This is the part that separates a media model from a language model: these are separate networks, and whether they sit in memory together is your choice, not the model's.
| Component | Parameters | Published as | On disk | What it does |
|---|---|---|---|---|
| model | 3.3B | F32 | 12.29 GB | Auxiliary encoder. |
These figures were stated by us, because the weights ship in a format with no readable header.
Memory by precision
Everything resident, at the model's native output size, on a card large enough that precision is the only constraint.
| Precision | Peak VRAM | Quality | What it costs you |
|---|---|---|---|
| BF16 | 7.55 GB | 100% | How the weights are published. No loss, and the largest footprint. |
| FP8 | 4.47 GB | 97% | Halves the denoiser with a small, usually invisible cost. Needs Ada or newer. |
| GGUF Q8_0 | 4.67 GB | 98% | Works on any card, unlike FP8. Slightly slower than native precision. |
| GGUF Q5_K_M | 3.59 GB | 95% | A middle step when Q8 will not fit. |
| GGUF Q4_K_M | 3.26 GB | 91% | The usual way a 12B image model gets onto an 8 GB card. Detail softens. |
| NF4 | 3.09 GB | 89% | Aggressive 4-bit. Fast to load, noticeably looser on fine detail. |
Memory by arrangement
The other lever: the same weights at the same precision, moved around differently. On this model quantising is the stronger lever: the denoiser is 100% of the pipeline and stays resident whatever you rearrange.
| Arrangement | Peak VRAM | Time | How it works |
|---|---|---|---|
| Everything resident | 7.55 GB | 5 min 2 s | All components stay on the GPU. Fastest, and needs the most memory. |
| Text encoder on CPU | 7.55 GB | 5 min 2 s | The prompt is encoded once on the processor, so the encoder never touches the GPU at all. Standard practice on video models, whose encoders are often larger than their denoisers. |
| Component offload | 7.55 GB | 5 min 2 s | One component on the GPU at a time. The text encoder runs, then makes way for the denoiser. Costs a few seconds per generation. |
| Sequential offload | 2.14 GB | 5 min 2 s | Weights stream layer by layer from system RAM. Runs almost anything on almost anything, and is many times slower. |
Which hardware runs MusicGen Large
118 of 118 consumer devices run it in some arrangement. Each row shows the fastest arrangement that fits on that device.
| Device | Memory | Verdict | Arrangement | Peak | Time |
|---|---|---|---|---|---|
| GeForce RTX 5090 NVIDIA | 32 GB | Slow | BF16 everything resident | 7.55 GB | 1.8x real time |
| RTX PRO 6000 Blackwell NVIDIA | 96 GB | Slow | BF16 everything resident | 7.55 GB | 1.8x real time |
| GeForce RTX 3070 Ti NVIDIA | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 1.1x real time |
| GeForce RTX 4090 NVIDIA | 24 GB | Slow | BF16 everything resident | 7.55 GB | 1x real time |
| GeForce RTX 3090 Ti NVIDIA | 24 GB | Slow | BF16 everything resident | 7.55 GB | 1x real time |
| Arc A750 Intel | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 1x real time |
| GeForce RTX 5080 NVIDIA | 16 GB | Slow | BF16 everything resident | 7.55 GB | 1x real time |
| RTX 6000 Ada Generation NVIDIA | 48 GB | Slow | BF16 everything resident | 7.55 GB | 1x real time |
| GeForce RTX 3090 NVIDIA | 24 GB | Slow | BF16 everything resident | 7.55 GB | 0.9x real time |
| GeForce RTX 3080 Ti NVIDIA | 12 GB | Slow | BF16 everything resident | 7.55 GB | 0.9x real time |
| GeForce RTX 3080 12GB NVIDIA | 12 GB | Slow | BF16 everything resident | 7.55 GB | 0.9x real time |
| GeForce RTX 5070 Ti NVIDIA | 16 GB | Slow | BF16 everything resident | 7.55 GB | 0.9x real time |
| GeForce RTX 5090 Laptop NVIDIA | 24 GB | Slow | BF16 everything resident | 7.55 GB | 0.9x real time |
| GeForce RTX 5060 Ti 8GB NVIDIA | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 0.8x real time |
| GeForce RTX 5060 NVIDIA | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 0.8x real time |
| GeForce RTX 3070 NVIDIA | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 0.8x real time |
| GeForce RTX 3060 Ti NVIDIA | 8 GB | Slow | GGUF Q8_0 everything resident | 4.67 GB | 0.8x real time |
| Radeon RX 7900 XTX AMD | 24 GB | Slow | BF16 everything resident | 7.55 GB | 0.8x real time |
| Apple M3 Ultra 96GB Apple | 96 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M3 Ultra 256GB Apple | 256 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M3 Ultra 512GB Apple | 512 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| GeForce RTX 5080 Laptop NVIDIA | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.8x real time |
| RTX A6000 NVIDIA | 48 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.8x real time |
| RTX A5000 NVIDIA | 24 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.8x real time |
| Apple M1 Ultra 64GB Apple | 64 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M1 Ultra 128GB Apple | 128 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M2 Ultra 64GB Apple | 64 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M2 Ultra 128GB Apple | 128 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| Apple M2 Ultra 192GB Apple | 192 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.8x real time |
| GeForce RTX 3080 10GB NVIDIA | 10 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.8x real time |
| Radeon PRO W7900 AMD | 48 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.8x real time |
| GeForce RTX 4080 SUPER NVIDIA | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.7x real time |
| GeForce RTX 4080 NVIDIA | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.7x real time |
| Radeon RX 7900 XT AMD | 20 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.7x real time |
| GeForce RTX 5070 NVIDIA | 12 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.7x real time |
| GeForce RTX 4070 Ti SUPER NVIDIA | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.7x real time |
| GeForce GTX 1660 SUPER NVIDIA | 6 GB | Barely usable | GGUF Q8_0 everything resident | 4.67 GB | 0.6x real time |
| GeForce RTX 2080 Ti NVIDIA | 11 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.6x real time |
| GeForce RTX 4090 Laptop NVIDIA | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.6x real time |
| Radeon RX 9070 XT AMD | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.6x real time |
| Radeon RX 9070 AMD | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.6x real time |
| Arc A770 16GB Intel | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.6x real time |
| Radeon RX 7800 XT AMD | 16 GB | Barely usable | BF16 everything resident | 7.55 GB | 0.5x real time |
| GeForce RTX 4060 Ti 8GB NVIDIA | 8 GB | Barely usable | GGUF Q8_0 everything resident | 4.67 GB | 0.5x real time |
| Apple M4 Max 36GB Apple | 36 GB | Barely usable | BF16 everything resident | 7.05 GB | 0.5x real time |
Source: facebook/musicgen-large . Downloaded 0.1 million times in the last month. See how we calculate, or browse every audio model.