Can You Run Gemma 4 on an RTX 5090 (32GB VRAM)?
Yes — best fit is Gemma 4 31B at Q4_0 (4-bit) (~17.4GB) ✅ Comfortable, using 31GB of usable memory on a 32GB RTX 5090.
Every Gemma 4 model on this configuration
| Model | Best quant | Size | Status |
|---|---|---|---|
| Gemma 4 E2B | BF16 (16-bit) | ~8 GB | ✅ Comfortable |
| Gemma 4 E4B | BF16 (16-bit) | ~15 GB | ✅ Comfortable |
| Gemma 4 12B | SFP8 (8-bit) | ~13.4 GB | ✅ Comfortable |
| Gemma 4 26B A4B | Q4_0 (4-bit) | ~15.6 GB | ✅ Comfortable |
| Gemma 4 31B | Q4_0 (4-bit) | ~17.4 GB | ✅ Comfortable |
How we got this verdict
32 - 1 = 31GB usable VRAM (the driver and OS reserve roughly 1 GB).
31B lands on Q4_0 here, so grab the official QAT (quantization-aware training) checkpoint instead of a generic Q4_0 quant. Google shipped a full-quality QAT build for 31B at essentially the same ~19GB download size, with noticeably better output quality than a naive 4-bit quant. See the QAT guide for exact file names.
The fastest lever on a fixed budget isn't more VRAM on the same card — it's switching cards. Compare this tier against neighboring cards below, or check the full checker for any GPU model.
Try a different configuration
Best fit: Gemma 4 31B · Q4_0 (4-bit) · ~17.4 GB ✅ Comfortable
FAQ
Can a RTX 5090 (32GB VRAM) run the 31B flagship model?
Yes, comfortably — Gemma 4 31B fits at Q4_0 (4-bit) (~17.4GB), using 31GB of usable memory on this configuration.
What's the best Gemma 4 model for a RTX 5090 (32GB VRAM)?
Gemma 4 31B at Q4_0 (4-bit) (~17.4GB) is the best fit — it's the largest model that runs comfortably within the 31GB of usable memory here.
What if I have more or less memory than 32GB VRAM?
This is the top configuration in this family — every Gemma 4 model already fits at its best available quant. With less (e.g. RTX 4090), the best-fitting model gets smaller or tighter — see the RTX 4090 page.