Can You Run Gemma 4 on an RTX 3080 Ti (12GB VRAM)?
Yes — best fit is Gemma 4 12B at Q4_0 (4-bit) (~6.7GB) ✅ Comfortable, using 11GB of usable memory on a 12GB RTX 3080 Ti.
Every Gemma 4 model on this configuration
| Model | Best quant | Size | Status |
|---|---|---|---|
| Gemma 4 E2B | BF16 (16-bit) | ~8 GB | ✅ Comfortable |
| Gemma 4 E4B | SFP8 (8-bit) | ~7.5 GB | ✅ Comfortable |
| Gemma 4 12B | Q4_0 (4-bit) | ~6.7 GB | ✅ Comfortable |
| Gemma 4 26B A4B | — | — | ❌ Won't fit |
| Gemma 4 31B | — | — | ❌ Won't fit |
How we got this verdict
12 - 1 = 11GB usable VRAM (the driver and OS reserve roughly 1 GB).
12B lands on Q4_0 here, so grab the official QAT (quantization-aware training) checkpoint instead of a generic Q4_0 quant. Google shipped a full-quality QAT build for 12B at essentially the same ~7.2GB download size, with noticeably better output quality than a naive 4-bit quant. See the QAT guide for exact file names.
Switching to the RTX 4070 (12GB VRAM) doesn't change the top recommendation — 12B is still the best fit, just with more headroom.
The fastest lever on a fixed budget isn't more VRAM on the same card — it's switching cards. Compare this tier against neighboring cards below, or check the full checker for any GPU model.
Real-world notes for this configuration
The RTX 3080 Ti standardizes at 12GB of GDDR6X on a 384-bit bus — unlike the plain 3080, which shipped in two different VRAM configurations under one name, the Ti variant never had that ambiguity. It sits just below the 3090's 24GB, using the same GA102 die family with fewer cores enabled and less memory.
- ⚠Very high power draw (350W board power) — worth checking PSU headroom, especially if this is going into a small always-on local-AI box.
- ⚠Unlike the plain 3080, there's no VRAM-variant confusion to check for when buying used — 12GB is the only configuration this card was ever sold in.
Try a different configuration
Best fit: Gemma 4 12B · Q4_0 (4-bit) · ~6.7 GB ✅ Comfortable
FAQ
Can a RTX 3080 Ti (12GB VRAM) run the 31B flagship model?
No — the 31B model needs at least 17.4GB usable memory even at Q4_0, and this configuration only has 11GB usable. Gemma 4 12B is the largest model that fits here.
What's the best Gemma 4 model for a RTX 3080 Ti (12GB VRAM)?
Gemma 4 12B at Q4_0 (4-bit) (~6.7GB) is the best fit — it's the largest model that runs comfortably within the 11GB of usable memory here.
What if I have more or less memory than 12GB VRAM?
With more (e.g. RTX 4070), you can run larger models or the same model more comfortably — see the RTX 4070 page. With less (e.g. RTX 3060), the best-fitting model gets smaller or tighter — see the RTX 3060 page.
Is the RTX 3080 Ti's VRAM ever ambiguous like the plain 3080's?
No — the 3080 Ti was only ever sold with 12GB of GDDR6X memory, unlike the plain 3080, which exists in both 10GB and 12GB variants under the same name.
How much more VRAM does the RTX 3090 have over the 3080 Ti?
Double — the 3090 has 24GB versus the 3080 Ti's 12GB, despite both using the same GA102 die family. That extra headroom is the main reason the 3090 remains a popular used pick specifically for local AI.
Related
Gemma 4 on RTX 3070/3080
The full setup guide for this device family.
Can You Run Gemma 4 on an RTX 3060 (12GB VRAM)?
NVIDIA GPU — 12GB VRAM
Can You Run Gemma 4 on an RTX 4070 (12GB VRAM)?
NVIDIA GPU — 12GB VRAM
Gemma 4 Hardware Requirements
The full memory table this checker is built on.
Full Hardware Checker
Every device type, every model, one tool.