Can You Run Gemma 4 on a CPU-Only PC (64GB RAM)?
Yes — best fit is Gemma 4 31B at Q4_0 (4-bit) (~17.4GB) ✅ Comfortable, using 32GB of usable memory on a 64GB CPU-Only PC.
Every Gemma 4 model on this configuration
| Model | Best quant | Size | Status |
|---|---|---|---|
| Gemma 4 E2B | BF16 (16-bit) | ~8 GB | ✅ Comfortable |
| Gemma 4 E4B | BF16 (16-bit) | ~15 GB | ✅ Comfortable |
| Gemma 4 12B | SFP8 (8-bit) | ~13.4 GB | ✅ Comfortable |
| Gemma 4 26B A4B | SFP8 (8-bit) | ~25 GB | ✅ Comfortable |
| Gemma 4 31B | Q4_0 (4-bit) | ~17.4 GB | ✅ Comfortable |
How we got this verdict
64 ÷ 2 = 32GB usable (the RAM-÷-2 rule: one copy of the weights, one copy for working memory).
31B lands on Q4_0 here, so grab the official QAT (quantization-aware training) checkpoint instead of a generic Q4_0 quant. Google shipped a full-quality QAT build for 31B at essentially the same ~19GB download size, with noticeably better output quality than a naive 4-bit quant. See the QAT guide for exact file names.
Going up to 128GB (128GB RAM) doesn't change the top recommendation — 31B is still the best fit, just with more headroom. One tier down, at 32GB RAM, the best fit drops to 12B (comfortable) instead.
Adding a discrete NVIDIA GPU is usually a bigger lever than more system RAM — it unlocks GPU offload and much faster tokens/sec, on top of whatever the RAM ÷ 2 rule already gives you here. See the GPU pages if a card is on the table.
Real-world notes for this configuration
64GB system RAM configurations are commonly sold as matched 2x32GB kits for dual-channel operation, though some budget boards and prebuilts ship 4x16GB instead — both hit the same total capacity, but populating all four DIMM slots on many consumer motherboards forces a lower maximum memory clock speed than a 2-DIMM kit, directly affecting CPU inference bandwidth.
- ⚠Check whether your 64GB is 2x32GB or 4x16GB — many consumer motherboards derate the maximum supported memory frequency when all four DIMM slots are populated, which matters for CPU-bound inference speed even at identical total capacity.
Try a different configuration
Best fit: Gemma 4 31B · Q4_0 (4-bit) · ~17.4 GB ✅ Comfortable
FAQ
Can a CPU-Only PC (64GB RAM) run the 31B flagship model?
Yes, comfortably — Gemma 4 31B fits at Q4_0 (4-bit) (~17.4GB), using 32GB of usable memory on this configuration.
What's the best Gemma 4 model for a CPU-Only PC (64GB RAM)?
Gemma 4 31B at Q4_0 (4-bit) (~17.4GB) is the best fit — it's the largest model that runs comfortably within the 32GB of usable memory here.
What if I have more or less memory than 64GB RAM?
With more (e.g. 128GB), you can run larger models or the same model more comfortably — see the CPU-Only PC page. With less (e.g. 32GB), the best-fitting model gets smaller or tighter — see the CPU-Only PC page.
Is 2x32GB or 4x16GB better for CPU-only Gemma 4 at 64GB total?
2x32GB is generally preferable on consumer motherboards — populating all four DIMM slots (4x16GB) commonly forces a lower maximum memory clock on mainstream platforms, which can measurably slow CPU-bound inference despite identical total capacity.
At 64GB RAM, is CPU-only viable for the larger Gemma 4 models?
See the model table above for the exact verdict — capacity-wise this tier reaches well into the lineup by the RAM-÷-2 rule, though this site's guide still puts the larger models under 1 token/sec on CPU, better suited to batch or non-interactive use than live chat.
Related
Gemma 4 CPU-Only Setup
The full setup guide for this device family.
Can You Run Gemma 4 on a CPU-Only PC (32GB RAM)?
CPU-Only PC — 32GB RAM
Can You Run Gemma 4 on a CPU-Only PC (128GB RAM)?
CPU-Only PC — 128GB RAM
Gemma 4 Hardware Requirements
The full memory table this checker is built on.
Full Hardware Checker
Every device type, every model, one tool.