Best hardware for Gemma 4 12B
Light assistant for mid-range GPUs.
12B parameters · dense · gemma family
Our picks for this model
Cheapest
$300 on eBay ↗ · ~29.8 tok/s est.
one card, no extra setup
Intel card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 10.05 GB (file 7.66 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.071/hr (rtx-a4000 class, on-demand) · cloud prices checked 2026-09-08.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× Intel Arc B580
our pick · Cheapest Intel card: more setup than NVIDIA for local AI
| $300 | 12 GB | 265 W (well below this table's average) | ~29.8 | 87 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3060 12GB our pick · Easiest | $305 | 12 GB | 245 W (well below this table's average) | ~23.5 (well below this table's average) | 67 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick · Fastest | $5,500 | 32 GB | 650 W | ~117 (well above this table's average) | never | Cloud always cheaper at this power price |
| 1× Intel Arc A770 16GB Intel card: more setup than NVIDIA for local AI
| $320 | 16 GB | 300 W (well below this table's average) | ~36.5 | 104 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3080 10GB
fits with a small context window (about 4k)
| $375 | 10 GB | 395 W | ~49.6 | 292 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3080 12GB | $435 + shipping unknown | 12 GB | 425 W | ~59.6 | 1,064 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5060 Ti 16GB | $706 | 16 GB | 255 W (well below this table's average) | ~29.2 (well below this table's average) | 209 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4070 Ti SUPER | $800 | 16 GB | 360 W | ~43.9 | 471 mo | Renting wins at 40 h/mo |
| 1× AMD Radeon RX 7900 XTX AMD card: more setup than NVIDIA for local AI
| $945 | 24 GB | 430 W | ~62.7 | 3,660 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A4000 | $950 | 16 GB | 215 W (well below this table's average) | ~29.2 (well below this table's average) | 206 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 | $1,000 | 16 GB | 395 W | ~46.8 | 722 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5070 Ti | $1,100 | 16 GB | 375 W | ~58.5 | 589 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 SUPER | $1,195 | 16 GB | 395 W | ~48.1 | 1,149 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5080 | $1,490 | 16 GB | 435 W | ~62.7 | 8,926 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 | $1,500 | 24 GB | 425 W | ~61.1 | 4,663 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 Ti | $1,610 | 24 GB | 525 W | ~65.8 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A5000 | $2,200 | 24 GB | 305 W | ~50.1 | 653 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090 | $2,950 | 24 GB | 525 W | ~65.8 | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W (well above this table's average) | ~61.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W | ~50.1 | 2,295 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~50.1 | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W (well above this table's average) | ~65.8 | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~61.1 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~50.1 | never | Cloud always cheaper at this power price |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $308 | 24 GB | 325 W | ~22.7 (well below this table's average) | 128 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~22.7 (well below this table's average) | never | Cloud always cheaper at this power price |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $899 | 24 GB | 355 W | ~43.9 | 386 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W (well above this table's average) | ~22.7 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Gemma 4 12B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 15.31 GB (file 12.67 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.071/hr (rtx-a4000 class, on-demand) · cloud prices checked 2026-09-08.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× Intel Arc A770 16GB Intel card: more setup than NVIDIA for local AI
| $320 | 16 GB | 300 W (well below this table's average) | ~22.1 | 104 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5060 Ti 16GB | $706 | 16 GB | 255 W (well below this table's average) | ~17.7 (well below this table's average) | 209 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4070 Ti SUPER | $800 | 16 GB | 360 W | ~26.5 | 471 mo | Renting wins at 40 h/mo |
| 1× AMD Radeon RX 7900 XTX AMD card: more setup than NVIDIA for local AI
| $945 | 24 GB | 430 W | ~37.9 | 3,660 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A4000 | $950 | 16 GB | 215 W (well below this table's average) | ~17.7 (well below this table's average) | 206 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 | $1,000 | 16 GB | 395 W | ~28.3 | 722 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5070 Ti | $1,100 | 16 GB | 375 W | ~35.4 | 589 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 SUPER | $1,195 | 16 GB | 395 W | ~29.1 | 1,149 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5080 | $1,490 | 16 GB | 435 W | ~37.9 | 8,926 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 | $1,500 | 24 GB | 425 W | ~36.9 | 4,663 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 Ti | $1,610 | 24 GB | 525 W | ~39.8 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A5000 | $2,200 | 24 GB | 305 W (well below this table's average) | ~30.3 | 653 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090 | $2,950 | 24 GB | 525 W | ~39.8 | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W (well above this table's average) | ~36.9 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W | ~30.3 | 2,295 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~30.3 | never | Cloud always cheaper at this power price |
| 1× NVIDIA GeForce RTX 5090 | $5,500 | 32 GB | 650 W | ~70.7 (well above this table's average) | never | Cloud always cheaper at this power price |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W (well above this table's average) | ~39.8 | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~36.9 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~30.3 | never | Cloud always cheaper at this power price |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $308 | 24 GB | 325 W | ~13.7 (well below this table's average) | 128 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~13.7 (well below this table's average) | never | Cloud always cheaper at this power price |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $899 | 24 GB | 355 W | ~26.5 | 386 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W (well above this table's average) | ~13.7 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Gemma 4 12B (Q8_0) in the buy-vs-rent calculator →