RigPrice.

Best hardware for Gemma 4 26B-A4B

Fast everyday assistant.

25.2B parameters · MoE (4B active per token, so it runs faster than its size suggests) · gemma family

Our picks for this model

Cheapest

1× AMD Radeon RX 7900 XTX

$945 on eBay ↗ · ~179 tok/s est.

one card, no extra setup

AMD card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.

Easiest

1× NVIDIA GeForce RTX 3090

$1,500 on eBay ↗ · ~174 tok/s est.

one card, no extra setup

Fastest

1× NVIDIA GeForce RTX 5090

$5,500 on eBay ↗ · ~334 tok/s est.

one card, no extra setup

Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.

Q4_K_M

Minimum VRAM: 19.75 GB (file 16.9 GB + context/runtime headroom at 8k ctx; see methodology).

Cheapest fitting cloud offer: Vast.ai $0.122/hr (rtx-3090 class, on-demand) · cloud prices checked 2026-09-09.

ConfigScreened priceTotal VRAMEst. wattsEst. tokens/secBreakeven vs cloudVerdict
1× AMD Radeon RX 7900 XTX our pick · Cheapest
AMD card: more setup than NVIDIA for local AI
$94524 GB430 W~179138 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 3090 our pick · Easiest$1,50024 GB425 W~174242 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5090 our pick · Fastest$5,50032 GB650 W~334 (well above this table's average)2,641 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 3090 Ti$1,61024 GB525 W~188415 moRenting wins at 40 h/mo
1× NVIDIA RTX A5000$2,20024 GB305 W (well below this table's average)~143196 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 4090$2,95024 GB525 W~188632 moRenting wins at 40 h/mo
2× NVIDIA GeForce RTX 3090$3,00048 GB775 W~174neverCloud always cheaper at this power price
1× NVIDIA RTX A6000$4,27048 GB375 W (well below this table's average)~143399 moRenting wins at 40 h/mo
2× NVIDIA RTX A5000$4,69948 GB535 W~143990 moRenting wins at 40 h/mo
2× NVIDIA GeForce RTX 4090$5,95048 GB975 W (well above this table's average)~188neverCloud always cheaper at this power price
4× NVIDIA GeForce RTX 3090$6,05096 GB1475 W (well above this table's average)~174neverCloud always cheaper at this power price
2× NVIDIA RTX A6000$8,57096 GB675 W~1433,544 moRenting wins at 40 h/mo
1× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$30824 GB325 W (well below this table's average)~64.7 (well below this table's average)34 moRenting wins at 40 h/mo
2× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$61848 GB575 W~64.7 (well below this table's average)162 moRenting wins at 40 h/mo
1× NVIDIA TITAN RTX
Turing-era card: much slower than an RTX 3090 for similar money
$89924 GB355 W (well below this table's average)~12583 moRenting wins at 40 h/mo
4× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$1,25696 GB1075 W (well above this table's average)~64.7 (well below this table's average)neverCloud always cheaper at this power price

Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.

  • Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
  • Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
  • Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
  • Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
  • Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
  • Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.

Run your own numbers for Gemma 4 26B-A4B (Q4_K_M) in the buy-vs-rent calculator →

Q8_0

Minimum VRAM: 30.25 GB (file 26.9 GB + context/runtime headroom at 8k ctx; see methodology).

Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-09-09.

ConfigScreened priceTotal VRAMEst. wattsEst. tokens/secBreakeven vs cloudVerdict
2× NVIDIA GeForce RTX 3090$3,00048 GB775 W~110127 moRenting wins at 40 h/mo
1× NVIDIA RTX A6000$4,27048 GB375 W (well below this table's average)~89.992 moRenting wins at 40 h/mo
2× NVIDIA RTX A5000$4,69948 GB535 W~89.9148 moRenting wins at 40 h/mo
1× NVIDIA GeForce RTX 5090$5,50032 GB650 W~210 (well above this table's average)213 moRenting wins at 40 h/mo
2× NVIDIA GeForce RTX 4090$5,95048 GB975 W~118285 moRenting wins at 40 h/mo
4× NVIDIA GeForce RTX 3090$6,05096 GB1475 W (well above this table's average)~110572 moRenting wins at 40 h/mo
2× NVIDIA RTX A6000$8,57096 GB675 W~89.9227 moRenting wins at 40 h/mo
2× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$61848 GB575 W~40.6 (well below this table's average)21 moBuying wins by month 21
4× NVIDIA Tesla P40
Passive datacenter card: needs a cooling mod and is far slower than modern cards
$1,25696 GB1075 W~40.6 (well below this table's average)65 moRenting wins at 40 h/mo

Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.

  • Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
  • Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
  • Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
  • Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
  • Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
  • Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.

Run your own numbers for Gemma 4 26B-A4B (Q8_0) in the buy-vs-rent calculator →