Best hardware for GLM-4.7 Flash 30B-A3B
Quick chat and agent tool use.
30B parameters · MoE (3B active per token, so it runs faster than its size suggests) · glm family
Our picks for this model
Cheapest
$945 on eBay ↗ · ~262 tok/s est.
one card, no extra setup
AMD card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 21.22 GB (file 18.3 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.122/hr (rtx-3090 class, on-demand) · cloud prices checked 2026-09-09.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× AMD Radeon RX 7900 XTX
our pick · Cheapest AMD card: more setup than NVIDIA for local AI
| $945 | 24 GB | 430 W | ~262 | 138 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 our pick · Easiest | $1,500 | 24 GB | 425 W | ~256 | 242 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick · Fastest | $5,500 | 32 GB | 650 W | ~490 (well above this table's average) | 2,641 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 Ti | $1,610 | 24 GB | 525 W | ~275 | 415 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A5000 | $2,200 | 24 GB | 305 W (well below this table's average) | ~210 | 196 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090 | $2,950 | 24 GB | 525 W | ~275 | 632 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W | ~256 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W (well below this table's average) | ~210 | 399 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~210 | 990 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W (well above this table's average) | ~275 | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~256 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~210 | 3,544 mo | Renting wins at 40 h/mo |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $308 | 24 GB | 325 W (well below this table's average) | ~94.8 (well below this table's average) | 34 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~94.8 (well below this table's average) | 162 mo | Renting wins at 40 h/mo |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $899 | 24 GB | 355 W (well below this table's average) | ~184 | 83 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W (well above this table's average) | ~94.8 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for GLM-4.7 Flash 30B-A3B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 35.39 GB (file 31.8 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-09-09.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W | ~147 | 127 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W (well below this table's average) | ~121 | 92 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~121 | 148 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W | ~158 | 285 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~147 | 572 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~121 | 227 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~54.6 (well below this table's average) | 21 mo | Buying wins by month 21 |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W | ~54.6 (well below this table's average) | 65 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for GLM-4.7 Flash 30B-A3B (Q8_0) in the buy-vs-rent calculator →