Best hardware for Qwen3.6 35B-A3B
Fast all-round assistant.
35B parameters · MoE (3B active per token, so it runs faster than its size suggests) · qwen family
Our picks for this model
Cheapest
$898 on eBay ↗ · ~251 tok/s est.
one card, no extra setup
Runs with a small context window (about 4k): fine for chat, tight for long documents. Step down a quant for more room.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 25.41 GB (file 22.29 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-08-02.
| Config | Total VRAM | Cheapest screened hw cost | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards
fits with a small context window (about 4k)
| 24 GB | $270
report
400 feedback (100%) · no returns · listed 5 days | 325 W | ~90.8 | 6.4 mo | Buying wins by month 7 |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | 48 GB | $550 2 screened listings: $270 + $280 + shipping unknown on 1 of 2 listings
report
400 feedback (100%) · no returns · listed 5 days | 575 W | ~90.8 | 16 mo | Buying wins by month 16 |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money
fits with a small context window (about 4k)
| 24 GB | $850
report
12,733 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 355 W | ~176 | 18 mo | Buying wins by month 18 |
| 1× AMD Radeon RX 7900 XTX
our pick
AMD card: more setup than NVIDIA for local AI
fits with a small context window (about 4k)
| 24 GB | $898
report
2,597 feedback (99.6%) · Top Rated seller · returns accepted · listed 3 days | 430 W | ~251 | 28 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | 96 GB | $1,140 4 screened listings: $270 + $280 + $290 + $300 + shipping unknown on 2 of 4 listings
report
400 feedback (100%) · no returns · listed 5 days | 1075 W | ~90.8 | 54 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090
fits with a small context window (about 4k)
| 24 GB | $1,180
report
164,757 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 425 W | ~245 | 32 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 Ti
fits with a small context window (about 4k)
| 24 GB | $1,350
report
5,792 feedback (99.8%) · Top Rated seller · returns accepted · listed 2 days | 525 W | ~264 | 45 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A5000
fits with a small context window (about 4k)
| 24 GB | $2,140
report
39,358 feedback (99.7%) · Top Rated seller · returns accepted · listed 23 days | 305 W | ~201 | 65 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | 48 GB | $2,420 2 screened listings: $1,180 + $1,240
report
164,757 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 775 W | ~245 | 89 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090
fits with a small context window (about 4k)
| 24 GB | $2,497
report
4,757 feedback (99.4%) · returns accepted · listed 23 days | 525 W | ~264 | 75 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 our pick | 48 GB | $3,520
report
164,767 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 375 W | ~201 | 85 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick | 32 GB | $4,220
report
17,037 feedback (99.9%) · Top Rated seller · returns accepted · listed 23 days | 650 W | ~469 | 148 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | 96 GB | $4,915 4 screened listings: $1,180 + $1,240 + $1,245 + $1,250
report
164,757 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 1475 W | ~245 | 409 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | 48 GB | $4,996 2 screened listings: $2,497 + $2,499
report
4,757 feedback (99.4%) · returns accepted · listed 23 days | 975 W | ~264 | 213 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | 96 GB | $7,140 2 screened listings: $3,520 + $3,620
report
164,767 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 675 W | ~201 | 218 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of median ask); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Qwen3.6 35B-A3B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 41.71 GB (file 37.81 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-08-02.
| Config | Total VRAM | Cheapest screened hw cost | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | 48 GB | $550 2 screened listings: $270 + $280 + shipping unknown on 1 of 2 listings
report
400 feedback (100%) · no returns · listed 5 days | 575 W | ~53.6 | 16 mo | Buying wins by month 16 |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | 96 GB | $1,140 4 screened listings: $270 + $280 + $290 + $300 + shipping unknown on 2 of 4 listings
report
400 feedback (100%) · no returns · listed 5 days | 1075 W | ~53.6 | 54 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | 48 GB | $2,420 2 screened listings: $1,180 + $1,240
report
164,757 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 775 W | ~144 | 89 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 | 48 GB | $3,520
report
164,767 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 375 W | ~118 | 85 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | 96 GB | $4,915 4 screened listings: $1,180 + $1,240 + $1,245 + $1,250
report
164,757 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 1475 W | ~144 | 409 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | 48 GB | $4,996 2 screened listings: $2,497 + $2,499
report
4,757 feedback (99.4%) · returns accepted · listed 23 days | 975 W | ~156 | 213 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | 96 GB | $7,140 2 screened listings: $3,520 + $3,620
report
164,767 feedback (99.6%) · Top Rated seller · returns accepted · listed 23 days | 675 W | ~118 | 218 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 2× NVIDIA RTX A5000 (48 GB) needs 2 screened cards and the market has 1; 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 1. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of median ask); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Qwen3.6 35B-A3B (Q8_0) in the buy-vs-rent calculator →