Best hardware for Devstral Small 2 24B
Coding assistant and code agents.
24B parameters · dense · mistral family
Our picks for this model
Cheapest
$320 on eBay ↗ · ~19.6 tok/s est.
one card, no extra setup
Runs with a small context window (about 4k): fine for chat, tight for long documents. Step down a quant for more room.
Intel card: runs well with Vulkan builds, but expect more setup than an NVIDIA card.
Rule-based picks at Q4_K_M, from live screened prices (how we estimate speed and choose picks). Every option is ranked below.
Q4_K_M
Minimum VRAM: 17.02 GB (file 14.3 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: Vast.ai $0.122/hr (rtx-3090 class, on-demand) · cloud prices checked 2026-09-09.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 1× Intel Arc A770 16GB
our pick · Cheapest Intel card: more setup than NVIDIA for local AI
fits with a small context window (about 4k)
| $320 | 16 GB | 300 W (well below this table's average) | ~19.6 | 32 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 our pick · Easiest | $1,500 | 24 GB | 425 W | ~32.7 | 242 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 our pick · Fastest | $5,500 | 32 GB | 650 W | ~62.7 (well above this table's average) | 2,641 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5060 Ti 16GB
fits with a small context window (about 4k)
| $706 | 16 GB | 255 W (well below this table's average) | ~15.7 (well below this table's average) | 77 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4070 Ti SUPER
fits with a small context window (about 4k)
| $800 | 16 GB | 360 W | ~23.5 | 96 mo | Renting wins at 40 h/mo |
| 1× AMD Radeon RX 7900 XTX AMD card: more setup than NVIDIA for local AI
| $945 | 24 GB | 430 W | ~33.6 | 138 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A4000
fits with a small context window (about 4k)
| $950 | 16 GB | 215 W (well below this table's average) | ~15.7 (well below this table's average) | 85 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080
fits with a small context window (about 4k)
| $1,000 | 16 GB | 395 W | ~25.1 | 93 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5070 Ti
fits with a small context window (about 4k)
| $1,100 | 16 GB | 375 W | ~31.3 | 102 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4080 SUPER
fits with a small context window (about 4k)
| $1,195 | 16 GB | 395 W | ~25.7 | 148 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5080
fits with a small context window (about 4k)
| $1,490 | 16 GB | 435 W | ~33.6 | 206 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 3090 Ti | $1,610 | 24 GB | 525 W | ~35.2 | 415 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A5000 | $2,200 | 24 GB | 305 W (well below this table's average) | ~26.9 | 196 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 4090 | $2,950 | 24 GB | 525 W | ~35.2 | 632 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W (well above this table's average) | ~32.7 | never | Cloud always cheaper at this power price |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W | ~26.9 | 399 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~26.9 | 990 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W (well above this table's average) | ~35.2 | never | Cloud always cheaper at this power price |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~32.7 | never | Cloud always cheaper at this power price |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~26.9 | 3,544 mo | Renting wins at 40 h/mo |
| 1× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $308 | 24 GB | 325 W | ~12.1 (well below this table's average) | 34 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~12.1 (well below this table's average) | 162 mo | Renting wins at 40 h/mo |
| 1× NVIDIA TITAN RTX Turing-era card: much slower than an RTX 3090 for similar money | $899 | 24 GB | 355 W | ~23.5 | 83 mo | Renting wins at 40 h/mo |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W (well above this table's average) | ~12.1 (well below this table's average) | never | Cloud always cheaper at this power price |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Devstral Small 2 24B (Q4_K_M) in the buy-vs-rent calculator →
Q8_0
Minimum VRAM: 28.36 GB (file 25.1 GB + context/runtime headroom at 8k ctx; see methodology).
Cheapest fitting cloud offer: RunPod $0.33/hr (rtx-a6000 class, on-demand) · cloud prices checked 2026-09-09.
| Config | Screened price | Total VRAM | Est. watts | Est. tokens/sec | Breakeven vs cloud | Verdict |
|---|---|---|---|---|---|---|
| 2× NVIDIA GeForce RTX 3090 | $3,000 | 48 GB | 775 W | ~18.6 | 127 mo | Renting wins at 40 h/mo |
| 1× NVIDIA RTX A6000 | $4,270 | 48 GB | 375 W (well below this table's average) | ~15.3 | 92 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A5000 | $4,699 | 48 GB | 535 W | ~15.3 | 148 mo | Renting wins at 40 h/mo |
| 1× NVIDIA GeForce RTX 5090 | $5,500 | 32 GB | 650 W | ~35.7 (well above this table's average) | 213 mo | Renting wins at 40 h/mo |
| 2× NVIDIA GeForce RTX 4090 | $5,950 | 48 GB | 975 W | ~20.1 | 285 mo | Renting wins at 40 h/mo |
| 4× NVIDIA GeForce RTX 3090 | $6,050 | 96 GB | 1475 W (well above this table's average) | ~18.6 | 572 mo | Renting wins at 40 h/mo |
| 2× NVIDIA RTX A6000 | $8,570 | 96 GB | 675 W | ~15.3 | 227 mo | Renting wins at 40 h/mo |
| 2× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $618 | 48 GB | 575 W | ~6.9 (well below this table's average) | 21 mo | Buying wins by month 21 |
| 4× NVIDIA Tesla P40 Passive datacenter card: needs a cooling mod and is far slower than modern cards | $1,256 | 96 GB | 1075 W | ~6.9 (well below this table's average) | 65 mo | Renting wins at 40 h/mo |
Fits, but not currently buyable as a set: 4× NVIDIA RTX A5000 (96 GB) needs 4 screened cards and the market has 2. We price a multi-card build from that many separate listings, so it stays unpriced until they exist.
- Our pick rows sit on top, then clean configs by price; cards with a caveat (cooling mod needed, or poor value for AI) group last, whatever their price. Watts and tokens/sec turn green or red when they sit far from this table's average (roughly 40% either way).
- Prices link straight to the screened eBay listing; hover for the seller summary. Seller details and the report option live on each card's own page.
- Configs without a screened purchase option are hidden until listings clear the trust screen (a multi-day survival window on eBay).
- Multi-card costs are the sum of that many separate screened listings; the cheapest listing counted twice is not a price anyone can pay. Prices exclude the PSU, board and case a multi-card build usually also needs.
- Tokens/sec are estimates, not benchmarks: spec-sheet memory bandwidth at 50% utilization divided by the model bytes read per token (how and why). Real speeds vary with the runtime and context length.
- Breakeven and verdict use default assumptions (40 h/mo, $0.16/kWh, 24-month horizon, resale at 65% of the going rate); adjust them in the calculator. Hardware links are screened, not guaranteed.
Run your own numbers for Devstral Small 2 24B (Q8_0) in the buy-vs-rent calculator →