- assumptions.resale_fraction
- 0.65
- assumptions.kwh_price
- 0.16
- assumptions.hours_per_month
- 40
- assumptions.horizon_months
- 24
- assumptions.power_model
- power_per_hr = est_system_watts / 1000 * kwh_price
- assumptions.excluded_costs
- cloud storage/egress excluded
- assumptions.hw_cost_basis
- sum of the gpu_count cheapest active buy_pick listings (used bucket); a config whose GPU has fewer screened listings than it needs is left unpriced
- assumptions.resale_basis
- resale_fraction x current median_ask x gpu_count
- assumptions.bandwidth_utilization
- 0.5
- assumptions.speed_model
- est_toks_per_s = mem_bandwidth_gbps x bandwidth_utilization / (file_gb x active_params_b / params_b); per-card bandwidth (layer-split decode is sequential), spec-sheet numbers, estimate not benchmark
- assumptions.recommend_rule
- three personas from the first priced quant, caveat/cooling-mod cards never win: cheapest = lowest cost clean config (tight fits and multi-GPU allowed); easiest = cheapest comfortable single clean card at >= 10 est tok/s, NVIDIA preferred, with fastest/multi-GPU fallbacks; fastest = highest est tok/s among comfortable fits, ties to cheaper
- assumptions.tight_fit_basis
- min_vram_tight_gb = file_gb + ~1.5 GB (about 4k context with a quantized KV cache); configs between the tight and comfortable bars publish as tight_fit