Economics

Buy vs Rent

Two break-even questions for any accelerator: how many GPU-hours until owning beats renting it by the hour, and — the one people actually ask — how many tokens of hosted OpenRouter inference until the card pays for itself. Pure capital break-even: power, cooling, networking, and support are deliberately excluded.

Tokens until this card is paid for

127.7Btokens

Buy H100 SXM5 ($30K) instead of renting Llama 3.3 70B Instruct on OpenRouter, and you break even after serving 127,659,574,468.085 tokens.

7.0 yr at 50.0M/day$0.235 / M tokensOpen-weight — self-hostable
Vs renting the card hourly

7.9K GPU-hrs

10.8 months at 24/7 vs a ~$3.80/hr cloud rate.

Purchase price

$30K

Illustrative street/list price for one H100 SXM5.

Throughput & memory fit — Llama 3.3 70B Instruct @ INT8

Weights

70 GB

of 80 GB

Fits one card?

No

needs ~2× or lower quant

Decode speed

~38 tok/s

single stream

Capacity

3.3M/day

1 stream, 24/7

At full single-stream utilization, one H100 SXM5 would take 105.7 yr to serve the 127.7B tokens needed to break even. Your 50.0M/day demand exceeds one card’s single-stream capacity — batching or extra cards raise aggregate throughput well beyond this floor. At INT8 the weights do not fit one card — drop to a lower precision or shard across cards.

Tokens to pay off vs Llama 3.3 70B Instruct — every accelerator

Fewest tokens = pays for itself fastest. Cheaper cards clear their price with far less usage.

AcceleratorBuy priceTokens to pay offAt 50.0M/day
$2501.1B21 days
$3301.4B28 days
$5002.1B43 days
$9994.3B2.8 mo
$1K5.5B3.7 mo
$1K6.0B4.0 mo
$3K10.6B7.1 mo
$4K17.0B11.3 mo
$5K21.3B1.2 yr
$6K23.4B1.3 yr
$7K28.9B1.6 yr
$7K29.8B1.6 yr
$8K34.0B1.9 yr
$9K38.3B2.1 yr
$10K40.9B2.2 yr
$10K42.6B2.3 yr
$11K46.8B2.6 yr
$11K48.0B2.6 yr
$12K49.0B2.7 yr
$12K51.1B2.8 yr
$15K63.8B3.5 yr
$17K72.0B3.9 yr
$18K76.6B4.2 yr
$20K85.1B4.7 yr
$25K106.4B5.8 yr
$25K106.4B5.8 yr
$30K127.7B7.0 yr
$32K136.2B7.5 yr
$34K144.7B7.9 yr
$35K147.1B8.1 yr
$36K153.2B8.4 yr
$38K161.7B8.9 yr
$40K170.2B9.3 yr
$85K361.7B19.8 yr
Capital break-even only. This is purchase price ÷ per-token rent — it excludes power, cooling, networking, hosting, and support, and assumes the card can actually serve the model at the throughput you need. OpenRouter prices are illustrative (they move often and vary by routed provider) — confirm current rates at openrouter.ai. Hourly rates for Intel Arc and Tenstorrent are nominal, since those are sold as cards, not cloud rentals.