Kinetic Alpha Compute Cost Navigator
companion to "Every Token Has an Address" · prices as of 2026-09-26 to 28
reference tool · not a live feed
H100 rental index
$2.54 – $2.70
Ornn OCPI · Silicon Data, Sep 26–27
Owner cost, H100-hour
$1.07 / $1.53
full utilisation / 70%
Cheapest million output tokens
$0.092
owned B200, DeepSeek R1, measured
Dearest million output tokens
$50.00
Claude Fable 5.1 list · 543× the cheapest
Largest lever
34×
workload shape by architecture · buyer-side

Build an address: one code from each of seven layers

taxonomy · v1 · 2026-09-28

The seven layers

taxonomy
#LayerUnit of priceSet by
7Commercial terms CTdiscount or premium to listcontract
6Delivery path DP$ per seat or per taskcontract; largest rent
5Workload shape WStokens per taskarchitecture; buyer
4Task type TT$ per run or per taskbinding constraint
price is set below this line, quantity above it
3Model class MC$ per million tokensvendor power or host competition
2Capacity form CF$ per chip-hourmarket; indexed, hedgeable
1Silicon SI$ per chip; wattsmarket; indexed, hedgeable
Modifiers ride with the workload shape: latency LAT-INT / LAT-BAT; input-to-output ratio IO-IN (above 10 to 1) / IO-BAL / IO-OUT; context bracket CTX-S (under 32k) / CTX-M (32k to 200k) / CTX-L (above 200k, where some vendors charge more).

How to read an address

  1. Classify the workload with one code from each layer.
  2. Read its cost floor from the price snapshot: the cheapest silicon and setup that can run that model class at that latency.
  3. Read its rent from the end-user map: what the buyer type actually pays and on what basis.
  4. Look up the levers for what closes the gap, ranked by size.
Cost means the resources consumed to produce the output. Rent means what is charged above that by whoever owns the layer. All prices are United States list, September 26 to 28, 2026.

Chip-hour prices, $ per chip-hour

vendor pages · read 2026-09-26 to 28

Token prices, $ per million tokens

vendor pages · read 2026-09-26 to 28
Batch is 50 percent off input and output at OpenAI, Anthropic and Google. Cache writes cost 25 percent (5-minute) to 100 percent (1-hour) over input list at Anthropic. Anthropic charges one rate across its full 1 million token context; OpenAI bills about 2 times input above 272,000 tokens and Google 2 times input above 200,000.

Measured serving throughput and cost

SemiAnalysis InferenceX · Nvidia

Other inputs

cited pages

Step 1 · From chip to chip-hour: an owner's cost against the market

model · taxonomy inputs
Owner's inputs
Utilisation
Cost at full utilisation
Cost at chosen utilisation
Rent share at the index ($2.62 mid)
Rent share at a hyperscaler ($11.68 mid)

Step 2 · From chip-hour to token

model
Tokens an hour
per chip, at utilisation
Cost per million tokens
chip-hour price ÷ tokens an hour
Hosted DeepSeek V4 Pro, blended 4:1
$1.85
$1.32 in / $3.96 out, Together
Hosted price over this cost
the host's gap: software, utilisation, redundancy, margin
cost per million tokens = price per chip-hour ÷ (tokens per second per chip × 3,600 × utilisation) × 106. The taxonomy's reference: $0.24 rented, about $0.10 owned at 70 percent, against $1.85 hosted (8 to 19 times).

Self-host break-even

model
Node cost a day
chips × rate × 24
Break-even utilisation
below this, hosted API is cheaper
Break-even tokens a day
the volume that justifies the node
Cost per M at full use
floor for this node
The taxonomy's reference: a rented eight-chip H200 node at $4.50 ($864 a day) beats hosted open-weight prices above about 13 percent utilisation, roughly 470 million tokens a day; on owned chips ($1.80) about 5 percent, or 190 million a day. Below those volumes, or for any frontier model, the choice is terms, not a chip.

Steps 3 and 4 · The token ladder: nine ways to buy a million output tokens

InferenceX · vendor list pages
Self-hosted cost (owned or rented chips)Hosted open-weight listFrontier list
Frontier vendors do not publish cost, so their rent is bounded by open-weight hosting prices, not measured: flagship output at $20 to $30 and the top tier at $50 is 5 to 13 times a hosted open-weight large model. A seat converts the token bill into a flat $20 a month with a usage cap; no vendor publishes tokens per seat, so the seat rent cannot be measured from outside.

End-user map: seventeen buyer types by dominant use case

taxonomy
Three patterns. Only three rows own the chip; every other buyer pays rent to at least two layers above it. The rows with the thinnest margins, the application companies, run the most multiplicative workloads, so their cost of goods is the least forecastable. The enterprise adopter is the only buyer paying rent at all four upper layers at once, which is why its total cost of AI is three to four times its token bill.

Most economic chip and setup, by workload class

taxonomy
Chips are chosen for training, batch and small-model work; terms are chosen for everything else. For every frontier workload the lever is layer seven, commercial terms. For every open-weight workload it is utilisation. Three rules of thumb: a newer generation is cheaper per token even at a higher hourly price (B200 is 2.2 times cheaper per token than H100 at index prices); a rented eight-chip H200 node beats hosted open-weight prices above about 13 percent utilisation; below that, or for any frontier model, the choice is terms.

Twelve levers, ranked by size

taxonomy · September 2026 prices
Buyer-sideBelongs to the chip ownerMoves variance, not the mean
Size is the multiple between worst and best case on that lever alone; bar length is on a log scale to 40×. The levers compound. Levers 1 to 5 and 8 to 9 need nothing from the market except measurement. Levers 6, 7, 10 and 11 pass through to buyers only where hosts compete. Lever 12 is the one no one yet provides at scale.

Pricer A · Enterprise bill ledger

model · illustrative inputsMC-FR-M | TT-INF-BA + TT-INF-ON | WS-T1 → WS-TK | DP-CLOUD | CT-COMMIT + CT-CACHE + CT-BATCH
A regulated enterprise budgeted a linear workload on a frontier mid model through a cloud commitment. The consumption multiplier is the shape drift; the floors are the same tokens on a hosted open-weight mid model and self-hosted.
The workload as budgeted
What actually happened
The terms
Token bill as budgeted
Token bill as it stands
Against the commitment
Seat bill
Same tokens, hosted open-weight mid
Same tokens, self-hosted
Rent multiple, bill over owned floor
Bill after termsBudgetCommitment, monthly share

Pricer B · Seat margin ledger

model · illustrative inputsMC-FR-M | TT-AGT | WS-TA | IO-IN | CTX-L | DP-API | CT-COMMIT + CT-CACHE
An application company sells $20 seats and buys frontier tokens for an agent loop (a fifty-turn session: about 1.0M input, 40k output tokens). Cache, routing and commitment are the three buyer-side levers; the heavy user is the tail the usage cap exists to stop.
The product
The model
The levers
Token cost per session
Token cost per mean seat
Gross margin per mean seat
Heavy seat
No routing25% routed to small tier50% routedBreak-even

Pricer C · Dealer's per-task ledger

model · stylisedMC-FR-M + MC-OW-L | TT-INF-BA | WS-TK | DP-GW | CT-PT → fixed $ per task
A dealer quotes a fixed price per completed task (lever 12). The cost splits into a chip leg (self-hosted share; hedgeable on a listed chip-hour index), a list leg (frontier share; vendor resets, no index) and a quantity leg (tokens per task; architecture and a contractual cap). Stylised: one cost per token on the self-hosted leg, a uniform shift in tokens per task, an index tracking the H200 rate one for one.
The task class
The routing
The contract
The stress
Expected cost per task
Fixed price offered
Hedgeable share of cost
Monthly revenue
Stressed cost, unhedged, no cap
With the index hedge
With hedge and cap
Monthly P&L under stress, hedge and cap
Unhedged, no capIndex hedge on the chip legHedge and capZero