KINETIC ALPHA
Research · Energy & Compute
Compute · Cost Structure · Market Structure

A Taxonomy of Compute Cost. The Bill Is Seven Rents Deep.

One physical cost, electricity through a chip, is resold at every layer on its way to the person who pays for an AI output. A seven-part address for each unit of compute makes the resale chain readable: what the output should cost at the bottom, what it is charged at the top, and where the gap opens. Three worked examples put the address to work for a finance team, an application company and a dealer.

Companion tool: the Compute Cost Navigator carries the full price snapshot, all seventeen buyer types, eleven workload classes, the twelve levers, the three pricers and a self-host break-even calculator, with every dial exposed. Builds on The Reserved Contract, Unbundled and Refiner and Merchant.
Layers in the address, each priced in its own unit
7
Silicon, capacity form, model class, task type, workload shape, delivery path, commercial terms.
An owner’s all-in cost of one H100-hour, run flat out
$1.07
Against $2.54 to $2.70 on the rental indices and $11 to $12 on demand at a hyperscaler.
A million output tokens, owned Blackwell versus the top frontier tier
9¢ → $50
Measured serving cost at one end, list price at the other. About 543 times.
The largest single lever, and it belongs to the buyer
34×
The same task at four architectures: $3.04 to $0.09. Layer five, workload shape.

Ask what a unit of AI compute costs and you get an answer in the wrong unit. The chip is priced by the hour. The model is priced by the million tokens. The product is priced by the seat. Nobody who pays one of those prices can see the other two, and the gap between them is where the money is.

This piece sets out a taxonomy we built to read that gap. It gives every unit of compute a seven-part address, one code from each of seven layers, running from the silicon that did the work up to the commercial terms it was sold on. Each layer has its own unit of price. Each layer inherits the full cost of the layer beneath it and adds a margin of its own. Once a workload has an address, three numbers become readable: what it should cost at the lowest layer, what it is being charged at the highest, and which layer in between is taking the rent.

The taxonomy was built for our own use, to understand the cost structure of a market that has indices at the bottom, list prices in the middle and flat fees at the top, with nothing joining them. We are publishing it because the same problem faces anyone who has to budget for, price, sell or hedge AI consumption. Sections 01 to 03 set out the address and the margin stack it exposes. Section 04 works three examples end to end: a finance team reconciling a bill, an application company pricing a seat, and a dealer pricing a fixed-price contract. Every list price is a United States list price read between September 26 and 28, 2026, and every measured cost carries its date, from the pages in the register at the end. Five of the figures are live, and the companion Navigator carries the rest.

01 · The address

Seven layers, one code from each. Price is decided below the line; quantity above it.

The address answers seven questions in order. What silicon ran it. How that silicon was held. Which model ran. What kind of task it was. How consumption grows with use. Who delivered it. On what terms it was sold.

The layers build on the Tokenomics Foundation’s five-layer stack (silicon, capacity, serving software, model, routing), which sets where price is decided, and on its consumption classes, which set how quantity scales; that keeps a classification lined up with the tools buyers already use. We add two layers the foundation leaves out, delivery path and commercial terms, because that is where most of the rent sits. The dashed line in Figure 1 marks the foundation’s own split: the price of a token is decided below it, and the number of tokens is decided above it.

Figure 1 · The seven layers, what each names, and its unit of price
Taxonomy
Read from the bottom up. Every layer inherits the cost of the one beneath and adds its own margin.
7
Commercial terms · CT
List, cached input, batch, long-context premium, provisioned throughput, committed spend, seat subscription, prepaid credits, spot.
Discount or premium to listSet by contract
6
Delivery path · DP
Direct from the model vendor, cloud-brokered, gateway or router, application seat, self-hosted.
$ per seat or per taskThe largest and least visible rent
5
Workload shape · WS
How consumption grows with use: constant, linear, multiplicative, agent-multiplicative, unbounded. Modifiers for latency, input-to-output ratio and context length.
Tokens per taskPartly under the buyer’s control
4
Task type · TT
Pretraining, post-training, evaluation, online inference, batch inference, retrieval, agent loop, media generation.
$ per run or per taskSets the binding constraint
price is set below this line, quantity above it
3
Model class · MC
Frontier flagship, mid and small; open-weight large, mid and small; fine-tuned derivatives; task models for embedding, speech, image and video.
$ per million tokensVendor pricing power or host competition
2
Capacity form · CF
Owned and hosted, reserved or committed, on-demand, spot or marketplace, serverless per token, bundled in a seat.
$ per chip-hourIndexed and hedgeable
1
Silicon · SI
Accelerator family and generation (Blackwell, Hopper, Ampere), inference cards, consumer cards, AMD Instinct, Google TPU, inference chips, CPU.
$ per chip; wattsIndexed and hedgeable
Blue layers are the compute complex, where prices are already indexed and hedgeable. Purple layers are where a buyer has partial control over quantity. Orange layers are where the largest margins sit, because their prices are set by contract rather than by a market. The address builds on the Tokenomics Foundation’s five-layer stack and consumption classes; the two upper layers, and the seven-part address itself, are ours.

An address is written as one code per layer, separated by bars. Figure 2 builds one. Pick a code from each layer and it returns the address, the unit each layer is priced in, and the price range the September snapshot gives for the layers that carry a price. The presets are the buyers this piece keeps returning to.

Figure 2 · Address builder
Live
One code from each layer. Prices are United States list, September 26 to 28, 2026, from the register.

Three conventions hold for the rest of the piece. Cost means the resources consumed to produce the output. Rent means what is charged above that cost by whoever owns the layer. And a price without a layer is not a price: “$2.70 an hour” is layer two, “$10 a million tokens” is layer three, “$20 a month” is layer six, and none of them can be compared until the same output has been given the same address.

02 · The margin stack

One kilowatt through one chip, resold at every layer. The rent grows on the way up.

Electricity is the only input that is physically consumed, and it is about 4 percent of the price of a rented chip-hour. Everything above it is capital, hosting, software, capability and contract, and each of those is somebody’s margin.

Figure 3 puts nine ways of buying the same million output tokens on one scale. The three self-hosted rows are measured throughput from SemiAnalysis InferenceX priced at an owner’s cost of about $1.20 a chip-hour, or at the $4.50 on-demand H200 rate for the rented row. The rest are list prices. Input tokens, caching and batch change the blend; they do not change the order.

Figure 3 · The token ladder: a million output tokens costs 9 cents on owned Blackwell and $50 at the top frontier tier
Snapshot
Cost or list price per million output tokens, US dollars, log scale, September 2026. Hover a bar.
Self-hosted cost (owned or rented chips)Hosted open-weight listFrontier list
Self-hosted rows: DeepSeek R1 at 45 tokens per second per user on B200 and H100, DeepSeek V4 Pro at 100 tokens per second per user on H200 (SemiAnalysis InferenceX), priced at owner cost or the Nebius on-demand H200 rate. Hosted and frontier rows are list output prices from the vendor pages in the register.

Step 1 · From chip to chip-hour

An owner’s all-in cost for one H100-hour is three lines: capital, electricity and hosting. Figure 4 builds it from the September inputs and sets it against the prices the same hour fetches on the market. At full utilisation the owner’s cost is about $1.07. At 70 percent, because idle hours cost the same, it is $1.53. The rental indices print $2.54 to $2.70, a neocloud charges $3.85 to $4.25 on demand, and a hyperscaler $11 to $12. The rent share of price is therefore about 40 to 60 percent at the index, 60 to 75 percent at a neocloud, and 85 to 90 percent at a hyperscaler. Spot at $0.79 sits below the owner’s own cost, which is what unsold capacity looks like.

Figure 4 · Chip-hour cost build: what an H100-hour costs its owner, and what it sells for
Live
Move the inputs. Market prices on the right are fixed at the September snapshot.
Owner’s inputs
Utilisation
Cost at full utilisation
Cost at chosen utilisation
Rent share at the index ($2.62 mid)
Rent share at a hyperscaler ($11.68 mid)

Step 2 · From chip-hour to token

An H200 serving DeepSeek V4 Pro produces about 5,100 tokens a second at 100 tokens a second per user, or 18.5 million tokens an hour. At the $4.50 on-demand rate that is $0.24 a million tokens; at an owner’s cost of about $1.80 a productive hour (70 percent utilisation) it is about $0.10. The same model is sold hosted at $1.32 a million input tokens and $3.96 a million output. On a four-to-one input-to-output blend the hosted price is about $1.85, roughly 8 times the rented cost and 19 times the owned cost. Serving software, utilisation, redundancy and the host’s margin sit in that gap.

Step 3 · From open-weight token to frontier token

Frontier vendors do not publish cost, so their rent can only be bounded, not measured. Frontier flagship output lists at $20 to $30 a million tokens, and the top tier at $50, which is 5 to 13 times a hosted open-weight large model. Frontier small tiers list at $1.20 to $5.00, on top of and sometimes below open-weight mid models, which is where price competition is real. The frontier premium is a price for a capability with no substitute. It is the one layer where a buyer has no cost-side lever, only routing and terms.

Step 4 · From token to seat

A seat converts a variable token bill into a flat fee: $20 a month for a prosumer plan at OpenAI, Anthropic and Cursor, $100 and up for heavy plans, $20 a seat plus usage at API rates for an enterprise plan. The vendor sets a usage cap so that the fixed fee covers the expected tokens. The cap is the vendor’s hedge, and the heavy user who hits it is the vendor’s tail risk. No vendor publishes tokens per seat, so the seat rent cannot be measured from outside. What can be seen is that the prosumer plans at OpenAI, Anthropic and Cursor all carry the same twenty-dollar price point, which says the price is set by the market for seats, not by the cost of tokens.

What this means for a buyer

The cost of a given output is set in the bottom three layers and is knowable to within a factor of two. The price paid is set in the top four and can be anywhere from 3 to 500 times that cost. The address is the tool for putting a number on that ratio for a particular buyer and a particular use, and section 04 does so three times.

03 · What the address tells a buyer

Who pays rent at which layer, and the twelve levers that close the gap, ranked by size.

Classify seventeen buyer types by their dominant use and three patterns fall out. Only three of them own the chip. The buyers with the thinnest margins run the least forecastable workloads. And the enterprise adopter is the only buyer paying rent at all four upper layers at once.

Figure 5 is the condensed map; the Navigator carries all seventeen rows with the revenue basis, cost basis and main cost risk for each. Rows run from the bottom of the stack, the chip owners, to the top, the seat buyers.

Figure 5 · Who pays rent where: seven buyer types, condensed
Taxonomy
Address gives the layers that matter most for that buyer. Full seventeen-row map in the Navigator.
BuyerDominant useAddress, main codesWhat it pays, and at which layerMargin shapeMain cost risk
Frontier labServing its own APISI-GPU-B/H · CF-OWN · MC-FR · TT-INF-ON · DP-API · CT-LISTOwned chip-hours at about $1 to $1.50 per H100-hour; serving software; idle capacity off-peakHighest rent in the stack; margin not disclosedDemand moving to open weights at the small tier; utilisation swings
NeocloudRenting chip-hoursSI-GPU-H/B · CF-RES sold as CF-ODAbout $1.07 an hour at full use, $1.53 at 70 percentRent 40 to 75 percent of price; spot below costThe index falling below owner cost on the older generation
Inference hostSelling tokens on open weightsSI-GPU-H/B · CF-OWN/RES · MC-OW · CF-SVLChip-hours at owner or reserved cost; $0.07 to $0.24 per million tokens producedPrice 8 to 19 times production cost, competed down by the number of hostsUtilisation below break-even; price war on popular models
AI-native applicationCoding, legal or support product on frontier modelsMC-FR-L/M · TT-AGT · WS-TA · DP-API/GW · CT-COMMIT+CACHEFrontier tokens at list less commitment discount; input-heavy at 25 to 1 in agent loopsThin and variable; agent depth sets the gross marginToken cost per seat exceeding seat price; vendor price reset; model retirement
Enterprise, regulatedDocument processing, service, internal copilotsMC-FR-M · TT-INF-BA/ON · WS-T1/TK · DP-CLOUD · CT-COMMITTokens through a cloud commitment; seats for knowledge workers; some self-hosting for data residencyPays rent at four layers: chip, model, cloud, seatUnused commitment; consumption drifting from linear to multiplicative; model retirement mid-contract
Enterprise, generalCoding assistants and productivity seatsMC-FR · TT-AGT · DP-APP · CT-SEAT/CREDIT$10 to $100 a seat a month plus credit overageRent to the application vendor; cheapest frontier tokens per outcome for light usersSeat sprawl; heavy users on credits
Developer or prosumerBuilding and using toolsMC-FR/OW · TT-AGT · DP-APP/API · CT-SEAT/CREDIT$20 to $200 a month in seats and creditsPays full retail rent; most tokens per dollar on open weightsCredit burn in agent loops
Highlighted rows are the two buyers worked in section 04. The dealer of the third example does not yet appear in the map, because no one sells that product at scale; that is lever 12 below.

The second thing the address returns is a ranked list of what to do about it. Figure 6 orders twelve levers by how much each can move the cost of a given output on the September prices, using the multiple between the worst and best case on that lever alone. The levers compound. The largest three sit in the upper layers and belong to the buyer. The chip-level levers are real but small by comparison, and they belong to whoever owns the chip.

Figure 6 · Twelve levers, ranked by size
Taxonomy
Size is the multiple between worst and best case on that lever alone. Bar length is on a log scale. Blue: buyer-side. Orange: belongs to the chip owner. Purple: moves variance, not the mean.
#LeverLayerSizeWho captures it todayEvidence
Levers 1 to 5 and 8 to 9 need nothing from the market except measurement, which is why the Tokenomics Foundation’s programme is built around them. Levers 6, 7, 10 and 11 pass through to buyers only where hosts compete. Lever 12 is the one no one yet provides at scale: the market has indices for the chip-hour and list prices for the token, but no instrument that fixes the cost of an outcome for the enterprise that has to budget for it.
Where the trading interest is, and is not

Only the two bottom layers are hedgeable on a listed instrument today. The chip-hour has two published H100 indices and B200, H200, A100 and MI300X series, and NYMEX has filed futures on two of them (see The Contract Lost Its Date). Nothing above layer two is indexed. The model-class price is a vendor list that resets a few times a year; the quantity of tokens is decided by architecture; the seat is a contract. So for a desk the address is a decomposition: it says what share of a client’s cost can be hedged on the screen (the chip leg of any self-hosted share), what has to be written as a private instrument (the frontier list leg and the quantity leg), and what has to be held as capital. Example C in the next section runs that decomposition on a stylised contract.

04 · Three worked examples

The same address, read by a finance team, an application company and a dealer.

Each example starts with an address, prices it from the snapshot, and then moves the dials that matter for that buyer. The inputs are round illustrative numbers. The prices are not.

All three ledgers use the same arithmetic for a token bill: input tokens at the list rate, less the cached share at 10 percent of list, plus output tokens at the list rate, with batch at half price where it applies. The Navigator’s margin-stack tab adds a self-host break-even calculator for the operator side.

Example A · Financial management
A regulated enterprise reconciles its AI bill against the budget
MC-FR-M | TT-INF-BA + TT-INF-ON | WS-T1 → WS-TK | DP-CLOUD | CT-COMMIT + CT-CACHE + CT-BATCH

The finance team budgeted a document-processing and customer-service workload as linear: one model call per request, on a frontier mid model, bought through a cloud commitment. Six months in, the product team has added a reasoning step and a retrieval loop, and the bill does not match the budget. The address shows why: the workload moved from WS-T1 to WS-TK, and the layer that moved is the one nobody was watching.

The ledger prices the tokens at list, applies the terms the enterprise actually has, and sets the result against the annual commitment. The floor line is what the same tokens would cost on a hosted open-weight mid model, and self-hosted. The gap between the bill and the floor is the rent the enterprise pays at four layers.

Figure 7 · The enterprise bill ledger
Live
Monthly figures. Prices are vendor list for the selected tier; DeepInfra list for the open-weight floor; the taxonomy’s self-hosted cost for the owned floor.
The workload as budgeted
What actually happened
The terms
Token bill as budgeted
Token bill as it stands
Against the commitment
Seat bill
Same tokens, hosted open-weight mid
Same tokens, self-hosted
Rent multiple, bill over owned floor
Bill after termsBudgetCommitment, monthly share

Three things the finance team can now do that it could not do from the invoice. It can name the layer that moved, and ask the product team for the consumption class of each feature before it ships rather than after. It can size the commitment to the class it expects, because an unused commitment is lost and an overrun is billed at list. And it can put the rent on paper: the same tokens would cost a fraction on an open-weight model, and the difference is the price of frontier capability plus the price of buying it through a cloud. Whether that is worth paying is a product question; until the address is written down, it is not a question at all.

Example B · Pricing a product
An AI-native application company prices a twenty-dollar seat on a frontier model
MC-FR-M | TT-AGT | WS-TA | IO-IN | CTX-L | DP-API | CT-COMMIT + CT-CACHE

A coding assistant sells seats at the market price for seats, $20 a month, and buys tokens at the market price for tokens. Its cost of goods is an agent loop: a fifty-turn coding session replays about a million input tokens and produces about forty thousand output tokens, an input-to-output ratio near 25 to 1. The workload shape is agent-multiplicative, the least forecastable class in the taxonomy.

The ledger prices a session, multiplies by sessions per seat, and sets it against the seat. Then it moves the three buyer-side levers that matter for this address: cache the replayed context, route sub-tasks to the small tier, and buy on a commitment. The heavy-user dial shows the tail: the seat that uses several times the mean, which the usage cap exists to stop.

Figure 8 · The seat margin ledger
Live
Per seat per month. Session shape from Vantage’s April measurement of a fifty-turn agentic coding session; cache read rates from the vendor price pages (10 percent of input at OpenAI and for Sonnet 5 and Haiku 4.5; 5 percent for Opus 5.5; 2.5 percent for Fable 5.1).
The product
The model
The levers
Token cost per session
Token cost per mean seat
Gross margin per mean seat
Heavy seat
No routing25% routed to small tier50% routedBreak-even

At the defaults the seat loses money: a session costs about $1.50 in tokens and twenty of them cost more than the $20 the seat brings in. That is the taxonomy’s reading of the application company’s row: thin and variable, with agent depth setting the gross margin. The way out is not a cheaper chip, because the company does not buy chips. It is the three upper-layer levers, and they compound. Restructure prompts so the stable prefix comes first and the cache hit rate rises from single digits to 80 percent; route the sub-tasks that do not need the frontier model to the small tier; then buy the remainder on a commitment. The ledger shows the seat crossing from loss to a margin that a software business can live with, and it shows the heavy user turning the mean back into a loss. The cap is not a product decision. It is the hedge.

Example C · Trading and intermediation
A dealer prices a fixed price per task for an enterprise client
MC-FR-M + MC-OW-L | TT-INF-BA | WS-TK | DP-GW | CT-PT → fixed $ per task

Lever 12 is the one no one provides at scale: converting a variable token cost into a fixed price per outcome. An enterprise wants to pay a set price for each completed contract review, for a year. A dealer that will write that contract has to know what a review costs, how much of that cost can be hedged on the screen, and how much capital it takes to hold the rest.

The address answers the first question and decomposes the other two. A review at the defaults runs about 200,000 input tokens and 8,000 output. The dealer routes part of the work to a frontier mid model at list and part to an open-weight large model it serves on rented or owned H200s. That splits the cost into three legs. A chip leg, the self-hosted share, which moves with the H200 rental rate and can be hedged on the listed indices. A list leg, the frontier share, which moves when the vendor resets its price and has no index. And a quantity leg, tokens per task, which moves when the client’s documents or the dealer’s prompts change shape, and which only architecture and a contractual cap can control.

Figure 9 · The dealer’s per-task ledger
Live
Stylised. Per task and per month. Chip-hour rate from the Nebius H200 on-demand price; serving throughput from SemiAnalysis InferenceX; frontier list from the vendor pages.
The task class
The routing
The contract
The stress
Expected cost per task
Fixed price offered
Hedgeable share of cost
Monthly revenue
Stressed cost, unhedged, no cap
With the index hedge
With hedge and cap
Monthly P&L under stress, hedge and cap
Unhedged, no capIndex hedge on the chip legHedge and capZero

The ledger’s first finding is uncomfortable for a desk. At the defaults the chip leg, the only part of the cost that can be hedged on a listed index, is around a fifth of the dealer’s cost. The frontier list leg is bigger and has no screen. The quantity leg is bigger still and has no screen either. So the index hedge, correctly sized, removes a small slice of the stress, and it is the contractual cap on tokens per task that does most of the work. That matches the taxonomy’s reading of lever 12: what it needs is a task-cost reference, a usage record standard and capital to hold the tail. The listed chip-hour future is the smallest of those three, and it is the only one that exists.

What the dealer ends up holding

A book of fixed-price tasks is short the vendor’s list price, short the client’s prompt discipline and long a small chip-hour position it can hedge. The first two exposures are why the seat vendors cap usage and why no one sells this to enterprises yet. The taxonomy does not make them hedgeable. It makes them nameable, which is the step before a price.

05 · Choosing a chip, or choosing terms

For open-weight work the lever is utilisation. For frontier work it is the contract.

For any self-hosted workload, the cost of a million tokens is the chip-hour price divided by what the chip produces in an hour. That one formula sorts the eleven workload classes into two groups: the ones where you choose a chip and the ones where you choose terms.

Three rules of thumb follow from the September numbers. A newer generation is cheaper per token even at a higher hourly price: on DeepSeek R1 a B200 produces 7 times the tokens of an H100 at 3.2 times the index price, so it is 2.2 times cheaper per token. Self-hosting an open-weight model on a rented eight-chip H200 node, about $864 a day, beats hosted open-weight prices once the node runs above about 13 percent utilisation, roughly 470 million tokens a day; on owned chips the break-even falls to about 5 percent, or 190 million tokens a day. Below those volumes, or for any frontier model, the economic choice is not a chip but a set of terms.

Figure 10 · The most economic setup, by workload class
Taxonomy
Condensed to six of eleven classes; the Navigator has all eleven with the switch condition for each.
Workload classAddressMost economic setupIndicative costWhen to switch
Batch extraction, summarisation, classification at volumeTT-INF-BA · MC-OW-L/M · WS-T1 · LAT-BATOwned or reserved B200 running open weights; spot H200 when interruption is tolerable$0.06 to $0.17 per million tokens self-hosted; $0.60 to $3.96 hostedBelow about 470 million tokens a day per rented node, use a hosted API at batch rates
Interactive assistants on open weightsTT-INF-ON · MC-OW-M · WS-T1 · LAT-INTHosted open-weight API until volume is steady; then H200 or B200 self-hosting at 100 tokens a second per userHosted $0.10 to $1.04 per million input; self-hosted $0.07 to $0.24Speed-critical products pay Groq or Cerebras rates for tokens per second, not per dollar
Small-model tasks: classification, extraction, embedding, rerankingTT-RET · MC-OW-S/TASK · WS-T1L4, L40S or consumer cards; Arm CPU for embeddings at low volume$0.49 to $1.09 a chip-hour; $0.05 to $0.15 per million tokens hostedNever a Hopper or Blackwell chip; the memory is wasted
Agentic coding and research on frontier modelsTT-AGT · MC-FR-L/M · WS-TA · IO-IN · CTX-LNo chip choice. Cheapest terms: no long-context premium, cache reads at 10 percent of input, committed spend, sub-tasks routed to the small tierInput is 25 tokens per output token; at 80 percent cache hits the effective input rate is about 28 percent of listWhen a task class is definable, buy per task from a router or dealer rather than per token
Long-document reasoningTT-INF-ON · MC-FR-L · WS-TK · CTX-LFlat context pricing or batch mode; open-weight large hosted when quality allowsFrontier $2 to $5 per million input, doubling above 200,000 tokens at two of three vendorsChunk and retrieve before paying long-context rates
Evaluation and regression suitesTT-EVAL · LAT-BATBatch API at half price, or spot chips for open weightsHalf of list; $0.79 an H100-hour spotNever interactive rates
The pattern: chips are chosen for training, batch and small-model work; terms are chosen for everything else. For every frontier workload the lever is layer seven. For every open-weight workload it is utilisation.
06 · What the address cannot yet see

Two rents are bounded, not measured. Both are the largest in the stack.

The taxonomy is only as good as the prices that fill it, and two of the seven layers publish none.

Frontier serving cost. No vendor publishes it, so the frontier rent in step 3 of the margin stack is bounded by open-weight hosting prices, not measured. The best proxy is the price of the largest open-weight model served on Blackwell at scale. Tokens per seat. No application vendor publishes it, so the seat rent in step 4 cannot be measured from outside. A gateway operator, or a large enterprise’s own usage export, is the only source. Tokens per watt across silicon. InferenceX covers Nvidia; the AMD, TPU, Groq and Cerebras figures are vendor claims until someone runs one model across all of them. Index coverage. Compute Desk’s index levels are not public; when they are, they belong beside the Silicon Data and Ornn rows. Training cost. Third-party estimates of frontier runs are modelled, not disclosed, and should carry that label wherever they are reused; DeepSeek V3 at $5.58 million is the one disclosed figure in the snapshot.

What would change the read

A published tokens-per-seat figure from any application vendor would turn the seat rent from a bound into a number, and would probably move lever 12 from “not yet provided” to “priceable”. A frontier vendor publishing serving cost, or a Blackwell-scale open-weight host publishing its unit economics, would do the same for the model rent. Neither is likely soon, and the taxonomy is built to work without them: the bounds are wide, but they are bounds, and a buyer who knows the floor can negotiate against it.

07 · Verdict

The address is the missing join between the chip-hour index, the token list and the seat.

Every unit of AI compute already has a price in three units that cannot be compared. The seven-layer address makes them comparable, and once they are comparable the rent at each layer becomes a number a buyer can budget against, a vendor can defend and a dealer can quote.

For the finance team, the address turns a surprise on the invoice into a named layer that moved. For the application company, it turns a thin margin into three levers with sizes. For the desk, it decomposes a client’s cost into the slice a listed future can hedge, the slice that needs a private instrument, and the slice that needs capital. In every case the address does not lower the cost by itself. It says where the cost is, which is the thing none of the three prices did.

Three things follow. First, the cost of an output is set in the bottom three layers and is knowable to within a factor of two from public prices; the price of an output is set in the top four and is not. Second, the largest levers belong to the buyer and need nothing from the market except measurement, which is why the Tokenomics Foundation’s work is built around them. Third, the instrument the market lacks is not another chip-hour contract. It is a fixed price per outcome, and the address is the object it would be written on.

Watch 1
A tokens-per-seat disclosure
Any application vendor, gateway operator or large enterprise publishing tokens consumed per seat per month. It turns the largest rent in the stack from a bound into a number.
Watch 2
A per-task price for enterprises
A router or dealer quoting a fixed price per completed task class, with a cap and a term. Lever 12, which nobody provides at scale today.
Watch 3
The frontier small tier against open-weight mid
The one place prices overlap ($1.20 to $5.00 against $0.32 to $3.00 output). Where competition is real, the frontier rent can be measured.
Watch 4
A cross-silicon throughput table
One model, one latency target, measured on H100, H200, B200, MI300X, TPU, Groq and Cerebras by someone other than the vendor. It fixes layer one.
Watch 5
Consumption class on the invoice
A vendor or gateway reporting reasoning tokens and agent depth per request. Without it, lever 1, the largest, cannot be measured by the buyer who owns it.
Sources & claims register

What this piece rests on

The taxonomy

  • Kinetic Alpha, Compute Consumption Taxonomy, working document, September 28, 2026: seven layers and code tables; margin stack (steps 1 to 4); price snapshot of September 26 to 28, 2026; end-user map (seventeen rows); most economic setup by workload class (eleven rows); optimisation register (twelve levers); open questions. Every table, range and multiple in this piece is taken from it. A spreadsheet companion carries the full snapshot of about ninety rows.
  • The five-layer stack (silicon, capacity, serving software, model, routing) and the consumption classes are the Tokenomics Foundation’s. Lever 1’s 34-times figure (one task, four architectures, $3.04 to $0.09) is the foundation’s worked example. Layers six and seven are ours.

Chip-hour prices, $ per chip-hour, read September 26 to 28

  • Indices: H100 rental index $2.70 (Silicon Data, SDH100RT, September 27); H100 SXM $2.54, H200 $4.78, B200 $8.04, A100 80 GB $0.95 (Ornn OCPI, September 26); MI300X $2.62 (Silicon Data).
  • Neoclouds and marketplaces: H100 SXM $3.99 on demand, B200 $6.69 on demand and $8.87 to $9.86 reserved cluster (Lambda); H100 $3.85 on demand and $0.79 spot, H200 $4.50 on demand and $0.79 spot (Nebius); H100 $3.00 reserved and $4.29 resold (SF Compute); A100 80 GB $1.59, L40S $1.09, L4 $0.49, RTX 4090 $0.74, RTX 5090 $0.99 (RunPod); MI300X $2.39 to $7.86 across six providers (GetDeploying).
  • Hyperscalers: H100 capacity block $5.19, B200 $12.36 and GB200 NVL72 $10.58 (AWS); H100 $11.06 and B200 $8.06 on demand (Google Cloud); H100 $12.29 on-demand list (Azure via Vantage); TPU v5e $1.20, v6e $2.70, v7 $12.00 (Google Cloud TPU); CPU $0.045 per vCPU-hour (AWS c7i via Vantage) and $0.01 per Arm core-hour (Oracle Cloud A1).

Token prices, $ per million tokens, input / output

  • Frontier: GPT-5.6 Sol 5.00 / 30.00, Terra 2.00 / 12.00, Luna 0.20 / 1.20, cached input at 10 percent (OpenAI); Claude Fable 5.1 and Mythos 5.1 10.00 / 50.00, Opus 5.5 4.00 / 20.00, Sonnet 5 2.00 / 10.00, Haiku 4.5 1.00 / 5.00, cache reads at 10 percent of input, cache writes 25 to 100 percent over input, one rate across the full context (Anthropic); Gemini 3.1 Pro 2.00 / 12.00 rising to 4.00 / 18.00 above 200,000 tokens, 3.5 Flash 1.50 / 9.00, Flash-Lite 0.30 / 2.50 (Google); Grok 4.7 2.00 / 6.00 (xAI). Batch is 50 percent off at OpenAI, Anthropic and Google; OpenAI bills about 2 times input above 272,000 tokens.
  • Open weights: DeepSeek V4 Pro 1.32 / 3.96, V4.1 Flash 0.30 / 1.20, Kimi K3 3.00 / 15.00, Qwen3.8 Max 2.00 / 6.00, GLM-5.3 1.40 / 4.40 (Together, Fireworks); DeepSeek V4 Pro 1.30 / 2.60, Llama 3.3 70B 0.10 / 0.32 (DeepInfra); Llama 3.3 70B 0.59 / 0.79, gpt-oss 120B 0.15 / 0.60, gpt-oss 20B 0.075 / 0.30, Llama 3.1 8B 0.05 / 0.08 (Groq); Mistral Large 3 0.50 / 1.50, Small 4 0.15 / 0.60 (Mistral).

Throughput, cost inputs and consumption

  • Measured serving: DeepSeek V4 Pro on H200 at 75 / 100 / 150 tokens a second per user gives 5,771 / 5,139 / 2,648 tokens a second per chip and $0.059 / $0.066 / $0.13 a million tokens (SemiAnalysis InferenceX, September 16); DeepSeek R1 on B200 against H100 at 45 tokens a second per user: 5,215 against 740 tokens a second per chip, $0.092 against $0.439 (InferenceX, July); gpt-oss 120B on B200 about 60,000 tokens a second per chip, vendor claim (Nvidia, October 2025).
  • Owner cost inputs: colocation $196 per kW per month, North America average, up 6.5 percent on the year (ColoPrice, September 26); US industrial electricity 8.83 cents per kWh, May 2026 (Data Center Scope); H100 $25,000 to $40,000 a chip, B200 $30,000 to $50,000 (IntuitionLabs, September 14).
  • Training: DeepSeek V3 $5.58 million and 2.79 million H800-hours, disclosed (arXiv 2412.19437); Grok 4 about $490 million, third-party estimate (Epoch AI); LoRA fine-tune of a 70B model $15 to $30 of GPU time, $5,000 to $15,000 all-in (Stratagem).
  • Consumption: a fifty-turn agentic coding session uses about 1,000,000 input and 40,000 output tokens (Vantage, April 15); cache hit rate 7 percent before and 74 to 93 percent after restructuring prompts, a 59 to 70 percent cost cut (ProjectDiscovery, April 10); enterprise generative AI spend $37 billion in 2025, applications 51 percent and infrastructure 49 percent, model APIs $12.5 billion (Menlo Ventures, December 2025); tokens about one quarter of the AI bill (Broadcom ValueOps, September 24).
  • Seats: ChatGPT Plus $20 and Pro $200; Claude Pro $20, Team $20 to $100, Enterprise $20 plus usage; Cursor Pro $20, Teams $40; Copilot Pro $10, Max $100 a month (Anthropic, Cursor, GitHub, OpenAI).

Claims register

  • $1.07 and $1.53 per H100-hour are the taxonomy’s owner-cost build: $30,000 a chip over 5 years and 8,760 hours ($0.68), 1.0 kW at 1.3 power usage effectiveness and 8.83 cents ($0.11), $196 per kW-month colocation ($0.27); at 70 percent utilisation the same cost over fewer hours. Figure 4 reproduces the arithmetic and lets the inputs move.
  • 9 cents to $50 are the ends of the token ladder: measured owned-B200 serving of DeepSeek R1 at owner cost, against the Claude Fable 5.1 list output price. The 543 multiple in the tile is 50 / 0.092. Different models at the two ends; the ladder compares ways of buying a million output tokens, not one model.
  • 34 times is the Tokenomics Foundation’s worked example of one task under four architectures, $3.04 to $0.09. It is the foundation’s figure, not ours.
  • 8 to 19 times (hosted open-weight price over production cost) and 3 to 500 times (price paid over cost) are the taxonomy’s ranges from steps 2 and 4 of the margin stack.
  • 13 percent and 5 percent break-evens are the taxonomy’s: an eight-chip H200 node at $4.50 an hour ($864 a day) against hosted open-weight list, and the same node at owner cost. The token-per-day figures (470 million and 190 million) follow from the 5,139 tokens a second per chip throughput.
  • Cache read rates in Figures 7 and 8 follow the vendor table, not the taxonomy’s round “10 percent” rule: 10 percent of input at OpenAI, Google and for Sonnet 5 and Haiku 4.5; 5 percent for Opus 5.5 (0.20 on 4.00); 2.5 percent for Fable 5.1 (0.25 on 10.00). Figure 9 uses Sonnet 5 at 10 percent.
  • Figures 7, 8 and 9 are Kinetic Alpha’s own arithmetic on round illustrative inputs and snapshot prices; the commitment discount, the heavy-user multiple, the dealer’s margin and the stresses are assumptions, marked as such on the dials. Figure 9 is stylised: it treats input and output tokens on the self-hosted leg at one cost per token, applies a uniform shift to tokens per task, and assumes the index hedge tracks the H200 rental rate one for one. None of that is a measurement.
  • The code example in the working document reads SI-GPU-H | CF-API | MC-FR-L | TT-INF-ON | WS-AGT | DP-GW | CT-TOK-LIST. Three of those codes are not in its own tables; this piece uses the table codes (CF-SVL, TT-AGT, WS-TA, CT-LIST) throughout.
  • Related Kinetic Alpha work: The Reserved Contract, Unbundled (layer two), Refiner and Merchant (who owns the chip-to-token margin), The Contract Lost Its Date and The Contract Has a Date. It Needs a Dealer. (what the listed chip-hour future does and does not hedge), The Token Price Index (layer three as an index).