Named in one sentence by the instrument builders, studied at length by people who never say “hedge”
In The Compute Crack Spread we wrote the producer’s margin as C = T·p − r: tokens per GPU-hour, times the price of a token, less the rent on the GPU. Everything that has happened since has been on the right-hand side of that equation. CME and Silicon Data list H100 and B200 rental-index futures on October 5, pending regulatory review, each contract “a month’s worth of rent.” ICE and Ornn, the GPU leg of Architect and Ornn, and Kalshi’s event contracts on Ornn prints all reference r. Bandi and Su estimated the first risk premium on r. Our own benchmark work is about how r is measured. The token price p has a research literature and no contract. And the term that is not in the producer’s equation at all, the quantity of tokens the buyer needs to get a job done, has neither.
The buyer does show up. CME’s release has Pete Keavey saying “AI builders and hyperscalers need to hedge as they grow,” and its body copy promising to allow “companies, including AI developers and hyperscalers, to lock in their costs”; Carmen Li’s own quote is that “together, that turns compute from something enterprises negotiate blindly into a market they can actually plan around.” Silicon Data’s cash-settlement note (Yu and Hou, August 24) names “AI developers and enterprises” on the demand side against “cloud providers and the broader compute supply chain” on the supply side. Dave Friedman’s market primer, the best independent map of the space, lists enterprises among the buyers of compute and then spends its analytical weight on fleet financing, observing that “the natural longs may be the more important source of sell-side hedging flow if they are financing fleets with debt.” His “longs” there are the fleet owners rather than the buyers — the reverse of the convention used here, and a taxonomy worth pinning down before the two literatures can talk to each other at all. Ornn, profiled by Axios in July, names lenders and “buyers and sellers of compute.” Nobody names a buyer.
The people who do study the buyer come from a different tradition. Figure 1 sorts them.
| Camp | Representative work | What it says about the end user | What it does not say |
|---|---|---|---|
| Exchanges and index providers | CME/Silicon Data (Aug 11); ICE/Ornn; Architect/Ornn; Kalshi; Silicon Data’s cash-settlement note | “AI developers and enterprises” will hedge rising rental costs | Which enterprises rent GPUs; what an API buyer’s exposure to a rental index is; basis |
| Independent market structure | Friedman’s primer and basis-risk notes; Chen (remio, Aug 20) on the CFTC review | Enterprises are among the buyers; Chen notes larger cloud companies “may already manage costs through long-term supply agreements” and that those private contracts “could reduce their need for exchange-traded hedges” | No decomposition of the buyer’s exposure. Chen tells buyers to check the index against their own rental contracts, which presumes they have rental contracts |
| Strategy consulting | BCG, Return on AI: How CFOs and CIOs Can Manage the Token Meter (Jul 1); BCG, Is AI Computing Power Becoming a Commodity? (Jul 23); Gartner “Inference Paradox” (Aug 17); EY on agentic token costs | Governance, attribution, routing, return on AI. BCG Jul 23 is the exception: separate baseload from flexible demand, throttle the flexible, and hedge AI compute “the same way they treat energy or foreign exchange exposure” | Jul 1, Gartner and EY: no price risk, no contracts, no transfer. Jul 23: a paragraph, not an analysis, and its hedge is on rent |
| FinOps practice | FinOps Foundation token-economics working group (Jun 3); CloudZero on AI gross margin (Aug 31); State of FinOps 2026 | The best empirical picture of the buyer: spikes, tier spreads, commitment break-evens, margin compression | The word “hedge” does not appear. Risk management means measurement and routing inside one purchasing relationship |
| Academic | Xing, AI Token Futures Market (arXiv 2603.21690, Mar 2026); Bandi & Su (arXiv 2607.12156); Das et al. on cloud forward pricing | Xing is the only paper that treats application-layer AI companies as the primary buy-side hedgers, claiming a 62–78% reduction in cost volatility (standard deviation) against a token index it proposes itself | Xing is a simulation with an assumed index and hedge efficiency = ρ² on an assumed correlation of 0.85; basis is not worked. Bandi & Su price rent, not tokens |
| Contract practice | Gupta on price-protection clauses; vendor pricing pages | Gupta proposes shared-savings mechanisms passing “typically 30-50%” of savings to customers, and a model-efficiency clause under which a version cutting inference cost by more than 20% passes “at least 50%” through; deprecation windows and caps (vendor terms) | Not analysed as a risk book; no link to any market |
Read across the rows and the pattern is a gap, not a disagreement. The market people know the buyer exists and have not modelled it. The buyer’s own people have modelled it and do not know a market is coming. The one paper that connects the two is a simulation by an independent researcher. That is the whole competitive landscape on the demand side, as of this week.
The producer is long the token price. The buyer is short it. The contracts are on neither.
Write the buyer’s side the same way we wrote the producer’s. A workflow, an agent, a support bot, a document pipeline, whatever the unit of business output is, consumes h tokens per outcome at a blended price p per token, and the enterprise needs N outcomes a month. Spend is S = p·h·N. If the output is sold on, the buyer’s own spread is M = V − p·h per outcome, with V the price the buyer charges its customer.
| Producer | Buyer | |
|---|---|---|
| Token price p | Long Revenue per GPU-hour is T·p | Short Cost per outcome is p·h |
| Heat rate | T, tokens per GPU-hour. Utilisation, batching, serving stack. Physics. | h, tokens per outcome. Prompt design, agent loops, model generation. Not physics. |
| Capacity term | r, the rent. The thing the futures settle on. | None unless self-hosting. The API buyer pays p, not r. |
| Volume | Fleet size, fixed on a lease. | N, the business itself. Not a risk to hedge. |
| What the announced contracts cover | r and, through H* = r/p, an implied token floor. | Nothing except for the minority who rent. |
The producer’s heat rate T is physics and utilisation; we measured it at 2.88 million tokens per GPU-hour idle to 11.16 million saturated on the NLR sweep, against a market heat rate H* = r/p of about 3.57 million at a $2.53 H100 rent and a $0.708 median Llama-3.3-70B price. The buyer’s heat rate h is not physics. It is the number of tokens a model generation, a prompt, an agent framework and a retry policy decide to spend on a task, and it is set by engineering choices the buyer only partly controls. Gartner’s number for what happens to it is the important one: cost per agentic workflow up more than fivefold through 2028 while per-token prices fall, because, in the words of Gartner’s Will Sommer, “each successive generation of AI capability will necessitate more, and often more expensive, tokens.”
So the sign structure is: the producer is long p and long T; the buyer is short p and long-exposed to h. A token price index would have natural counterparties on both sides, which is exactly what the CFTC’s request for comment asked for and exactly what a rental index lacks on the demand side, because an enterprise buying tokens from an API has no rent. In the crack spread piece we said the token leg fails the Commission’s tests for fixable reasons and the rent leg fails them structurally. The demand side is the reason the fixable one is worth fixing.
Price is falling. Quantity is exploding. The discontinuities are contractual.
p is deflating. a16z measured the cost of a token at constant capability falling 10× a year: $60 per million for GPT-3 in November 2021 to $0.06 three years later for a 3-billion-parameter open-weight model scoring equivalently — a like-for-like capability comparison, not frontier-to-frontier. Xing’s figure for GPT-4-class inference is $60 to under $1.50 in two years. A buyer short a commodity with that drift does not fear the level of p. It fears two other things: a jump, when a provider re-prices or retires the model a workflow was tuned on; and the dispersion of p, which the FinOps working group puts at 50–100× between a provider’s frontier and smallest models. Every routing decision is a price decision, and routing is the operational lever every consultancy recommends, so the buyer’s effective p is a portfolio weight, not a market print.
h is the risk. The FinOps working group’s plainest sentence is that “token consumption can spike dramatically based on user behavior, prompt design, or application bugs.” CloudZero’s analysis says two accounts on identical plans generate order-of-magnitude different inference costs. Output tokens, priced 3–5× input, are set by the model at runtime, not by the buyer. And the trend is structural: an agentic reasoning workflow costs at least 5× a chatbot before task complexity grows. This is a quantity risk with the shape of a volumetric exposure in power: the buyer does not know how much it will consume, and the not-knowing is correlated with the thing it wants, a more capable agent.
N is the business. More outcomes means more revenue or more work done. No treasurer hedges demand for their own product. It belongs in the equation because it multiplies the other two.
The discontinuities are where money has actually been lost. Two examples, both from the last fifteen months, both from the top of the stack.
The Pro plan went from 500 requests a month to “$20 of frontier model usage per month at API pricing,” with “unlimited” reserved for the Auto tier and additional usage “at cost.” The company wrote on July 4 that “we didn’t handle this pricing rollout well, and we’re sorry,” and offered refunds for June 16 to July 4. Read as a risk event: an AI application company holding a fixed-price liability against a floating token cost re-priced to pass-through. Its customers became the holders of p and h.
Claude Code’s weekly limits on Pro, Max, Team and seat-based Enterprise plans change on September 14: 25% above the original baseline, 17% below the temporary level users had been running on. Not a price change. A quantity change on a fixed-price plan, which is the same thing to a buyer whose h just went up.
A stylised buyer makes the ranking concrete. The sliders are there so you can disagree with the base case.
Base case
Shocks
Some of it is hedgeable. Most of it is manageable. The part that is neither, the buyer is paying for already.
| Exposure | Nature | Managed by | Transferred by | Available today |
|---|---|---|---|---|
| p, level | Market price, deflating | Routing, tier selection, caching, batch (50% discount) | A token price index forward or cap, on baseload volume | No No listed contract; OTC only in principle |
| p, jump (re-price, deprecation) | Contractual, discontinuous | Multi-provider architecture, model abstraction layer | Price-protection and model-efficiency clauses, deprecation notice periods, most-favoured pricing | Yes if negotiated. Gupta’s shared-savings and model-efficiency clauses are the template — prescriptive drafting, not observed contract data |
| h, tokens per outcome | Engineering quantity, trending up | Prompt discipline, context limits, kill switches, a named owner per endpoint, a baseline before a budget | Not insurable in any current form. The nearest analogue is a quantity cap in the contract | Operational only |
| Availability (rate limits, provisioned throughput) | Capacity option | Committed capacity, break-even modelled | Provisioned throughput is the option; its premium is embedded in the price | Yes and expensive: “for several current frontier and reasoning models, provisioned capacity costs more per token than pay-as-you-go even at 100% utilization” (FinOps WG) |
| r, rent (self-hosters only) | Market price | Term leases | CME/Silicon Data, ICE/Ornn, Architect futures; Kalshi | Oct 5 |
| N, volume | Business | – | – | – |
Two rows deserve a second look.
The availability row is the buyer already paying an option premium without calling it one. When provisioned capacity costs more per token than pay-as-you-go at full utilisation, the difference is the price of a capacity-locking option, and the FinOps working group’s advice, “before pursuing commitments, model the break-even arithmetic explicitly,” is a request to price the option. BCG found the physical version of the same behaviour on the rental side: enterprise GPU clusters running at utilisation “as low as 5%” (a Cast.AI figure relayed by BCG) because undercommitting risks not having capacity later. That is a buyer over-hedging a quantity risk with idle silicon, at 95% carrying cost, because no cheaper instrument exists. It is the same reason a utility keeps peakers.
The h row is where the consultancies are right and the markets are silent. Tokens per outcome is a quantity the buyer’s own engineers set, so the first-order tools are internal: the FinOps working group’s 30–60 days of measurement before a budget, the CIO playbook’s named owner per endpoint and automated kill switches, Gartner’s tiering and orchestration. But h is also correlated with the buyer’s success, because a better agent spends more tokens, and it is correlated across buyers, because a new model generation moves everyone’s h at once. That second correlation is the one that eventually makes it a market: an “agentic upgrade” risk that every buyer holds on the same date is the raw material of an index.
The producer is paying to hedge. The buyer would be paid to take the other side. The buyer’s biggest risk is the term nobody would pay it for.
Bandi and Su report preliminary evidence “consistent with a positive compute risk premium, suggesting hedging pressure on the part of compute providers” — estimated from synthetic futures built out of term rental contracts, which the paper itself calls likely upper bounds on true futures prices: futures below expected spot because the sellers need to sell. The reason they need to, in our reading rather than theirs, is that the fleets are financed with debt and the lenders want forward revenue. That is Keynes’s normal backwardation, in GPU-hours: the side with the hedging need pays, and here the side with the need is the producer. If the same pressure holds on the token leg, and the fleets generating tokens are the same fleets, then a buyer who commits to baseload volume at a forward price is not paying for insurance. It is being paid to provide it.
That is an unusual position for a corporate hedger. In the commodities a treasurer already knows, the sign of the premium is contested and the buyer usually ends up paying for certainty, through option premium or a contango it has to roll: the airline fixing jet fuel, the utility fixing gas. The inference buyer, short a deflating commodity whose producers are long and levered, would be the seller of certainty in the token market, collecting the premium and taking a forward price below today’s spot. BCG’s July 23 instinct, treat AI compute like an energy or FX exposure and hedge it where material, arrives at the right desk with the wrong sign.
Every dollar of premium available on p is small against the exposure on h. Which is why, when the demand side does show up in these markets, it will not be to lock a price. It will be to write cheap floors for producers on baseload volume it was going to buy anyway, and to spend the proceeds on the engineering that controls h.
That is also why the risk ends where it does. Look at the chain: the host indexes its toll to tokens; the model provider prices per token; the application company, when the fixed-price plan blows up, re-prices to pass-through, which is exactly what Cursor did. Every layer above the enterprise has found a way to be short h to the layer below. The enterprise end user is the last holder of both p and h, has no customer to pass them to, and is the only participant in the chain without a contract, a desk, or an analyst.
Six things, most of which need no market
- Own it like fuel. The line is now material: 98% of State of FinOps respondents manage AI spend, up from 31% two years ago. Fuel desks at airlines sit between procurement, operations and treasury; the inference book needs the same three owners, with FinOps as operations.
- Measure the three terms per workflow, not the bill. p, h and N separately, with h baselined over the 30–60 days the working group recommends before any budget is set. Attribution is the difference between “the bill went up” and “h went up on the support agent after the model upgrade.”
- Split baseload from flexible. BCG’s frame is right even where its instrument is wrong. Baseload is the volume you will consume in any state of the world; flexible is what gets throttled when p or h moves. Only baseload is ever a hedging candidate.
- Negotiate the jumps. Model-efficiency clauses (a stated share of provider cost declines passes through), deprecation notice periods long enough to re-tune h on a successor, and a cap on re-pricing within the term. These are the only instruments that address the second-largest row in Figure 3, and they cost nothing but negotiating time.
- Price the capacity option you already own. Compute the provisioned-throughput break-even. If it never crosses at 100% utilisation, you are paying for availability, which may be right, but say so, and size it to baseload.
- When a token index exists, sell certainty on baseload and buy nothing on h. Write the floor the producer needs, collect the premium, and treat the proceeds as the budget for the h engineering in step 2. Do not buy a rental future unless you rent.
Four things that would change this piece
- October 5. First CME prints. The question is not the price; it is whether any open interest is held by a firm that consumes tokens rather than selling GPU-hours. Large-trader categories will tell.
- A token index with a specification. The demand side has natural counterparties only on p. Whoever publishes a model-class-stratified token price with a methodology that survives the RFC’s fungibility and disclosed-price tests owns the buyer’s leg.
- The next “quantity change on a fixed-price plan.” September 14 is one. Each one moves h for a population of buyers on the same day, which is what an index needs and what a buyer’s programme has to survive.
- A consultancy or a FinOps working group that uses the word “forward.” The first one to connect the measurement tradition to a market will have written the buyer’s side properly. On the evidence of Figure 1, it has not happened yet.
Releases, papers and practice notes
- Exchanges and indices: CME Group / Silicon Data, compute futures launch October 5, pending regulatory review (Aug 11, 2026) · Silicon Data, Cash-settled compute futures (Yu, Hou, Aug 24, 2026) · ICE and Ornn · Architect and Ornn · Axios, Ornn profile (Jul 6, 2026)
- Independent market structure: Dave Friedman, Compute Derivatives Market Primer · Basis Risk Primer · The GPU Risk No One Is Managing (Dec 28, 2025) · Martin Chen, CFTC reviews compute futures, but GPU capacity is not oil (Aug 20, 2026)
- Consulting: BCG, Return on AI: How CFOs and CIOs Can Manage the Token Meter (Bijlsma, Kleine, Scognamiglio, Jul 1, 2026) · BCG, Is AI Computing Power Becoming a Commodity? (Belt, Thomas, Jul 23, 2026) · Gartner, AI inference costs per agentic workflow will increase more than fivefold through 2028 (Aug 17, 2026) · EY, Agentic AI token costs · CIO, The inference bill nobody budgeted for (Apr 28, 2026)
- FinOps practice: FinOps Foundation, Tokenomics: managing AI value in SaaS model token costs (working group, updated Jun 3, 2026) · State of FinOps 2026 · CloudZero, AI gross margin (Aug 31, 2026)
- Academic and measurement: Yicai Xing, AI Token Futures Market (arXiv 2603.21690, Mar 2026) · Bandi & Su, (Early) AI Compute Asset Pricing (arXiv 2607.12156) · Appenzeller, a16z, LLMflation (Nov 12, 2024)
- Repricing events and contract practice: Cursor, Clarifying our pricing (Jul 4, 2025) · BleepingComputer, Anthropic is cutting Claude Code’s current weekly limits by 17% (Aug 29, 2026) · Akhil Gupta, Should AI companies offer price protection clauses?
- Kinetic Alpha: The Compute Crack Spread · The Other Side of the Spark Spread · The compute risk premium · Five Indices, One Price · Token Price Index