KINETIC ALPHA
Research
Compute · Risk Frameworks · Market Structure

The Other Side of the Crack Spread

Every GPU-compute derivative announced this year is built for the supply chain. The buyer whose bill is the reason to have a market does not rent GPUs, appears in no one’s analysis beyond a single noun, and carries exposures the contracts do not touch. This piece writes down what those exposures are, who is studying them (almost nobody), why the risk ends up with the enterprise, and what a buyer’s programme looks like. It is the demand-side companion to The Compute Crack Spread.

Enterprise end users named as hedgers in the CME and Ornn launch communications
None
“Enterprises” appears as a category; no firm is named.
Gartner: inference cost per agentic workflow through 2028
>5×
While per-token prices fall. The quantity term is the risk.
a16z: cost of a token at constant capability
÷10 / yr
The buyer is short a deflating commodity.
Frontier vs smallest model, same provider (FinOps WG)
50–100×
The routing decision is the price decision.
01 · Where the buyer is in the literature

Named in one sentence by the instrument builders, studied at length by people who never say “hedge”

In The Compute Crack Spread we wrote the producer’s margin as C = T·p − r: tokens per GPU-hour, times the price of a token, less the rent on the GPU. Everything that has happened since has been on the right-hand side of that equation. CME and Silicon Data list H100 and B200 rental-index futures on October 5, pending regulatory review, each contract “a month’s worth of rent.” ICE and Ornn, the GPU leg of Architect and Ornn, and Kalshi’s event contracts on Ornn prints all reference r. Bandi and Su estimated the first risk premium on r. Our own benchmark work is about how r is measured. The token price p has a research literature and no contract. And the term that is not in the producer’s equation at all, the quantity of tokens the buyer needs to get a job done, has neither.

The buyer does show up. CME’s release has Pete Keavey saying “AI builders and hyperscalers need to hedge as they grow,” and its body copy promising to allow “companies, including AI developers and hyperscalers, to lock in their costs”; Carmen Li’s own quote is that “together, that turns compute from something enterprises negotiate blindly into a market they can actually plan around.” Silicon Data’s cash-settlement note (Yu and Hou, August 24) names “AI developers and enterprises” on the demand side against “cloud providers and the broader compute supply chain” on the supply side. Dave Friedman’s market primer, the best independent map of the space, lists enterprises among the buyers of compute and then spends its analytical weight on fleet financing, observing that “the natural longs may be the more important source of sell-side hedging flow if they are financing fleets with debt.” His “longs” there are the fleet owners rather than the buyers — the reverse of the convention used here, and a taxonomy worth pinning down before the two literatures can talk to each other at all. Ornn, profiled by Axios in July, names lenders and “buyers and sellers of compute.” Nobody names a buyer.

The people who do study the buyer come from a different tradition. Figure 1 sorts them.

Figure 1 · Who is writing about the inference buyer, and what they leave out
Six camps, sorted by how close they get to the buyer’s actual exposure. As of the week of September 3, 2026.
CampRepresentative workWhat it says about the end userWhat it does not say
Exchanges and index providersCME/Silicon Data (Aug 11); ICE/Ornn; Architect/Ornn; Kalshi; Silicon Data’s cash-settlement note“AI developers and enterprises” will hedge rising rental costsWhich enterprises rent GPUs; what an API buyer’s exposure to a rental index is; basis
Independent market structureFriedman’s primer and basis-risk notes; Chen (remio, Aug 20) on the CFTC reviewEnterprises are among the buyers; Chen notes larger cloud companies “may already manage costs through long-term supply agreements” and that those private contracts “could reduce their need for exchange-traded hedges”No decomposition of the buyer’s exposure. Chen tells buyers to check the index against their own rental contracts, which presumes they have rental contracts
Strategy consultingBCG, Return on AI: How CFOs and CIOs Can Manage the Token Meter (Jul 1); BCG, Is AI Computing Power Becoming a Commodity? (Jul 23); Gartner “Inference Paradox” (Aug 17); EY on agentic token costsGovernance, attribution, routing, return on AI. BCG Jul 23 is the exception: separate baseload from flexible demand, throttle the flexible, and hedge AI compute “the same way they treat energy or foreign exchange exposure”Jul 1, Gartner and EY: no price risk, no contracts, no transfer. Jul 23: a paragraph, not an analysis, and its hedge is on rent
FinOps practiceFinOps Foundation token-economics working group (Jun 3); CloudZero on AI gross margin (Aug 31); State of FinOps 2026The best empirical picture of the buyer: spikes, tier spreads, commitment break-evens, margin compressionThe word “hedge” does not appear. Risk management means measurement and routing inside one purchasing relationship
AcademicXing, AI Token Futures Market (arXiv 2603.21690, Mar 2026); Bandi & Su (arXiv 2607.12156); Das et al. on cloud forward pricingXing is the only paper that treats application-layer AI companies as the primary buy-side hedgers, claiming a 62–78% reduction in cost volatility (standard deviation) against a token index it proposes itselfXing is a simulation with an assumed index and hedge efficiency = ρ² on an assumed correlation of 0.85; basis is not worked. Bandi & Su price rent, not tokens
Contract practiceGupta on price-protection clauses; vendor pricing pagesGupta proposes shared-savings mechanisms passing “typically 30-50%” of savings to customers, and a model-efficiency clause under which a version cutting inference cost by more than 20% passes “at least 50%” through; deprecation windows and caps (vendor terms)Not analysed as a risk book; no link to any market
Sources in the list at the end. “Does not say” is a reading of each document, not a criticism of its purpose; the consulting and FinOps pieces were not written to address risk transfer.

Read across the rows and the pattern is a gap, not a disagreement. The market people know the buyer exists and have not modelled it. The buyer’s own people have modelled it and do not know a market is coming. The one paper that connects the two is a simulation by an independent researcher. That is the whole competitive landscape on the demand side, as of this week.

02 · Whose crack spread is it

The producer is long the token price. The buyer is short it. The contracts are on neither.

Write the buyer’s side the same way we wrote the producer’s. A workflow, an agent, a support bot, a document pipeline, whatever the unit of business output is, consumes h tokens per outcome at a blended price p per token, and the enterprise needs N outcomes a month. Spend is S = p·h·N. If the output is sold on, the buyer’s own spread is M = V − p·h per outcome, with V the price the buyer charges its customer.

Figure 2 · Two spreads, one price, opposite signs
The producer’s equation from the crack spread piece, and the buyer’s written to match.
Producer · neocloud, inference host
C = T·p − r
Tokens per GPU-hour, times token price, less rent. Long p, long T, short r. Break-even heat rate H* = r/p.
Buyer · enterprise, AI application
M = V − p·h
Value per outcome, less tokens per outcome times token price. Short p, exposed to h. No r unless self-hosting.
ProducerBuyer
Token price pLong Revenue per GPU-hour is T·pShort Cost per outcome is p·h
Heat rateT, tokens per GPU-hour. Utilisation, batching, serving stack. Physics.h, tokens per outcome. Prompt design, agent loops, model generation. Not physics.
Capacity termr, the rent. The thing the futures settle on.None unless self-hosting. The API buyer pays p, not r.
VolumeFleet size, fixed on a lease.N, the business itself. Not a risk to hedge.
What the announced contracts coverr and, through H* = r/p, an implied token floor.Nothing except for the minority who rent.

The producer’s heat rate T is physics and utilisation; we measured it at 2.88 million tokens per GPU-hour idle to 11.16 million saturated on the NLR sweep, against a market heat rate H* = r/p of about 3.57 million at a $2.53 H100 rent and a $0.708 median Llama-3.3-70B price. The buyer’s heat rate h is not physics. It is the number of tokens a model generation, a prompt, an agent framework and a retry policy decide to spend on a task, and it is set by engineering choices the buyer only partly controls. Gartner’s number for what happens to it is the important one: cost per agentic workflow up more than fivefold through 2028 while per-token prices fall, because, in the words of Gartner’s Will Sommer, “each successive generation of AI capability will necessitate more, and often more expensive, tokens.”

So the sign structure is: the producer is long p and long T; the buyer is short p and long-exposed to h. A token price index would have natural counterparties on both sides, which is exactly what the CFTC’s request for comment asked for and exactly what a rental index lacks on the demand side, because an enterprise buying tokens from an API has no rent. In the crack spread piece we said the token leg fails the Commission’s tests for fixable reasons and the rent leg fails them structurally. The demand side is the reason the fixable one is worth fixing.

03 · The three terms, ranked by what moves the bill

Price is falling. Quantity is exploding. The discontinuities are contractual.

p is deflating. a16z measured the cost of a token at constant capability falling 10× a year: $60 per million for GPT-3 in November 2021 to $0.06 three years later for a 3-billion-parameter open-weight model scoring equivalently — a like-for-like capability comparison, not frontier-to-frontier. Xing’s figure for GPT-4-class inference is $60 to under $1.50 in two years. A buyer short a commodity with that drift does not fear the level of p. It fears two other things: a jump, when a provider re-prices or retires the model a workflow was tuned on; and the dispersion of p, which the FinOps working group puts at 50–100× between a provider’s frontier and smallest models. Every routing decision is a price decision, and routing is the operational lever every consultancy recommends, so the buyer’s effective p is a portfolio weight, not a market print.

h is the risk. The FinOps working group’s plainest sentence is that “token consumption can spike dramatically based on user behavior, prompt design, or application bugs.” CloudZero’s analysis says two accounts on identical plans generate order-of-magnitude different inference costs. Output tokens, priced 3–5× input, are set by the model at runtime, not by the buyer. And the trend is structural: an agentic reasoning workflow costs at least 5× a chatbot before task complexity grows. This is a quantity risk with the shape of a volumetric exposure in power: the buyer does not know how much it will consume, and the not-knowing is correlated with the thing it wants, a more capable agent.

N is the business. More outcomes means more revenue or more work done. No treasurer hedges demand for their own product. It belongs in the equation because it multiplies the other two.

The discontinuities are where money has actually been lost. Two examples, both from the last fifteen months, both from the top of the stack.

Cursor, June–July 2025

The Pro plan went from 500 requests a month to “$20 of frontier model usage per month at API pricing,” with “unlimited” reserved for the Auto tier and additional usage “at cost.” The company wrote on July 4 that “we didn’t handle this pricing rollout well, and we’re sorry,” and offered refunds for June 16 to July 4. Read as a risk event: an AI application company holding a fixed-price liability against a floating token cost re-priced to pass-through. Its customers became the holders of p and h.

Anthropic, August 29, 2026

Claude Code’s weekly limits on Pro, Max, Team and seat-based Enterprise plans change on September 14: 25% above the original baseline, 17% below the temporary level users had been running on. Not a price change. A quantity change on a fixed-price plan, which is the same thing to a buyer whose h just went up.

A stylised buyer makes the ranking concrete. The sliders are there so you can disagree with the base case.

Figure 3 · A stylised buyer: one million outcomes a month
Spend = p · h · N. Five shocks, one at a time, against the base case. Round numbers chosen for arithmetic.
Base case
Shocks
Base spend / month
Largest shock
Adverse market-price shock
Of the two rows that are market price moves, the adverse one, the provider re-price, is typically the smallest. The largest is the one Gartner is forecasting for everybody. The deprecation jump sits between them and is contractual, not market. A rental-index future addresses none of the five rows for a buyer who does not rent, and only the first, through basis, for one who does.
04 · What each term can be transferred to

Some of it is hedgeable. Most of it is manageable. The part that is neither, the buyer is paying for already.

Figure 4 · Buyer exposure → nature of risk → where it can go
ExposureNatureManaged byTransferred byAvailable today
p, levelMarket price, deflatingRouting, tier selection, caching, batch (50% discount)A token price index forward or cap, on baseload volumeNo No listed contract; OTC only in principle
p, jump (re-price, deprecation)Contractual, discontinuousMulti-provider architecture, model abstraction layerPrice-protection and model-efficiency clauses, deprecation notice periods, most-favoured pricingYes if negotiated. Gupta’s shared-savings and model-efficiency clauses are the template — prescriptive drafting, not observed contract data
h, tokens per outcomeEngineering quantity, trending upPrompt discipline, context limits, kill switches, a named owner per endpoint, a baseline before a budgetNot insurable in any current form. The nearest analogue is a quantity cap in the contractOperational only
Availability (rate limits, provisioned throughput)Capacity optionCommitted capacity, break-even modelledProvisioned throughput is the option; its premium is embedded in the priceYes and expensive: “for several current frontier and reasoning models, provisioned capacity costs more per token than pay-as-you-go even at 100% utilization” (FinOps WG)
r, rent (self-hosters only)Market priceTerm leasesCME/Silicon Data, ICE/Ornn, Architect futures; KalshiOct 5
N, volumeBusiness

Two rows deserve a second look.

The availability row is the buyer already paying an option premium without calling it one. When provisioned capacity costs more per token than pay-as-you-go at full utilisation, the difference is the price of a capacity-locking option, and the FinOps working group’s advice, “before pursuing commitments, model the break-even arithmetic explicitly,” is a request to price the option. BCG found the physical version of the same behaviour on the rental side: enterprise GPU clusters running at utilisation “as low as 5%” (a Cast.AI figure relayed by BCG) because undercommitting risks not having capacity later. That is a buyer over-hedging a quantity risk with idle silicon, at 95% carrying cost, because no cheaper instrument exists. It is the same reason a utility keeps peakers.

The h row is where the consultancies are right and the markets are silent. Tokens per outcome is a quantity the buyer’s own engineers set, so the first-order tools are internal: the FinOps working group’s 30–60 days of measurement before a budget, the CIO playbook’s named owner per endpoint and automated kill switches, Gartner’s tiering and orchestration. But h is also correlated with the buyer’s success, because a better agent spends more tokens, and it is correlated across buyers, because a new model generation moves everyone’s h at once. That second correlation is the one that eventually makes it a market: an “agentic upgrade” risk that every buyer holds on the same date is the raw material of an index.

05 · The price of certainty may be negative

The producer is paying to hedge. The buyer would be paid to take the other side. The buyer’s biggest risk is the term nobody would pay it for.

Bandi and Su report preliminary evidence “consistent with a positive compute risk premium, suggesting hedging pressure on the part of compute providers” — estimated from synthetic futures built out of term rental contracts, which the paper itself calls likely upper bounds on true futures prices: futures below expected spot because the sellers need to sell. The reason they need to, in our reading rather than theirs, is that the fleets are financed with debt and the lenders want forward revenue. That is Keynes’s normal backwardation, in GPU-hours: the side with the hedging need pays, and here the side with the need is the producer. If the same pressure holds on the token leg, and the fleets generating tokens are the same fleets, then a buyer who commits to baseload volume at a forward price is not paying for insurance. It is being paid to provide it.

That is an unusual position for a corporate hedger. In the commodities a treasurer already knows, the sign of the premium is contested and the buyer usually ends up paying for certainty, through option premium or a contango it has to roll: the airline fixing jet fuel, the utility fixing gas. The inference buyer, short a deflating commodity whose producers are long and levered, would be the seller of certainty in the token market, collecting the premium and taking a forward price below today’s spot. BCG’s July 23 instinct, treat AI compute like an energy or FX exposure and hedge it where material, arrives at the right desk with the wrong sign.

The buyer would be paid to hedge the term that is falling, and cannot hedge the term that is rising.

Every dollar of premium available on p is small against the exposure on h. Which is why, when the demand side does show up in these markets, it will not be to lock a price. It will be to write cheap floors for producers on baseload volume it was going to buy anyway, and to spend the proceeds on the engineering that controls h.

That is also why the risk ends where it does. Look at the chain: the host indexes its toll to tokens; the model provider prices per token; the application company, when the fixed-price plan blows up, re-prices to pass-through, which is exactly what Cursor did. Every layer above the enterprise has found a way to be short h to the layer below. The enterprise end user is the last holder of both p and h, has no customer to pass them to, and is the only participant in the chain without a contract, a desk, or an analyst.

06 · The buyer’s programme

Six things, most of which need no market

  1. Own it like fuel. The line is now material: 98% of State of FinOps respondents manage AI spend, up from 31% two years ago. Fuel desks at airlines sit between procurement, operations and treasury; the inference book needs the same three owners, with FinOps as operations.
  2. Measure the three terms per workflow, not the bill. p, h and N separately, with h baselined over the 30–60 days the working group recommends before any budget is set. Attribution is the difference between “the bill went up” and “h went up on the support agent after the model upgrade.”
  3. Split baseload from flexible. BCG’s frame is right even where its instrument is wrong. Baseload is the volume you will consume in any state of the world; flexible is what gets throttled when p or h moves. Only baseload is ever a hedging candidate.
  4. Negotiate the jumps. Model-efficiency clauses (a stated share of provider cost declines passes through), deprecation notice periods long enough to re-tune h on a successor, and a cap on re-pricing within the term. These are the only instruments that address the second-largest row in Figure 3, and they cost nothing but negotiating time.
  5. Price the capacity option you already own. Compute the provisioned-throughput break-even. If it never crosses at 100% utilisation, you are paying for availability, which may be right, but say so, and size it to baseload.
  6. When a token index exists, sell certainty on baseload and buy nothing on h. Write the floor the producer needs, collect the premium, and treat the proceeds as the budget for the h engineering in step 2. Do not buy a rental future unless you rent.
07 · What to watch

Four things that would change this piece

  • October 5. First CME prints. The question is not the price; it is whether any open interest is held by a firm that consumes tokens rather than selling GPU-hours. Large-trader categories will tell.
  • A token index with a specification. The demand side has natural counterparties only on p. Whoever publishes a model-class-stratified token price with a methodology that survives the RFC’s fungibility and disclosed-price tests owns the buyer’s leg.
  • The next “quantity change on a fixed-price plan.” September 14 is one. Each one moves h for a population of buyers on the same day, which is what an index needs and what a buyer’s programme has to survive.
  • A consultancy or a FinOps working group that uses the word “forward.” The first one to connect the measurement tradition to a market will have written the buyer’s side properly. On the evidence of Figure 1, it has not happened yet.
Sources

Releases, papers and practice notes