KINETIC ALPHA
Research · Compute Markets
Compute Markets · Benchmark Design · Index Construction

The Compute Crack Spread

The compute futures that list on October 5 settle on the one price in this complex that does not float. The price that does float — what a token sells for, on a hundred-odd endpoints, repriced in hours — has no index at all. We build the series that should exist: reference token revenue per GPU-hour minus the rent, and its quotient, an implied inference heat rate, from a live order book and a published throughput sweep. The result reads like a refining margin, behaves like one, and points to where the first index-linked physical compute contract will actually come from — the token side, not the rent side. It also answers, with disclosed and executable prices, most of what the CFTC asked the market on August 19.

Market heat rate at the median active endpoint
3.57M
Tokens per GPU-hour needed to cover $2.53 of H100 rent at $0.708/M blended
Coverage at saturation
3.1×
A saturated H100 serving a 70B model produces 11.16M tokens per GPU-hour against a 3.57M requirement — the crack is +$5.37
Dispersion across active endpoints
1.8 → 8.8M
Five-fold, for one model, on the same screen, at zero search cost
Physical contracts referencing any rental index
0
Token-indexed revenue, by contrast, already exists at every inference host — involuntarily
01 · The premise

Index the leg that floats

Our companion piece on the October listing went looking for the balance sheet that carries a floating exposure to a GPU-hour and found that every commercial participant with a real exposure prices it fixed: the neocloud sells two-to-five-year take-or-pay, the lender sizes term debt against that contracted revenue and excludes the merchant tail, the buyer signs the same contract from the other side. No lease resets to Silicon Data, no covenant references Ornn, no ABS trigger keys to an OCPI print. The listing therefore runs the benchmark film backwards: oil and gas got deep because the physical market floated off a screen first — PEMEX formula pricing in 1986, Henry Hub spot liquidity before the 1990 contract — and compute is being asked to list first and hope the physical market reprices to it.

There is, however, a part of the compute complex that already runs the film forwards. The token market is unbundled, transaction-based, continuously repriced and disclosed at the endpoint level. OpenRouter routes on the order of 25 trillion tokens a week across 400-plus models and roughly a hundred providers, publishes a per-endpoint order book for every model — price, context, quantization, uptime — routes stochastically with weights proportional to the inverse square of price, and reprices in hours (we mapped that microstructure in August). That is a cash market in exactly the sense the contract-success literature requires: Brorsen and Fofana found an active cash market “perfectly predicts” whether a commodity gets a futures market.

So the administrators are indexing the leg that does not float and ignoring the leg that does. This is not a complaint about SKU-specific on-demand rent as a settlement object — it is the cleanest thing to settle on, and we have said so. It is a question about what the index is for. A take-or-pay buyer has no exposure to the on-demand rate. It has a fixed cost and a floating output. Its renewal price is negotiated off its own token economics; its lender’s exposure is to whether those token economics cover the fixed rent — the contracted project converts technology deflation into tenant credit, and tenant credit for an inference tenant is its token margin. The variable the whole physical market actually faces is the spread between the fixed compute price it signed and the floating token price it sells into. The rent index is a proxy for that spread, and the numbers below say it is a poor one.

The Commission asked this question five weeks before the listing

On August 19 the CFTC issued a request for comment on the listing of compute derivatives (PR 9286-26; 91 FR 54259, RIN 3038-AF77), saying it “preliminarily believes” compute “may not yet exhibit certain of the characteristics of commodities that typically underlie a commodity derivatives market, including fungibility, standardization, and sufficient liquidity,” and that “price formation primarily occurs in opaque bilateral transactions.” Comments close October 20 — fifteen days after the CME contracts are intended to begin trading.

Every one of those objections is true of the rent leg and false of the token leg. That asymmetry is the argument of this piece, and §05 maps the Commission’s questions one by one against what each series can actually answer.

02 · Construction

A crack spread, a heat rate, and why refining is the right analogy

A refiner does not hedge crude and hedge gasoline; it hedges the crack — the margin of converting one into the other at reference yields, the 3:2:1 being the convention. A merchant generator does not hedge power and hedge gas; it hedges the spark spread, power minus gas times heat rate, and the market heat rate (power over gas) tells it which units run. The compute version was set out in our spark-spread work as a three-leg chain with two heat rates: power to GPU-hour through a rate fixed by physics, GPU-hour to token through a rate fixed by nothing but utilization. Power is 2.7% of the GPU-hour price. The interesting conversion is the second one.

Definitions

Let p be the blended token price in dollars per million tokens at a reference input share (we use 0.422, the NLR trace’s own 0.73:1 input-to-output mix), T the reference throughput in millions of tokens per GPU-hour for the graded model on the graded SKU, and r the rent in dollars per GPU-hour. Then the compute crack spread is C = T·p − r, in dollars per GPU-hour; the market heat rate is H* = r / p, the tokens per GPU-hour a fleet must produce to break even at that venue’s price; and the coverage ratio is T / H*. A host whose fleet runs above H* “runs,” exactly as a CCGT more efficient than the market heat rate runs. Throughput enters the crack as a reference yield, not as a measurement of anyone’s fleet — that is what makes it a standard rather than telemetry.

Three inputs, three sources, and each is already published somewhere. The token price comes from the endpoint book — not an expenditure-weighted aggregate like Silicon Data’s SDLLMTK, which by its author’s own description is a usage-weighted average “irrespective of models” and falls when users switch models rather than when prices fall, but a grade-standardised price for a named model class. The reference throughput comes from a published benchmark sweep, held constant across venues so the series isolates price. The rent comes from the index the futures settle on. For the worked series we use Llama-3.3-70B-Instruct, because it is the one model class for which a deposited, reproducible throughput sweep exists on H100s (Llama-3.1-70B, vLLM, 1,024 runs: 2.88M total tokens per GPU-hour at 10 requests per second, 9.87M at 50, a median of 11.16M saturated), and because thirteen endpoints served it on the day of the pull.

03 · The series

From a live book

Figure 1 · Market heat rate by venue, Llama-3.3-70B-Instruct
Tokens per GPU-hour required to cover $2.53 of H100 rent at each endpoint’s blended price (0.422 input share). Achievable throughput from the NLR sweep: 2.88M at 10 req/s, 9.87M at 50 req/s, 11.16M saturated. Grey = flagged degraded by the router at the time of the pull.
DeepInfra fp8
11.14M · $0.227/M
Nebius fp8
8.84M · $0.286/M
Novita bf16
8.78M · $0.288/M
Parasail fp8
6.63M · $0.382/M
AkashML fp8
6.57M · $0.385/M
Crusoe bf16
4.69M · $0.539/M
Groq
3.59M · $0.706/M
SambaNova
3.56M · $0.710/M
CoreWeave fp16
3.56M · $0.710/M
Google (×2)
3.51M · $0.720/M
Together
2.43M · $1.040/M
Cloudflare fp8
1.77M · $1.426/M
Green clears at 50 req/s; amber clears only near saturation; red does not clear at any measured load. Grey bars are endpoints the router flagged (DeepInfra status −2, Nebius status −5).

Read the figure as a merit order. Two thirds of the active book prices tokens at a market heat rate of 3.5–3.6 million tokens per GPU-hour — SambaNova, Groq, CoreWeave, Google and their neighbours cluster within a few percent of one another at $0.71–0.72 per million blended — which a saturated H100 covers three times over. Cloudflare at $1.43 and Together at $1.04 sit at the top of the book with heat rates of 1.8 and 2.4 million, the kind of price that clears even at 10 requests per second and that survives, on OpenRouter’s routing rule, only because buyers direct orders to them. At the bottom, AkashML and Parasail need 6.6 million; Novita, the cheapest active endpoint, needs 8.8 million; and the two degraded endpoints price at 11.1 and 8.8 million — DeepInfra at $0.227 blended sits almost exactly on the saturated median throughput of the sweep. The bottom of the book is priced at the marginal cost of a saturated H100, and the venues there are the ones the router is currently flagging for uptime. That is what a market clearing at its own heat rate looks like.

The book, priced
Crack in $/GPU-hour at the three NLR load levels, against $2.53 of rent
VenueInput $/MOutput $/MBlended $/MQuant.Uptime 30mH*, M tok/GPU-hrCrack, sat.Crack, 50 req/sCrack, 10 req/s
DeepInfra0.1000.3200.227fp892% (status −2)11.14+0.01−0.29−1.88
Nebius0.1300.4000.286fp852% (status −5)8.84+0.66+0.29−1.71
Novita0.1350.4000.288bf1697%8.78+0.69+0.31−1.70
Parasail0.2200.5000.382fp895%6.63+1.73+1.24−1.43
AkashML0.2000.5200.385fp898%6.57+1.77+1.27−1.42
Crusoe0.2500.7500.539bf16100%4.69+3.49+2.79−0.98
Groq0.5900.7900.706unk100%3.59+5.34+4.43−0.50
SambaNova0.4500.9000.710unk100%3.56+5.39+4.48−0.48
CoreWeave0.7100.7100.710fp1699%3.56+5.39+4.48−0.49
Google (×2)0.7200.7200.720unk100%3.51+5.51+4.58−0.46
Together1.0401.0401.040unk98%2.43+9.08+7.73+0.47
Cloudflare0.2932.2531.426fp899%1.77+13.38+11.54+1.58
Source: OpenRouter public endpoints API for meta-llama/llama-3.3-70b-instruct, August 21, 2026 (Google appears twice at identical price; shown once). Rent $2.53 = Silicon Data SDH100RT product-page value, same day. Quantization is self-declared by the venue and unverified; the sweep’s precision is not stated in the paper text, which is one of the grade problems discussed in §05.

Three features of the series are the findings. First, the crack is large and positive for almost everyone at saturation — a median +$5.37 per GPU-hour against $2.53 of rent — and negative for almost everyone at 10 requests per second. The sign of a host’s P&L is set by utilization, not by price, which is the conclusion our spark-spread work reached from the cost side and is now visible from the revenue side. Second, the dispersion is five-fold across active venues for an identical declared good on one screen, a spread Stigler’s search-cost model cannot explain; the residual is unverified quality (fp8 versus bf16 versus “unknown”), capacity, and deliberately non-best-execution routing. A crack-spread series has to carry a grade, or it carries that dispersion as noise. Third — and this is the point of the exercise — the series moves with the hedger. Rent has been nearly still: seven-day changes under 1% across Silicon Data’s SKUs, the hyperscaler series unchanged on 39 of 50 trading days. The token leg repriced 5× in a day when OpenAI cut to its Luna tier against a disclosed 1.44× efficiency gain, DeepSeek cut V4-Pro 75%, Tencent raised Hunyuan up to 400%. A hedger who locks rent on GPU1 has hedged the quiet leg.

04 · Deflation

What the spread does as both legs fall, and who is short it

The crack is a difference of two deflating series, and their deflation rates are not the same. Rent deflates on the technology curve — ln 2 over a 2.5-year performance-per-dollar doubling, about 27.7% a year, the assumption we have used throughout. Token prices for a fixed model class deflate faster and more erratically, because they carry the technology curve and the host-layer margin compression that a 1/p² router enforces. Holding the reference throughput fixed — a constant-grade series — the surface below shows the saturated crack one year ahead.

Figure 2 · Saturated crack one year ahead, $/GPU-hour
From today’s median: $0.708/M blended, $2.53 rent, 11.16M tokens per GPU-hour
Token −0%Token −20%Token −40%Token −60%Token −80%
Rent −0%+5.37+3.79+2.21+0.63−0.95
Rent −15%+5.75+4.17+2.59+1.01−0.57
Rent −27% (one year of ln2/2.5)+6.07+4.49+2.91+1.33−0.25
Rent −40%+6.38+4.80+3.22+1.64+0.06
Rows: rent decline over the year; columns: token-price decline. The crack stays positive at saturation unless token prices fall more than ~70% while rent barely moves — the Luna-style event. Every cell is at saturation; at 10 req/s the whole surface is negative.

Who is short this spread — who loses when token prices fall faster than rent? Any balance sheet that has bought compute forward at a fixed price and sells tokens at a floating one: the inference host with a take-or-pay, the lab that signed a five-year cluster lease and prices its API against competitors, and, one step removed, the lender to either. The companion piece ranked participants by who needs a hedge and found no natural two-sided flat-price demand. The crack is where the two-sided demand is. The host is short the crack by construction. The neocloud that sells take-or-pay is long it — it has sold rent forward to someone whose ability to pay depends on the crack — and its lender is long it through tenant credit. A crack-spread instrument has a buyer and a seller on day one. A rent instrument has sellers and speculators.

05 · The Commission’s questions

Read the request for comment as a specification

The August 19 request for comment is usually read as a threat to the October listing. It is more useful read as a specification. The Commission is not asking whether compute should have a derivatives market — Chairman Pham’s cover line is that “America cannot win the AI race without a robust derivatives market for compute.” It is asking what a cash market has to look like before a contract can settle on it, and its questions are the same ones the contract-design literature has asked since Silber. Set them against the two candidate legs.

Figure 3 · The Commission’s objections, applied to each leg
Questions as posed in 91 FR 54259 §§1–2; assessment ours
What the RFC asksThe rent legThe token legVerdict
“Fungibility, standardization”A GPU-hour is not fungible across cluster quality, interconnect or tier; the tiers print 2.6–3× apartA token of a named model class at a declared precision is closer to fungible than a GPU-hour, but precision is self-declared — the grade has to be attested, not assertedBoth constrained
“Sufficient liquidity”No public volume at any tier; the exchange’s own cash-market exhibit is confidential~25 trillion tokens a week routed across ~100 providers, continuously; volumes per venue still unpublished, which is the honest gapBoth constrained
“Price formation primarily occurs in opaque bilateral transactions”True and central: two-to-five-year take-or-pay, bundled, negotiated, undisclosedFalse for tokens: prices are posted per endpoint, executable, and repriced in hours at zero search costRent leg fails
“What proportion of transactions occur at disclosed prices”Unknown, and unknowable from outside the administratorEffectively all routed volume transacts at a posted price — that is what the router isRent leg fails
“Dominant market participants may wield significant pricing power”Acute: hyperscalers set a posted rate the index samples and benefit from opacityReal but visible: concentration is observable endpoint by endpoint, and a 1/p² router mechanically punishes a lone price riseBoth constrained
“Manipulating a cash settlement index by adjusting a posted rate”The exposure the Commission names; a posted rate is precisely what SD samplesA venue can post a price it will not fill — the same attack, mitigable by uptime, status and minimum-fill tests the router already runsBoth constrained
The pattern is not that the token leg is clean — it has a real grade problem and no published volumes. It is that the token leg fails the Commission’s tests for fixable reasons, and the rent leg fails them for structural ones. Opaque bilateral price formation is a description of how compute is sold; it is not a description of how tokens are sold.

Which leaves an awkward implication for the comment file. The Commission has asked whether any cash price series satisfies its Appendix C criteria for a settlement object. On the rent side the honest answer is not yet, and not observably — the cash-market exhibit supporting the CME filing was submitted under separate cover with confidential treatment requested, so no commenter outside the exchange can evaluate it. On the token side the answer is the data exists and is public, but nobody has constructed the series. Those are very different failures, and only one of them is closed by anybody writing a comment letter.

06 · What an administrator would have to publish

Six requirements, four of them inherited from the rent-index critique

The series above is a demonstration, not a benchmark. For it to settle anything it needs the four things the rent indices have been criticised for lacking, and two more that are specific to tokens.

RequirementWhat the rent indices doWhat a crack-spread series must doDifficulty
Grade standardSKU-specific (H100, B200); cluster quality normalised with a proprietary, undisclosed frameworkA capability grade for the model class — the SIT-threshold construction from our token-index work (MMLU / HumanEval / GSM8K floors anchored to a reference model) plus a declared and attested precision (fp8 vs bf16 moves benchmarks by up to 16.6 points)Hard
Reference throughputNot applicableA published sweep per SKU per grade, reproducible, held constant within an index version and re-based on a governed schedule — the Brent model: add grades as the old one declinesModerate
Transaction basisSilicon Data: posted / quote-based; Ornn: invoice-verified prints, methodology unpublishedEndpoint prices are disclosed and executable; per-venue volumes are not published at any tier, so weighting is by count or by a governed panel, not by volume — and disclosed as suchModerate
Mix held constantn/aFixed input share per grade, so the series cannot fall because users switched models — the SDLLMTK failure modeEasy, and decisive
Manipulation surfaceA provider adjusting a posted rate — the Commission’s own questionA venue posting a price it will not fill; mitigated by uptime and status filters (the router already flags them) and a minimum-fill testModerate
Rent legThe settlement object itselfInherits SD-H100 or OCPI as published; the crack is long the token index and short the rent index, so the rent index’s restatement history (−4 to −7% on provider additions) passes straight throughInherited
The grade problem is the hard one and it is not hypothetical: precision is self-declared at four of the twelve endpoints in Figure 1 and simply “unknown” at five more.
07 · Adoption

The physical market can reprice on the token side instead

The adoption argument in the companion piece was that a benchmark gets used when either a dealer warehouses bilateral exposure and lays the residual off in futures (the WTI route) or the physical market reprices to the index so end users have something to hedge (the Henry Hub route). The second route looked closed for compute, because no operator will re-paper a five-year take-or-pay to float on Silicon Data — the contract exists precisely to remove that exposure, and the lender sized the loan against it.

But the physical market does not have to reprice on the rent side. It can reprice on the token side, and it already has: every inference host’s revenue floats on token prices today, involuntarily, through the router. The step from “my revenue floats on tokens” to “my compute cost has a component indexed to tokens” is small, and it is the step a toll takes.

The token-indexed toll

A host or lab contracts for capacity under a two-part tariff: a fixed capacity charge per GPU-month that covers the owner’s capital recovery and the lender’s debt service, and a conversion fee per GPU-hour that floats with a grade-standardised token price — rising when the crack widens, falling when it compresses. The owner (or a merchant standing between owner and host) is now long the token index and can sell it; the host has converted part of a fixed cost into a cost that moves with its revenue, which is the hedge it actually wanted. The lender’s senior claim sits on the fixed charge, which is exactly where it sits today.

This is the gas-era netback contract — crude priced off the products it became — and it is the first physical compute contract with an indexed element that anyone in this market has a reason to sign. Once it exists, the index it references is a benchmark in the only sense that matters: something physical settles on it.

From there the dealer route and the cash-market route converge. A merchant writing token-indexed tolls for hosts is long the crack; it lays the rent leg off on GPU1/GPU2 — the futures finally have a commercial short with a reason to be there — and carries the token leg until a token index is listable, warehousing it the way the gas marketers warehoused basis after Order 636. The rent futures get their dealer, the token index gets its physical reference, and the utilization that no administrator can see becomes the thing the toll’s conversion fee is paid on — which means someone has to meter it. Attested, independent metering, published under a methodology, is the other half of the product, and it is a business in its own right.

08 · Falsification

What would show this to be wrong

Three things. If the rent indices begin to move with the token market — if SDH100RT’s weekly changes start tracking the median endpoint price for the SKU’s reference model class — then the rent leg is carrying the information after all and the crack is redundant. If a major operator re-papers a term contract to float on a rent index before any host signs a token-indexed toll, adoption has arrived through the rent side and this piece has the order wrong. And if a disclosed-constituent, mix-held-constant token index with a published grade appears from an existing administrator, the “nobody has built it” claim is dead, and the remaining question is only who settles on it first. We would regard any of the three as good news for the market, and would say so.

Sources

Primary data and corpus references