Advisory research · Compute · Inference economics · 2 August 2026

Hedging the inference book

An inference provider — the Perplexity and Poe class of business, and every AI-native application above a certain scale — runs a book that would look familiar to any commodity retailer: fixed-price revenue downstream, floating-price input upstream. Subscriptions clear at $20 or $200 a month regardless of what the tokens underneath them cost; the tokens are bought at whatever the model APIs and the GPU rental market charge that week. Two instrument families now exist to hedge that book, and they settle in different units — token forwards in dollars per million tokens, compute futures in dollars per GPU-hour. The spread between them is the inference-efficiency curve, and deciding how much of each to hold is a genuine optimization, worked through here end to end. Two complications sharpen the problem: licensed open-weight models — the Kimi K3 class, open weights with separate commercial licensing above revenue or scale thresholds — add a third, unhedgeable leg to the self-hosted channel; and the Silicon Data LLM Token Expenditure Index shows why a price hedge alone does not cap the bill.

This extends the practice's hedge-program work — the specialty-lender design — one layer up the stack, and draws on the index methodology catalogue, the implied forward curve work, and the measured GPU power findings that make the efficiency basis quantifiable at all.

2 units
$/M-token vs $/GPU-hour
the same cost base quotes in both — the hedge decision is which unit, in what mix
91%
variance reduction, optimized two-leg hedge
vs 57% for GPU futures alone — the worked example below
2× / −90%
token spend up, token price down
Silicon Data’s LLM Token Expenditure Index: spending doubled since 2025 while $/token fell >90% — the quantity problem in one index
$234k
monthly cost base, worked example
100B tokens/month: 77% routed through APIs, 23% self-hosted on rented H100s

01The book, decomposed

Before instruments: what an inference provider is actually long and short. The answer differs by architecture, and the hedge follows the architecture.

The revenue side is mostly fixed-price. Perplexity's public ladder runs six SKUs — free, $20 Pro, $200 Max, $40 and $325 enterprise seats — and Poe sells points bundles; in both cases the subscriber pays a flat rate for usage the provider cannot perfectly meter in advance. The API side (usage-priced resale) passes some cost through, but the flagship consumer products are, in commodity terms, full-requirements contracts sold at a fixed tariff.

The cost side floats, through two different channels:

The two upstream architectures. Most providers at scale run both — routed for frontier quality, self-hosted for volume — which is exactly what makes the hedge a mix rather than a single instrument.
ChannelWhat floatsNatural hedge unitReference prices today
Routed — tokens bought from model APIs (OpenAI, Anthropic, Sonar-class ladders)The $/M-token rate card, repriced at the vendor's discretion; $1–15/M across current laddersToken forward, $/M-tokenOrnn's token price indices (realized $/M-token, by model family); Silicon Data's LLM token expenditure index
Self-hosted, permissive — MIT/Apache-class open weights served on rented GPU capacityThe GPU rental rate, plus the power underneath itCompute future, $/GPU-hrKalshi ladders (live), ICE/OCPI and CME/Silicon Data futures (pending), Compute Exchange forward auctions, AX perps
Self-hosted, licensed open-weight — the Kimi K3 class: weights free to download and serve, commercial deployment licensed above revenue or scale thresholdsGPU rental and power as above, plus a contingent license cost that switches on when the business succeedsGPU leg hedgeable as above; the license leg has no instrumentNone — the license term is bilateral, private, and indexed to the provider's own revenue, not to any published price
The trend is the trap

Inference costs fell more than 90% from 2024 to 2026, so the unhedged short has been a profitable position — every quarter the tokens got cheaper under fixed subscriptions. That history is why almost nobody in this sector hedges, and it is exactly backwards as risk logic: the position is short a price that has been falling, which means the book's loss scenario is the one nobody has lived through — a capacity crunch, a new model cycle that resets rate cards upward, or an export-control shock repricing the GPU layer. Hedging an inference book is not betting the trend breaks; it is refusing to be the counterparty who funds it if it does.

02Two units, one exposure — and the basis between them

A token forward hedges the bill. A compute future hedges the machine. The difference between them is a real, drifting, hedgeable-in-neither quantity: tokens per GPU-hour.

The token forward is the exact unit of the routed channel: it settles in $/M-token against a realized-price index, so the hedge and the invoice move together up to the index basis (the provider's model mix versus the index's — real, but second-order). The instrument class is young: Ornn publishes realized token price indices, the FalconX OTC precedent shows forwards on Ornn indices are executable, Compute Exchange runs forward auctions on capacity, and contract-design work on listed token futures is active. Nothing is deep yet. That matters for sizing, not for structure.

The compute future is the exact unit of the self-hosted channel's dominant line — but between the GPU-hour and the token sits the efficiency term: tokens served per GPU-hour, which improves with batching software, quantization, and load. Our measured-power work found its energy-side twin directly in the NLR traces: as request rate rises, power flattens while throughput keeps climbing, so energy per token falls even when the machine's draw barely moves. The same shape holds for cost per token. Hedging a token exposure with a GPU instrument therefore leaves the hedger short the efficiency curve — if serving efficiency jumps 40% on a software release, the GPU hedge is suddenly oversized against a cost base that just shrank.

cost = R·T·ptok  +  (1−R)·T·(g / E)·k R = routed share of tokens, T = tokens served, ptok = $/M-token, g = $/GPU-hr, E = tokens per GPU-hour, k = self-host overhead multiplier (power, idle capacity, orchestration). The token forward hedges ptok; the compute future hedges g; nothing listed hedges E — it enters every GPU-side hedge as basis. Under a licensed open-weight model the self-host term gains a third component, L(revenue) — zero below the license threshold, contractual above it.

The licensing layer — what the Kimi K3 class changes

Licensed open-weight releases occupy a deliberate middle ground: anyone can download, self-host, fine-tune and experiment, but commercial deployment above stated revenue or scale thresholds requires a separate license. For hedge design this does three specific things. First, it adds a cost leg with no instrument: the license term is bilateral and private, indexed to the provider's own revenue rather than to any published price, so neither token forwards nor GPU futures reach it. Second, its structure is wrong-way by construction — the obligation switches on precisely when the provider scales, which is a short call on the provider's own success, written to the model creator. It is the same shape as the threshold-triggered royalty a commodity producer grants in a streaming deal, and it should be valued that way rather than carried at zero because today's usage sits below the threshold.

Third, and least obvious: it erodes the routing option. A provider running both channels holds a real option to shift traffic toward whichever is cheaper, and that option is itself the book's best hedge — it is what disciplines the API vendors' rate cards. Permissive open weights keep the self-host strike clean at g/E; a licensed model raises and blurs that strike, which lowers the option's value and, at the margin, restores pricing power to the closed-model vendors. A hedging program should therefore treat the license terms of its self-host stack as part of its risk inventory: threshold distance, repricing rights, and audit provisions belong on the same monitoring sheet as hedge ratios.

03The worked example — evaluating the hedge

A Perplexity-class book, in round numbers, with every assumption stated.

Base case. Prices anchored to current public levels — Sonar-class API ladders, OCPI-region H100 spot near $1.70/hr — with the self-host line carried at a 2× all-in multiplier for power, idle, and orchestration per our facility-side work. The example assumes permissively-licensed weights; a licensed open-weight stack adds the unhedgeable L(revenue) term of §02 on top of every figure below.
ParameterValueNote
Tokens served100B / monthconsumer + API combined
Routed share R60%frontier-quality traffic on vendor APIs
Blended routed rate$3.00 / Mmid-ladder blend of input/output pricing
H100 rental$1.70 / hr× 2.0 all-in → effective $1.36 / M at 2.5M tokens/GPU-hr
Monthly cost base$234krouted $180k (77%) + self-hosted $54k (23%); ≈ $2.8M/yr
Vols (annualized)σtok 30% · σgpu 35% · σeff 20% · σbasis 10%stated, not fitted — the module lets you move them
Token–GPU correlation ρ0.60both load on the same scarcity factor, imperfectly

Unhedged, the cost base carries roughly 29.0% annualized volatility — on a $2.8M annual spend, a one-sigma year moves the bill by more than most inference businesses' entire operating margin. Now evaluate the three candidate programs the instrument set allows:

Variance-minimizing hedge ratios and effectiveness, base-case parameters. Ratios are notional as a share of the total cost base.
ProgramHedge ratiosVariance reductionWhat's left
GPU futures only
the liquid leg — Kalshi now, ICE/CME at listing
63% of cost base in GPU forwards57%The routed leg is hedged only through ρ, and the efficiency basis rides on everything — the cross-hedge oversizes the GPU position to reach the token exposure and picks up efficiency risk doing it
Token forwards only
the exact leg — thin, partly OTC
~69% in token forwards81%The self-hosted leg unhedged; index (model-mix) basis on the hedged leg
Optimized mix69% token forwards + 23% GPU futures91%Efficiency risk on the self-hosted slice and token-index basis — the two residuals no listed instrument reaches
The result generalizes

Match units first, optimize second. The mix wins not because of a clever weighting but because each leg is hedged in its own unit: token forwards sized to the routed share, GPU futures sized to the self-hosted share. Every attempt to reach the whole book with one instrument converts unit mismatch into basis risk — and the largest single residual in every program is the efficiency term, which argues for hedging less than fully on the GPU leg and treating serving-efficiency improvements as the book's natural, welcome hedge slippage.

The hedge optimizer

Move the assumptions; the optimizer re-solves the variance-minimizing ratios and effectiveness live. Routed share shifts the book between units; correlation decides how well the liquid GPU leg substitutes for the thin token leg.

Routed share of tokens
Token–GPU correlation ρ
Efficiency vol σeff
token-forward ratio (share of cost base)
GPU-future ratio
variance reduction, optimized mix
GPU-only alternative, for comparison

04The other side of the book — end users, and the quantity problem

Everything above hedges price. The Silicon Data LLM Token Expenditure Index is the public evidence that price is not what is breaking corporate AI budgets.

The index tracks realized LLM spending, and its 2026 reading is the whole argument in one pair of numbers: spending has doubled since 2025 even though cost per token has fallen more than 90% over three years — corporate token bills are up roughly 320% since 2023, and the Atlanta Fed expects firms to raise AI spend another 50% in 2026, to $280B. The bill is P × Q, the hedgeable instruments all live on P, and Q has been growing faster than P has been falling. A perfectly executed price hedge would have capped none of that.

For the end user — the enterprise deploying rather than reselling inference — this reorders the program. The price legs above still matter, but the exposures that dominate a corporate AI budget are quantity and contract structure: usage growth, vendor rate-card resets, and single-provider dependence. The emerging corporate playbook — restructure contracts so model providers carry outcome risk rather than pure usage pricing, multi-source across vendors, hold an open-weight (increasingly a licensed open-weight) fallback — is hedging by procurement, and the licensed open-weight class is precisely what makes the fallback credible below threshold and priced above it.

The instrument gap this exposes

An expenditure index that doubles while its price index falls 90% is an advertisement for a contract nobody lists yet: the expenditure-linked structure — a cap or swap on realized token spend rather than token price, the direct analogue of the quantity-contingent products power markets built when they hit the same P×Q problem. Silicon Data's expenditure index is, structurally, the settlement benchmark such a contract needs; per our methodology catalogue, it would face the same disclosure bar as every other index in the complex before anything could settle on it. Until then, the honest statement to an end user is: the market can hedge your price; your quantity is a budgeting problem wearing a derivative's clothes.

05What breaks the hedge

Four things the optimizer cannot see, in the order they will actually bite.


06The program, in one paragraph

Decompose the book by channel and hedge each in its own unit: token forwards against the routed share, GPU futures against the self-hosted share, sized by the variance optimizer but shaded toward under-hedging on the GPU leg — the efficiency curve is a tailwind you should let run. Prefer capped structures to fixed strips where the venue can quote them, because the exposure is asymmetric. Validate every reference index against your own vendor invoices, monthly, exactly as a lender validates against its servicing tape. And if any capacity is self-hosted, remember the book extends one commodity further down than tokens: the inference load shape is a power position, and it is the expensive-hours one.