Advisory research · Compute · Inference economics · 2 August 2026

Hedging the inference book

An inference provider — the Perplexity and Poe class of business, and every AI-native application above a certain scale — runs a book that would look familiar to any commodity retailer: fixed-price revenue downstream, floating-price input upstream. Subscriptions clear at $20 or $200 a month regardless of what the tokens underneath them cost; the tokens are bought at whatever the model APIs and the GPU rental market charge that week. Two instrument families now exist to hedge that book, and they settle in different units — token forwards in dollars per million tokens, compute futures in dollars per GPU-hour. The spread between them is the inference-efficiency curve, and deciding how much of each to hold is a genuine optimization, worked through here end to end.

This extends the practice's hedge-program work — the specialty-lender design — one layer up the stack, and draws on the index methodology catalogue, the implied forward curve work, and the measured GPU power findings that make the efficiency basis quantifiable at all.

2 units
$/M-token vs $/GPU-hour
the same cost base quotes in both — the hedge decision is which unit, in what mix
91%
variance reduction, optimized two-leg hedge
vs 57% for GPU futures alone — the worked example below
>90%
inference cost decline, 2024–2026
the short leg has been a tailwind — which is precisely why the risk is asymmetric now
$234k
monthly cost base, worked example
100B tokens/month: 77% routed through APIs, 23% self-hosted on rented H100s

01The book, decomposed

Before instruments: what an inference provider is actually long and short. The answer differs by architecture, and the hedge follows the architecture.

The revenue side is mostly fixed-price. Perplexity's public ladder runs six SKUs — free, $20 Pro, $200 Max, $40 and $325 enterprise seats — and Poe sells points bundles; in both cases the subscriber pays a flat rate for usage the provider cannot perfectly meter in advance. The API side (usage-priced resale) passes some cost through, but the flagship consumer products are, in commodity terms, full-requirements contracts sold at a fixed tariff.

The cost side floats, through two different channels:

The two upstream architectures. Most providers at scale run both — routed for frontier quality, self-hosted for volume — which is exactly what makes the hedge a mix rather than a single instrument.
ChannelWhat floatsNatural hedge unitReference prices today
Routed — tokens bought from model APIs (OpenAI, Anthropic, Sonar-class ladders)The $/M-token rate card, repriced at the vendor's discretion; $1–15/M across current laddersToken forward, $/M-tokenOrnn's token price indices (realized $/M-token, by model family); Silicon Data's LLM token expenditure index
Self-hosted — open-weight models served on rented GPU capacityThe GPU rental rate, plus the power underneath itCompute future, $/GPU-hrKalshi ladders (live), ICE/OCPI and CME/Silicon Data futures (pending), Compute Exchange forward auctions, AX perps
The trend is the trap

Inference costs fell more than 90% from 2024 to 2026, so the unhedged short has been a profitable position — every quarter the tokens got cheaper under fixed subscriptions. That history is why almost nobody in this sector hedges, and it is exactly backwards as risk logic: the position is short a price that has been falling, which means the book's loss scenario is the one nobody has lived through — a capacity crunch, a new model cycle that resets rate cards upward, or an export-control shock repricing the GPU layer. Hedging an inference book is not betting the trend breaks; it is refusing to be the counterparty who funds it if it does.

02Two units, one exposure — and the basis between them

A token forward hedges the bill. A compute future hedges the machine. The difference between them is a real, drifting, hedgeable-in-neither quantity: tokens per GPU-hour.

The token forward is the exact unit of the routed channel: it settles in $/M-token against a realized-price index, so the hedge and the invoice move together up to the index basis (the provider's model mix versus the index's — real, but second-order). The instrument class is young: Ornn publishes realized token price indices, the FalconX OTC precedent shows forwards on Ornn indices are executable, Compute Exchange runs forward auctions on capacity, and contract-design work on listed token futures is active. Nothing is deep yet. That matters for sizing, not for structure.

The compute future is the exact unit of the self-hosted channel's dominant line — but between the GPU-hour and the token sits the efficiency term: tokens served per GPU-hour, which improves with batching software, quantization, and load. Our measured-power work found its energy-side twin directly in the NLR traces: as request rate rises, power flattens while throughput keeps climbing, so energy per token falls even when the machine's draw barely moves. The same shape holds for cost per token. Hedging a token exposure with a GPU instrument therefore leaves the hedger short the efficiency curve — if serving efficiency jumps 40% on a software release, the GPU hedge is suddenly oversized against a cost base that just shrank.

cost = R·T·ptok  +  (1−R)·T·(g / E)·k R = routed share of tokens, T = tokens served, ptok = $/M-token, g = $/GPU-hr, E = tokens per GPU-hour, k = self-host overhead multiplier (power, idle capacity, orchestration). The token forward hedges ptok; the compute future hedges g; nothing listed hedges E — it enters every GPU-side hedge as basis.

03The worked example — evaluating the hedge

A Perplexity-class book, in round numbers, with every assumption stated.

Base case. Prices anchored to current public levels — Sonar-class API ladders, OCPI-region H100 spot near $1.70/hr — with the self-host line carried at a 2× all-in multiplier for power, idle, and orchestration per our facility-side work.
ParameterValueNote
Tokens served100B / monthconsumer + API combined
Routed share R60%frontier-quality traffic on vendor APIs
Blended routed rate$3.00 / Mmid-ladder blend of input/output pricing
H100 rental$1.70 / hr× 2.0 all-in → effective $1.36 / M at 2.5M tokens/GPU-hr
Monthly cost base$234krouted $180k (77%) + self-hosted $54k (23%); ≈ $2.8M/yr
Vols (annualized)σtok 30% · σgpu 35% · σeff 20% · σbasis 10%stated, not fitted — the module lets you move them
Token–GPU correlation ρ0.60both load on the same scarcity factor, imperfectly

Unhedged, the cost base carries roughly 29.0% annualized volatility — on a $2.8M annual spend, a one-sigma year moves the bill by more than most inference businesses' entire operating margin. Now evaluate the three candidate programs the instrument set allows:

Variance-minimizing hedge ratios and effectiveness, base-case parameters. Ratios are notional as a share of the total cost base.
ProgramHedge ratiosVariance reductionWhat's left
GPU futures only
the liquid leg — Kalshi now, ICE/CME at listing
63% of cost base in GPU forwards57%The routed leg is hedged only through ρ, and the efficiency basis rides on everything — the cross-hedge oversizes the GPU position to reach the token exposure and picks up efficiency risk doing it
Token forwards only
the exact leg — thin, partly OTC
~69% in token forwards81%The self-hosted leg unhedged; index (model-mix) basis on the hedged leg
Optimized mix69% token forwards + 23% GPU futures91%Efficiency risk on the self-hosted slice and token-index basis — the two residuals no listed instrument reaches
The result generalizes

Match units first, optimize second. The mix wins not because of a clever weighting but because each leg is hedged in its own unit: token forwards sized to the routed share, GPU futures sized to the self-hosted share. Every attempt to reach the whole book with one instrument converts unit mismatch into basis risk — and the largest single residual in every program is the efficiency term, which argues for hedging less than fully on the GPU leg and treating serving-efficiency improvements as the book's natural, welcome hedge slippage.

The hedge optimizer

Move the assumptions; the optimizer re-solves the variance-minimizing ratios and effectiveness live. Routed share shifts the book between units; correlation decides how well the liquid GPU leg substitutes for the thin token leg.

Routed share of tokens
Token–GPU correlation ρ
Efficiency vol σeff
token-forward ratio (share of cost base)
GPU-future ratio
variance reduction, optimized mix
GPU-only alternative, for comparison

04What breaks the hedge

Four things the optimizer cannot see, in the order they will actually bite.


05The program, in one paragraph

Decompose the book by channel and hedge each in its own unit: token forwards against the routed share, GPU futures against the self-hosted share, sized by the variance optimizer but shaded toward under-hedging on the GPU leg — the efficiency curve is a tailwind you should let run. Prefer capped structures to fixed strips where the venue can quote them, because the exposure is asymmetric. Validate every reference index against your own vendor invoices, monthly, exactly as a lender validates against its servicing tape. And if any capacity is self-hosted, remember the book extends one commodity further down than tokens: the inference load shape is a power position, and it is the expensive-hours one.