Advisory research · Compute · Inference economics · 2 August 2026
An inference provider — the Perplexity and Poe class of business, and every AI-native application above a certain scale — runs a book that would look familiar to any commodity retailer: fixed-price revenue downstream, floating-price input upstream. Subscriptions clear at $20 or $200 a month regardless of what the tokens underneath them cost; the tokens are bought at whatever the model APIs and the GPU rental market charge that week. Two instrument families now exist to hedge that book, and they settle in different units — token forwards in dollars per million tokens, compute futures in dollars per GPU-hour. The spread between them is the inference-efficiency curve, and deciding how much of each to hold is a genuine optimization, worked through here end to end.
This extends the practice's hedge-program work — the specialty-lender design — one layer up the stack, and draws on the index methodology catalogue, the implied forward curve work, and the measured GPU power findings that make the efficiency basis quantifiable at all.
Before instruments: what an inference provider is actually long and short. The answer differs by architecture, and the hedge follows the architecture.
The revenue side is mostly fixed-price. Perplexity's public ladder runs six SKUs — free, $20 Pro, $200 Max, $40 and $325 enterprise seats — and Poe sells points bundles; in both cases the subscriber pays a flat rate for usage the provider cannot perfectly meter in advance. The API side (usage-priced resale) passes some cost through, but the flagship consumer products are, in commodity terms, full-requirements contracts sold at a fixed tariff.
The cost side floats, through two different channels:
| Channel | What floats | Natural hedge unit | Reference prices today |
|---|---|---|---|
| Routed — tokens bought from model APIs (OpenAI, Anthropic, Sonar-class ladders) | The $/M-token rate card, repriced at the vendor's discretion; $1–15/M across current ladders | Token forward, $/M-token | Ornn's token price indices (realized $/M-token, by model family); Silicon Data's LLM token expenditure index |
| Self-hosted — open-weight models served on rented GPU capacity | The GPU rental rate, plus the power underneath it | Compute future, $/GPU-hr | Kalshi ladders (live), ICE/OCPI and CME/Silicon Data futures (pending), Compute Exchange forward auctions, AX perps |
Inference costs fell more than 90% from 2024 to 2026, so the unhedged short has been a profitable position — every quarter the tokens got cheaper under fixed subscriptions. That history is why almost nobody in this sector hedges, and it is exactly backwards as risk logic: the position is short a price that has been falling, which means the book's loss scenario is the one nobody has lived through — a capacity crunch, a new model cycle that resets rate cards upward, or an export-control shock repricing the GPU layer. Hedging an inference book is not betting the trend breaks; it is refusing to be the counterparty who funds it if it does.
A token forward hedges the bill. A compute future hedges the machine. The difference between them is a real, drifting, hedgeable-in-neither quantity: tokens per GPU-hour.
The token forward is the exact unit of the routed channel: it settles in $/M-token against a realized-price index, so the hedge and the invoice move together up to the index basis (the provider's model mix versus the index's — real, but second-order). The instrument class is young: Ornn publishes realized token price indices, the FalconX OTC precedent shows forwards on Ornn indices are executable, Compute Exchange runs forward auctions on capacity, and contract-design work on listed token futures is active. Nothing is deep yet. That matters for sizing, not for structure.
The compute future is the exact unit of the self-hosted channel's dominant line — but between the GPU-hour and the token sits the efficiency term: tokens served per GPU-hour, which improves with batching software, quantization, and load. Our measured-power work found its energy-side twin directly in the NLR traces: as request rate rises, power flattens while throughput keeps climbing, so energy per token falls even when the machine's draw barely moves. The same shape holds for cost per token. Hedging a token exposure with a GPU instrument therefore leaves the hedger short the efficiency curve — if serving efficiency jumps 40% on a software release, the GPU hedge is suddenly oversized against a cost base that just shrank.
A Perplexity-class book, in round numbers, with every assumption stated.
| Parameter | Value | Note |
|---|---|---|
| Tokens served | 100B / month | consumer + API combined |
| Routed share R | 60% | frontier-quality traffic on vendor APIs |
| Blended routed rate | $3.00 / M | mid-ladder blend of input/output pricing |
| H100 rental | $1.70 / hr | × 2.0 all-in → effective $1.36 / M at 2.5M tokens/GPU-hr |
| Monthly cost base | $234k | routed $180k (77%) + self-hosted $54k (23%); ≈ $2.8M/yr |
| Vols (annualized) | σtok 30% · σgpu 35% · σeff 20% · σbasis 10% | stated, not fitted — the module lets you move them |
| Token–GPU correlation ρ | 0.60 | both load on the same scarcity factor, imperfectly |
Unhedged, the cost base carries roughly 29.0% annualized volatility — on a $2.8M annual spend, a one-sigma year moves the bill by more than most inference businesses' entire operating margin. Now evaluate the three candidate programs the instrument set allows:
| Program | Hedge ratios | Variance reduction | What's left |
|---|---|---|---|
| GPU futures only the liquid leg — Kalshi now, ICE/CME at listing | 63% of cost base in GPU forwards | 57% | The routed leg is hedged only through ρ, and the efficiency basis rides on everything — the cross-hedge oversizes the GPU position to reach the token exposure and picks up efficiency risk doing it |
| Token forwards only the exact leg — thin, partly OTC | ~69% in token forwards | 81% | The self-hosted leg unhedged; index (model-mix) basis on the hedged leg |
| Optimized mix | 69% token forwards + 23% GPU futures | 91% | Efficiency risk on the self-hosted slice and token-index basis — the two residuals no listed instrument reaches |
Match units first, optimize second. The mix wins not because of a clever weighting but because each leg is hedged in its own unit: token forwards sized to the routed share, GPU futures sized to the self-hosted share. Every attempt to reach the whole book with one instrument converts unit mismatch into basis risk — and the largest single residual in every program is the efficiency term, which argues for hedging less than fully on the GPU leg and treating serving-efficiency improvements as the book's natural, welcome hedge slippage.
The hedge optimizer
Move the assumptions; the optimizer re-solves the variance-minimizing ratios and effectiveness live. Routed share shifts the book between units; correlation decides how well the liquid GPU leg substitutes for the thin token leg.
Four things the optimizer cannot see, in the order they will actually bite.
Decompose the book by channel and hedge each in its own unit: token forwards against the routed share, GPU futures against the self-hosted share, sized by the variance optimizer but shaded toward under-hedging on the GPU leg — the efficiency curve is a tailwind you should let run. Prefer capped structures to fixed strips where the venue can quote them, because the exposure is asymmetric. Validate every reference index against your own vendor invoices, monthly, exactly as a lender validates against its servicing tape. And if any capacity is self-hosted, remember the book extends one commodity further down than tokens: the inference load shape is a power position, and it is the expensive-hours one.