Advisory research · Compute · Inference economics · 2 August 2026

The inference spark spread

On 30 July a frontier lab cut the published price of one model tier by 80% in three weeks. Over the past year the neocloud GPU rental rate that tier is served on roughly doubled off its trough. Those two facts are in direct tension, and the tension has an exact form: it is a spark spread — the structure a power trader uses to price a generator's margin as the power price less the gas price times the heat rate. Written that way, the inference version reduces to a single unobservable term, tokens per GPU-hour, and the arithmetic inverts to a threshold: the serving efficiency a rate card requires in order to clear.

The heat rate is the whole difficulty. In power it is published per unit and regulated; in inference nobody publishes it. This piece computes it directly — 1,024 measured vLLM serving runs on Llama-3.1-70B across four H100s, joined sample-by-sample to the NVML power traces recorded alongside them, released under CC-BY by a national laboratory. That gives a measured denominator to put underneath the July rate cards. This is a required-efficiency bound computed from published prices. It is not an accounting of any company's margins, and nothing here is a claim about any firm's financial condition.

11.6M
measured tokens per GPU-hour
Llama-3.1-70B bf16, TP=4 on H100, saturated online serving — median of 1,024 runs
−80% / +115%
token price vs GPU rental
Luna card cut 30 Jul 2026; US East neocloud H100 ~$1.35 (Q3 2025 trough) → ~$2.90 (Q3 2026)
5.0×
rise in required efficiency
Luna break-even threshold, old vs new rate card, holding GPU price and overhead fixed
51 Wh
per million tokens, saturated
down from 175 Wh at light load — measured, and ~0.9% of the new Luna rate at $0.06/kWh

01The divergence, stated with verified numbers

Both legs move, in opposite directions, and both are tiered. The tiering is the first finding.

On 30 July 2026 OpenAI repriced the GPT-5.6 family, three weeks after its 9 July launch. The cut was not uniform. Luna fell 80%, from $1.00/$6.00 to $0.20/$1.20 per million input/output tokens. Terra fell 20%, from $2.50/$15.00 to $2.00/$12.00. Sol, the flagship, was left unchanged at $5.00/$30.00. Batch and Flex processing are reported at half standard price, putting Luna's batch tier at $0.10/$0.60.

The stated driver is competitive. A CNBC investigation found Chinese-origin models reaching a weekly peak of 46% of US enterprise token volume on OpenRouter, and at least 30% of it every week since 8 February 2026 — against 11% averaged over the prior twelve months and 4.5% in the first half of 2025. OpenRouter's own attribution is price: open-weight Chinese models running 60–90% cheaper than the US frontier — and at the extreme, well beyond that band: as of June 2026 DeepSeek V4 Flash listed at $0.14/M input against the then-current GPT-5.5 flagship at $5.00, a 97% gap. OpenAI framed the cut as passing through efficiency gains in its serving stack.

The demand side is moving the other way, which is what makes the divergence interesting rather than merely deflationary. Silicon Data's LLM Token Expenditure Index shows aggregate LLM spending has doubled since 2025 even as cost per token fell more than 90% over three years, and the Federal Reserve Bank of Atlanta expects firms to raise AI spend roughly 50% in 2026, to $280bn. Falling unit price with rising total spend is the P×Q problem the inference hedging piece works in its §04: every hedgeable instrument lives on price, while quantity outran it. A spark spread is a per-unit margin — it says nothing about volume, and volume is where the money has actually gone.

The tiering is informative on its own

80% at the bottom, 20% in the middle, zero at the top. Price pressure is concentrated exactly where open-weight and Chinese models compete on substitutable work, and absent where they do not. Because tiers run on different model sizes and different hardware, the spread has to be computed per tier; an aggregate blend of the three would average away the only place the arithmetic binds.

The compute leg is tiered too — and "GPU prices are rising" is only half true

The convenient framing is that GPU rentals rose while token prices fell. The verified picture is more specific, and the specificity matters because a lab's input cost depends on which tier it buys. Silicon Data's Q3 2026 regional cut shows hyperscaler H100 rates falling 26–54% by region since Q1 2025, while non-hyperscaler (neocloud) rates bottomed in mid-2025 and recovered strongly — US East from roughly $1.35 to about $2.90, the Nordics from about $1.75 to roughly $3.30, now above where the series started. Silicon Data computes the average premium across those seven regions as falling from roughly 250% to about 120% — a per-region average of quarterly medians, which is why the range midpoints in the table below do not reproduce it directly.

LegQ1 2025 / troughQ3 2026DirectionSource
H100, neocloud (US East)~$1.35/hr (Q3 2025)~$2.90/hrup ~115%Silicon Data H100 index, regional
H100, neocloud (range, 7 regions)$1.35–$2.70$2.00–$3.30up, dispersedSilicon Data, Q3 2026 partial
H100, hyperscaler (range)$7.25–$10.40$4.80–$6.10down 26–54%Silicon Data, Q3 2026 partial
B200, neocloud (SDB200RT index)4.40 (1 Jan 2026)5.48 (30 Mar 2026)up 24.4% YTD to MarSilicon Data SDB200RT; March mean price $5.09/hr, median $4.96
GPT-5.6 Luna, blended 3:1$2.25/M$0.45/Mdown 80%Published rate card, 30 Jul 2026
GPT-5.6 Terra, blended 3:1$5.63/M$4.50/Mdown 20%Published rate card, 30 Jul 2026
GPT-5.6 Sol, blended 3:1$11.25/M$11.25/MunchangedPublished rate card
Blended token rates use a 3:1 input:output mix throughout unless stated; sensitivity to that assumption is worked in §04. GPU figures are quarterly or monthly medians of quoted rental prices, not transacted contract prices — the same quote-basis limitation the practice's index methodology catalogue raises against every published GPU index.

So the divergence is real but it is a neocloud divergence. A lab renting marginal capacity from the specialist tier has watched its input price roughly double off the trough while its output price on one tier fell 80%. A lab buying hyperscaler on-demand has seen its input price fall — but no lab at scale buys marginal inference capacity there. The relevant leg for this analysis is the neocloud tier, and it is the leg that rose.

One reported figure we could not verify

The framing that H100 forwards are in contango — forward above spot, the market pricing a future crunch — is reported but we could not confirm a live forward quote against a contemporaneous spot from a public source. What the corpus does establish independently is a $1.23/GPU-hour convenience yield (23.6% of spot) between the financial forward and provider reserved-tier term sheets at the one-year point, measured in the implied forward curve work. Contango in GPU forwards is that same quantity seen from the other side. We use the measured convenience yield and leave the contango claim as reported rather than verified.


02The spark spread, constructed

The analogy is exact in form and breaks in one specific place. Both halves matter.

A gas generator's margin is the spark spread: the price it sells power at, less the price it buys gas at, scaled by how much gas it burns per unit of power. The scaling term is the heat rate, in MMBtu per MWh, and it is the whole reason the spread is tradeable — it is published per unit, it is an engineering property of the turbine, and it changes slowly.

spark spreadpower = Ppower − Pgas × HR Ppower in $/MWh · Pgas in $/MMBtu · HR (heat rate) in MMBtu/MWh

The inference version has the identical shape. A model lab sells tokens at a published $/M-token rate and buys GPU-hours at a rental rate. The scaling term is the number of tokens the stack produces per GPU-hour — which enters as a divisor rather than a multiplier, because it is expressed as output per input rather than input per output.

spark spreadinference = ptok − (g × k ÷ E) × 106 ptok = blended token price, $/M tokens · g = GPU rental rate, $/GPU-hr · E = tokens served per GPU-hour · k = self-host overhead multiplier (power, idle, orchestration, replication, KV-cache headroom)

This is the cost equation from the inference hedging piece — cost = R·T·ptok + (1−R)·T·(g/E)·k — restricted to the self-hosted channel and rearranged as a margin rather than a cost. E is the heat rate of inference. 1/E is tokens' "gas burn": GPU-hours consumed per token delivered.

Where the analogy breaks

PropertyPower: heat rateInference: tokens per GPU-hour
Determined byTurbine thermodynamics. A physical property of the plant.Software. Batching policy, quantization, speculative decoding, KV-cache management, scheduler, framework version.
StabilityDegrades slowly and predictably over the asset's life.Improves discontinuously with every serving-stack release. A framework upgrade can move it 30% overnight.
Load dependenceModest. Part-load heat rate is worse, and the curve is published.Severe and unpublished. Measured here: a 4× swing between light load and saturation on identical hardware.
DisclosurePublished per generating unit; used in dispatch, regulation and settlement.Published by nobody. No index administrator, no lab, no cloud provider publishes a serving-efficiency reference.
Consequence for the spreadSpark spread is directly observable and directly tradeable.Two of three terms are observable; the third is private. The spread is computable only where someone has measured E.

That last row is why this analysis exists here and essentially nowhere else. The measured GPU power work established that the denominator of any compute-normalised metric has to be measured rather than assumed — nameplate TDP overstates device energy by 12–21%, and duty factors run 0.795–0.880. The same discipline applies one layer up: the efficiency term in a token-versus-GPU spread has to be measured, and the measurement has to be stated with its load condition attached.


03The heat rate of inference, measured

1,024 vLLM serving runs, joined to 715,989 NVML power samples over a 19.9-hour window. Same hardware, same model, same weights — only the offered load changes.

The NLR generative-AI power dataset (Vercellino et al., arXiv:2604.07345, CC-BY-4.0) was published to support data-centre load modelling, and the practice used it previously for duty factors and capacity-factor derates. It contains something the power analysis did not need: the raw vLLM benchmark logs that ran alongside the wattmeter traces. Those logs carry total_token_throughput, total_input_tokens and total_output_tokens per run. Dividing throughput by the tensor-parallel width gives measured tokens per GPU-hour directly, with no modelling step.

Measurement basisValue
Model / precisionLlama-3.1-70B-Instruct, bfloat16, KV cache auto
Hardware / parallelism4 × H100 on one node, tensor_parallel_size = 4
Serving stackvLLM, OpenAI-compatible completions endpoint
WorkloadInstructCoder prompts, 1,000 per run, request rate swept 10 → 1,000 req/s
Observed prompt shapemean 175 input / 158 output tokens — a 1.1:1 input:output mix
Runs / power samples1,024 benchmark runs; 715,989 four-GPU NVML samples (GPU sockets only — CPU, DRAM and facility overhead excluded) across 19.9 h

Measured throughput, power and energy per token vs offered load

The finding the whole spread rests on: above roughly 50 req/s, GPU power flattens while throughput keeps climbing — so energy per token keeps falling even though the machine draws the same watts. Tokens per GPU-hour varies across this sweep on identical silicon.

Read. Four-GPU NVML draw rises 17% from light load to the saturation plateau (2,020 W → ~2,370 W), peaking at 2,384 W. Throughput over the same range rises 304% (2.88M → 11.63M tokens per GPU-hour). Energy per million tokens therefore falls 71%, from 175 Wh to 51 Wh. At $0.06/kWh and PUE 1.3 the saturated figure is about $0.0040 per million tokens — 0.9% of the new Luna blended rate. That is a floor: the wattmeter reads GPU sockets only, so CPU, DRAM, fans and PSU losses sit outside it and the true facility figure is higher. Even several times higher it does not change the conclusion — electricity is not the binding cost of inference at these token prices. The rental rate is.

The corpus assumption was conservative by about 4.6×

The inference hedging worked example used E = 2.5M tokens per GPU-hour. Against this measurement, 2.5M sits close to a lightly loaded server at 10 req/s, measured at 2.88M — not a saturated one. A well-utilised 70B-class deployment runs 11.6M at the median and 15.9M at the 90th percentile. That is not an error in the prior work — it is the point that piece made, that E is unhedgeable basis — but it materially changes the break-even arithmetic, and in the lab's favour.

What this measurement is and is not


04Break-even efficiency: what the rate card requires

Set the spread to zero and solve for E. The result is the number of tokens per GPU-hour a stack must produce for a given rate card to cover a given rental rate.

Ebreak-even = (g × k × 106) ÷ ptok tokens per GPU-hour required to break even · g in $/GPU-hr · ptok in $/M tokens · k dimensionless

This is the inference analogue of the market-implied heat rate a power trader computes as the power price divided by the gas price: the efficiency the market is implicitly demanding at prevailing prices. Cut the numerator's price and the required efficiency rises proportionally. The 80% Luna cut therefore raises the required efficiency by exactly , mechanically, holding the compute leg fixed.

Required efficiency by tier and GPU price

Tier / rate cardBlended 3:1 g = $2.00g = $2.50g = $2.90g = $3.30 vs measured 11.6M
Luna — old card ($1.00/$6.00)$2.25/M1.8M2.2M2.6M2.9Mclears 4.0–6.5×
Luna — new card ($0.20/$1.20)$0.45/M8.9M11.1M12.9M14.7M1.3× → 0.8×
Luna — batch/flex ($0.10/$0.60)$0.225/M17.8M22.2M25.8M29.3Mfar above
Terra — old card ($2.50/$15.00)$5.63/M0.71M0.89M1.03M1.17Mclears 9.9–16.4×
Terra — new card ($2.00/$12.00)$4.50/M0.89M1.11M1.29M1.47Mclears 7.9–13.1×
Sol — unchanged ($5.00/$30.00)$11.25/M0.36M0.44M0.52M0.59Mclears 19.8–31.5×
Required tokens per GPU-hour at k = 2.0 (self-host overhead multiplier from the inference hedging worked example: power, idle, orchestration, replication). GPU price range spans the Q3 2026 neocloud H100 band across the seven regions Silicon Data tracks ($2.00–$3.30). Coverage badges give the range across the four GPU-price columns (worst case at g = $3.30, best at g = $2.00). The measured comparison is the saturated median (11.63M) for a 70B-class dense model in bf16 on H100; smaller or more aggressively quantized models achieve multiples of it.

Luna's new card, across overhead and GPU price

Overhead kg = $2.00g = $2.50g = $2.90g = $3.30Interpretation at k
1.2 — near bare-metal, own orchestration5.3M6.7M7.7M8.8MClears on 70B-class throughput in every region
1.5 — lean operator6.7M8.3M9.7M11.0MClears at the median, thin at the top of the range
2.0 — the corpus assumption8.9M11.1M12.9M14.7MAt or above measured median across most of the range
2.5 — replication-heavy, latency-SLA serving11.1M13.9M16.1M18.3MAbove the measured median throughout; above p90 only at the top two rental rates
All at a 3:1 input:output blend ($0.45/M). Measured reference: 11.6M median, 15.9M p90, 16.6M maximum observed.

Sensitivity to the input:output blend

Input and output tokens carry different prices and different costs, and the blend assumption does real work. Luna's 6:1 price ratio between output and input means an output-heavy workload is better for the lab per token; an input-heavy one is worse on price but cheaper to serve.

Input:outputBlended $/MRequired E at g=$2.50, k=2.0Typical of
1:1$0.7007.1MChat turns, short prompts
2:1$0.5339.4MMixed API traffic
3:1$0.45011.1MBase case used throughout
5:1$0.36713.6MRAG, document QA
10:1$0.29117.2MLong-context agents, tool-calling chains
Heavier input mixes lower the blended price and raise the required efficiency — but they also raise achievable efficiency, since prefill is more compute-efficient per token than decode. The two effects offset by an amount this dataset cannot separate, because its prompt mix is fixed at 1.1:1. That unresolved offset is the largest single uncertainty in the calculation and we flag it rather than assume it away.

Break-even calculator — required efficiency against measured throughput

Set a rate card, a GPU rental rate, an overhead multiplier and a prompt mix. The calculator returns the required tokens per GPU-hour, and compares it against the measured 70B-class reference from §03. The bar is the requirement; the marker is what was measured.

Rate card
GPU rental rate, $/GPU-hr
Self-host overhead multiplier k
Input : output blend
Measured reference (70B dense, bf16, H100)


05What the answer means, in three cases

The arithmetic supports a conclusion about model class, not about margins. That is a narrower claim than the headline invites, and it is the one the data will carry.

CaseConditionReading
1 — Margin give-backRequired E comfortably below measured throughputThe cut surrenders margin from a very high base. Competitive, not existential. This is where the Terra and Sol rate cards sit against this denominator, and by a wide margin: Sol requires 0.44M at g=$2.50 against 11.63M measured — roughly 26× coverage, or 19.8× at the top of the rental range — even on a 70B-class dense denominator that is far more expensive to serve than Sol's own class is likely to be.
2 — Compression toward zeroRequired E near measured throughputThe rate card would price at or near cost if served on this denominator. This is where Luna's new card sits against a 70B-class dense denominator — 11.1M required at the mid-range rental rate against 11.63M measured, and the coverage flips from 1.31× to 0.79× across the rental range. It is a statement about the denominator, not about the tier's actual economics, which depend on a serving stack nobody outside the lab has seen.
3 — Below cost to serveRequired E far above measured throughputTokens sold below the compute cost of producing them. The arithmetic does not support placing any tier here on the evidence available, and we decline to assert it. The batch/flex tier's 22M requirement would qualify against a 70B denominator — but batch tiers exist precisely to be filled with deferrable work at high batch sizes, where achievable E is far above the online figure measured here.
What the calculation actually establishes

Read case 2 carefully. It does not say Luna's margins are near zero — it says Luna cannot be a 70B-class dense bf16 model served on rented neocloud H100 at k = 2.0 and still clear. The break-even threshold is a constraint on the joint choice of model size, quantization, hardware generation and capacity procurement. Any one of several things resolves it: a smaller or sparser model, aggressive quantization, Blackwell-class hardware, owned or reserved capacity below the neocloud spot rate, or a leaner overhead multiplier. All of these are ordinary engineering and procurement choices, and the lab publicly attributed the cut to exactly such improvements.

The finding, stated precisely: the July rate card is a public statement about the serving stack. It sets an upper bound on the model class that can economically sit behind that tier. A price cut of this size is not a claim about willingness to lose money — it is a claim about efficiency, and the arithmetic says the claim has to be large: on the order of a 5× improvement, or an equivalent reduction in the effective cost of the compute underneath.

The spread, valued

Running the spread rather than the threshold, at the measured 70B-class median and a 3:1 blend:

Rate cardRevenue $/MCost $/M @ g=$2.50Spread $/MImplied margin on this denominator
Luna — old2.2500.4301.82080.9%
Luna — new0.4500.4300.0204.5%
Terra — new4.5000.4304.07090.4%
Sol — unchanged11.2500.43010.82096.2%
Cost = g·k/E with g = $2.50/GPU-hr, k = 2.0, E = 11.63M measured. These are spreads against a 70B-class dense denominator on rented neocloud capacity — an upper bound on any frontier lab's true input cost, and therefore a lower bound on its true spread. They are not estimates of anyone's realised margins.

06The hedge, and the answer to Pirrong

Work the sign convention and the two sides of the market fall out — including the one Pirrong argued does not exist.

Who is long the spread and who is short it

ParticipantEconomicsExposureHedge
Model lab (serves its own weights)Margin = Q × [ptok − g·k/E]. Earns the token price, pays the GPU rate.LONG the spread — hurt by falling token prices and rising GPU pricesSell token forwards, buy GPU futures. Structurally identical to a generator hedging a spark spread by selling power forward and buying gas forward.
Inference reseller (the Perplexity / Poe class)Margin = S − Q × ptok. Fixed subscription revenue, floating token cost.SHORT tokens — hurt by rising token pricesBuy token forwards. Worked end to end in the inference hedging piece: 91% variance reduction for the optimized two-leg mix vs 57% for GPU futures alone.
Neocloud / GPU ownerRevenue is g; collateral value is g capitalised.LONG GPU price twice overSell GPU futures. The specialty lender programme is the financing-side version of the same short.
The structural point

The lab is long the token leg. The reseller is short it. They are natural counterparties on the same instrument — and the July divergence is precisely the event that gives both of them a reason to trade it at the same time. The lab has just watched its output price fall 80% on a tier; the reseller has spent two years watching token prices fall and is structurally reluctant to hedge a position that keeps winning. When token prices next turn, both sides want the same contract from opposite directions.

Pirrong's objection, and what survives it

Craig Pirrong's May 2026 critique — catalogued in the index methodology work — makes three arguments against compute futures: compute prices trend (tech-cycle driven) rather than mean-revert; the value chain is concentrated and vertically integrated, so there is weak natural two-sided hedging demand; and chip-specific contracts risk orphaning on 18–24 month silicon cycles, as DRAM futures did.

ObjectionAnswer from this structureStatus
Weak two-sided hedging demand. Concentration and vertical integration mean too few natural counterparties.The spark-spread decomposition manufactures the two sides. Labs are long the token leg and short nothing; resellers are short the token leg and long nothing. Every dollar the lab wants to lock, the reseller wants to lock from the other side. The token leg — not the GPU leg — is where the natural two-sided book lives.Answered, on the token leg
Prices trend rather than mean-revert. Trending prices make hedging a one-sided loser and deter liquidity.Nothing in this structure addresses it. Both legs have trended hard — tokens down >90% over three years, neocloud GPU up ~115% off the trough. A spread between two trending series is not obviously stationary either.Stands
Silicon-cycle orphaning. Chip-specific contracts die with the chip generation.Nothing in this structure addresses it, and the efficiency term makes it worse: E is chip-generation-specific, so a spread contract inherits the orphaning risk of both legs plus the basis between them.Stands, arguably sharpened

One objection of three is answered. That is worth stating plainly rather than claiming the critique has been disposed of — Pirrong's remaining two are the harder ones, and the efficiency basis this piece measures is an argument for his orphaning concern, not against it.


07What would have to exist

A tradeable inference spark spread needs three published inputs. Two exist in some form. The third exists nowhere.

LegWhat exists todayWhat is missing
Token price indexSilicon Data's LLM Token Expenditure Index (expenditure-weighted $/M tokens across the market); Ornn publishes realized token price indices by model family.No published rulebook, per the methodology catalogue's headline finding on all four administrators. And an expenditure-weighted index measures P×Q, not P — it moves when the mix shifts to cheaper models even if no price changes, which is exactly what its own CEO has attributed recent moves to.
GPU price indexFour administrators — Silicon Data, Ornn, Kalshi's derived curve, Compute Desk. Live ladders at Kalshi; ICE/OCPI and CME/Silicon Data futures pending.Same problem: none has published a methodology. All are quote-based rather than transaction-anchored, which overstates dispersion and is vulnerable to the same critique as a rate card.
Efficiency referenceNothing. No administrator, lab, cloud or standards body publishes a serving-efficiency benchmark tied to a settlement date.Everything. The spread is computable today only because a national laboratory released measured throughput and power under CC-BY for an unrelated purpose — data-centre load planning. A durable market needs that measurement institutionalised: a versioned reference workload, a published model class, a stated load condition, and a settlement calendar.

This is the same conclusion the methodology catalogue reached about GPU indices, one layer up and one degree more acute. There, the missing thing was disclosure of a methodology for measuring a price everyone agrees exists. Here, the missing thing is the measurement itself — and unlike a price, an efficiency figure cannot be inferred from quotes, because no one quotes it. Someone has to run the benchmark and publish the number, with the load condition attached, on a schedule. Until then the inference spark spread is computable in research and not settleable in a contract.


08What this analysis is not

Repeated deliberately, because the arithmetic is more inviting than it is authoritative.

The one-paragraph version. Token prices and GPU rental prices are moving apart, and the gap between them is a spark spread with a missing heat rate. Measure the heat rate — 11.6M tokens per GPU-hour for a 70B-class dense model at saturation, 4× better than the same hardware lightly loaded, at 51 Wh per million tokens — and the July rate cards resolve into a threshold. Sol and Terra clear that threshold by more than an order of magnitude and the cut on Terra gives back a slice of a very wide implied spread. Luna's new card lands on it, which is not a statement about anyone's margins but a constraint on what model can sit behind that tier at that price. And the same decomposition that produces the threshold produces the two-sided hedging demand Pirrong argued compute lacked: labs long the token leg, resellers short it, natural counterparties manufactured by exactly the divergence that prompted the question. His other two objections are untouched and still stand.