Advisory research · Compute · Inference economics · 2 August 2026
On 30 July a frontier lab cut the published price of one model tier by 80% in three weeks. Over the past year the neocloud GPU rental rate that tier is served on roughly doubled off its trough. Those two facts are in direct tension, and the tension has an exact form: it is a spark spread — the structure a power trader uses to price a generator's margin as the power price less the gas price times the heat rate. Written that way, the inference version reduces to a single unobservable term, tokens per GPU-hour, and the arithmetic inverts to a threshold: the serving efficiency a rate card requires in order to clear.
The heat rate is the whole difficulty. In power it is published per unit and regulated; in inference nobody publishes it. This piece computes it directly — 1,024 measured vLLM serving runs on Llama-3.1-70B across four H100s, joined sample-by-sample to the NVML power traces recorded alongside them, released under CC-BY by a national laboratory. That gives a measured denominator to put underneath the July rate cards. This is a required-efficiency bound computed from published prices. It is not an accounting of any company's margins, and nothing here is a claim about any firm's financial condition.
Both legs move, in opposite directions, and both are tiered. The tiering is the first finding.
On 30 July 2026 OpenAI repriced the GPT-5.6 family, three weeks after its 9 July launch. The cut was not uniform. Luna fell 80%, from $1.00/$6.00 to $0.20/$1.20 per million input/output tokens. Terra fell 20%, from $2.50/$15.00 to $2.00/$12.00. Sol, the flagship, was left unchanged at $5.00/$30.00. Batch and Flex processing are reported at half standard price, putting Luna's batch tier at $0.10/$0.60.
The stated driver is competitive. A CNBC investigation found Chinese-origin models reaching a weekly peak of 46% of US enterprise token volume on OpenRouter, and at least 30% of it every week since 8 February 2026 — against 11% averaged over the prior twelve months and 4.5% in the first half of 2025. OpenRouter's own attribution is price: open-weight Chinese models running 60–90% cheaper than the US frontier — and at the extreme, well beyond that band: as of June 2026 DeepSeek V4 Flash listed at $0.14/M input against the then-current GPT-5.5 flagship at $5.00, a 97% gap. OpenAI framed the cut as passing through efficiency gains in its serving stack.
The demand side is moving the other way, which is what makes the divergence interesting rather than merely deflationary. Silicon Data's LLM Token Expenditure Index shows aggregate LLM spending has doubled since 2025 even as cost per token fell more than 90% over three years, and the Federal Reserve Bank of Atlanta expects firms to raise AI spend roughly 50% in 2026, to $280bn. Falling unit price with rising total spend is the P×Q problem the inference hedging piece works in its §04: every hedgeable instrument lives on price, while quantity outran it. A spark spread is a per-unit margin — it says nothing about volume, and volume is where the money has actually gone.
80% at the bottom, 20% in the middle, zero at the top. Price pressure is concentrated exactly where open-weight and Chinese models compete on substitutable work, and absent where they do not. Because tiers run on different model sizes and different hardware, the spread has to be computed per tier; an aggregate blend of the three would average away the only place the arithmetic binds.
The convenient framing is that GPU rentals rose while token prices fell. The verified picture is more specific, and the specificity matters because a lab's input cost depends on which tier it buys. Silicon Data's Q3 2026 regional cut shows hyperscaler H100 rates falling 26–54% by region since Q1 2025, while non-hyperscaler (neocloud) rates bottomed in mid-2025 and recovered strongly — US East from roughly $1.35 to about $2.90, the Nordics from about $1.75 to roughly $3.30, now above where the series started. Silicon Data computes the average premium across those seven regions as falling from roughly 250% to about 120% — a per-region average of quarterly medians, which is why the range midpoints in the table below do not reproduce it directly.
| Leg | Q1 2025 / trough | Q3 2026 | Direction | Source |
|---|---|---|---|---|
| H100, neocloud (US East) | ~$1.35/hr (Q3 2025) | ~$2.90/hr | up ~115% | Silicon Data H100 index, regional |
| H100, neocloud (range, 7 regions) | $1.35–$2.70 | $2.00–$3.30 | up, dispersed | Silicon Data, Q3 2026 partial |
| H100, hyperscaler (range) | $7.25–$10.40 | $4.80–$6.10 | down 26–54% | Silicon Data, Q3 2026 partial |
| B200, neocloud (SDB200RT index) | 4.40 (1 Jan 2026) | 5.48 (30 Mar 2026) | up 24.4% YTD to Mar | Silicon Data SDB200RT; March mean price $5.09/hr, median $4.96 |
| GPT-5.6 Luna, blended 3:1 | $2.25/M | $0.45/M | down 80% | Published rate card, 30 Jul 2026 |
| GPT-5.6 Terra, blended 3:1 | $5.63/M | $4.50/M | down 20% | Published rate card, 30 Jul 2026 |
| GPT-5.6 Sol, blended 3:1 | $11.25/M | $11.25/M | unchanged | Published rate card |
So the divergence is real but it is a neocloud divergence. A lab renting marginal capacity from the specialist tier has watched its input price roughly double off the trough while its output price on one tier fell 80%. A lab buying hyperscaler on-demand has seen its input price fall — but no lab at scale buys marginal inference capacity there. The relevant leg for this analysis is the neocloud tier, and it is the leg that rose.
The framing that H100 forwards are in contango — forward above spot, the market pricing a future crunch — is reported but we could not confirm a live forward quote against a contemporaneous spot from a public source. What the corpus does establish independently is a $1.23/GPU-hour convenience yield (23.6% of spot) between the financial forward and provider reserved-tier term sheets at the one-year point, measured in the implied forward curve work. Contango in GPU forwards is that same quantity seen from the other side. We use the measured convenience yield and leave the contango claim as reported rather than verified.
The analogy is exact in form and breaks in one specific place. Both halves matter.
A gas generator's margin is the spark spread: the price it sells power at, less the price it buys gas at, scaled by how much gas it burns per unit of power. The scaling term is the heat rate, in MMBtu per MWh, and it is the whole reason the spread is tradeable — it is published per unit, it is an engineering property of the turbine, and it changes slowly.
The inference version has the identical shape. A model lab sells tokens at a published $/M-token rate and buys GPU-hours at a rental rate. The scaling term is the number of tokens the stack produces per GPU-hour — which enters as a divisor rather than a multiplier, because it is expressed as output per input rather than input per output.
This is the cost equation from the inference hedging piece — cost = R·T·ptok + (1−R)·T·(g/E)·k — restricted to the self-hosted channel and rearranged as a margin rather than a cost. E is the heat rate of inference. 1/E is tokens' "gas burn": GPU-hours consumed per token delivered.
| Property | Power: heat rate | Inference: tokens per GPU-hour |
|---|---|---|
| Determined by | Turbine thermodynamics. A physical property of the plant. | Software. Batching policy, quantization, speculative decoding, KV-cache management, scheduler, framework version. |
| Stability | Degrades slowly and predictably over the asset's life. | Improves discontinuously with every serving-stack release. A framework upgrade can move it 30% overnight. |
| Load dependence | Modest. Part-load heat rate is worse, and the curve is published. | Severe and unpublished. Measured here: a 4× swing between light load and saturation on identical hardware. |
| Disclosure | Published per generating unit; used in dispatch, regulation and settlement. | Published by nobody. No index administrator, no lab, no cloud provider publishes a serving-efficiency reference. |
| Consequence for the spread | Spark spread is directly observable and directly tradeable. | Two of three terms are observable; the third is private. The spread is computable only where someone has measured E. |
That last row is why this analysis exists here and essentially nowhere else. The measured GPU power work established that the denominator of any compute-normalised metric has to be measured rather than assumed — nameplate TDP overstates device energy by 12–21%, and duty factors run 0.795–0.880. The same discipline applies one layer up: the efficiency term in a token-versus-GPU spread has to be measured, and the measurement has to be stated with its load condition attached.
1,024 vLLM serving runs, joined to 715,989 NVML power samples over a 19.9-hour window. Same hardware, same model, same weights — only the offered load changes.
The NLR generative-AI power dataset (Vercellino et al., arXiv:2604.07345, CC-BY-4.0) was
published to support data-centre load modelling, and the practice used it previously for
duty factors and capacity-factor derates.
It contains something the power analysis did not need: the raw vLLM benchmark logs that ran
alongside the wattmeter traces. Those logs carry total_token_throughput,
total_input_tokens and total_output_tokens per run. Dividing throughput by
the tensor-parallel width gives measured tokens per GPU-hour directly, with no modelling step.
| Measurement basis | Value |
|---|---|
| Model / precision | Llama-3.1-70B-Instruct, bfloat16, KV cache auto |
| Hardware / parallelism | 4 × H100 on one node, tensor_parallel_size = 4 |
| Serving stack | vLLM, OpenAI-compatible completions endpoint |
| Workload | InstructCoder prompts, 1,000 per run, request rate swept 10 → 1,000 req/s |
| Observed prompt shape | mean 175 input / 158 output tokens — a 1.1:1 input:output mix |
| Runs / power samples | 1,024 benchmark runs; 715,989 four-GPU NVML samples (GPU sockets only — CPU, DRAM and facility overhead excluded) across 19.9 h |
Measured throughput, power and energy per token vs offered load
The finding the whole spread rests on: above roughly 50 req/s, GPU power flattens while throughput keeps climbing — so energy per token keeps falling even though the machine draws the same watts. Tokens per GPU-hour varies 4× across this sweep on identical silicon.
Read. Four-GPU NVML draw rises 17% from light load to the saturation plateau (2,020 W → ~2,370 W), peaking at 2,384 W. Throughput over the same range rises 304% (2.88M → 11.63M tokens per GPU-hour). Energy per million tokens therefore falls 71%, from 175 Wh to 51 Wh. At $0.06/kWh and PUE 1.3 the saturated figure is about $0.0040 per million tokens — 0.9% of the new Luna blended rate. That is a floor: the wattmeter reads GPU sockets only, so CPU, DRAM, fans and PSU losses sit outside it and the true facility figure is higher. Even several times higher it does not change the conclusion — electricity is not the binding cost of inference at these token prices. The rental rate is.
The inference hedging worked example used E = 2.5M tokens per GPU-hour. Against this measurement, 2.5M sits close to a lightly loaded server at 10 req/s, measured at 2.88M — not a saturated one. A well-utilised 70B-class deployment runs 11.6M at the median and 15.9M at the 90th percentile. That is not an error in the prior work — it is the point that piece made, that E is unhedgeable basis — but it materially changes the break-even arithmetic, and in the lab's favour.
Set the spread to zero and solve for E. The result is the number of tokens per GPU-hour a stack must produce for a given rate card to cover a given rental rate.
This is the inference analogue of the market-implied heat rate a power trader computes as the power price divided by the gas price: the efficiency the market is implicitly demanding at prevailing prices. Cut the numerator's price and the required efficiency rises proportionally. The 80% Luna cut therefore raises the required efficiency by exactly 5×, mechanically, holding the compute leg fixed.
| Tier / rate card | Blended 3:1 | g = $2.00 | g = $2.50 | g = $2.90 | g = $3.30 | vs measured 11.6M |
|---|---|---|---|---|---|---|
| Luna — old card ($1.00/$6.00) | $2.25/M | 1.8M | 2.2M | 2.6M | 2.9M | clears 4.0–6.5× |
| Luna — new card ($0.20/$1.20) | $0.45/M | 8.9M | 11.1M | 12.9M | 14.7M | 1.3× → 0.8× |
| Luna — batch/flex ($0.10/$0.60) | $0.225/M | 17.8M | 22.2M | 25.8M | 29.3M | far above |
| Terra — old card ($2.50/$15.00) | $5.63/M | 0.71M | 0.89M | 1.03M | 1.17M | clears 9.9–16.4× |
| Terra — new card ($2.00/$12.00) | $4.50/M | 0.89M | 1.11M | 1.29M | 1.47M | clears 7.9–13.1× |
| Sol — unchanged ($5.00/$30.00) | $11.25/M | 0.36M | 0.44M | 0.52M | 0.59M | clears 19.8–31.5× |
| Overhead k | g = $2.00 | g = $2.50 | g = $2.90 | g = $3.30 | Interpretation at k |
|---|---|---|---|---|---|
| 1.2 — near bare-metal, own orchestration | 5.3M | 6.7M | 7.7M | 8.8M | Clears on 70B-class throughput in every region |
| 1.5 — lean operator | 6.7M | 8.3M | 9.7M | 11.0M | Clears at the median, thin at the top of the range |
| 2.0 — the corpus assumption | 8.9M | 11.1M | 12.9M | 14.7M | At or above measured median across most of the range |
| 2.5 — replication-heavy, latency-SLA serving | 11.1M | 13.9M | 16.1M | 18.3M | Above the measured median throughout; above p90 only at the top two rental rates |
Input and output tokens carry different prices and different costs, and the blend assumption does real work. Luna's 6:1 price ratio between output and input means an output-heavy workload is better for the lab per token; an input-heavy one is worse on price but cheaper to serve.
| Input:output | Blended $/M | Required E at g=$2.50, k=2.0 | Typical of |
|---|---|---|---|
| 1:1 | $0.700 | 7.1M | Chat turns, short prompts |
| 2:1 | $0.533 | 9.4M | Mixed API traffic |
| 3:1 | $0.450 | 11.1M | Base case used throughout |
| 5:1 | $0.367 | 13.6M | RAG, document QA |
| 10:1 | $0.291 | 17.2M | Long-context agents, tool-calling chains |
Break-even calculator — required efficiency against measured throughput
Set a rate card, a GPU rental rate, an overhead multiplier and a prompt mix. The calculator returns the required tokens per GPU-hour, and compares it against the measured 70B-class reference from §03. The bar is the requirement; the marker is what was measured.
The arithmetic supports a conclusion about model class, not about margins. That is a narrower claim than the headline invites, and it is the one the data will carry.
| Case | Condition | Reading |
|---|---|---|
| 1 — Margin give-back | Required E comfortably below measured throughput | The cut surrenders margin from a very high base. Competitive, not existential. This is where the Terra and Sol rate cards sit against this denominator, and by a wide margin: Sol requires 0.44M at g=$2.50 against 11.63M measured — roughly 26× coverage, or 19.8× at the top of the rental range — even on a 70B-class dense denominator that is far more expensive to serve than Sol's own class is likely to be. |
| 2 — Compression toward zero | Required E near measured throughput | The rate card would price at or near cost if served on this denominator. This is where Luna's new card sits against a 70B-class dense denominator — 11.1M required at the mid-range rental rate against 11.63M measured, and the coverage flips from 1.31× to 0.79× across the rental range. It is a statement about the denominator, not about the tier's actual economics, which depend on a serving stack nobody outside the lab has seen. |
| 3 — Below cost to serve | Required E far above measured throughput | Tokens sold below the compute cost of producing them. The arithmetic does not support placing any tier here on the evidence available, and we decline to assert it. The batch/flex tier's 22M requirement would qualify against a 70B denominator — but batch tiers exist precisely to be filled with deferrable work at high batch sizes, where achievable E is far above the online figure measured here. |
Read case 2 carefully. It does not say Luna's margins are near zero — it says Luna cannot be a 70B-class dense bf16 model served on rented neocloud H100 at k = 2.0 and still clear. The break-even threshold is a constraint on the joint choice of model size, quantization, hardware generation and capacity procurement. Any one of several things resolves it: a smaller or sparser model, aggressive quantization, Blackwell-class hardware, owned or reserved capacity below the neocloud spot rate, or a leaner overhead multiplier. All of these are ordinary engineering and procurement choices, and the lab publicly attributed the cut to exactly such improvements.
The finding, stated precisely: the July rate card is a public statement about the serving stack. It sets an upper bound on the model class that can economically sit behind that tier. A price cut of this size is not a claim about willingness to lose money — it is a claim about efficiency, and the arithmetic says the claim has to be large: on the order of a 5× improvement, or an equivalent reduction in the effective cost of the compute underneath.
Running the spread rather than the threshold, at the measured 70B-class median and a 3:1 blend:
| Rate card | Revenue $/M | Cost $/M @ g=$2.50 | Spread $/M | Implied margin on this denominator |
|---|---|---|---|---|
| Luna — old | 2.250 | 0.430 | 1.820 | 80.9% |
| Luna — new | 0.450 | 0.430 | 0.020 | 4.5% |
| Terra — new | 4.500 | 0.430 | 4.070 | 90.4% |
| Sol — unchanged | 11.250 | 0.430 | 10.820 | 96.2% |
Work the sign convention and the two sides of the market fall out — including the one Pirrong argued does not exist.
| Participant | Economics | Exposure | Hedge |
|---|---|---|---|
| Model lab (serves its own weights) | Margin = Q × [ptok − g·k/E]. Earns the token price, pays the GPU rate. | LONG the spread — hurt by falling token prices and rising GPU prices | Sell token forwards, buy GPU futures. Structurally identical to a generator hedging a spark spread by selling power forward and buying gas forward. |
| Inference reseller (the Perplexity / Poe class) | Margin = S − Q × ptok. Fixed subscription revenue, floating token cost. | SHORT tokens — hurt by rising token prices | Buy token forwards. Worked end to end in the inference hedging piece: 91% variance reduction for the optimized two-leg mix vs 57% for GPU futures alone. |
| Neocloud / GPU owner | Revenue is g; collateral value is g capitalised. | LONG GPU price twice over | Sell GPU futures. The specialty lender programme is the financing-side version of the same short. |
The lab is long the token leg. The reseller is short it. They are natural counterparties on the same instrument — and the July divergence is precisely the event that gives both of them a reason to trade it at the same time. The lab has just watched its output price fall 80% on a tier; the reseller has spent two years watching token prices fall and is structurally reluctant to hedge a position that keeps winning. When token prices next turn, both sides want the same contract from opposite directions.
Craig Pirrong's May 2026 critique — catalogued in the index methodology work — makes three arguments against compute futures: compute prices trend (tech-cycle driven) rather than mean-revert; the value chain is concentrated and vertically integrated, so there is weak natural two-sided hedging demand; and chip-specific contracts risk orphaning on 18–24 month silicon cycles, as DRAM futures did.
| Objection | Answer from this structure | Status |
|---|---|---|
| Weak two-sided hedging demand. Concentration and vertical integration mean too few natural counterparties. | The spark-spread decomposition manufactures the two sides. Labs are long the token leg and short nothing; resellers are short the token leg and long nothing. Every dollar the lab wants to lock, the reseller wants to lock from the other side. The token leg — not the GPU leg — is where the natural two-sided book lives. | Answered, on the token leg |
| Prices trend rather than mean-revert. Trending prices make hedging a one-sided loser and deter liquidity. | Nothing in this structure addresses it. Both legs have trended hard — tokens down >90% over three years, neocloud GPU up ~115% off the trough. A spread between two trending series is not obviously stationary either. | Stands |
| Silicon-cycle orphaning. Chip-specific contracts die with the chip generation. | Nothing in this structure addresses it, and the efficiency term makes it worse: E is chip-generation-specific, so a spread contract inherits the orphaning risk of both legs plus the basis between them. | Stands, arguably sharpened |
One objection of three is answered. That is worth stating plainly rather than claiming the critique has been disposed of — Pirrong's remaining two are the harder ones, and the efficiency basis this piece measures is an argument for his orphaning concern, not against it.
A tradeable inference spark spread needs three published inputs. Two exist in some form. The third exists nowhere.
| Leg | What exists today | What is missing |
|---|---|---|
| Token price index | Silicon Data's LLM Token Expenditure Index (expenditure-weighted $/M tokens across the market); Ornn publishes realized token price indices by model family. | No published rulebook, per the methodology catalogue's headline finding on all four administrators. And an expenditure-weighted index measures P×Q, not P — it moves when the mix shifts to cheaper models even if no price changes, which is exactly what its own CEO has attributed recent moves to. |
| GPU price index | Four administrators — Silicon Data, Ornn, Kalshi's derived curve, Compute Desk. Live ladders at Kalshi; ICE/OCPI and CME/Silicon Data futures pending. | Same problem: none has published a methodology. All are quote-based rather than transaction-anchored, which overstates dispersion and is vulnerable to the same critique as a rate card. |
| Efficiency reference | Nothing. No administrator, lab, cloud or standards body publishes a serving-efficiency benchmark tied to a settlement date. | Everything. The spread is computable today only because a national laboratory released measured throughput and power under CC-BY for an unrelated purpose — data-centre load planning. A durable market needs that measurement institutionalised: a versioned reference workload, a published model class, a stated load condition, and a settlement calendar. |
This is the same conclusion the methodology catalogue reached about GPU indices, one layer up and one degree more acute. There, the missing thing was disclosure of a methodology for measuring a price everyone agrees exists. Here, the missing thing is the measurement itself — and unlike a price, an efficiency figure cannot be inferred from quotes, because no one quotes it. Someone has to run the benchmark and publish the number, with the load condition attached, on a schedule. Until then the inference spark spread is computable in research and not settleable in a contract.
Repeated deliberately, because the arithmetic is more inviting than it is authoritative.
The one-paragraph version. Token prices and GPU rental prices are moving apart, and the gap between them is a spark spread with a missing heat rate. Measure the heat rate — 11.6M tokens per GPU-hour for a 70B-class dense model at saturation, 4× better than the same hardware lightly loaded, at 51 Wh per million tokens — and the July rate cards resolve into a threshold. Sol and Terra clear that threshold by more than an order of magnitude and the cut on Terra gives back a slice of a very wide implied spread. Luna's new card lands on it, which is not a statement about anyone's margins but a constraint on what model can sit behind that tier at that price. And the same decomposition that produces the threshold produces the two-sided hedging demand Pirrong argued compute lacked: labs long the token leg, resellers short it, natural counterparties manufactured by exactly the divergence that prompted the question. His other two objections are untouched and still stand.