AI Efficiency Is Repricing The Compute Market | Steve Hou
Wednesday, 22 July 2026 · 4 min read · Listen to the episode ↗
Steve Hou of Silicon Data makes the case that futures contracts and derivatives for physical AI compute would let cloud providers hedge capacity risk currently embedded in equities and fixed income, enabling bolder infrastructure commitments while giving new entrants the revenue certainty lenders require. The GPU forward curve shifted from backwardation in late 2024 to a flatter contango shape by March 2025, with one-year contract prices rising monotonically even as headlines suggested oversupply.
Silicon Data is building futures contracts and derivatives for the physical AI compute market so that cloud providers and buyers can hedge risk currently held in equity and fixed income form. Steve Hou argues that futures contracts would let cloud providers commit more boldly to capacity acquisition without overexposure, and that new cloud entrants and their lenders specifically want the revenue certainty that forward contracts provide, analogous to any commodity market. Hou acknowledges cloud providers resist the commodity label and notes substitutability in compute will exist but will not be perfect.
The GPU forward curve shifted from backwardation in November 2024 to a flatter or slightly contango shape by March 2025, meaning cloud providers are no longer offering long-term contract discounts and are instead looking to raise prices. One-year term compute contract prices have been monotonically increasing even amid news suggesting AI supply excess, and multiple providers have raised prices, indicating firming demand rather than a one-off event. A short-term H100 spot price dip around June 2025 caused market concern, but one-year contract prices were largely unchanged. When H100 rates came down, B200 and H200 rental rates continued rising significantly. A100 rental rates have held steady, which Hou reads as consistent with inference demand absorbing older chip capacity while training workloads shift to newer chips, though he flags this as consistent with the data but not confirmed. Hou directly counters Michael Burry's view that chips are fast-depreciating two-to-three-year assets and predicts Hopper-class chips will become the new workhorse for inference as newer training chips come online.
The Copper AI inference price index blends input and output token prices weighted by usage volume from public routing platforms and tilts toward independent developers and small-to-medium enterprises rather than large enterprise or hyperscaler customers, making it a leading indicator rather than a fully representative market sample. The index appeared to plateau around early June 2025, which Hou attributes to users becoming more rational about cost and substituting quality for cost efficiency. This plateau coincided with a turn in AI-related equity markets including the SMH and AI beneficiary indices. Hou frames the current AI capex trade as priced on second derivatives, meaning that even if growth continues at a slower pace, valuations can sell off. The index movement since early March was partly driven by token maxing, where US corporate users maximized frontier model usage regardless of cost, and the subsequent turn reflects substitution toward efficiency rather than an outright decline in token demand.
Hou argues that the transition from token maxing to smart model routing and token efficiency is now becoming the norm rather than a hobbyist practice. He identifies a key routing challenge in that a question can appear simple but actually be complex, making quality and cost assessment difficult. He identifies a risk of an air pocket where capex ROI from the current paradigm does not materialize quickly enough and large enterprise adoption is slow to ramp. Productivity gains in aggregate are partly suppressed because the fastest AI adopters are small enterprises that are small in dollar-weighted terms, and enterprise adoption overall remains vanishingly small with enterprises still nervous about cost.
Memory is the primary input constraint for GPUs, and AI models are highly memory-hungry, which has caused demand to outpace supply and driven prices up sharply. Algorithmic innovations from Chinese models including Kimi are improving memory efficiency so that memory demand no longer has to grow linearly with context length. Hou says the Kimi efficiency moment felt very similar to the DeepSeek moment and that such improvements will keep recurring. Despite falling memory prices, Hou applies a price-times-quantity framework and argues memory makers will still do well because volume growth will more than offset margin compression, though he would be very shocked if memory makers are still running gross margins of around 85 percent in two years. Multimodal voice AI such as ChatGPT Live creates significantly more data and has been largely overlooked as a demand driver, and video as a modality has barely been scratched, representing further future demand.
On the US-China dynamic, Hou draws a parallel to Tesla and BYD, arguing Chinese models are very capable but geopolitical decoupling limits their direct market access in the United States. Chinese AI models are being released free and open-sourced to the US partly because the Chinese domestic economy is weak and Chinese companies historically do not pay for SaaS. He predicts efficiency gains will continue regardless of regulatory outcomes, and that both the US and China doubling down on compute investment means roughly double the overall demand for compute and hardware.
Hou sees genuine enterprise adoption and demonstrable ROI as the bigger near-term dynamic to watch. He does not expect any single lab to win outright as the dominant platform and believes an orchestration layer will accrue significant value as enterprises make models more substitutable, routing low-value tasks to cheaper open models and high-value tasks to frontier intelligence. He predicts coming quarters will show more visible ROI from AI adoption but frames the productivity payoff as J-shaped, meaning it will still be slower than many optimists expect before the benefits become clearly evident.
This summary was generated from the episode transcript and can contain mistakes.