PodBrowser
Forward Guidance

AI Efficiency Is Repricing The Compute Market | Steve Hou

Wednesday, 22 July 2026 · 4 min read · Listen to the episode ↗

Steve Hou makes the case that AI compute risk is mispriced because it is currently embedded in equity and fixed income rather than traded directly, and that Silicon Data's futures and derivatives market would let cloud providers and data centers hedge that exposure efficiently.

Silicon Data is building futures contracts and derivatives for the physical AI compute market, targeting data center operators, cloud providers, and enterprises that need to hedge compute-related risks. Steve Hou argues that AI compute risk is currently held in equity and fixed income form and is not being priced efficiently, and that cloud providers have a natural incentive to use futures to lock in revenue certainty.

Silicon Data's token index is an expenditure-weighted price index aggregating data from over 300 models, blending input and output token prices weighted by usage volume drawn from public routing platforms including Open Router. Because it excludes direct usage data from OpenAI, Anthropic, and the major hyperscalers, the index tilts toward independent developers and small-to-medium enterprises, which are more price-sensitive than large enterprise buyers. Hou notes that most movement in the index comes from shifts in user behavior and substitution across models rather than from changes in per-token list prices at launch, and he describes it as functioning analogously to the PCE inflation measure for AI.

In early June Hou posted publicly that the index appeared to be plateauing, which he noted coincided with a turn in AI-beneficiary equities including SMH. He interprets the move since June as reflecting substitution toward cheaper and more efficiently routed models rather than an outright collapse in token demand. He argues that markets are pricing AI capital expenditure on the second derivative, meaning even a slower rate of growth can trigger selloffs at current valuations, and that cheaper capable models from Grok, Mistral, and others could challenge the margin structure of frontier models and the capex financing paradigm supporting them.

Hou's structural view is that no single AI lab will win the market outright and that an orchestration layer will accrue significant value as models become more substitutable, with low-value tasks routed to cheaper open-weight models and high-value reasoning tasks routed to frontier intelligence. He cites Open Router as having evolved from a hobbyist developer tool into an emerging enterprise norm for smart model routing, and names Palantir's ontology approach and Databricks as examples of the broader enterprise trend of making AI practically useful in complex workflows. He flags a potential transition-phase air pocket where return on investment on current capital expenditure does not materialize quickly enough while large enterprises are still identifying workflows, and he argues that sustainable adoption requires the model layer to become substitutable so that AI does not drain enterprise budgets through reliance on a single expensive frontier model. Using Jevons paradox logic, he suggests that if price competition compresses per-token costs but total adoption grows substantially, overall profit for leading providers can still work out.

On the GPU rental market, Hou says H100 one-year rental prices were monotonically increasing through July 20th despite fears about excess AI supply, with the forward curve shifting from backwardation in November to near-contango by late March, meaning long-term contract discounts are disappearing. When H100 rates experienced some recent downward fluctuation, B200 and H200 rates continued rising significantly. A100 rental rates have been holding steady rather than declining, which Hou interprets as evidence that inference demand is growing fast enough to sustain demand for a chip available for approximately five years. His cross-sectional reading is that inference demand is growing extremely fast while some training workflows shift from H100 to newer chips, and he predicts Hopper-generation chips will transition from being the leading training chips to becoming the new workhorse inference chips, following the pattern the A100 established.

Hou argues that memory supply is not keeping up with AI-driven demand, causing prices to rise sharply, with longer context windows making models highly memory-hungry. He describes Kimi's algorithmic improvements to memory efficiency as feeling similar to the DeepSeek moment, suggesting ongoing algorithmic gains will continue to compress memory demand growth. He predicts gross margins of memory makers are unlikely to remain near 85 percent in two years because efficiency-driven demand compression will follow, though he frames overall memory demand as a price-times-quantity dynamic where falling unit prices are offset by expanding usage from multimodal voice AI, video AI applications, and longer context windows.

Hou draws a parallel between the US-China AI competition and the Tesla-BYD dynamic, where Chinese models dominate outside the US while US models persist domestically. He argues that US-China decoupling drives bottleneck trades in compute hardware and predicts the US government could effectively restrict Chinese models from federal procurement through model safety rules that Chinese models would fail to meet. With both the US and China doubling down on capex and state intervention, he believes overall demand for compute and hardware could effectively double. He frames AI enterprise ROI as following a J-shaped curve, meaning returns will materialize more slowly than most people hope, and identifies genuine enterprise adoption rather than consumer use cases as the more important near-term dynamic to watch.

This summary was generated from the episode transcript and can contain mistakes.