PodBrowser
This Week Startups

Why Data Is the Next $1 Trillion Market

Wednesday, 8 July 2026 · 4 min read · Listen to the episode ↗

Nikhil Basu Trivedi makes the case that energy, compute, and data are the three foundational substrates for AI, and that while energy and compute are represented by massive public companies like Nvidia, no comparably large pure-play public data company yet exists, creating what he sees as a significant market opportunity.

Nikhil Basu Trivedi argues that the three fundamental substrates for AI models are energy, compute, and data, and that while energy and compute are represented by enormous public companies such as Nvidia, no comparably large pure-play public company focused on data yet exists. He predicts this absence of large incumbents creates meaningful opportunity in the data sector. The AI labs are currently focused on data, but activity on that side is more murky and less publicly discussed than their work on energy and compute. Protege, in its second year of business, operates in the data licensing space and is already generating hundreds of millions in revenue, with public deals like Reddit and the New York Times representing only a visible fraction of a much larger volume of such arrangements.

Michael argues that proprietary data is where economic value will concentrate in the AI race, and that the SaaS Apocalypse narrative is overblown because established SaaS companies hold proprietary data, customer relationships, and workflows that AI will help them leverage rather than destroy. He points to Salesforce growing at 10 to 12 percent but down 40 percent for the year, and Box trading at 3.4 times trailing price-to-sales, as examples of data-rich companies being systematically undervalued by public market investors chasing high-growth momentum names. HubSpot, with over three billion dollars in ARR, is valued at less than ten billion dollars in public markets, while some private companies with far less revenue carry valuations of ten billion dollars or more. The MeritTech SaaS index is trading at approximately 3.6 to 3.7 times revenue, and Monday.com is trading at approximately two times revenue despite being a rule-of-40 company.

Michael views current AI investment dynamics as resembling ZIRP-era 2021 conditions, with platform firms supplying effectively unlimited capital and revenue multiples returning to 2021 levels. He does not view Nvidia's forward PE of approximately 24 times as comparable to Cisco at 100 times in 1999, but notes that earnings growth assumptions are doing significant work in the valuation, and predicts the Nasdaq could fall 20 percent easily over the next 12 to 24 months if growth expectations are not met. He identifies three key correction catalysts: a semiconductor earnings miss, financing risk on large infrastructure projects such as the Oracle and OpenAI Stargate cluster, and any disruption to the circular capital flow supporting AI infrastructure buildout. Private credit funding data center buildouts is flagged as a poorly understood risk where any hiccup could create contagion, and leveraged ETFs on semiconductor companies, including three-times leveraged products, represent a potential contagion vector where retail losses could cascade to hedge funds.

Nikhil identifies a decline in the compute crunch at the hyperscaler level as the most important signal to watch for overall AI health, noting that every hyperscaler has so far said it is compute constrained and will be for quarters to come. He predicts that one major hyperscaler pulling back CapEx by at least 10 percent would put significant pressure on the upcoming earnings cycle. He views Nvidia's top line as having significant predictability but worries more about OpenAI's bottom line, citing reports that it needs to raise several hundred billion dollars. Anthropic hit approximately 64 billion dollars in annual run rate, with a prediction it could reach at least 100 billion dollars in revenue in 2027.

Two months ago many companies stopped paying AI subscriptions and shifted to usage-based pricing, with Cursor using Kimi K2.5 and Airbnb using Qwen cited as examples of shifting to open-weight Chinese models. There is concern that Nvidia Nebatron and Google Gemma model families may be insufficient replacements for DeepSeek, Kimi, and Qwen for startups seeking open alternatives, meaning many startups that do not want to pay OpenAI or Anthropic may have no adequate domestic open-source fallback. The Chinese government may also preclude the release of future open-weight models, mirroring accessibility debates in the US.

Windborn, founded in 2019 by four Stanford co-founders, collects proprietary atmospheric data via its own weather balloons and uses that data to build AI-powered weather forecasting models it describes as among the most accurate in the world. Cuts to the US National Weather Service have created a gap that Windborn is filling, and it sells weather data and forecasting to governments including NOAA and the Department of Defense as well as commercial customers. The company is cited as a concrete example of the proprietary data thesis, where owning the underlying data rather than the model layer is the durable competitive advantage.

Basu Trivedi notes that the pool of talent capable of fine-tuning, switching, or building their own models is finite, and that the number of companies and researchers that matter in AI appears to be concentrating even as it has become easier than ever to build something. There is an arms race among investors to find younger founders in an AI-native world, with first checks going to increasingly early-stage builders. The current environment is compared to the ZIRP era in that tourist money has been replaced by tourist founders chasing easy conditions, with the underlying observation that building a company remains hard and takes a very long time regardless of market conditions.

This summary was generated from the episode transcript and can contain mistakes.