SF Compute: Commoditizing Compute
Friday, 11 April 2025 · 4 min read · Listen to the episode ↗
The conversation centers on SF Compute's strategy to commoditize GPU compute resources, highlighting Core Weave's long-term contract approach amid price sensitivity in the GPU market. It critiques hyperscale cloud providers while suggesting that proprietary product development could enhance profitability. Additionally, the dialogue addresses the dynamics of GPU demand, pricing strategies, and the emergence of financialized GPU marketplaces, emphasizing the potential impact of AI and blockchain technologies on compute resource management and trading.
Alessio welcomes listeners and highlights Evan's influence in the San Francisco tech community. Evan discusses SF Compute's journey and Core Weave's recent IPO, emphasizing its long-term contract strategy in the GPU market. He explains that Core Weave's success is rooted in understanding customer behavior, where GPU customers are more price-sensitive due to the high costs of hardware. This sensitivity means that even minor price changes can have significant financial implications.
The conversation emphasizes the importance of long-term contracts for maximizing profits, suggesting that targeting low credit risk customers can lead to better financing terms. Core Weave is compared to banks or real estate companies, operating differently from traditional cloud providers. The discussion critiques hyperscalers like Microsoft and AWS, predicting they may face losses when reselling Nvidia GPUs due to low margins.
Challenges in the GPU market are highlighted, including the impact of small customers seeking discounts on profitability. The speaker suggests that hyperscalers could improve margins by developing proprietary products rather than reselling GPUs. The relationship between interest rates and GPU sales is analyzed, with lower rates allowing for aggressive pricing strategies.
The conversation critiques the traditional CPU cloud model for its inadequacy in the GPU market, warning that high pricing strategies can lead to customer loss. Many companies are shifting to low-priced short-term contracts, which is detrimental in the current financial climate of high interest rates.
CoreWeave's business model is central to the discussion, particularly regarding why major players like Nvidia and Microsoft haven't created similar services. Nvidia avoids competing with its customers to maintain profitable relationships, while Microsoft’s entry could lead to customer concentration issues.
The financial challenges faced by companies like Digital Ocean and Together are discussed, particularly the difficulties of coupling software and hardware. Long-term contracts and hardware investments can lead to significant debt, complicating the achievement of high margins. The conversation also emphasizes the risks of purchasing large GPU clusters and suggests treating GPU clouds as real estate businesses.
The discussion touches on the importance of distinguishing between trading and inference workloads, as well as the price sensitivity in GPU software. The origins of SF Compute reveal a focus on training models for music and audio, leading to a reactive strategy of subleasing unused capacity and creating a liquid GPU market.
Strategies for increasing cluster utilization by selling idle clusters are explored, with a focus on generating revenue through market pricing and long-term contracts. The conversation references a previous issue with the H100 GPU glut, highlighting the complexities of supply and demand dynamics influenced by VC funding.
Speaker 1 expresses skepticism about market speculators and predicts that test time inference will significantly increase compute usage, particularly in the bio and pharma sectors. They question whether increased demand will impact compute prices, depending on new chip rollouts.
The challenges faced by grad students as customers of traditional cloud services are noted, emphasizing the significance of grants that need to be spent quickly. The discussion includes VCs providing GPU clusters, addressing credit risk and the difficulties startups face in securing loans for hardware.
The conversation shifts to pricing dynamics, where one speaker inquires about the oddity of one-week pricing being higher than one-day and one-month options. The response clarifies that preemptible pricing is generally cheaper than non-preemptible options, allowing users to buy compute on an hourly basis.
A strategy for purchasing compute resources emphasizes continuous buying without interruptions, highlighting the benefits of preemptible pricing. The idea of creating a more tailored compute system beyond what hyperscale providers offer is explored, referencing OpenAI and Claude's batch API.
The potential for financial engineering in the compute market is discussed, including plans to establish a financialized GPU marketplace. The conversation also touches on de-risking technical and financial perspectives in computing, emphasizing the importance of auditing clusters and identifying unreliable hardware.
Speaker 1 discusses unique interactions in data centers, noting manageable common issues. They emphasize the importance of standardizing commodity contracts and describe the process of creating a "this or better" spec list for storage. Speaker 2 stresses the necessity for standard contracts in trading, particularly for GPUs.
The impact of inflated venture capital markets on startups is discussed, with pressure from VCs for high valuations potentially leading to failures. A proposed solution is the introduction of futures to enhance security in the economic system. Expectations management is a key theme, with the company intentionally setting low expectations for their product to ensure customer satisfaction.
Speaker 1 reflects on advice about perseverance and the challenges of developing an email client in a competitive landscape. They highlight hiring for key roles, emphasizing the meaningful impact of work at SF Compute, particularly in supporting significant projects such as cancer research.
This summary was generated from the episode transcript and can contain mistakes.