PodBrowser
Latent Space

Why Every Agent needs Open Source Cloud Sandboxes

Thursday, 24 April 2025 · 6 min read · Listen to the episode ↗

The discussion emphasizes the necessity of Open Source Cloud Sandboxes for AI agents, aiming to streamline the development process and enhance user experiences. Key insights include the expansion of AI applications beyond simple code execution into data analysis, supported by a flexible sandbox infrastructure that allows rapid environment changes. Additionally, the conversation highlights the importance of understanding the unique challenges and opportunities in utilizing sandboxes effectively for LLMs and the evolving landscape of cryptocurrencies and blockchain technology in this context.

Alessio, the CTO at Decibel, introduces co-host Swix and guest Váček from E2B, discussing their investment in E2B and Váček's history with DevBook, which aimed to enhance developer experience through interactive documentation and playgrounds. Váček explains the transition from DevBook to E2B, highlighting the creation of an AI agent capable of running code and managing integrations, which was open-sourced with a focus on sandbox technology.

The initial hypothesis for E2B was that code-generation agents required an environment for code execution. A test project deploying a small agent in a sandbox gained unexpected popularity, shifting the audience focus towards those interested in building projects with tools like Lovable. Váček reflects on the decline in model performance post-launch, noting discrepancies between initial demos and actual capabilities, which led to unrealistic user expectations. The go-to-market strategy emphasized code interpretation and data visualization in a headless Jupyter-like environment, with Python emerging as a strong language for data visualization.

The conversation highlights the growth trends in sandbox usage for AI agents, with applications expanding beyond simple code execution to include data analysis and deep research agents. The speakers suggest viewing the sandbox as a runtime environment for LLMs and agents, with various use cases like file creation and task management. Growth statistics reveal a dramatic rise in sandbox usage, from 40,000 in March 2024 to 15 million in March 2025.

The relationship between infrastructure and applications is discussed, noting that while 2024 focused on agents not fully utilizing the sandbox, 2025 will require more features to support LLMs. The importance of being agnostic to LLMs is emphasized, with aspirations to provide foundational technology akin to Kubernetes, but with an improved developer experience. Potential development strategies include browser emulation, virtual machines, or custom Python sandboxes.

The necessity of Open Source Cloud Sandboxes for agents is emphasized, particularly in simplifying the development process for AI engineers. The initial marketing approach faced challenges as potential users struggled to grasp the product's application, leading to a shift in focus towards "code interpreting" to resonate with familiar concepts from OpenAI. Educating the market is crucial, with team members working to demonstrate specific use cases that help developers envision practical applications.

The primary audience for these sandboxes consists of AI engineers and web developers, particularly those using JavaScript and TypeScript. The platform aims to streamline tools for product developers, allowing them to concentrate on building rather than managing infrastructure. The conversation also touches on the coexistence of GPT wrappers and ModelLabs, indicating a growing recognition of their potential profitability.

Usage statistics reveal a significant preference for Python in tasks like code interpreting, while JavaScript is favored for application development. The discussion raises critical questions about the need for an AI-focused cloud environment, highlighting the unique value proposition of a dedicated virtual machine that operates independently of traditional cloud models. This approach allows for rapid sandbox termination and replacement, enhancing security by isolating untrusted code execution.

The platform supports multiple programming languages and offers customizable server capabilities within sandboxes. Balancing speed, security, accessibility, and observability is a key challenge, particularly for AI engineers. The system's infrastructure is designed to be flexible, allowing users to switch runtimes without restarting sandboxes, with the ultimate goal of enabling LLMs to autonomously manage the infrastructure.

The conversation highlights the challenges faced in implementing a billing project, particularly the reluctance of engineers to take on the task, which ultimately required a senior engineer's involvement. Despite being a board-level objective, the project took a year to complete due to its complexity. The discussion emphasizes that the difficulties in pricing and billing stem from integration and implementation issues rather than the business model itself, with concerns about customer satisfaction when enforcing billing limits.

Insights into agent usage reveal that early-stage startups may run agents for extended periods, complicating billing without aligning with product-market fit. The CTO of Orb's perspective on pricing includes the concept of reselling tokens, with market dynamics showing significant price variations among companies. The idea of "bring your own key" solutions has not gained as much traction as anticipated, and economic considerations are shifting towards evaluating the value delivered by agents rather than merely comparing costs.

Technical discussions focus on the development of forking and sandbox checkpoints, which allow for pausing code execution while retaining context. This feature enables parallel problem-solving with multiple agents, where each node represents a snapshot sandbox, facilitating forking and checkpointing. Interest in forking and merging processes is noted, with the team aiming to provide a toolkit for users to monitor successful paths.

The conversation also addresses Managed Control Points (MCPs), noting a lack of focus on them despite their relevance in AI infrastructure. There is uncertainty about their current utility and application, suggesting that users are still exploring effective uses, with a call for a clear strategy regarding MCPs and the importance of having an API-first approach.

The discussion emphasizes the necessity for agents to utilize Open Source Cloud Sandboxes, particularly through the ability to spin up E2B instances seamlessly. A dedicated domain is proposed to facilitate account creation and sandbox launching, supporting the development of complex applications. An API-first approach for LLMs is highlighted as crucial for aligning with new dashboard infrastructure.

Concerns are raised about the potential negative impact of adapting the human web for agents, suggesting that separate interfaces for human and agent interactions may be necessary to address monetization challenges. The conversation acknowledges that new technologies often replicate existing structures rather than innovate, with a recognition of the complexities of real-world applications.

The discussion also covers the evolution of technology, including the transition from mobile web domains to potential new formats for LLM experiences. The importance of understanding E2B applications, such as AI data analysis and generative UI, is emphasized, alongside the need for broader platform support beyond Linux. Future discussions regarding the Open R1 project and collaboration with academics for model training are mentioned, highlighting how E2B employs sandboxes during the reinforcement learning phase.

E2B's capability to run multiple sandboxes for training without expensive GPU clusters is recognized, aligning with the lifecycle of AI agents. The unexpected use case of E2B in model training is noted, along with plans for a startup and research program for universities. The conversation touches on the challenges in the GPU market and the potential for GPUs to enhance data analysis and machine learning model training.

The company aims to leverage LLMs for application development and deployment, envisioning a comprehensive platform tailored for LLMs. The team's relocation to San Francisco is motivated by the desire to engage directly with users in the AI hub, facilitating quicker feedback implementation. Their customer engagement strategy has evolved from direct interactions to a focus on automation and scaling.

The conversation highlights the significance of in-person interactions, particularly in San Francisco, where opportunities for user engagement are greater. Direct feedback from users is emphasized as crucial for accelerating growth and building relationships in the B2B sector. The preference for face-to-face meetings over virtual ones is noted, as they foster better engagement. The discussion also touches on ongoing hiring efforts, recognizing the talent pool in Europe, especially in the Czech Republic, and current hiring needs across various roles.

This summary was generated from the episode transcript and can contain mistakes.