PodBrowser
a16z AI

From Code Search to AI Agents: Inside Sourcegraph's Transformation with CTO Beyang Liu

Tuesday, 25 November 2025 · 2 min read · Listen to the episode ↗

The discussion centers on Sourcegraph's evolution from a code search engine to an AI agent, AMP, utilizing large language models to enhance coding efficiency. It highlights the integration of agentic LLMs for improved developer productivity and the critical balance between model performance and user interaction. Additionally, concerns about the reliability of AI models and the implications for market efficiency are addressed, alongside the need for clear regulations in AI and a focus on open-source development to foster innovation.

The conversation explores the current landscape of computer science and AI, noting that while developers report increased productivity, they find coding less enjoyable. Beyang Liu, co-founder of Sourcegraph, highlights the evolution of Sourcegraph from a code search engine to an AI agent, driven by the need to help developers understand code more efficiently. The integration of large language models (LLMs) into their technology has led to the development of AMP, a coding agent that improves pull request merge rates.

Sourcegraph's focus on large code bases has prompted the creation of AMP, which leverages agentic tool-use LLMs capable of robust reasoning. The decision to start from first principles for AMP aims to explore its disruptive potential, making it accessible to both professional developers and non-coders. The conversation also addresses the trade-offs between intelligence and latency in model performance, with a recognition of user preferences for interaction styles with coding agents.

The discussion emphasizes the critical components of AI models, including system prompts and feedback loops, and introduces the concept of an "agent" defined by user input and expected behaviors. Concerns about the reliability of AI models are raised, particularly regarding transitions between versions, but well-constructed agents can still deliver reliable outcomes. The dialogue touches on market efficiency, questioning whether speed or correctness is prioritized, and discusses AMP's two types of agents: a fast, ad-supported agent and a smart, usage-based pricing agent.

The conversation reviews early models in agentic tool use and current models like GPT-5, noting preferences for larger models for smart coding and smaller models for quick edits. The distinction between generalist and specialized models is made, highlighting that specialized agents improve efficiency for specific workloads. The future of software engineering is questioned, considering whether it will involve integrated development environments (IDEs) or agents operating on a command-line interface (CLI).

Both speakers reflect on the early narratives surrounding AI, critiquing the extreme framing of AGI and emphasizing that current models primarily perform pattern matching. They express concerns about the potential dependency on open-source models from China, especially if U.S. models do not improve. The conversation also addresses the cautious release of open-source models due to copyright and data usage concerns, noting that the U.S. has fallen behind in this area.

The discussion advocates for clear, nationwide regulations that target specific applications rather than broad existential risks, emphasizing the importance of maintaining competition at the model layer to prevent anti-competitive behavior. There is a sense of optimism that academia and industry will prioritize openness in AI development, fostering a collaborative environment for innovation.

This summary was generated from the episode transcript and can contain mistakes.