PodBrowser
a16z AI

From Vector Databases to Knowledge Engines: The Next Layer of AI

Tuesday, 5 May 2026 · 3 min read · Listen to the episode ↗

The discussion highlights the shift toward knowledge engines in AI, exemplified by Pinecone's evolution from a vector database, which lacks context, to a knowledge engine that enhances efficiency in data retrieval. The introduction of NoQL streamlines agent communication, increasing task completion rates significantly. Additionally, there is an emphasis on the importance of trust and security in AI deployments, with plans for democratizing access and standardizing tools, paving the way for a Cambrian explosion of vertical AI applications.

Peter Levine notes a significant shift in user interaction from human users to agents, with 85% of agents focused on knowledge retrieval. Ash Ashratush emphasizes that systems originally designed for human interaction create inefficiencies for agents, which often lack context and issue multiple queries, leading to a bottleneck in data retrieval. Agents frequently fail to complete tasks, with completion rates below 50%.

The conversation introduces Pinecone's evolution from a vector database to a knowledge engine, crucial for addressing agents' needs. While a vector database functions like a library, a knowledge engine acts as an expert, providing context tailored to specific tasks. This distinction is vital, as agents using vector databases struggle with context and efficiency. Ashratush describes the process of setting up a knowledge engine, where humans provide context and expected answers, akin to a compiler refining data into new artifacts. The complexity of organizing data and its foundational role in training a knowledge engine is highlighted, along with a new protocol enabling agents to specify how they want responses.

The transformative impact of large language models (LLMs) and knowledge engines on task completion and efficiency is discussed, with task completion rates increasing from 50-60% to over 90% and knowledge retrieval time reduced by 85%. This efficiency is reflected in a 40-90% reduction in token usage, resulting in cost savings and improved performance. Pinecone's development of an operations agent exemplifies the shift from traditional ETL pipelines to a dynamic context compiling approach, enhancing user experience.

The introduction of NoQL (Knowledge Query Language) facilitates communication between agents and knowledge engines, structured around intent, time, and governance to ensure purposeful, timely, and trustworthy queries. The economics of building knowledge retrieval stacks are discussed, with Pinecone's infrastructure offering an 85% reduction in code requirements, simplifying application development. Pinecone aims to foster a developer-centric community, with plans to make NoQL publicly available and standardize it for agent applications.

Looking ahead, there is anticipation for a Cambrian explosion of vertical AI applications driven by standardization, allowing developers to focus on specific business cases. Knowledge retrieval is positioned as a crucial component of agent work, underscoring its importance in the evolving landscape of AI. Speaker 1 emphasizes the critical role of trust in deploying AI within large enterprises, advocating for a trusted knowledge engine that ensures traceability and citations for data sources, essential for explainable AI.

Despite the rapid pace of AI demos, users encounter barriers such as ETL complexities and security concerns, prompting a call for simplified processes. Speaker 1 notes that pricing strategies will focus on knowledge curation, extraction, and task completion rather than infrastructure costs, suggesting that pricing could be linked to token savings. They also mention plans to expand the interface for agents, aiming for a greater number of agents than human users.

A new pricing structure for the core database is announced, aimed at democratizing access and broadening economic opportunities. Speaker 2 expresses surprise at the knowledge engine's impact on token usage, while Speaker 1 reflects on historical technological shifts, likening current advancements to past transitions in CPU and networking optimizations. The conversation underscores the importance of efficiently moving data and optimizing processes in AI, acknowledging that the industry is still developing with numerous opportunities ahead. Trust and security remain critical areas for improvement, alongside historical themes of process offloading, security, and data governance. The mention of MCP interfaces as a standard for data access reflects ongoing issues related to token optimization, fostering excitement about the evolving landscape of AI and technology.

This summary was generated from the episode transcript and can contain mistakes.