PodBrowser
This Day in AI

Gemini 3 Flash, GPT-Image-1.5, Skills vs MCPs, and Our 2025 Model Reviews - EP99.29

Monday, 22 December 2025 · 5 min read · Listen to the episode ↗

The podcast discusses the launch of Gemini 3 Flash, showcasing its advanced reasoning capabilities and improved output even with a cost increase. It also reviews GPT-Image-1.5, emphasizing its design limitations compared to the more reliable Nano Banana Pro. Additionally, the distinction between skills and Model Control Protocols (MCPs) is explored, underscoring their potential synergy for enterprise AI applications and the shift toward specialized capabilities in AI workflows by 2025.

Chris discusses the release of Gemini 3 Flash, which features a reasoning-focused version and a standard model. He notes that while Gemini 2.5 Flash was effective for tasks like summarization, Gemini 3 Flash has demonstrated impressive benchmarks despite a price increase. The input cost is 50 cents per million tokens, and the output cost is $3 per million tokens. Chris finds Gemini 3 Flash to have a tight output that accurately meets requests, suggesting it feels like a new model altogether.

The conversation shifts to user experiences with Gemini Flash 3, where one user switched from Opus 4.5 for quicker issue resolution. Both Gemini Flash 3 and Gemini Pro 3 are praised for their intelligence and efficiency, although reliability concerns persist, with users frequently switching models due to performance issues. Complaints about Opus 4.5 indicate a decline in capability over time, and GPT 5.2 is also viewed as ineffective, leading some users to revert to Sonnet 4.5 or Grok 4.1 for more dependable performance.

The podcast also covers the recent release of GPT-Image-1.5, seen as a response to the success of Nano Banana Pro. While GPT-Image-1.5 is slightly cheaper, Nano Banana Pro is regarded as more reliable for character consistency and infographics. Chris speculates that OpenAI's release may be reactionary and questions whether users will prefer it over Nano Banana Pro, which has made significant advancements in text accuracy and prompt flow.

Participants discuss the performance of image models, particularly comparing Nano Banana Pro and GPT-Image-1.5. One participant notes that their vision affects their evaluation of these models, but they agree that GPT-Image-1.5 appears poorly designed. While a holiday card created with GPT-Image-1.5 is better than expected, it still fails to accurately represent the participants. In contrast, Nano Banana Pro receives praise for its quality, especially in editing images to meet specific themes.

The conversation shifts to recent releases, including the Firecrawl agent, which effectively addresses past limitations in data extraction, allowing users to compile histories of model releases. The speaker highlights its reliability for research, emphasizing its ability to perform tailored queries and verify data, enhancing the quality of information retrieved compared to traditional search methods. Trust in the Firecrawl agent has grown, making it a valuable tool for research and code writing.

Speaker 1 praises the Firecrawl agent for its efficiency in quick searches and documentation reading, noting its upcoming integration into SimTheory as an MCP. Speaker 2 introduces the Gemini Deep Research Agent, which excels in structured research processes, gathering information and synthesizing findings into comprehensive reports with citations. They conclude that building context is vital for effective AI task execution, underscoring the importance of powerful research tools in planning.

A key discussion point is the distinction between skills and MCPs, questioning whether enterprises should prioritize building skills, MCPs, or both. An article from VentureBeat is referenced, noting Anthropic's launch of Enterprise Agent Skills, which challenge OpenAI in the workplace AI space. Skills are described as having a "front matter" that provides context for their use, making them suitable for repeatable processes and enhancing accuracy.

The philosophical shift in AI is highlighted, with skills representing a new way to conceptualize AI capabilities. Skills are expected to drive tool use, suggesting a model with a general-purpose agent equipped with specialized capabilities. The conversation also touches on the potential merging of skills and MCPs due to their overlap, emphasizing the ecosystem around MCPs, including secure hosting and enterprise data integration.

Enterprise-level teams are increasingly sharing assistants tailored for specific tasks, leveraging multiple assistants to enhance efficiency. The discussion highlights the development of skills using Model Control Protocol (MCP), including a theme maker for simulation theory. A new skill has been created to streamline theme creation, resulting in more concise prompts, while document skills allow models to generate files in a Python sandbox with precise instructions.

The conversation shifts to AI model releases, with Mike expressing surprise at the rapid pace of developments. The Gemini 2.5 Pro was confirmed to have been released in March 2025, marking a significant moment for Google. The year showcased impressive advancements in model intelligence, with Gemini 2.5 Pro emerging as the best model, while Llama was deemed the worst. Predictions for 2026 include expectations for all model providers to support skills by early January.

The podcast discusses the computing burden faced by model providers and the potential for longer-running processes in AI models. There is a debate on whether model providers will assume more computing responsibilities or if products like Sim Theory will manage processes independently. Concerns are raised about the implications of these longer processes, suggesting they may favor model companies over users.

Looking ahead, predictions for 2026 include the rise of "agentic workflows" that extend beyond chat interactions, with organizations expected to develop their own AI strategies rather than relying on generic subscriptions. There is skepticism about the future of AI businesses, alongside a desire for organizations to maintain control over their workflows. The discussion touches on potential competition from companies like Anthropic and anticipates significant improvements in open-source models by 2026.

Anticipation for a shift in workflows by 2025 is expressed, focusing on planning phases with agents capable of executing tasks. The rise of adaptable workers managing multiple tasks simultaneously is expected to enhance productivity. The future of OpenAI is debated, with skepticism about its competitive edge and a negative outlook on a predicted new app store. The podcast critiques the integration store for web apps, suggesting it lacks value compared to MCPs.

As the podcast wraps up, there is gratitude expressed to listeners and reflections on the year in AI, with a sentiment that the year's advancements have been average. The closing remarks offer a light-hearted take on the year's AI developments and a hopeful outlook for the future.

This summary was generated from the episode transcript and can contain mistakes.