PodBrowser
The AI Daily Brief

How to Decide What Work AI Should Do for You: The AI Deputization Audit

Friday, 14 August 2026 · 4 min read · Listen to the episode ↗

This episode introduces the AI Deputization Audit, a five-dimension scoring framework for deciding which recurring tasks to hand off to AI entirely, keep with the human, or split into a duet where AI handles part of the work but the human stays closely involved. Most knowledge work today lands in that duet category, and the framework identifies specific blockers preventing fuller handoff, matching each to tools like computer-use agents or show-rather-than-tell recorders such as Grokbot.

The AI Deputization Audit is a framework for deciding which recurring work to hand off to AI, with the word deputization chosen deliberately over automation to reflect a relationship where humans remain selectively involved rather than fully removed. The audit scores tasks across five dimensions: how worth automating the task is based on frequency and time cost, how teachable it is, how quickly output can be checked, how serious errors would be, and whether the task must be done by that specific person. Scores of eight to ten qualify for full deputization, scores of zero to three should stay with the human, and scores of four to seven are duets where AI handles part of the work but the human stays highly involved. Most knowledge work today falls in the duet category.

Step four of the audit requires naming the specific blocker preventing handoff. Processes that span websites or legacy software with no API are solved by computer-use agents that click screens. Processes that are easy to demonstrate but hard to explain in a prompt are solved by show-rather-than-tell tools like Grokbot, which lets users record a browser task once and have the bot replicate it as a repeatable routine. Processes where AI lacks ongoing context such as account history or personal formats may be addressed by ambient observation tools like ChatGPT's computer history feature, which watches activity across apps and builds context without deliberate user action. Blockers involving taste, judgment, costly or irreversible mistakes, and human relationships are not solved by any of these new capabilities, and recording screens introduces a new security concern that may itself become a blocker where privacy was already a consideration.

Google released Gemini 3.7 Flash, running at 340 tokens per second, more than twice as fast as GPT-56 Luna, with a coding benchmark score improved from 48.6 to 65.3 percent over 3.6 Flash and prices cut by half. At 40 cents per task it costs the same as MuSpark 1.2 and roughly eight times more than ultra-cheap models. Cognition noted it delivers Sonnet 5 coding performance for less than half the cost, though Sonnet 5 has not been widely adopted. Brandon Galang placed it on the Pareto Frontier despite being technically edged out by GPT-56 Luna, which he would not use for coding tasks given its smaller size.

OpenAI introduced Ultra Fast Mode for GPT-56 Sol, claiming frontier intelligence at 14 times the speed and 750 tokens per second, more than twice the pace of Gemini 3.7 Flash. The mode uses optimizations built to run models on Cerebras hardware, making available infrastructure the limiting factor. It is currently API-only for select customers and targets latency-sensitive workflows including real-time voice, customer support, coding, financial research, and security response. Pricing was not disclosed.

An AlphaSense study tested US and Chinese models on real-world financial analysis tasks using earnings call transcripts, SEC filings, and news articles. GPT-56 Sol completed tasks roughly 13 percent cheaper than Kimi K3 with quality scores about 20 percent higher. Opus 5 delivered lower quality than Opus 48 while costing more than five times as much. Gemma 4 and Inkling matched GLM 5.2 quality for less than one fifth the cost. One analyst noted that some more expensive models ended up cheaper overall because they used tokens more efficiently.

OpenAI chief revenue officer Denise Dresser departed after nine months, days after former COO Brad Lightcap also left, and former CEO of Apps and AGI deployment Fiji Simo left in July for health reasons. Dresser and CFO Sarah Friar had been seen as a business-focused team preparing OpenAI for an IPO now delayed until next year. Her replacement is Dolly Rodgic, former president and COO of Wiz. Sources told Axios that president Greg Brockman has been building his own leadership team, which may explain the pattern of exits.

Microsoft's Windows Recall, announced around early 2024, took encrypted local screenshots every few seconds and faced significant backlash over privacy concerns before being relaunched as an opt-in feature with finer controls. ChatGPT's computer history is seen as overcoming prior limitations by recording interaction events rather than constant screenshots, and its value proposition of having AI perform work is stronger than Recall's original pitch of simply helping users find things they had previously seen. John O Neill, a plumbing company owner with no engineers, went from zero to automated dispatch and office chores within 24 hours using Grokbot. High-efficiency workers often find that pausing to learn new AI tools feels like a short-term loss even when long-term savings are significant, and duet tasks are expected to migrate into the fully deputized category as AI improves at learning individual workflows.

This summary was generated from the episode transcript and can contain mistakes.