Hermes Agent App Clearly Explained (and how to use it)
Saturday, 6 June 2026 · 4 min read · Listen to the episode ↗
In this episode, Alex Finn walks through the Hermes desktop app and explains why he considers it the best AI agent experience available, surpassing Claude and replacing earlier setups built around Telegram, Signal, and iMessage bots. He addresses the most common complaint, monthly costs often reported around one thousand dollars, tracing the problem to single-thread conversations that inflate token usage and arguing that the app's automatic session creation alone can cut costs three to four times.
Hermes desktop app is described by Alex Finn as the moment Hermes overtook Claude as the best AI agent experience, replacing prior setups through Telegram, Signal, and iMessage that required confusing separate group chats with bots. Finn considers it the single best way to use Hermes and frames the architectural choices in the app as directly addressing the most common user complaints.
The most cited complaint about Hermes is cost, often reported at around one thousand dollars per month. Finn attributes this primarily to users running everything inside a single conversation thread, which forces every new message to carry the entire prior conversation as context, inflating token usage substantially. Hermes desktop automatically creates a new session each time, and Finn argues this alone can reduce costs by three to four times.
The app supports multiple profiles, each carrying its own model, skills, personality file, memories, and session history. Finn runs a Default Hermes profile on Claude Opus 4.8 for high-level thinking and strategy, a profile called GP Teamies on GPT-4.5 or GPT-5 for coding because he considers it stronger than Opus 4.8 for that task and it carries higher usage limits, and a profile called Quinn running a local model on an Nvidia DGX Spark for quick web research at no cost. Finn recommends organizing profiles by model strengths rather than by job role personas, arguing that creating fifty specialized role-based profiles generates cognitive overhead and that a sufficiently capable model can handle diverse tasks without a dedicated agent for each.
Hermes ships with over 150 skills installed by default, and each skill adds context that raises cost per message. The app also generates new skills automatically in the background based on user behavior. A tool sets feature allows grouping multiple skills for complex tasks. The artifacts feature automatically organizes all links, images, media, and files shared with the agent into one searchable location without manual filing instructions, effectively productizing a second brain use case that users were previously building manually through custom instructions.
The cron section lets users view and create scheduled tasks with a single click, replacing a CLI-based setup that frequently failed silently. Finn describes cron jobs and routines as essential to extracting value from Hermes and recommends a reverse prompt approach where the user provides personal context and asks the agent to generate the best possible prompt, noting that a generic morning brief request often returns headlines from months or even a year ago due to model memory cutoffs. Finn runs a cron job called Daily AI Business Opportunity Scan that executes every twenty minutes using the local Qwen 3.7 model, reading Reddit and X to surface other people's challenges and business opportunities. Output includes a custom dashboard with challenges, source threads, founder comments, why Finn is positioned to solve the problem, and a suggested first move, with an option for the agent to automatically build a micro-SaaS prototype. Running this on a local model makes the every-twenty-minute cadence free, whereas the same job on a cloud model would incur significant ongoing costs.
Sub agents in Hermes are copies of the main agent sharing the same skills, personality, and settings, suited for running the same capability across multiple parallel strands simultaneously, such as building several features of a micro-SaaS at once. Profiles carry distinct skill sets, context, and memories and are appropriate when a fundamentally different capability is needed. A YouTube workflow example uses Qwen for research, Hermes with Opus 4.8 for scripting, and a GPT profile for thumbnail image generation. Hermes is also architected to swap models and thinking settings dynamically, so new model releases can be substituted immediately without waiting for a platform update, which Finn cites as a primary reason for switching from OpenClaw, where models are described as hard-coded and updates as unfocused and unreliable.
The current hardware recommendation for local inference is the Nvidia DGX Spark, a plug-and-play device with 128 gigabytes of unified memory priced at approximately four thousand eight hundred dollars after a recent price increase. The Mac Studio was previously recommended but is largely sold out with remaining units carrying smaller memory configurations. Finn recommends proving Hermes's value in daily use before committing to hardware at that price. His broader argument is that the highest-value application of Hermes agents is deploying them to solve real challenges for other people rather than personal experimentation, and that AI tools should be evaluated as investments capable of producing returns rather than compared to entertainment subscriptions.
This summary was generated from the episode transcript and can contain mistakes.