A demonstration of twelve agents handing work down a chain is easy to build and easy to film. A product that keeps a user past the second week is harder to make, and it usually runs on one agent behind one screen.
The revenue evidence
Revenue is already sorting the two groups. OpenAI reported 900 million weekly ChatGPT users in February 2026, up from 800 million in October 2025, with about $25 billion in annualized revenue. Claude Code, a terminal program that launched in May 2025, was reported past $1 billion in annualized revenue about six months later, and near $2.5 billion by mid-2026. Cursor passed $2 billion in reported annual recurring revenue in 2026 with a few hundred employees. Each of those products keeps its machinery out of sight. None of them puts the agent count on the pricing page.
One number gives a conference talk away: the agent count. A team runs forty agents across twelve tools; the room nods; nobody asks what the forty do at nine on Monday. Ask a customer what they open in the morning and the answer is one thing: a terminal, a chat box, a browser panel, an inbox. The agents are plumbing. The count is an internal metric, and it has never appeared on an invoice.
The console era
The 2025 enterprise budget mostly went to consoles: agent inventories, run histories, permission matrices, cost dashboards. On 25 June 2025, Gartner forecast that more than 40% of agentic AI projects would be canceled by the end of 2027, naming cost, unclear value, and weak risk controls. MIT's GenAI Divide study that August found 95% of enterprise pilots delivered no measurable impact on profit and loss. A console asks the user to operate the machine, which is the labor the user came to avoid.
Coordination burden compounds with the count. Each added agent brings a prompt, a credential, a retry policy, and a fresh decision about who owns which failure. The pilot stalls for a plain reason: the console is complete the day it is built and never finished being configured.
Public enthusiasm for swarms has cooled too. Berkeley's paper "Why Do Multi-Agent LLM Systems Fail?" (March 2025) found that gains from multi-agent systems over single agents were often minimal on popular benchmarks, and catalogued how the systems break. Cognition argued in June 2025 that long-running production work favors a single-threaded agent, since parallel swarms compound errors and lose context across handoffs. Anthropic published a multi-agent research system the same month that beat a single agent by 90.2% on an internal evaluation, at roughly 15 times the token cost of a chat turn. That result is documented, and so is its bill. It landed where the question was open-ended and parallel search paid for itself. Research is one of the few tasks shaped that way.
This is where the Avengers model turns expensive. Assemble a specialist for every threat: a planner, a researcher, a coder, a critic, each with a narrow brief and deep expertise. On a slide it reads as velocity. In production it behaves as a single node. Lose the critic and nothing gets evaluated. Lose the researcher and the planner works from stale memory. Teams reproduce the shape whenever they port an org chart into a topology diagram.
Products that retained users shipped small surfaces for large systems. Claude Code is a command in a terminal. ChatGPT is a text field. Cursor is an editor with a panel. GitHub launched Agent HQ on 28 October 2025 as one platform where agents from several vendors run inside the existing workflow, with a single control layer for administrators, and by February 2026 both Claude and Codex ran inside it. Anthropic's 2026 post on Managed Agents describes a hosted service that exposes a small set of interfaces built to outlast any particular implementation, and splits the agent into a brain and hands so the parts can be swapped underneath. What a person touches is the stable layer.
Buyers test that layer during a trial. A procurement team in 2026 asks for a demo account, a pricing page, and a name for the support call. The vendor's org chart never appears in the evaluation. What gets forwarded inside the company is a screen that produced a result in one sitting.
What to ask first
Three questions separate the two piles, and they are cheap to ask before the next sprint.
- Can a new user reach a useful result without configuring a single agent? If the first session demands a topology decision, the product has not been designed yet.
- Remove any one agent, model, or vendor. Does the system keep working? A stack that holds together only when every specialist is present is a demo with a maintenance contract.
- Does the customer ever need to know how many agents are running?
The count belongs to the operator. The buyer keeps the row that appeared in the ledger.
Scale stays invisible from the outside, which is why it never sold anything on its own. The interface is what ships, what gets reviewed, and what a buyer can hold. Build the surface first. The machinery underneath can change every quarter while the customer keeps pressing the same button.