Cohere Labs counted 696,291 tools from 123,069 public Model Context Protocol server listings, then classified 18,058 of 694,411 analyzed tools as capable of performing their nearest occupational task. That is 2.6005%.
The result is a sharp check on agent-directory abundance. It is not evidence that those tools completed work in production. The released Agentic Task Ecosystem dataset compares documentation text with occupational task statements. It does not run the tools.
That distinction changes what the number can support. The study measures the public supply of callable integrations. It does not measure adoption, reliability, workflow completion or job loss.
The denominator is documentation, not deployed software
Cohere collected listings on 7 and 8 May 2026 from MCP World, MCP Store, Glama, Smithery, the official MCP registry, mcp.so and mcp-repository. A tool is one named capability parsed from a server’s README or documentation, such as listing issues or running a query.
The researchers say they deduplicated servers republished across directories. The public package contains the resulting 123,069 server rows, not the pre-deduplication lineage, so that stage cannot be replayed from the release.
Tool duplication remains. The raw table has 696,291 rows. The analysis excludes 1,876 duplicate tool IDs, three empty names and one embedding miss, leaving 694,411 rows. A Clarqo count also finds 76,103 repetitions beyond the first among normalized, nonempty name-description pairs across server IDs. That does not prove every repeat is a cloned capability. Generic operations often share names and descriptions. It does mean 696,291 is not a count of 696,291 unique things an agent can do.
What the 2.6% test actually measures
Each analysis-set tool was paired with its nearest task that Cohere had marked software-performable in O*NET. A language model then assigned good, partial or bad. A good match means the described tool performs the named action. Partial includes information retrieval or one step in work that a person still coordinates. Bad means the selected task is unrelated.
The released labels contain 18,058 good matches, 591,826 partial matches and 84,522 bad matches. Five rows have no label.
The arithmetic is reproducible. The classification is not yet independently reproducible. As of 10 September, the dataset card does not name the evaluator model, publish its judging prompt or provide a validation sample for the match labels. It says a technical report is still coming.
Nor is there a disclosed cosine-similarity threshold that creates a good match. Every tool receives its nearest task before the model judgment. In the released rows, good labels appear at similarities as low as 0.3123, while bad labels reach 0.7296. The categories overlap. The operative threshold is therefore the undisclosed judge’s interpretation of the task and description.
Incomplete documentation contributes noise but does not explain the entire result. Descriptions are missing for 48,954 analyzed tools, about 7%. Their good-match rate is 2.18%, compared with 2.63% where a description exists.
The occupation figures use a broader map
The source task table contains [18,796 ONET 29.2 statements](https://www.onetcenter.org/dictionary/29.2/text/task_statements.html) across 923 US occupations. Cohere, not ONET, marks 9,423 of those tasks software-performable. Good tool matches touch 1,380 statements, or 14.65% of that researcher-defined subset.
Direct good matches reach 491 occupations. Reproducing Cohere’s claim that 419 occupations have no agentic activity requires a broader union: occupations reached by a good task match plus occupations assigned through the separate cluster taxonomy. That union reaches 504 occupations. Subtracting 504 from 923 produces 419.
Both findings are valid descriptions of the released tables. They answer different questions. The 2.6% figure asks how many public tool descriptions pass the strict task-match judge. The 419 figure asks whether any public tool activity appears through either task matching or the broader cluster assignment.
O*NET adds another boundary. It describes US occupations, and the package uses release 29.2 from February 2025. The current O*NET release archive shows newer versions existed when the MCP listings were collected. The benchmark is not a global or live inventory of work.
Integration supply is not workflow completion
An MCP directory can grow quickly because a single server exposes many small functions. Complete work needs more: authentication, sequencing, state, policy, exception handling, retries and a test that the intended outcome occurred. Tool count measures the available verbs. Automation depends on whether a system can compose them into a reliable sentence.
Private enterprise servers are absent. Cohere argues that this omission makes 2.6% a floor because private tools may skew toward back-office processes. The missing tier makes both numerator and denominator incomplete. Without its match rate, it cannot mathematically establish a floor.
The commercial context also belongs in the interpretation. Cohere sells North Automations, a product for orchestrating enterprise workflows. That does not invalidate the dataset. It raises the value of publishing the evaluator, prompt and validation evidence behind its most quoted result.
For buyers, the practical lesson is narrower than an employment forecast and more useful. A directory listing proves that someone documented an interface. It says nothing about permissions, success rates or recovery when the fifth step fails. Procurement should count completed, audited runs, not tool names.
The 2.6% result is best read as a reported classification of public MCP supply. It shows why abundant integrations do not yet equal abundant automation. Until the judging method is fully released, it does not show that 97.4% of tools failed a real task.
Discussion
Sign in to join the discussion.
No comments yet. Be the first to share your thoughts.