Robin Saige we check the tools AI agents call

The state of the tool economy

a data essay · census 2026-08-14 · every number is a live door · grades ● observed ◐ derived ○ reported

One sentence: the supply side is sprinting, the demand side hasn't shown up, and the layer that would connect them — accountability — is almost entirely missing. Everything below is measured, clickable, and re-derivable.

What was counted — the denominators, up front

21,988 = every name the registry has ever listed (cumulative, ○ reported — not a health population). 11,094 = servers probed in the latest census — every share below uses THIS denominator unless it says otherwise. Subject classification covers ~76% of it ◐; multi-label applies at subject grain only. Counts printed here equal what their doors return.

1 · A land grab, half the flags on nothing

21,988 names ever listed (cumulative ○); 11,094 probed in the latest census; only 6,040 answer a real handshake. In one week, over a hundred servers died while more arrived. Listing is free; maintenance isn't — the registry behaves like early npm: publish-and-abandon is the dominant lifecycle.

2 · The abundance is partly a wardrobe

One gateway answers for 1,312 registry names — over a fifth of everything alive. The top five multi-server operators carry a similar share again. Any market-size claim built on listing counts is inflated before it is spoken. The concentration section →

3 · Mostly a read-layer, not an act-layer

By kind: 5,126 retrieve public facts · 1,906 take actions · 380 generate · 105 predict. This wave of MCPs overwhelmingly wraps existing APIs for agents to read. Even in media, wrappers outnumber generators three to one.

4 · Where it's being built — the sector map

sectorserversshare of probeddoor
agent infrastructure2,13219.2%browse ◐
finance & assets1,89217.1%browse ◐
work & personal1,0629.6%browse ◐
government & public data1,0299.3%browse ◐
knowledge & media9939.0%browse ◐
software & infra6666.0%browse ◐
commerce & industry4363.9%browse ◐
environment & place2712.4%browse ◐

Three readings. The machine economy's first citizen is crypto — the largest nameable sector, and the same community building agent-money (x402). The second-biggest sector is the ecosystem building for itself — memory, routing, LLM tooling: more picks and shovels than mines. And government & public data is a real belt1,029 servers of civic portals, GIS, law, and trade: the most answer-keyed corner of the whole economy, and where truth-checking can grow fastest.

5 · Beautifully formed, unaccountable

Most servers are perfectly legible to a harness — formation is solved. But 2,239 sit at take its word, and exactly 11 servers in the entire ecosystem have ever had an answer independently verified true. Capability outran accountability by two orders of magnitude. That gap is the defining fact of this world. The ladder →

6 · Where checking is possible, the tools are honest

69 of 75 decided truth checks agreed with the official source — one to the euro against the raw Eurostat file. The problem today isn't lying tools; it's that the honest and the hollow are indistinguishable, which rationally suppresses trust, and therefore use. Every check →

7 · The money is a hamlet; the trust language arrived early

The paid tool economy: 13 distinct operators. All tripwire signals read quiet. Yet 1,352 servers carry the register's trust & verification label ◐ (their name, description, or tools speak in verified / receipts / audit vocabulary) — a mix of genuine verifiers and trust-flavored marketing; the count is register membership and equals what the door returns. The market is pricing trust language before anyone measures trust. Trust-washing will arrive before verification does, unless verification gets there first.

8 · The historical rhyme: 1995, not 2005

Directories before search engines; HTTP before the padlock. Supply-side landgrab, demand-side vacuum, trust layer vacant — and platforms beginning to gate their own gardens while the open registry gates nothing. What resolved that era wasn't more supply; it was trust infrastructure making abundance navigable. That seat is still open.

Coverage, stated plainly. Kind classification covers the census with an evidence tier per server; subject classification covers ~76% (the rest describe themselves too thinly to place — itself a finding). The trust-vocabulary count mixes genuine verifiers with marketing and will be split when the register matures. The first run of the subject analysis inflated one bucket via a publisher-prefix bug — caught before publication, method in the open repo.

robin saige · the observatory checks itself: usage · corrections · every series with SQL on data