Robin Saige we check the tools AI agents call

Corrections

This log launches non-empty on purpose. A verifier that has never published a mistake has either never measured anything or is hiding the misses. Ours are here — and the numbers moved a long way; most were our own instrument catching itself.

10
corrections published
6
caught by our own instrument
<24h
fastest, from publish to fix
0
deleted, ever
2026-08-15T19:09:33Z — was wrong: The site's six-bucket health taxonomy (alive-and-usable / behind-a-login / gateway / costume-farm / dead / alive-no-tools) conflated two independent questions — does it answer (health) and whose host is it (company) — so a gateway that answered was subtracted from 'alive & usable', and charts crossed axes that should never share a bucket. The taxonomy shaped the home map, the spectrum, and the matrix for four days.
correction: Ground-up rebuild (2026-08-16): health is now purely Q1 (answers / walled / dead, with sub-outcomes), company purely Q5 (solo / shared / gateway / costume), and the type system makes a mixed-axis chart a build error. The census page re-scores every server's health from its actual probe outcome. Old URLs redirect; machine endpoints are unchanged.
where: robinsaige.com home, /spectrum and derived charts, 2026-08-12 to 2026-08-16
2026-08-15T18:09:40Z — was wrong: The home-page funnel's last stage, labeled 'whose answers we could verify as true', counted every server with ANY truth check on record — including the ~600 servers the wide lanes had merely ATTEMPTED (verdict unverifiable). It showed 609 where the honest number was 11. The count had equalled decided servers by coincidence until 2026-08-15's wide-lane runs; a reader caught the inflation the same evening.
correction: The funnel now counts only servers with at least one independently CONFIRMED answer (11), states the attempted count separately and explicitly ('reached for — attempts never counted as verification'), and its stages now follow the instrument's own scope doctrine: answered → public-fact retrieval → has an answer key → verified true.
where: robinsaige.com home page, hours on 2026-08-15
2026-08-15T15:57:44Z — was wrong: The first full-pool wide-lane truth run (2026-08-15, 868 checks) recorded 150 decided verdicts (37 agree / 105 disagree / 8 partial) that a pre-registered sample audit showed were dominated by harness errors, not server behaviour: the tariff keyword 'import' matched the word 'IMPORTANT' in unrelated servers, the USITC primary lookup returned no rate for one probe code, and payment-walled answers were scored as disagreement instead of unverifiable.
correction: All 150 decided verdicts of the run voided to unverifiable and all 87 V3W upgrades rolled back before any finding was published. The pre-registered falsifier ('>5% wrong-tool selection voids the harness') worked as designed. Harness fixed and re-run; disagreements remain human-review-gated.
where: robinsaige.com/truth (per-server checks, ~30 minutes)
2026-08-15T07:45:00Z — was wrong: The home-page "tool economy at a glance" treemap counted server profiles across ALL census runs. From the 2026-08-14 re-census onward every health bucket was roughly doubled: "alive & usable" showed 7,652 (census figure: 3,898), "behind a login" 6,016 (3,062) — the map summed 21,907 servers over an 11,094-server census. The /spectrum page was scoped correctly, so the two pages contradicted each other.
correction: Every population-wide profile aggregate now scopes to the latest census run via one shared helper (latest_profile_run, chosen by recency, never by row count), regression-tested. Home and /spectrum now show identical bucket counts. Same sweep also fixed: /find duplicating each result once per census run, drill-down headers stating the capped 300 instead of the true bucket size, and a hardcoded "~10,813-server census" in the rating text (now derived from the actual cohort).
where: robinsaige.com home page, 2026-08-14T20:06Z to 2026-08-15
2026-08-14T20:22:00Z — was wrong: Finding RS-F-002 on /atlas framed pipeworx as one operator wearing 1,312 costumes — implying deception/duplication.
correction: Corrected: gateway.pipeworx.io is a multi-tenant gateway/aggregator. Of its 1,312 registry names, 1,266 serve DISTINCT tool inventories — concentration, not duplication, and no intent asserted. The finding is now a classification of where endpoints live, not a costume/deception claim.
where: /atlas (finding RS-F-002)
2026-08-13T17:20:00Z — was wrong: llms.txt and the home page stated Robin Saige is "listed in the registry it measures" / "one of the servers in this registry".
correction: It is not yet in the official MCP registry — verifiable with our own instrument: a registry search for "robinsaige" returns zero. Wording corrected to "an MCP server, registry publication pending"; the claim will be restored (and made checkable) once published. Caught by an MCP-architecture audit of our own server.
where: robinsaige.com/llms.txt + home page (2026-08-13)
2026-08-13T15:10:00Z — was wrong: Prior framing held that A2A agent cards barely exist — a landscape probe of 168 A2A-partner domains found 1, and a build note called publishing one "roughly the second genuine card on the internet".
correction: That measured the wrong population. A2A's own advertised partners do not publish cards, but MCP server operators do: an unbiased sample of 200 live MCP domains found ~9.5% publish a valid A2A agent card (2026-08-13). Cards exist in real numbers; "second card on the internet" was false.
where: agent-network landscape memo (2026-08-11) / internal
2026-08-13T15:10:00Z — was wrong: The traffic tripwire's first A2A-card-density reading was 16.5% (33 of 200).
correction: That sample was ordered alphabetically, which over-represents demo-host platforms (nip.io, *.workers.dev sort early) — 37 of the first 200 were on demo hosts. An unbiased hash-ordered sample gives 9.5% (19 of 200). The sampler was fixed to hash-order; the census reading is corrected.
where: robinsaige.com/signals (first run, 2026-08-13)
2026-08-12T06:00:00Z — was wrong: A 300-server random sample (2026-08-11, seed 7) suggested 27% of registry servers complete an MCP handshake.
correction: A full census of all 10,813 remote endpoints (2026-08-12) measured 55.0%. The census supersedes the sample; 27% must never be quoted as current. The gap is the finding: unversioned one-off measurements of this ecosystem cannot be trusted — including ours.
where: internal sample / draft post
2026-08-12T06:00:00Z — was wrong: The same sample recorded publisher pipeworx-io as "all 403" (dead).
correction: On 2026-08-12 all 1,312 pipeworx-io listings completed a handshake, cross-checked from a second vantage. It is one gateway wearing 1,312 costumes = 22% of live endpoints. Liveness alone would have scored it perfect.
where: internal sample