Method
The whole probe, stated so you can run it yourself. Recomputability is the credential; trust is not requested.
The probe
We enumerate the official registry, then send every server that lists a remote endpoint a real MCP initialize handshake, then tools/list. One pass. A 12-second deadline per stage. No retries inside a run — a server that fails once counts as failed, and we say so; the alive share is therefore a lower bound.
The politeness contract
- We identify ourselves in the User-Agent, with a contact address.
- One request chain per server. We do not crawl past the endpoint.
- We never bypass authentication. Auth-walled servers get their own class — alive but unverifiable from outside — never scored dead, never scored verified.
- We probe from a fixed vantage and disclose it; a second vantage cross-checks for servers that special-case the prober.
The grades
observed our probe recorded it · reported the registry claims it · derived we computed it from the above. A derived number is never quoted as a measurement.
Verifiability, before verification
Every server is also being classed by how checkable it even is — from opaque (no stated source) to publicly re-derivable against a named primary source. That distribution — how much of this world anyone could ever verify — is the measurement we think matters most.
What this instrument cannot see
Saying it plainly is what makes the rest credible. The census reaches initialize and tools/list and stops — by politeness it never invokes tools across the whole registry. Therefore the census cannot see:
- Behavior on call — what a tool actually does, including payment walls, which sit almost entirely behind tools/call (98 of 99 in the first paid-probe, 2026-08-14). Call-level questions get their own pre-registered probes on candidate subsets: the truth engine, and the paid-probe feeding the tripwire.
- Payment demands at the handshake were, before 2026-08-14, filed under http_error; they now carry their own class, payment_required. Earlier censuses undercount them — at zero.
- Description content before v2 — the first census stored description hashes only: enough to detect change, not to search text. Probes now store the text (capped 16 KB).
- Anything auth-walled (about a quarter of the registry): alive but unverifiable from outside — never scored dead, never scored verified.
- Tools needing real arguments — the paid-probe calls one conservatively-chosen tool with empty arguments; a validation failure says nothing about behavior on real input, so no_payment_observed is evidence about that call only, never "this server is free."
A blind spot named is a question queued, not a flaw hidden. This register grows whenever a new one is found — a sixth trap is always assumed.
The house rules
Pre-register before running · publish nulls · a finding carries a falsifier · never claim beyond coverage · corrections are public and never deleted · a baseline comes from a query, never from a quotation. These are not decoration; on an instrument whose product is honesty, they are the product.
Verification, by kind
“Is the answer true?” only applies to public-fact retrieval. Every kind gets the check that fits it — and a tool is never marked down for failing a test that cannot apply.
| kind | best-known check | never |
|---|---|---|
| retrieval | re-derivation against the published primary source, signed receipt | an LLM verdict — the judge stays deterministic and auditable |
| computational | known-answer tests + metamorphic relations (convert A→B→A must invert; a hash has one right answer) — re-computation makes us the authority | nothing to gate — this class is fully checkable |
| action | contract checks (declared destructive/read-only/idempotent vs behavior) + CONSENTED-TARGET checks: fire the action only where WE own the target (an email-sender at our own inbox, a webhook at our own receiver) | firing a real action at anyone else's target — under any methodology |
| generative | stability probes (same input twice → variance vs the disclosed nondeterminism) + model/version disclosure | quality judging — taste is not measurement; an LLM judging 'good output' is gameable, unauditable, and breaks the referee seat |
| predictive | vintage scoring: archive the prediction at T, score at T+horizon against the realized public outcome (Brier/MAE — proper scoring rules) | grading a forecast before its horizon arrives |
| private | seeded-fixture verification: WE supply a known dataset under written invitation, the operator loads it, we query — known answers on private infrastructure, verdict-only publishing | credentialed probing without written invitation; holding customer data |
| orchestration | identity/costume analysis + pass-through fidelity (ask the gateway and its public upstream the same question; compare) | treating concentration alone as deception (the pipeworx lesson) |
The verdict
allow · warn · block synthesizes Q1–Q4 (does it answer · what is it · can it be checked · did it hold). Q5 — what company it keeps — contextualizes and never sentences. A verdict never renders without its because-clause, and is a dated observation with a signed receipt, never a warranty.
Scope
Population = servers listed in the official MCP registry with a callable remote endpoint. Listed-only servers are a coverage statement. We never probe past authentication uninvited, never fire real actions, and say plainly what the instrument cannot see.
Legibility (bands A–D)
Can an agent’s harness read the name, description, and typed parameters well enough to call the tool right? Graded A–D from stored inventories (formerly “harness readiness”).