Robin Saige we check the tools AI agents call

Method

The whole probe, stated so you can run it yourself. Recomputability is the credential; trust is not requested.

The probe

We enumerate the official registry, then send every server that lists a remote endpoint a real MCP initialize handshake, then tools/list. One pass. A 12-second deadline per stage. No retries inside a run — a server that fails once counts as failed, and we say so; the alive share is therefore a lower bound.

The politeness contract

The grades

observed our probe recorded it · reported the registry claims it · derived we computed it from the above. A derived number is never quoted as a measurement.

Verifiability, before verification

Every server is also being classed by how checkable it even is — from opaque (no stated source) to publicly re-derivable against a named primary source. That distribution — how much of this world anyone could ever verify — is the measurement we think matters most.

What this instrument cannot see

Saying it plainly is what makes the rest credible. The census reaches initialize and tools/list and stops — by politeness it never invokes tools across the whole registry. Therefore the census cannot see:

A blind spot named is a question queued, not a flaw hidden. This register grows whenever a new one is found — a sixth trap is always assumed.

The house rules

Pre-register before running · publish nulls · a finding carries a falsifier · never claim beyond coverage · corrections are public and never deleted · a baseline comes from a query, never from a quotation. These are not decoration; on an instrument whose product is honesty, they are the product.

Verification, by kind

“Is the answer true?” only applies to public-fact retrieval. Every kind gets the check that fits it — and a tool is never marked down for failing a test that cannot apply.

kindbest-known checknever
retrievalre-derivation against the published primary source, signed receiptan LLM verdict — the judge stays deterministic and auditable
computationalknown-answer tests + metamorphic relations (convert A→B→A must invert; a hash has one right answer) — re-computation makes us the authoritynothing to gate — this class is fully checkable
actioncontract checks (declared destructive/read-only/idempotent vs behavior) + CONSENTED-TARGET checks: fire the action only where WE own the target (an email-sender at our own inbox, a webhook at our own receiver)firing a real action at anyone else's target — under any methodology
generativestability probes (same input twice → variance vs the disclosed nondeterminism) + model/version disclosurequality judging — taste is not measurement; an LLM judging 'good output' is gameable, unauditable, and breaks the referee seat
predictivevintage scoring: archive the prediction at T, score at T+horizon against the realized public outcome (Brier/MAE — proper scoring rules)grading a forecast before its horizon arrives
privateseeded-fixture verification: WE supply a known dataset under written invitation, the operator loads it, we query — known answers on private infrastructure, verdict-only publishingcredentialed probing without written invitation; holding customer data
orchestrationidentity/costume analysis + pass-through fidelity (ask the gateway and its public upstream the same question; compare)treating concentration alone as deception (the pipeworx lesson)

The verdict

allow · warn · block synthesizes Q1–Q4 (does it answer · what is it · can it be checked · did it hold). Q5 — what company it keeps — contextualizes and never sentences. A verdict never renders without its because-clause, and is a dated observation with a signed receipt, never a warranty.

Scope

Population = servers listed in the official MCP registry with a callable remote endpoint. Listed-only servers are a coverage statement. We never probe past authentication uninvited, never fire real actions, and say plainly what the instrument cannot see.

Legibility (bands A–D)

Can an agent’s harness read the name, description, and typed parameters well enough to call the tool right? Graded A–D from stored inventories (formerly “harness readiness”).