Robin Saige we check the tools AI agents call

InsideOut (Riley)

com.luthersystems.insideout/mcp

Designs, prices, and deploys AWS/GCP cloud infrastructure from plain-English requirements.

— the operator's own registry description reported

Publisher com.luthersystems.insideout · first seen 2026-08-12 · endpoint https://app.luthersystems.com/v1/insideout-mcp

allow
verdict
mid-solo
cluster
ok_tools
answers
24
tools served

Dependency rating derived

allow
allow — alive & usable · action ● · V n/a
allow — because answers ●, 24 tools legible (band A) ●, no confirmed drift in 2 probes ●; not yet checked: truth (0 runs)

Cluster mid-solo · distinctiveness 0.464 (moderate). Every figure is a percentile or class within the 11,094-server censussee the whole spectrum.

Does it answer? Q1 · observed
answersalive_tools
latency298 ms
vs populationtypical
protocolcurrent
authopen
Can it be used? Q4 · observed
tools24
vs populationmid
capabilities3
harnessA
Did it hold when checked? Q3–Q4 · derived
identitysolo
servers on its host1
verifiabilityunassessed
duplicate inventoryno
drift events0
answers trueno primary

🔏 Signed receipt rr_0650cb35101ae2a878a110e4 · ed25519 · key rs-rcpt-2026-08 — this rating is tamper-evident; an agent gets the full signature from check_server and can verify it against the published key.

Q1Does it answer?alive & usable — 2 probes ●evidence ↓
Q2What is it?action ● evidence ↓
Q3Can it be checked?V n/a for this kind ◐evidence ↓
Q4Did it hold?harness A · truth not yet run · no drift ●evidence ↓
Q5What company does it keep?mid-solo · 1 on its host (context, not verdict) ◐evidence ↓

● observed · ◐ derived · ○ reported — the verdict synthesizes Q1–Q4; Q5 is context. Methodology →

History observed

Every probe we made, oldest → newest (2 shown, 2026-08-12 → 2026-08-14). Green answered · amber answered-but-walled · red no useful answer. A gap in our cadence is a gap in coverage, not evidence about the server.

changelog (1 events)
datetypewhat changed
2026-08-12outcomefirst probe: ok_tools

What kind of tools these are derived

Verification only means something relative to a tool's NATURE — an email-sender has no "true answer", a generator has no answer key. Classification evidence: classified from tool names/descriptions/annotations WE OBSERVED. What checking applies to each kind (the full methodology →):

naturetoolsmeaning what checking applies
action10changes the worldno true answer exists; we check the CONTRACT (destructive/read-only/idempotent declarations, error legibility) and never fire real actions
retrieval-public6serves public factstruth-checkable against a public source — the V-ladder applies in full
generative5creates contentno answer key exists; disclosure + stability checks, never truth verdicts
computational2deterministic transformsself-checkable by re-computation and known-answer tests
predictive1claims about the futurescoreable only in retrospect — the archive scores yesterday's predictions against today's outcome

Tools last seen observed

awsinspect awsinspect_batch convoawait convoinspect convoopen convoreply convostatus credawait gcpinspect gcpinspect_batch help stackdiff stackrollback stackversions submit_feedback tfdeploy tfdestroy tfdrift tfgenerate tflogs tfoutputs tfplan tfruns tfstatus

Harness readiness observed

Can an agent's harness pick this tool and call it correctly? Graded on the three things an agent reads — name, description, typed parameters. Band A — an agent can pick and call these reliably.

legibility checktools passingok
name is clear & specific24/24
has a description24/24
description is substantive24/24
parameters are typed24/24
parameters are described24/24

24 of 24 tools graded · 24 with a captured input schema. RS-008, Tier-1 — no domain knowledge, applies to any tool. “Alive” is not the same as “usable”.

Truth checks observed

Not in the Tier-2 trade lane, or not yet checked. Tier-2 re-derives a server's answers against published primary sources — see truth checks.

Protocol conformance observed

2025-06-18
protocol version
session model
v2.0.0
server version insideout-agent

Capabilities: logging prompts tools

Drift observed

No confirmed drift on record. Baseline set 2026-08-14; changes appear here after human review.

raw probe records
probed (UTC) outcomehttpmsprotocol
2026-08-14T19:58:17Zok_tools2002982025-06-18
2026-08-12T05:33:32Zok_tools2002412025-06-18

Watch or embed this verdict

verdict badge

Operators: put the live verdict in your README — it updates with every census, links back here, and is a dated observation, never a warranty:

[![Robin Saige verdict](https://robinsaige.com/badge/com.luthersystems.insideout/mcp.svg)](https://robinsaige.com/s/com.luthersystems.insideout/mcp)

Depend on it? Subscribe to its change feed — outcome changes and confirmed drift, no account needed. Building on it? The full dossier as JSON — verdict, rating, truth checks, history — stable enough to gate CI on.

← overview