Robin Saige we check the tools AI agents call

retro

com.qretro/retro

QRetro retrospectives and planning poker over MCP: boards, action items, poker games, estimates.

— the operator's own registry description reported

Publisher com.qretro · first seen 2026-08-12 · endpoint https://mcp.qretro.com

allow
verdict
mid-slow-solo
cluster
ok_tools
answers
29
tools served

Dependency rating derived

allow
allow — alive & usable · retrieval-public ● · V3 has answer key ◐
allow — because answers ●, 29 tools legible (band A) ●, no confirmed drift in 2 probes ●; not yet checked: truth (3 attempts unverifiable, 0 disagreements)

Cluster mid-slow-solo · distinctiveness 0.623 (unusual). Every figure is a percentile or class within the 11,094-server censussee the whole spectrum.

Does it answer? Q1 · observed
answersalive_tools
latency575 ms
vs populationslow
protocolcurrent
authopen
Can it be used? Q4 · observed
tools29
vs populationmid
capabilities3
harnessA
Did it hold when checked? Q3–Q4 · derived
identitysolo
servers on its host1
verifiabilityunassessed
duplicate inventoryno
drift events0
answers true0/0 agree

🔏 Signed receipt rr_97d7891cb84d69cf8b317258 · ed25519 · key rs-rcpt-2026-08 — this rating is tamper-evident; an agent gets the full signature from check_server and can verify it against the published key.

Q1Does it answer?alive & usable — 2 probes ●evidence ↓
Q2What is it?retrieval-public ● evidence ↓
Q3Can it be checked?V3 has answer key ◐evidence ↓
Q4Did it hold?harness A · truth not yet run · no drift ●evidence ↓
Q5What company does it keep?mid-slow-solo · 1 on its host (context, not verdict) ◐evidence ↓

● observed · ◐ derived · ○ reported — the verdict synthesizes Q1–Q4; Q5 is context. Methodology →

History observed

Every probe we made, oldest → newest (2 shown, 2026-08-12 → 2026-08-14). Green answered · amber answered-but-walled · red no useful answer. A gap in our cadence is a gap in coverage, not evidence about the server.

changelog (1 events)
datetypewhat changed
2026-08-12outcomefirst probe: ok_tools

What kind of tools these are derived

Verification only means something relative to a tool's NATURE — an email-sender has no "true answer", a generator has no answer key. Classification evidence: classified from tool names/descriptions/annotations WE OBSERVED. What checking applies to each kind (the full methodology →):

naturetoolsmeaning what checking applies
retrieval-public16serves public factstruth-checkable against a public source — the V-ladder applies in full
action13changes the worldno true answer exists; we check the CONTRACT (destructive/read-only/idempotent declarations, error legibility) and never fire real actions

Tools last seen observed

poker.game.get poker.game.task.reveal poker.game.task.select poker.game.task.sync poker.game.tasks.add poker.game.tasks.import poker.game.tasks.list poker.games.create poker.games.list poker.iterations.list poker.sources.list retro.actions.complete retro.actions.create retro.actions.list retro.actions.update retro.board.actions.list retro.board.health.get retro.board.insights.list retro.board.messages.delete_own retro.board.messages.list retro.board.messages.update retro.board.roti.get retro.board.suggested_actions.promote retro.board.suggested_actions.reject retro.board.summary.get retro.boards.list retro.boards.search retro.team.members.list retro.teams.list

Harness readiness observed

Can an agent's harness pick this tool and call it correctly? Graded on the three things an agent reads — name, description, typed parameters. Band A — an agent can pick and call these reliably.

legibility checktools passingok
name is clear & specific29/29
has a description29/29
description is substantive29/29
parameters are typed29/29
parameters are described29/29

29 of 29 tools graded · 29 with a captured input schema. RS-008, Tier-1 — no domain knowledge, applies to any tool. “Alive” is not the same as “usable”.

Truth checks observed

Not yet confirmed — 3 attempt(s) currently unverifiable (rate limits, walls, or unmappable schemas — never counted against the server) · 0 disagreements.

claim classkeyverdictprimary source
tariff-duty6109.10.00unverifiableus-hts
tariff-duty8471.30.01unverifiableus-hts
tariff-dutyunmappedunverifiableus-hts

Protocol conformance observed

2025-06-18
protocol version
session model
1.3.0
server version QRetro.com

Capabilities: prompts resources tools

Drift observed

No confirmed drift on record. Baseline set 2026-08-14; changes appear here after human review.

raw probe records
probed (UTC) outcomehttpmsprotocol
2026-08-14T19:58:25Zok_tools2005752025-06-18
2026-08-12T05:33:40Zok_tools2005862025-06-18

Watch or embed this verdict

verdict badge

Operators: put the live verdict in your README — it updates with every census, links back here, and is a dated observation, never a warranty:

[![Robin Saige verdict](https://robinsaige.com/badge/com.qretro/retro.svg)](https://robinsaige.com/s/com.qretro/retro)

Depend on it? Subscribe to its change feed — outcome changes and confirmed drift, no account needed. Building on it? The full dossier as JSON — verdict, rating, truth checks, history — stable enough to gate CI on.

← overview