Robin Saige we check the tools AI agents call

OpenAI Tools MCP Server

ai.com.mcp/openai-tools

Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.

— the operator's own registry description reported

Publisher ai.com.mcp · first seen 2026-08-12 · endpoint https://openai-tools.run.mcp.com.ai/mcp

allow
verdict
mid-shared
cluster
ok_tools
answers
9
tools served

Dependency rating derived

allow — mid-shared — safe to depend on

Cluster mid-shared · distinctiveness 0.333 (typical). Every figure is a percentile or class within the 11,094-server censussee the whole spectrum.

Reach — can an agent get to it?
answersalive_tools
latency167 ms
vs populationtypical
protocolcurrent
authopen
Use — can it use it?
tools9
vs populationmid
capabilities5
harnessB
Trust — can it rely on it?
identityshared
servers on its host1
servers on its domain8
verifiabilityunassessed
duplicate inventoryno
drift events0
answers trueno primary

🔏 Signed receipt rr_72abae2a17d371f4fb52400c · ed25519 · key rs-rcpt-2026-08 — this rating is tamper-evident; an agent gets the full signature from check_server and can verify it against the published key.

What kind of tools these are derived

Verification only means something relative to a tool's NATURE — an email-sender has no "true answer", a generator has no answer key. Classification evidence: classified from tool names/descriptions/annotations WE OBSERVED. What checking applies to each kind (the full methodology →):

naturetoolsmeaning what checking applies
action4changes the worldno true answer exists; we check the CONTRACT (destructive/read-only/idempotent declarations, error legibility) and never fire real actions
unclassified4not classifiable from its textTier-1 only until its tools describe themselves
generative1creates contentno answer key exists; disclosure + stability checks, never truth verdicts

Tools last seen

createImage createImageEdit createImageVariation createModeration createTranscription createTranslation deleteModel listModels retrieveModel

Harness readiness observed

Can an agent's harness pick this tool and call it correctly? Graded on the three things an agent reads — name, description, typed parameters. Band B — usable, with rough edges.

legibility checktools passingok
name is clear & specific9/9
has a description9/9
description is substantive6/96/9
parameters are typed9/9
parameters are described0/90/9

9 of 9 tools graded · 9 with a captured input schema. RS-008, Tier-1 — no domain knowledge, applies to any tool. “Alive” is not the same as “usable”.

Truth checks observed

Not in the Tier-2 trade lane, or not yet checked. Tier-2 re-derives a server's answers against published primary sources — see truth checks.

Protocol conformance observed

2025-06-18
protocol version
session model
0.6.0
server version openai

Capabilities: logging prompts resources resourcesTemplates tools

Drift

No confirmed drift on record. Baseline set 2026-08-14; changes appear here after human review.

History observed

Every probe we made, oldest → newest (2 shown, 2026-08-12 → 2026-08-14). Green answered · amber answered-but-walled · red no useful answer. A gap in our cadence is a gap in coverage, not evidence about the server.

changelog (1 events)
datetypewhat changed
2026-08-12outcomefirst probe: ok_tools
raw probe records
probed (UTC) outcomehttpmsprotocol
2026-08-14T19:57:34Zok_tools2001672025-06-18
2026-08-12T05:32:52Zok_tools2001302025-06-18

Watch or embed this verdict

verdict badge

Operators: put the live verdict in your README — it updates with every census, links back here, and is a dated observation, never a warranty:

[![Robin Saige verdict](https://robinsaige.com/badge/ai.com.mcp/openai-tools.svg)](https://robinsaige.com/s/ai.com.mcp/openai-tools)

Depend on it? Subscribe to its change feed — outcome changes and confirmed drift, no account needed. Building on it? The full dossier as JSON — verdict, rating, truth checks, history — stable enough to gate CI on.

← overview