framebench
app.framebench/framebench
Estimated game fps for any GPU or Apple Silicon chip, with the limiter and tweaks.
— the operator's own registry description reported
Publisher app.framebench · first seen 2026-08-12 · endpoint https://framebench.app/mcp
Dependency rating derived
allow — alive & usable · kind unknown · V n/a
Cluster thin-solo · distinctiveness 0.243 (typical). Every figure is a percentile or class within the 11,094-server census — see the whole spectrum.
🔏 Signed receipt rr_947e965626989a73046d14bd · ed25519 · key rs-rcpt-2026-08 — this rating is tamper-evident; an agent gets the full signature from check_server and can verify it against the published key.
● observed · ◐ derived · ○ reported — the verdict synthesizes Q1–Q4; Q5 is context. Methodology →
History observed
Every probe we made, oldest → newest (2 shown, 2026-08-12 → 2026-08-14). Green answered · amber answered-but-walled · red no useful answer. A gap in our cadence is a gap in coverage, not evidence about the server.
changelog (1 events)
| date | type | what changed |
|---|---|---|
| 2026-08-12 | outcome | first probe: ok_tools |
What kind of tools these are derived
Verification only means something relative to a tool's NATURE — an email-sender has no "true answer", a generator has no answer key. Classification evidence: classified from tool names/descriptions/annotations WE OBSERVED. What checking applies to each kind (the full methodology →):
| nature | tools | meaning | what checking applies |
|---|---|---|---|
| unclassified | 2 | not classifiable from its text | Tier-1 only until its tools describe themselves |
| action | 1 | changes the world | no true answer exists; we check the CONTRACT (destructive/read-only/idempotent declarations, error legibility) and never fire real actions |
Tools last seen observed
Harness readiness observed
Can an agent's harness pick this tool and call it correctly? Graded on the three things an agent reads — name, description, typed parameters. Band A — an agent can pick and call these reliably.
| legibility check | tools passing | ok |
|---|---|---|
| name is clear & specific | 3/3 | ✓ |
| has a description | 3/3 | ✓ |
| description is substantive | 3/3 | ✓ |
| parameters are typed | 3/3 | ✓ |
| parameters are described | 2/3 | 2/3 |
3 of 3 tools graded · 3 with a captured input schema. RS-008, Tier-1 — no domain knowledge, applies to any tool. “Alive” is not the same as “usable”.
Truth checks observed
Not in the Tier-2 trade lane, or not yet checked. Tier-2 re-derives a server's answers against published primary sources — see truth checks.
Protocol conformance observed
Capabilities: tools
Drift observed
No confirmed drift on record. Baseline set 2026-08-14; changes appear here after human review.
raw probe records
| probed (UTC) | outcome | http | ms | protocol |
|---|---|---|---|---|
| 2026-08-14T19:57:45Z | ok_tools | 200 | 157 | 2025-06-18 |
| 2026-08-12T05:33:02Z | ok_tools | 200 | 146 | 2025-06-18 |
Watch or embed this verdict
Operators: put the live verdict in your README — it updates with every census, links back here, and is a dated observation, never a warranty:
[](https://robinsaige.com/s/app.framebench/framebench)
Depend on it? Subscribe to its change feed — outcome changes and confirmed drift, no account needed. Building on it? The full dossier as JSON — verdict, rating, truth checks, history — stable enough to gate CI on.