Official MCP Registry versus a live reliability scorecard
Last updated:
Use the official MCP Registry to discover published server metadata, not as proof that a server connects and completes your task today. Add a dated canary result for operational evidence, then review permissions and security separately; none of those layers replaces the others.
The registry policy, MCP trust model and CanaryIndex v0.1 method on this page were checked against their linked sources on October 4, 2026.
Discovery and reliability answer different questions
The official Registry's moderation policy calls the service permissive and tells consumers to assume minimal-to-no moderation. It removes illegal content, malware, spam and non-functioning servers when identified, but explicitly does not remove every low-quality, buggy or vulnerable server (official moderation policy).
MCP's security policy separately says users and administrators choose which servers to trust. A local stdio server runs with the process privileges it is given, and the protocol is not a sandbox (MCP security policy).
| Question | Official registry listing | Dated live scorecard | Security review |
|---|---|---|---|
| Can I discover the server and its published metadata? | Yes | Sometimes | Not its main job |
| Did one known task complete recently? | No continuous proof | Yes, for covered tools and fixtures | Not necessarily |
| Was output quality checked against a known answer? | No | Yes, if the method publishes an oracle and threshold | Not necessarily |
| What did that observed task cost and how long did it take? | Not a live measurement | Can record both | Not necessarily |
| Is the server safe for my credentials, files and permissions? | Not certified | Not certified | This is the review's job |
| Will it work for every input, account and location tomorrow? | No | No | No |
What a useful scorecard must publish
A reliability claim is reproducible only when it includes:
- the exact tool/build and category;
- the public or redistributable fixture and expected answer;
- the timestamp, account context and runner location when relevant;
- whether the call completed, whether its output passed, and why;
- the scorer version, threshold, latency and known cost;
- a separate failed to observe state for cases where the checker could not make the call.
Do not count an account-plan denial, missing credential or checker's network failure as a product failure. It is missing evidence, not a failed task.
What CanaryIndex actually measures
CanaryIndex is a narrow public scorecard, not a replacement MCP registry. Version 0.1 runs comparable Apify Actors in two categories: one-page PDF text-layer to Markdown and a direct public audio/video URL to transcript. Those Actors can be reached by agents through Apify's MCP server, but the current weekly canary calls the Actor runs directly; it does not claim to test arbitrary MCP-server initialization or every MCP transport.
Its current method:
- runs the same fixed public fixture for each comparable tool on Mondays at 08:17 UTC;
- accepts the PDF result only when all three expected tokens are present (configured threshold 0.80 over three tokens);
- scores transcription as
max(0, 1 − word error rate)against two published transcript interpretations, with a 0.70 threshold; - defines reliability as accepted observations divided by completed observations;
- records unknown cost as unknown, never zero;
- requires both a public listing and a latest passing canary before recommending a tool;
- excludes payment, sponsorship, affiliate status and publisher ownership from routing.
The monthly runner is capped at $5, with a configured worst case of $0.36 per weekly suite. Publisher-owned tools are visibly labeled, receive no ranking advantage, and use a labeled simulated FREE-tier price where owner event charges are exempt.
Read the current result instead of copying a score from this article: machine recommendations, latest raw observation, and full method.
A practical selection flow
- Discover: find the candidate in the official registry or a relevant subregistry and inspect its publisher metadata.
- Threat-model: review transport, command, requested permissions, data access, authentication and publisher trust before connecting.
- Observe: run one representative task with a known answer, or use a transparent current scorecard that covers the same category.
- Classify: record passed, failed, or failed to observe. Do not collapse the third state into failure.
- Repeat: a single pass is an observation, not a permanent reliability guarantee.
For the step-by-step protocol checks—connection, authentication, tools/list, tools/call, and expected-answer comparison—use the companion MCP reliability testing guide. This page is intentionally different: it explains which decision belongs to the registry, scorecard and security-review layers rather than duplicating that testing procedure.
FAQ
Does an official MCP Registry listing certify reliability?
No. The official policy is permissive and does not promise continuous task testing. Treat the listing as discovery metadata.
Does a passing live canary prove an MCP server is secure?
No. A canary reports one task observation. Permissions, credentials, code execution and data handling need a separate security review.
Why keep “failed” separate from “failed to observe”?
A failed result means the task ran and did not meet its contract. Failed to observe means the checker could not establish a result because of its own account, credentials, network or other measurement limit.
Does CanaryIndex test every MCP server?
No. Version 0.1 covers comparable Apify Actors in two narrow categories and calls their Actor runs directly. It publishes the limitation instead of generalizing those observations to the whole MCP ecosystem.
Related
- Procedure: which MCP server works today, and how to check
- Evidence before completion: passing tests need a negative control
- Package grounding: check a function against an exact package version
- Open CanaryIndex · method