Official MCP Registry versus a live reliability scorecard

Last updated:

Use the official MCP Registry to discover published server metadata, not as proof that a server connects and completes your task today. Add a dated canary result for operational evidence, then review permissions and security separately; none of those layers replaces the others.

The registry policy, MCP trust model and CanaryIndex v0.1 method on this page were checked against their linked sources on October 4, 2026.

Discovery and reliability answer different questions

The official Registry's moderation policy calls the service permissive and tells consumers to assume minimal-to-no moderation. It removes illegal content, malware, spam and non-functioning servers when identified, but explicitly does not remove every low-quality, buggy or vulnerable server (official moderation policy).

MCP's security policy separately says users and administrators choose which servers to trust. A local stdio server runs with the process privileges it is given, and the protocol is not a sandbox (MCP security policy).

QuestionOfficial registry listingDated live scorecardSecurity review
Can I discover the server and its published metadata?YesSometimesNot its main job
Did one known task complete recently?No continuous proofYes, for covered tools and fixturesNot necessarily
Was output quality checked against a known answer?NoYes, if the method publishes an oracle and thresholdNot necessarily
What did that observed task cost and how long did it take?Not a live measurementCan record bothNot necessarily
Is the server safe for my credentials, files and permissions?Not certifiedNot certifiedThis is the review's job
Will it work for every input, account and location tomorrow?NoNoNo

What a useful scorecard must publish

A reliability claim is reproducible only when it includes:

Do not count an account-plan denial, missing credential or checker's network failure as a product failure. It is missing evidence, not a failed task.

What CanaryIndex actually measures

CanaryIndex is a narrow public scorecard, not a replacement MCP registry. Version 0.1 runs comparable Apify Actors in two categories: one-page PDF text-layer to Markdown and a direct public audio/video URL to transcript. Those Actors can be reached by agents through Apify's MCP server, but the current weekly canary calls the Actor runs directly; it does not claim to test arbitrary MCP-server initialization or every MCP transport.

Its current method:

The monthly runner is capped at $5, with a configured worst case of $0.36 per weekly suite. Publisher-owned tools are visibly labeled, receive no ranking advantage, and use a labeled simulated FREE-tier price where owner event charges are exempt.

Read the current result instead of copying a score from this article: machine recommendations, latest raw observation, and full method.

A practical selection flow

  1. Discover: find the candidate in the official registry or a relevant subregistry and inspect its publisher metadata.
  2. Threat-model: review transport, command, requested permissions, data access, authentication and publisher trust before connecting.
  3. Observe: run one representative task with a known answer, or use a transparent current scorecard that covers the same category.
  4. Classify: record passed, failed, or failed to observe. Do not collapse the third state into failure.
  5. Repeat: a single pass is an observation, not a permanent reliability guarantee.

For the step-by-step protocol checks—connection, authentication, tools/list, tools/call, and expected-answer comparison—use the companion MCP reliability testing guide. This page is intentionally different: it explains which decision belongs to the registry, scorecard and security-review layers rather than duplicating that testing procedure.

FAQ

Does an official MCP Registry listing certify reliability?

No. The official policy is permissive and does not promise continuous task testing. Treat the listing as discovery metadata.

Does a passing live canary prove an MCP server is secure?

No. A canary reports one task observation. Permissions, credentials, code execution and data handling need a separate security review.

Why keep “failed” separate from “failed to observe”?

A failed result means the task ran and did not meet its contract. Failed to observe means the checker could not establish a result because of its own account, credentials, network or other measurement limit.

Does CanaryIndex test every MCP server?

No. Version 0.1 covers comparable Apify Actors in two narrow categories and calls their Actor runs directly. It publishes the limitation instead of generalizing those observations to the whole MCP ecosystem.