
Product
Introducing Socket Scanning for VS Code Marketplace Extensions
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.
mcp-toolpick
Advanced tools
Will an assistant actually pick your tool? Scores an MCP server's tool names, descriptions and annotations against Anthropic's published guidance and the MCP spec — every rule quoting the sentence it comes from — then puts your tools in a catalogue beside
Will an assistant actually pick your tool?
Point it at an MCP server, with a few of the questions your users really ask. It drops your tools into a catalogue beside a dozen other people's, ranks them for each question, and tells you how often yours reaches the shortlist a model actually sees. Then it checks your descriptions and annotations against the sentences Anthropic and the MCP spec publish.
npx mcp-toolpick https://your-server.example/mcp --prompts ./prompts.tsv
Run against our own server, which fails its own test:
referencesource (streamable-http, 4 tools)
list_datasets
! [description-states-limits] Description never says what the tool is not for or does not return.
Anthropic's own worked example ends: "It will not provide any other information about the stock or company."
search_records
! [description-states-limits] Description never says what the tool is not for or does not return.
get_record
! [description-states-limits] Description never says what the tool is not for or does not return.
ok verify_quote
0 errors, 3 warnings, 0 notes across 4 tools.
Why any of these? npx mcp-toolpick --why <rule>
Selection test (bm25-offline, your 4 tools against 12 others)
ranked first: 33% of 3 queries
in the top 5: 67% — below this the model never sees the tool
search_records misses the shortlist, beaten by github_search_code, verify_quote, web_search
The tool that missed is the right tool for the question that was asked. That is the finding worth having, and you cannot get it by re-reading your own descriptions. The next section explains how it is measured; the linting comes after.
No install, no key, no config file, no account. It works against a live URL, a
--stdio command, or a JSON file of tool definitions.
Scoring prose is the easy half. The question you actually care about is whether an assistant with thirty tools in front of it reaches for yours.
So the checker builds a catalogue of your tools plus a bank of well-written decoys — GitHub, Slack, Jira, Postgres, Stripe, Sentry, a filesystem, a calendar — ranks them with BM25 for each query, and reports two numbers:
BM25 is not a guess at the mechanism. The tool-search tool ships in a
tool_search_tool_bm25_20251119 variant and indexes "tool names, descriptions,
argument names, and argument descriptions" — the same four fields this indexes.
Running it offline means the test is free, deterministic, and can sit in CI.
With no arguments the queries are your own tool names said in plain words.
That is a floor, not a forecast: losing to github_search_code on the phrase
"search records" is damning, and winning proves very little. For a number worth
quoting, pass the requests your users actually type:
npx mcp-toolpick https://your-server/mcp --prompts ./prompts.tsv
search_records find me the reporting threshold for benzene
get_record open the record for CAS 71-43-2
The decoys are synthetic fixtures written to be good, because beating a badly
written decoy proves nothing. If you would rather compete against something
real, --decoys-from <url> uses another live server's tools, and
--decoys-file takes your own.
Our own server passes the name-baseline floor on all 4 queries. Given questions a user would actually ask, it does not — this is the run at the top of the file, one row per prompt:
| prompt | rank of search_records |
|---|---|
search the reference datasets for benzene | 3rd |
look up a published fact with its source quote | 1st |
find me the reporting threshold for benzene | 16th of 16 — last |
The third one is the point. search_records is exactly the right tool for that
question, and it loses to github_search_code — because its description says
"search by state name, chemical, certificate number" and never once says
threshold, limit, reporting or how much. The catalogue is not being
unfair to it. It simply does not contain the words the question is made of.
You cannot see that by reading your own descriptions, because you know what they mean.
The offline test is the default because it costs nothing. If you want the real thing:
export ANTHROPIC_API_KEY=...
npx mcp-toolpick https://your-server/mcp --model claude-opus-4-6,gpt-5.1
This sends each query to each model with the whole shuffled catalogue attached and records which tool it called. It spends your money, so it is opt-in and never runs by accident. Tool order is shuffled with a fixed seed per query, so what gets measured is the description rather than the position in the list.
--provider is required for a model id the tool does not recognise. It will
never guess a vendor for you — handing an unknown model to the wrong API gets
you one vendor's answer under another vendor's name.
$ npx mcp-toolpick --why destructive-by-default
destructive-by-default (warn, tool scope)
A tool that sets neither hint is advertised as possibly destructive, because
that is the spec default.
It rests on this sentence, read 2026-09-08:
"destructiveHint?: boolean If true, the tool may perform destructive updates
to its environment. If false, the tool performs only additive updates. (This
property is meaningful only when readOnlyHint == false) Default: true"
— Model Context Protocol, Specification 2025-06-18 — schema reference
https://modelcontextprotocol.io/specification/2025-06-18/schema
Thirteen rules, each carrying a publisher, a URL, a verbatim quote and the date
it was read. --list-rules prints all of them. If you disagree with a rule you
can read what it is claiming and argue with the document rather than with me.
One claim in the file has no fetchable link and says so on its face: the Claude
connectors directory's requirement that every tool carry a title and the
applicable hint sits behind a submission form, so annotations-title carries a
note recording when we read it instead of a URL we cannot give you.
Run 2026-09-15 against a random sample of public MCP servers, 70 of the 80 drawn having answered:
| share of the 808 tools | |
|---|---|
| never say what the tool is not for | 70% |
set neither readOnlyHint nor destructiveHint | 41% |
no human-readable title | 42% |
| at least one parameter documented nowhere at all | 10% |
A fifth of the 808 have a parameter with no description field — but for half of
those, the author did write one, in the tool's own description or as an enum or
a bound on the parameter itself. Those are notes saying where the text is and why
tool search cannot see it there. The 10% above is what the checker warns about.
No parameter name is exempt for looking self-explanatory; benchmark/results.md
has the split and the reason.
The second row is why destructive-by-default is in the rule set. Because
destructiveHint defaults to true, a tool that sets no hints is telling every
client that reads annotations that it may perform destructive updates. Almost
none of those 332 tools are destructive. Nobody had told them.
An earlier run of the same 80 URLs, on 2026-09-08, reached 77 servers and 1,045
tools and put those figures at 75%, 30%, 31% and 16%. Do not read the difference
as a trend: one host supplying eight of the servers stopped answering tools/list
in between, taking about a quarter of the sample with it, and the servers present
in both runs changed as well. benchmark/results.md shows the arithmetic. It is
the reason the benchmark now captures catalogues once and scores them offline.
Seventy servers is a small sample, and a much larger one points the same way from a different angle. PolicyLayer's MCP Security Audit (snapshot 1 July 2026, updated monthly) classified 517,973 tools across 32,820 servers:
3.6% of tools warn the model about what they do. The other 96.4% don't.
— PolicyLayer, MCP Security Audit — July 2026, read 2026-09-10, https://policylayer.com/research/state-of-mcp-2026
That is a different measurement, not a bigger version of this one: they searched tool descriptions for warning words — "irreversible", "permanent", "cannot be undone" — while this checker reads the annotations, a separate field with a published default. A tool can pass one and fail the other. Read theirs for the size of the problem across the ecosystem and this one for what your own server does, in about ten seconds, with no account.
Method, sample and full table: benchmark/results.md.
- run: npx mcp-toolpick https://your-server/mcp --max-warn 10
Exit 0 when everything passes, 1 on any error-severity finding or above
--max-warn, 2 if it could not connect.
Errors are reserved for things that are categorically broken: no description at
all, readOnlyHint and destructiveHint both true, duplicate tool names, more
than 50 tools in one server. On the 77-server sample exactly one server tripped
an error. Everything else is a warning, so a default run does not fail your
build for having room to improve.
Checks
--rules <a,b> only these rules
--list-rules every rule, with the document behind it
--why <rule> the sentence a rule rests on, and where it was read
Selection test
--decoys <n> how many competing tools (default 12, max 27)
--decoys-file <p> your own decoys, as a JSON tool list
--decoys-from <url> use a real MCP server's tools as the decoys
--prompts <p> real requests: JSON [{tool,prompt}] or TSV "tool<TAB>prompt"
--no-select skip the selection test
--model <id,...> run it against real models instead (needs your API key)
--provider <p> anthropic | openai, if the model id does not say
Output
--json machine-readable, everything
--quiet findings only
--max-warn <n> exit 1 above this many warnings (default: no limit)
--header k=v extra HTTP header, repeatable (e.g. auth)
import { loadTools, scoreCatalogue, selectionTest, decoySet } from "mcp-toolpick";
const { tools } = await loadTools("https://your-server/mcp");
const report = scoreCatalogue(tools);
assert.equal(report.counts.error, 0);
const sel = selectionTest(tools, decoySet(12), [
{ tool: "search_records", prompt: "find me the reporting threshold for benzene" },
]);
console.log(sel.rows[0].rank, sel.rows[0].beaten_by); // 16 [ 'github_search_code', ... ]
Pin shortlist_rate in a test once you know what yours actually is. Asserting a
number you have not measured is how a green suite ends up meaning nothing.
--json gives you the same structure from the CLI.
There is no score out of 100. A number invites you to optimise the number. The output is a list of findings, each of which names a document.
It does not call your tools. It reads tools/list and stops. Nothing here
executes anything on your server.
It cannot tell you whether your tool is any good. It can tell you that a model searching a crowded catalogue will not find it, which is a different and more tractable problem.
mcp-tool-lint also lints MCP
tool definitions, from a JSON file you produce yourself, with string heuristics.
If that fits how you work, use it. This one connects to the running server,
checks annotations against their spec defaults, cites a document for every rule,
and runs the selection test.
TDQS is Glama's open rubric for tool-definition quality, and
it is already scoring the ecosystem at a scale nothing here approaches — every
tool of every server in their registry. Run it. It grades a definition on its
own terms, against fixed anchors: "TDQS scores a tool definition, not tool
behavior." That is a different question from the one here, which is
comparative — put your tool in a catalogue beside other good tools and see
whether a search for your user's actual question surfaces it. A definition can
score well and still lose that. The server this was built for scores
A, 4.4/5.0 on TDQS
and its search_records tool still misses the shortlist in the example at the
top of this README.
Node 18+. No dependencies.
MIT. Built by referencesource.org, which runs an MCP server of its own — checked by this tool, with no special case, in the example at the top.
FAQs
Will an assistant actually pick your tool? Scores an MCP server's tool names, descriptions and annotations against Anthropic's published guidance and the MCP spec — every rule quoting the sentence it comes from — then puts your tools in a catalogue beside
We found that mcp-toolpick demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Product
Socket now scans VS Code extensions, giving teams early detection of risky behaviors, hidden capabilities, and supply chain threats in developer tools.

Research
/Security News
Socket uncovered two malicious VS Code themes in a GlassWorm-linked cluster with thousands of installs across VS Code Marketplace and Open VSX.

Security News
/Company News
Capital One is partnering with Socket to proactively secure its open source supply chain.