New:Introducing Socket Scanning for VS Code Marketplace Extensions.Learn more →
Get Started

mcp-toolpick

Package Overview
Dependencies
Maintainers
1
Versions
1
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

mcp-toolpick

Will an assistant actually pick your tool? Scores an MCP server's tool names, descriptions and annotations against Anthropic's published guidance and the MCP spec — every rule quoting the sentence it comes from — then puts your tools in a catalogue beside

latest
Source
npmnpm
Version
0.1.0
Version published
Maintainers
1
Created
Source

mcp-toolpick

Will an assistant actually pick your tool?

Point it at an MCP server, with a few of the questions your users really ask. It drops your tools into a catalogue beside a dozen other people's, ranks them for each question, and tells you how often yours reaches the shortlist a model actually sees. Then it checks your descriptions and annotations against the sentences Anthropic and the MCP spec publish.

npx mcp-toolpick https://your-server.example/mcp --prompts ./prompts.tsv

Run against our own server, which fails its own test:

referencesource (streamable-http, 4 tools)

list_datasets
  ! [description-states-limits] Description never says what the tool is not for or does not return.
      Anthropic's own worked example ends: "It will not provide any other information about the stock or company."
search_records
  ! [description-states-limits] Description never says what the tool is not for or does not return.
get_record
  ! [description-states-limits] Description never says what the tool is not for or does not return.
ok verify_quote

0 errors, 3 warnings, 0 notes across 4 tools.
Why any of these? npx mcp-toolpick --why <rule>

Selection test (bm25-offline, your 4 tools against 12 others)
  ranked first:      33% of 3 queries
  in the top 5:      67%  — below this the model never sees the tool
  search_records misses the shortlist, beaten by github_search_code, verify_quote, web_search

The tool that missed is the right tool for the question that was asked. That is the finding worth having, and you cannot get it by re-reading your own descriptions. The next section explains how it is measured; the linting comes after.

No install, no key, no config file, no account. It works against a live URL, a --stdio command, or a JSON file of tool definitions.

The selection test

Scoring prose is the easy half. The question you actually care about is whether an assistant with thirty tools in front of it reaches for yours.

So the checker builds a catalogue of your tools plus a bank of well-written decoys — GitHub, Slack, Jira, Postgres, Stripe, Sentry, a filesystem, a calendar — ranks them with BM25 for each query, and reports two numbers:

  • ranked first — your tool won.
  • in the top 5 — Anthropic's tool-search tool "returns up to 5 matching tools by default". Below the top 5 the model never sees your tool at all, so this is the number that matters.

BM25 is not a guess at the mechanism. The tool-search tool ships in a tool_search_tool_bm25_20251119 variant and indexes "tool names, descriptions, argument names, and argument descriptions" — the same four fields this indexes. Running it offline means the test is free, deterministic, and can sit in CI.

With no arguments the queries are your own tool names said in plain words. That is a floor, not a forecast: losing to github_search_code on the phrase "search records" is damning, and winning proves very little. For a number worth quoting, pass the requests your users actually type:

npx mcp-toolpick https://your-server/mcp --prompts ./prompts.tsv
search_records	find me the reporting threshold for benzene
get_record	open the record for CAS 71-43-2

The decoys are synthetic fixtures written to be good, because beating a badly written decoy proves nothing. If you would rather compete against something real, --decoys-from <url> uses another live server's tools, and --decoys-file takes your own.

What a real miss looks like

Our own server passes the name-baseline floor on all 4 queries. Given questions a user would actually ask, it does not — this is the run at the top of the file, one row per prompt:

promptrank of search_records
search the reference datasets for benzene3rd
look up a published fact with its source quote1st
find me the reporting threshold for benzene16th of 16 — last

The third one is the point. search_records is exactly the right tool for that question, and it loses to github_search_code — because its description says "search by state name, chemical, certificate number" and never once says threshold, limit, reporting or how much. The catalogue is not being unfair to it. It simply does not contain the words the question is made of.

You cannot see that by reading your own descriptions, because you know what they mean.

Against real models

The offline test is the default because it costs nothing. If you want the real thing:

export ANTHROPIC_API_KEY=...
npx mcp-toolpick https://your-server/mcp --model claude-opus-4-6,gpt-5.1

This sends each query to each model with the whole shuffled catalogue attached and records which tool it called. It spends your money, so it is opt-in and never runs by accident. Tool order is shuffled with a fixed seed per query, so what gets measured is the description rather than the position in the list.

--provider is required for a model id the tool does not recognise. It will never guess a vendor for you — handing an unknown model to the wrong API gets you one vendor's answer under another vendor's name.

Every rule cites the sentence it rests on

$ npx mcp-toolpick --why destructive-by-default

destructive-by-default  (warn, tool scope)

A tool that sets neither hint is advertised as possibly destructive, because
that is the spec default.

It rests on this sentence, read 2026-09-08:

  "destructiveHint?: boolean If true, the tool may perform destructive updates
   to its environment. If false, the tool performs only additive updates. (This
   property is meaningful only when readOnlyHint == false) Default: true"

  — Model Context Protocol, Specification 2025-06-18 — schema reference
    https://modelcontextprotocol.io/specification/2025-06-18/schema

Thirteen rules, each carrying a publisher, a URL, a verbatim quote and the date it was read. --list-rules prints all of them. If you disagree with a rule you can read what it is claiming and argue with the document rather than with me.

One claim in the file has no fetchable link and says so on its face: the Claude connectors directory's requirement that every tool carry a title and the applicable hint sits behind a submission form, so annotations-title carries a note recording when we read it instead of a URL we cannot give you.

What the rules found on 808 live tools

Run 2026-09-15 against a random sample of public MCP servers, 70 of the 80 drawn having answered:

share of the 808 tools
never say what the tool is not for70%
set neither readOnlyHint nor destructiveHint41%
no human-readable title42%
at least one parameter documented nowhere at all10%

A fifth of the 808 have a parameter with no description field — but for half of those, the author did write one, in the tool's own description or as an enum or a bound on the parameter itself. Those are notes saying where the text is and why tool search cannot see it there. The 10% above is what the checker warns about. No parameter name is exempt for looking self-explanatory; benchmark/results.md has the split and the reason.

The second row is why destructive-by-default is in the rule set. Because destructiveHint defaults to true, a tool that sets no hints is telling every client that reads annotations that it may perform destructive updates. Almost none of those 332 tools are destructive. Nobody had told them.

An earlier run of the same 80 URLs, on 2026-09-08, reached 77 servers and 1,045 tools and put those figures at 75%, 30%, 31% and 16%. Do not read the difference as a trend: one host supplying eight of the servers stopped answering tools/list in between, taking about a quarter of the sample with it, and the servers present in both runs changed as well. benchmark/results.md shows the arithmetic. It is the reason the benchmark now captures catalogues once and scores them offline.

Seventy servers is a small sample, and a much larger one points the same way from a different angle. PolicyLayer's MCP Security Audit (snapshot 1 July 2026, updated monthly) classified 517,973 tools across 32,820 servers:

3.6% of tools warn the model about what they do. The other 96.4% don't.

— PolicyLayer, MCP Security Audit — July 2026, read 2026-09-10, https://policylayer.com/research/state-of-mcp-2026

That is a different measurement, not a bigger version of this one: they searched tool descriptions for warning words — "irreversible", "permanent", "cannot be undone" — while this checker reads the annotations, a separate field with a published default. A tool can pass one and fail the other. Read theirs for the size of the problem across the ecosystem and this one for what your own server does, in about ten seconds, with no account.

Method, sample and full table: benchmark/results.md.

In CI

- run: npx mcp-toolpick https://your-server/mcp --max-warn 10

Exit 0 when everything passes, 1 on any error-severity finding or above --max-warn, 2 if it could not connect.

Errors are reserved for things that are categorically broken: no description at all, readOnlyHint and destructiveHint both true, duplicate tool names, more than 50 tools in one server. On the 77-server sample exactly one server tripped an error. Everything else is a warning, so a default run does not fail your build for having room to improve.

Options

Checks
  --rules <a,b>        only these rules
  --list-rules         every rule, with the document behind it
  --why <rule>         the sentence a rule rests on, and where it was read

Selection test
  --decoys <n>         how many competing tools (default 12, max 27)
  --decoys-file <p>    your own decoys, as a JSON tool list
  --decoys-from <url>  use a real MCP server's tools as the decoys
  --prompts <p>        real requests: JSON [{tool,prompt}] or TSV "tool<TAB>prompt"
  --no-select          skip the selection test
  --model <id,...>     run it against real models instead (needs your API key)
  --provider <p>       anthropic | openai, if the model id does not say

Output
  --json               machine-readable, everything
  --quiet              findings only
  --max-warn <n>       exit 1 above this many warnings (default: no limit)
  --header k=v         extra HTTP header, repeatable (e.g. auth)

As a library

import { loadTools, scoreCatalogue, selectionTest, decoySet } from "mcp-toolpick";

const { tools } = await loadTools("https://your-server/mcp");
const report = scoreCatalogue(tools);
assert.equal(report.counts.error, 0);

const sel = selectionTest(tools, decoySet(12), [
  { tool: "search_records", prompt: "find me the reporting threshold for benzene" },
]);
console.log(sel.rows[0].rank, sel.rows[0].beaten_by);   // 16 [ 'github_search_code', ... ]

Pin shortlist_rate in a test once you know what yours actually is. Asserting a number you have not measured is how a green suite ends up meaning nothing.

--json gives you the same structure from the CLI.

What it deliberately does not do

There is no score out of 100. A number invites you to optimise the number. The output is a list of findings, each of which names a document.

It does not call your tools. It reads tools/list and stops. Nothing here executes anything on your server.

It cannot tell you whether your tool is any good. It can tell you that a model searching a crowded catalogue will not find it, which is a different and more tractable problem.

mcp-tool-lint also lints MCP tool definitions, from a JSON file you produce yourself, with string heuristics. If that fits how you work, use it. This one connects to the running server, checks annotations against their spec defaults, cites a document for every rule, and runs the selection test.

TDQS is Glama's open rubric for tool-definition quality, and it is already scoring the ecosystem at a scale nothing here approaches — every tool of every server in their registry. Run it. It grades a definition on its own terms, against fixed anchors: "TDQS scores a tool definition, not tool behavior." That is a different question from the one here, which is comparative — put your tool in a catalogue beside other good tools and see whether a search for your user's actual question surfaces it. A definition can score well and still lose that. The server this was built for scores A, 4.4/5.0 on TDQS and its search_records tool still misses the shortlist in the example at the top of this README.

Requirements

Node 18+. No dependencies.

Licence

MIT. Built by referencesource.org, which runs an MCP server of its own — checked by this tool, with no special case, in the example at the top.

Keywords

mcp

FAQs

Package last updated on 20 Sep 2026

Related posts