
Company News
AWS Security Hub Adds Socket for Supply Chain Security
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.
rcsb-mcp
Advanced tools
An MCP server for interrogating PDB structures — search, inspect, and cross-reference across the RCSB Search, Data, and Sequence Coordinates APIs
An MCP server for interrogating Protein Data Bank structures — discover, inspect, and cross-reference — from LLM clients (Claude Desktop, MCP Inspector, Cursor, etc.). It spans three RCSB APIs:
| Tool | What it does |
|---|---|
rcsb_list_pdb_search_attributes | Discover searchable attribute paths, types, and operators. schema="structure" (default, ~677) or schema="chemical" (~57: chem_comp.*, drugbank_info.*, ...). |
rcsb_find_go_terms | Resolve a free-text molecular function / biological process / cellular component to Gene Ontology ids (via EBI QuickGO), annotated with PDB entry counts — then search by rcsb_polymer_entity_annotation.annotation_lineage.id. |
rcsb_find_interpro_domains | Resolve a free-text protein domain / family / fold to InterPro ids (via EBI InterPro API), annotated with PDB entry counts — then search by rcsb_polymer_entity_annotation.annotation_id. |
rcsb_find_enzyme_classes | Resolve a free-text enzyme / reaction to Enzyme Commission (EC) numbers (via EBI Search/IntEnz), annotated with PDB entry counts — then search by rcsb_polymer_entity.rcsb_ec_lineage.id (hierarchical). |
rcsb_find_disease_terms | Resolve a free-text disease / condition to MONDO ids (via EBI OLS), annotated with PDB entry counts — then search by rcsb_uniprot_annotation.annotation_lineage.id (hierarchical, UniProt-based). |
rcsb_find_organisms | Resolve a free-text organism / common name / clade to NCBI Taxonomy ids (via UniProt taxonomy), annotated with PDB entry counts — then search by rcsb_entity_source_organism.taxonomy_lineage.id (hierarchical: a clade id matches every organism beneath it). |
rcsb_search_fulltext | Free-text keyword search (e.g. "CRISPR Cas9"), optionally refined with structured attributes filters (AND/OR) and sort. |
rcsb_search_by_attribute | Structured search on one or more indexed attributes (resolution, organism, release date, ...) combined with a single AND/OR. Each AttributeFilter supports exists, negation, case_sensitive; chemical=True (text_chem). |
rcsb_search_by_sequence | MMseqs2 sequence-similarity search (BLAST-like). |
rcsb_search_by_chemical | Chemical search by SMILES/InChI descriptor (whole-molecule or substructure) or molecular formula. |
rcsb_search_by_structure | 3D shape-similarity search against a reference PDB assembly or chain. |
rcsb_search_by_seqmotif | Short sequence-motif search (PROSITE pattern, regex, or simple wildcards). |
rcsb_search_strucmotif | 3D structural-motif search: structures sharing a geometric arrangement of specific residues (e.g. a catalytic triad). |
The two text tools (rcsb_search_fulltext, rcsb_search_by_attribute)
also take group_by_identity (100/95/90/70/50/30) to return one representative
per sequence-identity cluster — i.e. non-redundant results. To search
chemical-component attributes, find the path with
rcsb_list_pdb_search_attributes(schema="chemical"), then pass chemical=True to
rcsb_search_by_attribute / rcsb_search_fulltext (usually with return_type="mol_definition").
Both catalogs (structure and chemical) are generated from the live metadata schemas by
scripts/generate_search_attributes.py.
Counting and faceting are output options on every rcsb_search_* tool, not separate
tools: each response includes total_count (the full match count — for "how many ..." run a
search with limit=1 and read it), and passing facets returns a breakdown
(terms/histogram/date_histogram/range/cardinality) instead of hits. The rcsb_search_by_*
service tools (sequence, chemical, structure, seq/struc-motif) also take optional attributes
filters, so e.g. a sequence search can be restricted to an organism in one call.
Sorting is likewise available on every rcsb_search_* tool via sort_by (an
attribute path) + sort_direction (asc/desc), replacing the default score ordering (for
the similarity searches this overrides the similarity-ranked order). Only attributes indexed
for sorting work — those exposing exact_match (strings) or equals (numbers/dates) in
rcsb_list_pdb_search_attributes; sorting is not available for return_type="mol_definition"
(chemical-component results are ranked by score only).
Paging. Every search tool that returns hits accepts limit (1–100, default
10) and offset (default 0). Each response reports total_count, has_more,
and next_offset; to fetch the next page, call the tool again with the same
query and offset set to the returned next_offset.
There is one tool per Data API GraphQL root field. Each takes a list of IDs
(singular lookups = a one-element list) plus an optional fields argument to
override the curated default selection with your own GraphQL sub-selection.
Unknown IDs are reported under not_found. Discover the paths to put in fields
with rcsb_describe_data_object — browse a level, drill into a nested object with
into=, or search the schema by keyword with query= + max_depth=. Every path it
returns is verified against the live schema, so don't guess field names.
| Tool | Object | Example ID |
|---|---|---|
rcsb_get_entries | PDB entries | "4HHB" |
rcsb_get_polymer_entities | Polymer entities (protein/NA) | "4HHB_1" |
rcsb_get_nonpolymer_entities | Ligand/cofactor entities | "4HHB_3" |
rcsb_get_branched_entities | Carbohydrate entities | "5FMB_2" |
rcsb_get_polymer_entity_instances | Polymer chains | "4HHB.A" |
rcsb_get_nonpolymer_entity_instances | Bound-ligand instances | "4HHB.E" |
rcsb_get_branched_entity_instances | Glycan chains | "5FMB.C" |
rcsb_get_assemblies | Biological assemblies | "4HHB-1" |
rcsb_get_interfaces | Assembly interfaces | "1BMV-1.1" |
rcsb_get_chem_comps | Chemical components / ligands | "HEM", "ATP" |
rcsb_get_entry_groups | Entry groups | "G_1002266" |
rcsb_get_polymer_entity_groups | Polymer entity groups (seq. clusters) | "85_70" |
rcsb_get_nonpolymer_entity_groups | Non-polymer entity groups | "ATP" |
rcsb_get_uniprot | UniProt record (single) | "P69905" |
rcsb_get_pubmed | PubMed record (single, integer) | 6726807 |
rcsb_get_group_provenance | Grouping provenance (single) | "provenance_sequence_identity" |
rcsb_describe_data_object | Introspect an object's live GraphQL schema to build a fields= selection: browse a level, drill into a nested object with into=, or search by keyword with query= + max_depth= (flat, incl. nested + cross-object paths). Returns verified dotted paths. The Data API analogue of rcsb_list_pdb_search_attributes. | — |
The Search API only returns identifiers, so a search is the first step: batch the
returned ids into the matching rcsb_get_* tool to fetch titles, organisms, and
other metadata (these tools query the GraphQL endpoint, batching every requested ID
into one request). All 16 typed tools are generated from a single registry in
queries.py (DATA_OBJECTS), so adding a field or
endpoint is a one-line change.
Maps alignments and positional annotations between sequence reference systems
(UNIPROT, NCBI_PROTEIN, NCBI_GENOME, PDB_ENTITY, PDB_INSTANCE). Each
tool takes an optional fields argument to override the default selection; use
rcsb_describe_seqcoord_object to discover what fields are available.
This is the only RCSB API that cross-references NCBI (RefSeq protein /
genome) — the Data API only knows UniProt. So "what NCBI proteins map to a PDB
structure?" is answered by rcsb_seqcoord_alignments, not the Data API. PDB query
ids must be entity-level (4HHB_1), not a bare entry (4HHB); for a whole
entry, query each polymer entity.
| Tool | What it does |
|---|---|
rcsb_seqcoord_alignments | Cross-reference a sequence across PDB / UniProt / NCBI with aligned ranges (e.g. 4HHB_1 → NCBI proteins NP_000508, NP_000549). |
rcsb_seqcoord_annotations | Positional features for one sequence, from one or more annotation sources (UNIPROT, PDB_ENTITY, PDB_INSTANCE, PDB_INTERFACE). |
rcsb_seqcoord_group_alignments | Alignments among members of a sequence group (MATCHING_UNIPROT_ACCESSION / SEQUENCE_IDENTITY). |
rcsb_seqcoord_group_annotations | Annotations across a group; summary=True returns a positional summary. |
rcsb_describe_seqcoord_object | Introspect the live schema to discover fields available on a seqcoord object (for use with fields=). |
# run the published package without installing (recommended for clients)
uvx rcsb-mcp
# or install it
pip install rcsb-mcp
rcsb-mcp is listed in the Official MCP Registry
as io.github.rcsb/rcsb-mcp, so registry-aware clients can discover it directly.
For local development, install from the project root instead:
pip install -e .
# or with uv
uv pip install -e .
# unit tests (no network)
hatch test # or: python tests/test_queries.py
# run the server over stdio
python -m rcsb_mcp.server
# or, after install:
rcsb-mcp
# inspect interactively
npx @modelcontextprotocol/inspector python -m rcsb_mcp.server
There is also an end-to-end evaluation suite (evals/) — 10
read-only, stable questions that measure how well an LLM can drive these tools to
answer real PDB questions. See evals/README.md to run it.
Edit claude_desktop_config.json:
~/Library/Application Support/Claude/claude_desktop_config.json%APPDATA%\Claude\claude_desktop_config.json{
"mcpServers": {
"rcsb-mcp": {
"command": "uvx",
"args": ["rcsb-mcp"]
}
}
}
For a local source checkout, point at the module instead:
{
"mcpServers": {
"rcsb-mcp": {
"command": "python",
"args": ["-m", "rcsb_mcp.server"],
"cwd": "/absolute/path/to/rcsb-mcp/src"
}
}
}
Restart Claude Desktop. The tools appear under the connectors (plug) icon.
rcsb_search_fulltext (keyword + attributes)rcsb_search_fulltext (keyword + attributes, sort_by)rcsb_search_by_sequencercsb_search_by_chemicalrcsb_search_by_structurercsb_search_by_seqmotifrcsb_find_go_terms → rcsb_search_by_attribute on rcsb_polymer_entity_annotation.annotation_lineage.idrcsb_find_interpro_domains → rcsb_search_by_attribute on rcsb_polymer_entity_annotation.annotation_idrcsb_find_enzyme_classes → rcsb_search_by_attribute on rcsb_polymer_entity.rcsb_ec_lineage.idrcsb_find_disease_terms → rcsb_search_by_attribute on rcsb_uniprot_annotation.annotation_lineage.idrcsb_find_organisms → rcsb_search_by_attribute on rcsb_entity_source_organism.taxonomy_lineage.idrcsb_search_fulltext with group_by_identity=90rcsb_search_by_attribute (read total_count)rcsb_search_fulltext with facetsrcsb_search_strucmotifrcsb_list_pdb_search_attributes(schema="chemical") + rcsb_search_by_attribute with chemical=Truercsb_get_entriesrcsb_get_polymer_entitiesrcsb_get_chem_compsrcsb_get_assembliesrcsb_get_uniprotrcsb_seqcoord_alignmentsrcsb_seqcoord_alignments per entity (4HHB_1, 4HHB_2), to_ref=NCBI_PROTEINrcsb_seqcoord_annotationsrcsb_describe_data_object to find the path, then the matching rcsb_get_* tool with fields=RCSB_MCP_REPORT_BASE_URL)rcsb_render_report returns a self-contained link, not the markup — so the
agent never has to reproduce the ~20 KB document (the single most expensive step
of a report turn). The whole report is gzip+base64url-packed into the URL, so the
server stores nothing and any replica renders any link on demand.
Set RCSB_MCP_REPORT_BASE_URL to the origin that serves this MCP (e.g.
https://rcsb-mcp.rcsb.org) and the tool returns
{ url: "<base>/r?d=<packed report>", html: null }. The agent hands the user that
link; opening it hits the stateless render endpoint:
GET /r?d=<gzip+base64url of the report JSON> → text/html
The endpoint decodes, validates against the report schema, and renders with the
fixed template. It is hardened for a public route: the d token and its
decompressed size are both capped (a gzip bomb is refused before it expands), the
page carries Content-Security-Policy: default-src 'none' + noindex, and every
value is escaped by the template — a crafted link can only ever produce an escaped
report.
Fallback. When RCSB_MCP_REPORT_BASE_URL is unset (e.g. local stdio dev with
no reachable endpoint) or a report is too large to pack into a URL, the tool
returns html instead of url, and the rcsb_search_assistant prompt tells the model to
write it to a .html file. Reports compress to ~1 KB even at 50 rows, so the
size fallback is rare.
https://search.rcsb.org/rcsbsearch/v2/query (POST, JSON body).https://data.rcsb.org/graphql (POST, GraphQL). It returns
HTTP 200 even for query errors, reporting them in an errors array.https://sequence-coordinates.rcsb.org/graphql
(POST, GraphQL; same HTTP-200-with-errors behavior).rcsb_find_* resolvers map free text to ontology ids via EBI services — the non-RCSB
dependencies: GO via QuickGO (.../QuickGO/services/ontology/go/search), InterPro
(.../interpro/api/entry/interpro/), EC via EBI Search/IntEnz (.../ebisearch/ws/rest/intenz),
and disease via OLS/MONDO (.../ols4/api/search?ontology=mondo). The resolved ids then drive
RCSB annotation searches (rcsb_polymer_entity_annotation.*, rcsb_polymer_entity.rcsb_ec_lineage.id,
rcsb_uniprot_annotation.annotation_lineage.id).rcsb_search_by_attribute is in the
Search API attribute reference;
the Data API schema is documented at
data.rcsb.org/index.html#gql-api.The server also exposes an MCP prompt, rcsb_search_assistant ("RCSB PDB search
assistant") — the full tool-routing guide, followed by the search requirements and the
HTML-report output format. Because it is served over the protocol's prompts
capability, any MCP client can list and invoke it (e.g. Claude Desktop surfaces server
prompts in the + / prompt menu); there's nothing to copy-paste. The policy half lives
in
src/rcsb_mcp/prompts/rcsb_search_assistant.md
and the guide half is
rcsb_mcp_guide.md, joined at request time so
the two never drift; both ship with the package.
Invoke it when you want answers formatted as a PDB report. It is self-sufficient: the
appended guide is the same text as the always-on server instructions, so the prompt
still works on clients that never inject those — and the tool descriptions' "see the
server instructions" cross-references resolve against it. A second prompt,
rcsb_mcp_guide, serves that guide on its own for sessions that want the routing
guidance without the report policy; load one or the other, not both.
FAQs
An MCP server for interrogating PDB structures — search, inspect, and cross-reference across the RCSB Search, Data, and Sequence Coordinates APIs
The pypi package rcsb-mcp receives a total of 426 weekly downloads. As such, rcsb-mcp popularity was classified as not popular.
We found that rcsb-mcp demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.
Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Company News
Socket is now in the AWS Security Hub Extended plan. Adopt it through AWS, apply committed spend, and block malicious open source packages.

Research
/Security News
Popular npm packages keyv and cacheable compromised.

Security News
A misconfiguration gave three Anthropic models internet access, and one, believing it was in a simulation, shipped a credential-stealing package to PyPI.