MCP Server for 3D protein structural data retrieval & analysis from RCSB PDB, PDBe, and UniProt.
Federated protein structure & annotation across experimental (PDB) and predicted (AlphaFold) models via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://protein.caseyjhand.com/mcp
Experimental (PDB) and predicted (AlphaFold) protein structures, federated behind one surface. Search, fetch, align, compare, and annotate structures and their ligands across RCSB, AlphaFold DB, 3D-Beacons, UniProt, InterPro, and Foldseek — all keyless. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|---|
protein_search_structures | Search experimental and predicted structures by free text, sequence, or organism/method/resolution filters, with optional facet breakdowns. |
protein_get_structure | Fetch metadata and coordinate-file URLs by ID — experimental (PDB), predicted (AlphaFold), or best-available — with batch partial success and optional coordinate inlining. |
protein_find_similar | Find sequence homologs (RCSB mmseqs2) or fold homologs (Foldseek) from a sequence, PDB ID, or UniProt accession. |
protein_track_ligands | Resolve ligand names/formulas to component IDs, find structures containing a ligand, or map binding-site residues. |
protein_compare_structures | Structurally align multiple structures (TM-align / jFATCAT) to a reference or as a full pairwise matrix. |
protein_analyze_collection | Profile the PDB into distributions and trends with server-side facets — counts, histograms, timelines, and cross-tabs. |
protein_get_annotations | Fetch UniProt features and natural variants plus InterPro domain/family memberships with GO terms. |
| Resource | Description |
|---|---|
pdb://{entry_id} | Experimental structure summary for a PDB entry — title, method, resolution, organism, bound ligands, and per-entity chain IDs in both the author (authAsymIds) and mmCIF label (labelAsymIds) namespaces. |
af://{uniprot} | Predicted-structure summary for a UniProt accession from AlphaFold DB — mean pLDDT, confidence-band fractions, model URLs, and version. |
All resource data is also reachable via tools — pdb://{entry_id} mirrors protein_get_structure for source: experimental, and af://{uniprot} mirrors it for source: predicted. Many MCP clients are tool-only and don't surface resources; the summaries remain reachable through the tools.
protein_search_structures toolcontent_type scopes the search to experimental, predicted, or all (default) — all is a genuine union, so computed models appear alongside PDB entriessource; sequence hits in either universe expose a chainable entry id plus the matched polymer entityId; experimental hits carry title, method, resolution, and organism enrichment, and AlphaFold models their parsed UniProt accessionstart and limit page through ranked results; nextStart is returned while another page remains, and an empty page past the end names the offset in notice rather than reporting no matchesfacets return a method / organism / release-year breakdown alongside the hits — each dimension may be listed once and reports how many matches carry no value for it; a capped dimension is named in notice, with protein_analyze_collection (larger bucket_limit) as the route to the long tailprotein_get_structureprotein_get_structure toolsource: experimental batches PDB entry IDs (also resolving computed-model IDs like AF_*/MA_* from search, tagged source: predicted with their provider); source: predicted takes UniProt accessions for AlphaFold models with pLDDT/PAE; source: best_available takes UniProt accessions and returns the top federated model (highest-resolution experimental if one exists, else the best prediction)failed[]; requested/processed disclose IDs dropped beyond the batch cap, and every advisory (cap, failure, overflow) joins into one noticesource: experimental, computed models included, also carry polymerEntities (both authAsymIds and labelAsymIds), ligands, molecularWeight, and releaseDatecoordinateUrls lists only files that exist: BinaryCIF comes from RCSB's ModelServer, the PDB format is omitted for large mmCIF-only entries, and a computed model's files come from its provider (all three formats from AlphaFold DB, mmCIF from ModelArchive) — an AlphaFold model whose provider lookup fails keeps only its RCSB BinaryCIF, named in noticeinclude_coords inlines coordinate content, subject to a response budget — an over-budget batch returns a per-structure size outline (re-call with sections: [ids]), and a single oversized file is withheld with a pointer to its coordinateUrlsattribution block naming upstream data licenses and citationsprotein_find_similar toolby: sequence runs a synchronous RCSB mmseqs2 search; by: structure runs an asynchronous Foldseek search against experimental and predicted databases — query from a raw sequence, a PDB ID, or a UniProt accessionstart/limit and report totalCount, echoing start and returning nextStart while another page remains; an empty page past the end names the offset in notice, distinct from a search with no matchespdb100 + afdb50; override via databases (e.g. afdb-swissprot, BFVD)status: computing with a ticketId — re-call with ticket_id to resume; a completed structure search returns the same ticket so a new start pages the finished jobquery, 0-based, default 0) and reports queryCount, with a notice naming the other queries; pass query with ticket_id to read another chain's hits from the same job. An out-of-range query is rejected (query_out_of_range), not answered with an empty listscore across every searched database (hits without a score last, ties by database then target) before start/limit pagingsequence, max_evalue, min_identity under by: sequence; ticket_id, databases, query under by: structure) — a field the selected mode can't consume is rejected, not ignoredprotein_track_ligands toolmode: find_ligand resolves a name or formula to chemical component IDs with formula, weight, SMILES, and InChIKey — ranked by deposition frequency, most-common match firsttotalCount and candidatesConsidered report how many components matched and how many were ranked; a broad name whose matches exceed the candidate pool gets a notice to narrow the queryquery matches on exact composition, spaced (C29 H31 N7 O) or unspaced; anything else (a component ID included) matches on name and synonymsmode: structures_with_ligand returns PDB entries containing a ligand by exact component ID, with start/limit paging and nextStart while another page remains; a page past the end names the offset in notice instead of reporting no entriesmode: binding_site returns the protein residues lining a ligand's pocket in a structure, with contact distances; ligand instances page with start/limit like structures_with_ligandasymId, seqId) and author numbering (authAsymId, authSeqId) — 1IEP's imatinib pocket lists label THR93 as author THR315; the ligand instance reports its own author chain and residue numberprotein_compare_structures tooltm-align, fatcat-rigid, or fatcat-flexible; optional per-structure chain restricts the alignment to a single mmCIF label chainreference: first aligns every structure to the first; reference: all_pairs computes the full pairwise matrix; a structure repeated in structures[] is compared oncestatus: computing with a job uuid; a failed pair degrades only its own row{ a, b, uuid } entry in resume[] to poll a computing pair instead of resubmitting; a resumed pair reports a/b in the order its job was submitted, whatever the current structures[] order, and a resume under a different method is rejectedmodeledResidues and 0–100 coverage, ordered [a, b]; TM-score is normalized by a's length, so the same pair scores differently when reversedprotein_analyze_collection toolmethod, organism, polymer_type, resolution, release_year, or molecular_weightgroup_by dimension for a breakdown, or two distinct dimensions for a cross-tab (the first nests the second); a repeated dimension is rejectedinterval sets a histogram bin width (a number, for resolution or molecular_weight) or date-histogram period (year, the only one RCSB accepts) — applies to whichever requested dimension can consume that type; rejected when neither canquery, organism, method, or max_resolution; content_type selects the structure universebucket_limit caps buckets per dimension level, not per response — a cross-tab applies it separately to the parent and each nested child, up to bucket_limit × (1 + bucket_limit) buckets; notice names every capped position and bucketsReturned gives the realized totalmissingValueCount — matches carrying no value for that attribute (e.g. a resolution breakdown excludes NMR entries; computed models have neither method nor resolution)protein_get_annotations toolambiguity — pass chain (an author chain ID) to select a specific oneinclude scopes which classes are fetched (features, domains, variants, all); limit caps each class independently (1–200, default 50), with a truncated class disclosed in noticeattribution block naming the upstream data licenses and citations (see Upstream data licensing)pdb://{entry_id} resourceapplication/json — title, method, resolution, organism, bound ligands, and per-entity chain IDs in both the author (authAsymIds) and mmCIF label (labelAsymIds) namespacesprotein_get_structure for source: experimental; entry_id is a PDB entry ID (e.g. 4HHB)af://{uniprot} resourceapplication/json — mean pLDDT, confidence-band fractions, model URLs (cif/pdb/bcif), and AlphaFold model versionuniprot accepts a UniProt accession or an AlphaFold DB entry ID (e.g. AF-P69905-F1); mirrors protein_get_structure for source: predictedBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
PDB / AlphaFold-specific:
ticketId / per-pair uuid) instead of blocking — re-call with ticket_id or a resume[] entry to poll the same job instead of resubmittingAgent-friendly output:
source (experimental / predicted), the engine and database that produced it, and effective-query / total-count echoes so agents can reason about coveragefailed[], per-pair status) instead of failing the whole request, each with actionable recovery textsource and status unions, computing results with resume tickets, and budget-overflow outlines let callers branch on data, not string parsingA public instance is available at https://protein.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"protein": {
"type": "streamable-http",
"url": "https://protein.caseyjhand.com/mcp"
}
}
}
Add the following to your MCP client configuration file. No API key is required — every upstream provider is keyless.
{
"mcpServers": {
"protein-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/protein-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"protein-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/protein-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"protein-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/protein-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/protein-mcp-server.git
cd protein-mcp-server
bun install
All upstream providers are keyless, so the server runs out of the box with no configuration. Every variable below is optional.
| Variable | Description | Default |
|---|---|---|
PROTEIN_ASYNC_POLL_TIMEOUT_MS | Max wall-clock to poll an async job (alignment / Foldseek) before returning a computing result. | 30000 |
PROTEIN_MAX_BATCH_IDS | Cap on IDs accepted by protein_get_structure in one batch (1–100). | 25 |
PROTEIN_MAX_COMPARE_STRUCTURES | Cap on structures per protein_compare_structures call (2–25). | 10 |
PROTEIN_FACET_BUCKET_CAP | Default cap on buckets per protein_analyze_collection dimension (1–500). | 50 |
PROTEIN_FANOUT_CONCURRENCY | Max concurrent upstream requests for per-ID / per-pair fan-out (1–16). | 5 |
RCSB_SEARCH_BASE_URL | Base URL for the RCSB Search API v2. | https://search.rcsb.org |
ALPHAFOLD_BASE_URL | Base URL for the AlphaFold Protein Structure Database API. | https://alphafold.ebi.ac.uk |
FOLDSEEK_BASE_URL | Base URL for the Foldseek structural-similarity search service. | https://search.foldseek.com |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for the HTTP server. | 3010 |
MCP_SESSION_MODE | HTTP session mode: stateless, stateful, or auto. The server declares stateless in code; set this to override it. | stateless |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of provider base-URL overrides and tuning limits.
Build and run:
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
Run checks and tests:
bun run devcheck # Lint, format, typecheck, security
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
docker build -t protein-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=http -p 3010:3010 protein-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/protein-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools/resources and inits the provider services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). |
src/mcp-server/resources | Resource definitions (*.resource.ts). |
src/services | Provider service layer — RCSB (search, data, facets), AlphaFold, 3D-Beacons (best-available), UniProt (incl. InterPro/GO), Structural Comparison alignment, Foldseek, and shared HTTP/identifier/concurrency helpers. |
tests/ | Unit and integration tests mirroring src/. |
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped logging, ctx.state for tenant-scoped storagesrc/mcp-server/*/definitions/index.tsStructure and annotation data comes from public upstream databases, each under its own license. protein_get_structure and protein_get_annotations carry an attribution block on every response — the license, citation, and homepage for each source that contributed to that specific response — so the attribution obligation travels with the data to downstream consumers rather than living only here. CC BY / CC BY-SA sources require attribution on redistribution; CC0 sources are citation-only (attribution encouraged, not required).
| Source | Contributes to | License |
|---|---|---|
| RCSB PDB | protein_get_structure — experimental records | CC0 1.0 Universal |
| AlphaFold DB | protein_get_structure — predicted models | CC BY 4.0 |
| ModelArchive | protein_get_structure — MA_* computed models | CC BY 4.0 |
| SWISS-MODEL | protein_get_structure — best_available models | CC BY-SA 4.0 |
| BFVD | protein_get_structure — best_available models | CC BY 4.0 |
| UniProt | protein_get_annotations | CC BY 4.0 |
| InterPro | protein_get_annotations — domain/family data | CC0 1.0 Universal |
| GO | protein_get_annotations — GO terms | CC BY 4.0 |
best_available federates predicted models through 3D-Beacons, so the attribution block credits the actual contributing provider (AlphaFold DB, SWISS-MODEL, BFVD, …); a provider without a curated license entry carries a See provider terms fallback pointing back to 3D-Beacons rather than a fabricated license. InterPro's own domain/family classifications are CC0; the GO terms carried alongside them are separately CC BY 4.0, so each is credited independently only when it actually contributes. Full citations for each source travel in the attribution block of the relevant tool responses. This covers upstream data licensing — the server's own code is licensed separately (see License).
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y protein-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-protein-mcp-server": {
"command": "npx",
"args": [
"-y",
"protein-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/protein-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.