Back to Directory/Developer Tools

io.github.cyanheads/chembl-mcp-server

Link compounds to protein targets, rank bioactivity, and look up drug mechanisms and indications.

Developer ToolsTypeScriptv0.3.2

@cyanheads/chembl-mcp-server

Link compounds to protein targets, rank bioactivity (IC50/Ki/EC50), and look up drug mechanisms and indications over ChEMBL via MCP. STDIO or Streamable HTTP.

7 Tools (+1 opt-in) • 2 Resources

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://chembl.caseyjhand.com/mcp


Overview

Drug-discovery data over ChEMBL (EBI) — the curated link between compounds, protein targets, and measured bioactivity (IC50/Ki/EC50), plus drug mechanisms and indications. Search compounds by name, ID, or structure, resolve protein targets, rank bioactivity measurements, and look up drug mechanisms and indications from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

ToolDescription
chembl_search_moleculesFind compounds by name / ChEMBL ID / InChIKey, or run a structure search (exact | similarity | substructure) from a SMILES.
chembl_get_bioactivitiesThe flagship compound↔target bridge: bioactivity measurements for a molecule, a target, or both (the compound×target pair), ranked on pchembl_value, or the measurements without one via potency_view. Large sets spill to a canvas.
chembl_search_targetsResolve a protein / gene symbol / UniProt accession to the ChEMBL target ID chembl_get_bioactivities needs.
chembl_get_drug_infoDrug pharmacology — mechanism(s) of action, molecular target(s), action type, first-approval year, and clinical indications.
chembl_get_assayAssay provenance behind a bioactivity row — type, target, organism, and ChEMBL's 1–9 confidence score.
chembl_dataframe_queryRun a read-only SQL SELECT over the bioactivity rows spilled to a canvas — rank, group, dedupe, aggregate across the full set.
chembl_dataframe_describeList the tables and columns staged on a canvas, so you can write correct SQL before querying.
chembl_dataframe_dropDrop a named staged table from a canvas. Opt-in via CHEMBL_DATAFRAME_DROP_ENABLED=true — absent from tools/list when off, since TTL already reclaims staged tables.

Resources

ResourceDescription
chembl://molecule/{chemblId}A molecule record by ChEMBL ID — the same shape a chembl_search_molecules row carries.
chembl://target/{chemblId}A target record by ChEMBL target ID — preferred name, type, organism, and component UniProt accessions + gene symbols.

All resource data is also reachable via the tools, so tool-only MCP clients lose nothing. There are no prompts — the canonical workflows are short tool chains an agent composes directly, and the cross-server chain guidance ships as server-level instructions instead.

Capability reference

chembl_search_molecules tool

  • Default search_type=name matches drug names, synonyms, ChEMBL IDs, and InChIKeys in one query; a query that is exactly a ChEMBL ID or InChIKey routes to ChEMBL's single-record lookup instead of the fuzzy text index (totalCount: 1)
  • Structure search via search_type: exact, similarity (Tanimoto ≥ similarity_threshold, integer 40–100, default 70), or substructure — supply structure as a SMILES; max_phase_min (name search only) restricts to compounds at or above a max clinical phase
  • Every row carries max_phase, MW, AlogP, Lipinski rule-of-five violations, and QED; only search_type=similarity results carry a Tanimoto similarity percent — absent, not null, on other modes
  • Paginated via nextCursor / cursor, omitted (not null) on the last page — redeem a cursor with the same filters that minted it
  • Chain molecule_chembl_id into chembl_get_bioactivities or chembl_get_drug_info

chembl_get_bioactivities tool

  • Supply at least one of molecule_chembl_id or target_chembl_id; supplying both narrows to that compound–target pair — neither is a missing_filter error
  • Filter by standard_type (IC50/Ki/EC50/…), pchembl_value_min, assay_type, organism; ranked on pchembl_value — comparable only within one standard_type
  • potency_view selects potency_ranked (default, measurements with a derivable pchembl_value) or null_potency (the excluded rows); pchembl_value_min with null_potency is a contradictory_potency_filter error, and totalCount spans both views
  • Numerics are coerced to number | null at the service boundary — a missing potency reads as null, never 0
  • Large sets spill to a DataCanvas table per view (bioactivities / bioactivities_null_potency), capped at CHEMBL_MAX_SPILL_ROWS (default 50,000) and reported truncated: true + staged_row_count when hit; requires CANVAS_PROVIDER_TYPE=duckdb
  • The inline preview is always capped at limit (default 25) regardless of spill status; the optional canvas_id reuses a canvas, but re-querying the same view replaces its prior rows

chembl_search_targets tool

  • Supply at least one of accession (UniProt, e.g. P00533), gene_symbol, or query (free-text); narrow with organism and target_type — none supplied is a missing_input error
  • A UniProt accession is the most precise input — chain it from a uniprot/protein server
  • Each row carries target type, organism, and component UniProt accessions + gene symbols, flattened from ChEMBL's nested component synonyms
  • Paginated via nextCursor / cursor, the same contract as chembl_search_molecules
  • Chain target_chembl_id into chembl_get_bioactivities

chembl_get_drug_info tool

  • Supply molecule_chembl_id; returns mechanism(s) of action, molecular target(s), action type, first-approval year, and clinical indications with the max phase reached for each
  • Mechanisms and indications are fetched with Promise.allSettled, so a rejected list degrades to a disclosed partial result rather than failing the call
  • Each list carries its own mechanisms_status / indications_status (complete / truncated / failed) next to a *_total_count — an empty array is authoritative only when the status is complete
  • A mechanism's target_chembl_id chains into chembl_get_bioactivities

chembl_get_assay tool

  • Supply assay_chembl_id from a chembl_get_bioactivities row
  • Returns description, assay type (binding / functional / ADMET / toxicity), the target measured, organism, and ChEMBL's 1–9 confidence score (9 = direct assay on the protein target, lower = homologous or indirect)
  • Call it to judge whether two measurements are comparable before ranking them together

chembl_dataframe_query tool

  • Accepts a single read-only SELECT against a canvas_id from a spilled chembl_get_bioactivities call; writes, DDL, and non-SELECT statements are rejected by the framework SQL gate
  • Reference each staged table by the name chembl_get_bioactivities returned — bioactivities (potency_ranked) or bioactivities_null_potency (null_potency); discover columns with chembl_dataframe_describe first
  • Two independent bounds, each disclosed: truncated is the canvas engine's own query-result cap; rendered_rows is how many rows the content[] markdown table holds under its character budget — either can trip without the other; page past both with SQL LIMIT/OFFSET
  • structuredContent.rows always carries the full materialized result regardless of the render bound
  • Requires CANVAS_PROVIDER_TYPE=duckdb, else a canvas_disabled error

chembl_dataframe_describe tool

  • Supply a canvas_id from a spilled chembl_get_bioactivities call
  • Returns each staged table/view with its row count, kind, and column names + types
  • Requires CANVAS_PROVIDER_TYPE=duckdb, else a canvas_disabled error

chembl_dataframe_drop tool

  • Opt-in — registered only when CHEMBL_DATAFRAME_DROP_ENABLED=true; absent from tools/list when off, though it still appears in the server manifest carrying the enable hint
  • Drops a named staged table by canvas_id + table_name; returns dropped: true if it existed, false if already gone
  • Rarely needed — per-table and per-canvas TTL already reclaim staged tables; reach for it only to free a large table early in a long session
  • Requires CANVAS_PROVIDER_TYPE=duckdb

chembl://molecule/{chemblId} resource

  • Molecule record as application/json — the same shape a chembl_search_molecules row carries (ID, names, structures, properties, max clinical phase)
  • chemblId is validated against the CHEMBL\d+ pattern
  • Fully covered by the tool surface — a convenience injectable-context mirror of the per-record fetch

chembl://target/{chemblId} resource

  • Target record as application/json — preferred name, type, organism, and component UniProt accessions + gene symbols
  • chemblId is validated against the CHEMBL\d+ pattern
  • Fully covered by the tool surface — a convenience injectable-context mirror of the per-record fetch

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

ChEMBL-specific:

  • Bidirectional bioactivity bridge — one tool serves both compound→target and target→compound, ranked on pchembl_value
  • Structure search (exact / similarity / substructure) consolidated under one discovery tool via a search_type enum
  • String → number | null numeric coercion at the service boundary — a missing potency becomes null, never 0
  • DataCanvas spill on the flagship tool — tens-of-thousands-of-row activity sets stream to a DuckDB table you inspect with chembl_dataframe_describe and query via chembl_dataframe_query
  • Server-level instructions carry the cross-server chain guidance and the ChEMBL CC BY-SA 3.0 attribution requirement

Agent-friendly output:

  • Provenance on every response — total-found counts, applied-filter echo, and a spill notice so agents know whether the preview is the full set or a slice of a canvas table
  • Truncation disclosure — capped searches report shown / cap / totalCount, and spilled tables report truncated + staged_row_count, so a page or slice is never mistaken for the complete set
  • Typed, recoverable errors — missing_filter / missing_input / contradictory_potency_filter / canvas_disabled carry recovery hints, so callers correct the call without parsing prose
  • Never fabricates — normalization and format() preserve null potency / units; a missing measurement renders as "not reported", never 0

Getting started

Public Hosted Instance

A public instance is available at https://chembl.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "chembl-mcp-server": {
      "type": "streamable-http",
      "url": "https://chembl.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

ChEMBL is keyless — no API key or account is required.

Add the following to your MCP client configuration file:

{
  "mcpServers": {
    "chembl-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/chembl-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "chembl-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/chembl-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "chembl-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/chembl-mcp-server:latest"]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

To unlock the analytical SQL path (the bioactivities spill and the chembl_dataframe_* tools), add "CANVAS_PROVIDER_TYPE": "duckdb" to the env.

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • Optional: set CANVAS_PROVIDER_TYPE=duckdb to enable the DataCanvas SQL path for large bioactivity sets.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/chembl-mcp-server.git
  1. Navigate into the directory:
cd chembl-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment:
cp .env.example .env
# edit .env to override any defaults (all optional)

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts. ChEMBL is keyless, so every variable is optional.

VariableDescriptionDefault
CANVAS_PROVIDER_TYPESet to duckdb to enable the bioactivity spill and the chembl_dataframe_* SQL tools. When none, large sets inline a preview but never spill.none
CHEMBL_API_BASE_URLBase URL for the ChEMBL REST data API. Override for a private mirror or pinned host.https://www.ebi.ac.uk/chembl/api/data
CHEMBL_REQUEST_TIMEOUT_MSPer-request timeout in milliseconds for upstream ChEMBL fetches.30000
CHEMBL_MAX_PAGE_SIZEChEMBL per-page cap when streaming activity pages for the spill (max 1000).1000
CHEMBL_DEFAULT_LIMITDefault result limit applied when callers omit it.25
CHEMBL_MAX_SPILL_ROWSCeiling on rows chembl_get_bioactivities stages to a canvas table, and so on the upstream page drain behind it. Over the cap the response reports truncated: true.50000
CHEMBL_DATAFRAME_DROP_ENABLEDRegister the opt-in chembl_dataframe_drop tool (absent from tools/list when off).false
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTPort for the HTTP server.3010
MCP_AUTH_MODEAuth mode: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (RFC 5424).info
LOGS_DIRDirectory for log files (Node.js only).<project-root>/logs
OTEL_ENABLEDEnable OpenTelemetry instrumentation.false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security, changelog sync
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec
    

Docker

docker build -t chembl-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=stdio chembl-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/chembl-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them. The fully-resolved @duckdb native binary is copied from the build stage so CANVAS_PROVIDER_TYPE=duckdb works at runtime.

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point — registers tools/resources and inits the ChEMBL service + optional canvas.
src/configServer-specific environment variable parsing and validation with Zod.
src/mcp-server/tools/definitionsTool definitions (*.tool.ts). Five ChEMBL tools plus the three chembl_dataframe_* canvas tools.
src/mcp-server/resources/definitionsResource definitions (*.resource.ts). Molecule and target record mirrors.
src/services/chemblThe single ChEMBL upstream client — URL builder, pagination, numeric coercion, nested-structure flattening, activity page stream.
src/services/canvas-accessor.tsModule-level holder for the optional DataCanvas wired in createApp({ setup }).
tests/Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
  • Register new tools and resources in the createApp() arrays
  • Wrap the ChEMBL API: validate raw → normalize to the flat domain type → return the output schema; never fabricate missing fields (absent numerics become null, never 0)

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
npx -y @cyanheads/chembl-mcp-server

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-cyanheads-chembl-mcp-server": {
      "command": "npx",
      "args": [
        "-y",
        "@cyanheads/chembl-mcp-server"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

@cyanheads/chembl-mcp-servernpm

Compatible MCP Clients

io.github.cyanheads/chembl-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More