Search and retrieve bioRxiv and medRxiv preprints — by DOI, date interval, or keyword — via MCP.
Search and retrieve bioRxiv and medRxiv preprints — by DOI, date interval, or keyword — via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://biorxiv.caseyjhand.com/mcp
bioRxiv and medRxiv preprint metadata and full text, searchable via EuropePMC. Fetch preprints by DOI, browse by date interval or subject category, search by keyword and author, resolve journal-publication crosswalks, and extract full text from the rendered article page, from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|---|
biorxiv_get_preprint | Fetch full metadata, abstract, revision history, and journal crosswalk for one or more preprints by DOI |
biorxiv_list_recent | List preprints posted or updated within a date interval, with optional server, category, and funder filters |
biorxiv_search_preprints | Search preprints by keyword and/or author via EuropePMC for relevance ranking, enriched with bioRxiv/medRxiv metadata |
biorxiv_get_published_version | Resolve a preprint DOI to its journal publication record (journal DOI, name, published date) |
biorxiv_get_fulltext | Retrieve a preprint's full text as best-effort Markdown extracted from its rendered HTML article page |
biorxiv_list_categories | List valid subject category strings for bioRxiv and medRxiv |
biorxiv_get_preprint toolhttps://doi.org/…, doi:…, a biorxiv.org/medrxiv.org article URL, a vN or article-page suffix (.full, .full.pdf, .article-metrics, …) — and reports the bare DOIrevisions[] — one API call per DOI, no enumeration loopawards), JATS XML full-text link (jatsxmlUrl), and published journal DOI (publishedJournalDoi) once acceptedawards holds each award value as upstream records it (one value can run several grants together), deduplicated; funder names are left out because api.biorxiv.org attributes them to unrelated organizationsbiorxiv, medrxiv, or both (default both); when both, each DOI fans out in parallel and partial failures report per-DOI in failed[]failed[] entry carries a reason (not_found, invalid_doi_format, upstream_unavailable, rate_limited) and a retryable flag — a DOI is only reported not_found when every attempted server answeredreason: "rate_limited" rather than folding into upstream_unavailable, carrying retryAfter — the wait in seconds api.biorxiv.org asked forbiorxiv_list_recent toolbiorxiv_list_categories, in any case, with _, -, or a space (Cell Biology, cell_biology, cell-biology)invalid_category when no server applied itfunder filter by ROR ID, bare (021nxhr62) or as https://ror.org/021nxhr62, checked against the ROR pattern and checksum before any request; combines with categoryserver="both" queries bioRxiv alone with a notice, and server="medrxiv" raises invalid_funderapi.biorxiv.org has no funder record for raises invalid_funder, never an empty pagecursor (0, 30, 60, …)include_abstract: true adds them for the whole page, and biorxiv_get_preprint returns them for up to 10 DOIs per calltotal count per server; a cursor past the last page is marked exhausted: true rather than reading as zero results in the intervalserver="both" (default), per-server pagination state is independent ({ biorxiv: { cursor, total }, medrxiv: { cursor, total } }); one server not answering is named in failed[] while the other's page still returnsupstream_unavailable (or rate_limited) error instead of an empty pagebiorxiv_search_preprints toolAUTH:"…" field query, ANDed with the keyword query); optional date_from/date_to range and server scope (default both)cursor_mark pages through the same ranked list; a page past the last match comes back empty with a notice saying so, and a token EuropePMC does not recognize raises invalid_cursor_marksearch_unavailablebiorxiv_get_preprint, including type, license, awards, and authorCorrespondingInstitutionpartial_results and a per-record enrichment_error (service_error, rate_limited, or not_found); those records' abstracts come from one EuropePMC lookup keyed by their DOIs, and a failed lookup leaves them without one, with a noticeinclude_abstract: false drops them from every result, enriched and fallback alike, for a response about a third the size, keeping every other fieldrate_limited error carrying the origin's Retry-After wait — the search itself has no metadata to fall back on, unlike enrichmentbiorxiv_get_published_version tool/pubs/{server}/{doi} endpoint for richer metadata than the publishedJournalDoi field on biorxiv_get_preprintserver field names which server answered (never "both")biorxiv, medrxiv, or both (default both) — the two servers share their DOI prefixes, so a DOI alone doesn't identify one10.64898/ DOIs, which /pubs cannot look up by preprint DOI, resolve through the preprint's own journal DOI; when the crosswalk has no record either way, that journal DOI returns alone, without journal name or date, with a notice saying soupstream_unavailable, or rate_limited with the origin's wait on an HTTP 429 — never doi_not_found, which would assert an absence nothing establishedbiorxiv_get_fulltext toolwww.{server}.org/content/{doi}v{N}.full) and extracts Markdown — there is no keyless JATS sourceversion or a vN suffix on the DOI (the two must agree; a version the preprint lacks raises version_not_found), confirmed via the details API first; only DOI resolution fans out across biorxiv/medrxiv/both (default both) — the full-text fetch itself targets whichever server answered, named in the output server fieldoffset/limit character chunking (default limit 20,000, max 50,000); response reports totalChars, remainingChars, hasMore, and a full-article wordCount counted from the same Markdown, and the extracted article is cached per version so paging costs one origin fetchfulltext_unavailable error routing to biorxiv_get_preprintapi.biorxiv.org during resolution — returns a retryable rate_limited error carrying the origin's Retry-After wait, with the recovery hint naming which origin is limitingbiorxiv_list_categories toolbiorxiv_list_recentBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
bioRxiv-specific:
BiorxivApiService wraps api.biorxiv.org — details, publications, and crosswalk endpoints with retry and exponential backoff; a 429 is classified as a retryable rate_limited error carrying the parsed Retry-After wait, with the upstream response body kept out of the payloadEuropePmcService wraps the EuropePMC search endpoint for relevance-ranked keyword and/or author results, classified the same way on a 429BiorxivFullTextService fetches and extracts Markdown from the rendered HTML article pages on www.biorxiv.org / www.medrxiv.org — a distinct origin from the JSON APIPromise.allSettled — both biorxiv and medrxiv queried in parallel when server="both", results merged and deduplicated by DOIUser-Agent header including a mailto address (BIORXIV_MAILTO env var) per Cold Spring Harbor Lab API guidelinesAgent-friendly output:
failed[] with a typed reason and retryable flag instead of aborting the whole batch or listing callreason: "rate_limited" carrying the origin's parsed retryAfter wait, distinguished from a generic upstream_unavailablebiorxiv_search_preprints results carry enriched plus a typed enrichment_error (service_error / rate_limited / not_found) so callers branch on data, not string parsingexhausted cursors are flagged as an out-of-range artifact rather than an empty interval, and biorxiv_get_fulltext reports totalChars / remainingChars / hasMore for chunked reads{beta}, {+/-}, [≥]) become their characters, structured-abstract headings read Results: …, list items read • …, and figure and table blocks are dropped. The Markdown in content[] escapes upstream text so it renders as writtenA public instance is available at https://biorxiv.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"biorxiv-mcp-server": {
"type": "streamable-http",
"url": "https://biorxiv.caseyjhand.com/mcp"
}
}
}
Add the following to your MCP client configuration file.
{
"mcpServers": {
"biorxiv-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/biorxiv-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"BIORXIV_MAILTO": "your@email.com"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"biorxiv-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/biorxiv-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"BIORXIV_MAILTO": "your@email.com"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"biorxiv-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "-e", "BIORXIV_MAILTO=your@email.com", "ghcr.io/cyanheads/biorxiv-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 BIORXIV_MAILTO=your@email.com bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/biorxiv-mcp-server.git
cd biorxiv-mcp-server
bun install
cp .env.example .env
# optionally set BIORXIV_MAILTO for polite API access
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
| Variable | Description | Default |
|---|---|---|
BIORXIV_MAILTO | Email address included in the User-Agent header for polite API access per Cold Spring Harbor Lab guidelines. Optional, but recommended. | — |
BIORXIV_API_BASE_URL | Override the bioRxiv API base URL. | https://api.biorxiv.org |
EUROPEPMC_API_BASE_URL | Override the EuropePMC base URL. | https://www.ebi.ac.uk/europepmc/webservices/rest |
BIORXIV_WEB_BASE_URL | Override the bioRxiv website base URL (full-text HTML source for biorxiv_get_fulltext). | https://www.biorxiv.org |
MEDRXIV_WEB_BASE_URL | Override the medRxiv website base URL (full-text HTML source for biorxiv_get_fulltext). | https://www.medrxiv.org |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | HTTP server port. | 3010 |
MCP_HTTP_ENDPOINT_PATH | HTTP endpoint path. | /mcp |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.). | info |
LOGS_DIR | Directory for log files (Node.js only). | <project-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Build and run:
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
Run checks and tests:
bun run devcheck # Lint, format, typecheck, security
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
docker build -t biorxiv-mcp-server .
docker run --rm -e BIORXIV_MAILTO=your@email.com -p 3010:3010 biorxiv-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/biorxiv-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools and initializes services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). Six tools across bioRxiv and medRxiv. |
src/services/biorxiv | BiorxivApiService — details, publications, and crosswalk endpoint wrappers with retry. |
src/services/biorxiv-fulltext | BiorxivFullTextService — rendered HTML article page fetch and Markdown extraction. |
src/services/europe-pmc | EuropePmcService — preprint keyword/author search endpoint wrapper. |
tests/ | Unit and integration tests mirroring the src/ structure. |
See CLAUDE.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped logging, ctx.state for tenant-scoped storagesrc/mcp-server/tools/definitions/index.tsIssues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @cyanheads/biorxiv-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-biorxiv-mcp-server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/biorxiv-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/biorxiv-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.