Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths.
Search CERN Open Data, fetch records, files, analysis environments, CMS good-run lists, HLT paths via MCP. STDIO or Streamable HTTP.
Particle-physics data from the CERN Open Data Portal: collision, simulated and derived datasets, analysis software, environments and documentation from ALICE, ATLAS, CMS, LHCb and other experiments. Search it with exact-vocabulary filters and live facet counts, open records with their license and citation, list the files that hold the data, assemble a record's analysis environment, and look up CMS good-run lists and trigger paths. Runs as a stdio process or a local Streamable HTTP server.
| Tool | Description |
|---|---|
cern_opendata_search_records | Search datasets, software, environments, documentation and supplementary records with exact-vocabulary filters and live facet counts |
cern_opendata_get_records | Fetch full metadata for 1–20 records by recid, DOI, CMS dataset path or documentation slug, with license and citation |
cern_opendata_list_files | Page through a record's file indexes and files: XRootD URIs, HTTPS URLs, sizes, checksums, tape availability |
cern_opendata_get_analysis_env | Assemble a record's analysis environment: container images, CMSSW release, global tag, linked environment and software records, guide sections |
cern_opendata_get_validated_runs | Get a CMS validated-run (good-run) list for a dataset, a list or a run period, with luminosity-section ranges |
cern_opendata_search_trigger_paths | Look up CMS High-Level Trigger paths by name or prefix, parsed into run ranges, versions and L1 seeds |
cern_opendata_list_reference | Decode the vocabulary the other tools accept: experiments, record types, energies, formats, identifiers, query syntax, licensing, run periods |
| Resource | Description |
|---|---|
cern-opendata://record/{recid} | One record's metadata, license and citation, in the cern_opendata_get_records record shape |
Tool-only clients get the same data from cern_opendata_get_records.
cern_opendata_search_records toolquery (an OpenSearch query_string, up to 500 characters) plus OR-list filters type, experiment, collision_energy, collision_type, file_type, availability and collection, each an array or a comma-separated string; year_from/year_to and min_events/max_events bound the data-taking year and the event countsort (bestmatch, mostrecent, title, title_desc), limit 1–50 (default 10) and page from 1; paging reaches the first 10,000 matches, and page × limit past that fails as page_window_exceededhits with recids, plus eight live facets that each ignore their own filter; applied_filters echoes what ran, with values outside the verified vocabulary listed under unrecognized_valuescern_opendata_get_records toolids: 1–20 recids, DOIs, CMS dataset paths (/Primary/Era/TIER) or documentation slugs, mixed in one array or comma-separated stringlicense with its basis (record, cern_terms_default, not_stated) and, when it has a DOI, a ready citation; documentation and news bodies are cut at 30,000 charactersmissing with interpreted_as and guidance instead of failing the call; file lists come from cern_opendata_list_filescern_opendata_list_files toolrecid required; without index, returns the record's file indexes and regular files, and with an index key, that index's fileslimit 1–500 (default 50), continued with next_cursor; each file carries xrootd_uri, https_url, size_in_bytes, checksum and availability, and each index a uri_list_url listing every XRootD URI in iton demand sit on tape and must be requested on the record's portal page first; an umbrella record with no files of its own returns its children recidscern_opendata_get_analysis_env toolrecid required; software carries the record's own container images, CMSSW release, global tag and environment recidenvironment_records (condition, VM, validation) for the record's run periods and example_software that declares it works with the record, up to 50 between them; guides quotes the linked section of the first two portal guides, each capped at 12,000 charactersseparately_licensed: true; linked records or guides that can't be read leave a notice instead of failing the callcern_opendata_get_validated_runs toolrecid (a CMS collision dataset or a validated-run list) or run_period (Run2012B; 2012B also matches); variant full or muons_only; run_min/run_max; limit 1–2000 (default 200)recid bounds the runs to the first and last run the dataset lists, echoed in run_bounds; when several lists match, matched_lists names them and no runs are readlumi_sections and lumi_ranges, and list.https_url downloads the whole list file; CMS only, so other records fail as no_validated_runscern_opendata_search_trigger_paths toolpath: an exact name (HLT_IsoMu24) or a prefix with one trailing * (HLT_IsoMu*); HLT_ is added when missing and a _v<n> version suffix dropped; optional year, limit 1–50 (default 10) and pagefirst_seen, last_seen, per-version run ranges with their l1_seed, and HLT menu record links; parsed: false marks a record to read from its abstract_htmlcern_opendata_list_reference tooltopic: experiments, record_types, collision_energies, collision_types, file_types, availability, identifiers, query_syntax, licensing or run_periods; omit it for every tablerun_periods is a dated snapshot, while cern_opendata_get_validated_runs reads the live list collectioncern-opendata://record/{recid} resourcerecid (leading zeros ignored) as application/json, in the cern_opendata_get_records record shape: metadata, license and citation, without file listsrecid comes from cern_opendata_search_records; reads carry a 15-minute public cache hintBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CERN Open Data-specific:
13 tev → 13TeV, lhcb → LHCb, Pb-Pb → PbPb, dataset/collision → Dataset::Collision); unknown values are sent as given and flaggedondemand) records included in every search and lookup, where the portal otherwise drops them silently; every hit and file states its availabilityAgent-friendly output:
portal_url on each hit and record, a license with its basis, a DOI citation, and applied_filters or effectiveQuery echoing what rancern_opendata_get_records returns unresolved ids under missing with guidance, and cern_opendata_get_analysis_env reports unreadable linked records or guides in a notice rather than failingkind, license.basis, interpreted_as, scope, variant, run_bounds.source and parsed let callers branch on data, not string parsingcontent[] and relayed as received (HTML in _html fields) in structuredContentPortal metadata and datasets are CC0 under the CERN Open Data Terms of Use. Software, container images, documentation and guide code are licensed separately, per record (software is commonly GPL). cern_opendata_get_records reports each record's license and its basis: record when the record states one, cern_terms_default for a dataset that states none (CC0 under the Terms of Use), and not_stated otherwise. cern_opendata_get_analysis_env marks container images, software and guide code as separately licensed.
CERN asks reusers to cite each dataset's DOI in applications and publications. cern_opendata_get_records returns a ready citation for every record with a DOI.
This server is an independent project and is not affiliated with or endorsed by CERN.
rate_limited with retryAfter. A hosted deployment shares that one budget across every user behind its egress IP. cern_opendata_get_analysis_env and cern_opendata_get_validated_runs cost 2–4 requests each.file_type up to 100), with the rest counted in other_count. A filter does not narrow its own facet, only the hits and the other facets.on demand must be requested on the record's portal page before download; staging them is a write and out of scope. A record whose availability is ondemand lists none of its files through the API, so cern_opendata_list_files reports only the count and size its metadata states.cern_opendata_list_files returns under children.Add the following to your MCP client configuration file.
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cern-opendata-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"cern-opendata-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cern-opendata-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/cern-opendata-mcp-server.git
cd cern-opendata-mcp-server
bun install
cp .env.example .env
# optional: adjust transport, logging, or telemetry settings
The server has no settings of its own; these framework variables apply.
| Variable | Description | Default |
|---|---|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | HTTP server port. | 3010 |
MCP_HTTP_HOST | HTTP server host. | 127.0.0.1 |
MCP_SESSION_MODE | HTTP session mode: stateless, stateful, or auto. | stateless |
MCP_AUTH_MODE | Authentication: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.). | info |
LOGS_DIR | Directory for log files (Node.js only). | <app-root>/logs |
OTEL_ENABLED | Enable OpenTelemetry. | false |
See .env.example for the common framework overrides.
Build and run the production version:
# One-time build
bun run rebuild
# Run the built server
bun run start:http
# or
bun run start:stdio
Run checks and tests:
bun run devcheck # Lints, formats, type-checks, and more
bun run test # Runs the test suite
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point: registers the tools and resource, sets the server instructions, starts the portal client. |
src/mcp-server/tools | Tool definitions (*.tool.ts), plus the shared input helpers and list enrichment. |
src/mcp-server/resources | Resource definitions. The cern-opendata://record/{recid} resource. |
src/mcp-server/record-schema.ts | The record output schema shared by cern_opendata_get_records and the resource. |
src/services/cern-opendata | Portal client (pacing, retries, per-call deadline, byte ceilings, caches), normalization, vocabulary tables, text rendering, trigger parsing. |
tests/ | Unit and tool tests, mirroring the src/ structure, run against fixture portal responses. |
docs/design.md | Tool surface, verified portal behavior, and design decisions. |
See CLAUDE.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for logging and ctx.enrich for notices and paging contextsrc/mcp-server/*/definitions/index.tsIssues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
This project is licensed under the Apache 2.0 License. See the LICENSE file for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @cyanheads/cern-opendata-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-cern-opendata-mcp-server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/cern-opendata-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/cern-opendata-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.