Search and query CDC public health data — mortality, vaccinations, surveillance, behavioral risk.
Search and query CDC public health data — mortality, vaccinations, surveillance, behavioral risk (Socrata SODA API) via MCP. STDIO or Streamable HTTP.
Public Hosted Server: https://cdc.caseyjhand.com/mcp
CDC public health data — the Socrata-based CDC Open Data portal, plus CDC WONDER, a separate CDC system for national mortality statistics. Search the catalog, inspect dataset schemas, and run SoQL queries across vaccination, surveillance, and behavioral-risk data, or query WONDER for deaths, population, and death rates by year, age, sex, and race. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
| Tool | Description |
|---|---|
cdc_discover_datasets | Search the CDC dataset catalog by keyword, category, or tag |
cdc_get_dataset_schema | Fetch column schema, row count, and metadata for a dataset |
cdc_list_catalog_vocabulary | List the catalog's category and tag values with their entry counts |
cdc_query_dataset | Execute SoQL queries — filter, aggregate, sort, full-text search, select fields |
cdc_query_wonder | Query CDC WONDER for national mortality, population, and death rates across five databases |
| Resource | Description |
|---|---|
cdc://datasets | Top 50 most-viewed catalog entries, for orientation |
cdc://datasets/{datasetId} | Dataset metadata and the first 100 columns for a specific dataset |
Both resources mirror data also reachable via cdc_discover_datasets and cdc_get_dataset_schema, for clients that surface resources but not tools.
| Prompt | Description |
|---|---|
analyze_health_trend | Guided workflow for investigating a public health question across CDC data |
cdc_discover_datasets tooldomain selects data.cdc.gov (default) or chronicdata.cdc.gov — both front the same catalog, so switching hosts neither widens nor narrows a searchquery, category, and tags filters — tags union (a dataset matches on any one tag, so each tag added widens the result set), while query and category intersect with the tag setoffset is capped at 9999, and offset + limit must not exceed 10,000 — Socrata's catalog ceilingorder: dataset_id (default) sorts deterministically for stable pagination; relevance ranks by best match but is not stably paginable across pagescolumnCount — a value of 0 marks a non-tabular asset (chart, map, story, file, or href) that yields no data from the other tools; assetType is descriptive onlydescription is converted to plain text before it is cut to 300 characters, so the budget buys visible text rather than the HTML tags catalog entries arrive wrapped intotalCount and appliedFilters; a notice distinguishes an offset past the end of the result set from a search that matched nothing, and resolves a category/tags value that matched nothing against the catalog vocabulary — category: "Vaccination" comes back naming "Vaccinations" (89 datasets) rather than advising a broader search. The vocabulary is read only on that branch, so an ordinary search costs no extra requestcdc_get_dataset_schema tooldatasetId (e.g. bi63-dtpu) and the same domain enum as the other Socrata toolscolumn_limit, max 500) — catalog schemas run 3 to 322 columns, so ordinary datasets arrive whole; wider ones report totalCount, truncated, and nextOffset to pass back as column_offsetcolumn_offset at or past the column count returns an empty window rather than an errornot_queryable when the ID names a non-tabular catalog asset, rather than returning an empty column listrowCount prefers a live count(*) fetched alongside the metadata; rowCountSource says live or cached, since Socrata's cached figure is built once and can understate an actively-updated dataset by a third or more. A failed count falls back to the cached figure and never fails the schema responsedescription is returned in full as plain text — markup stripped, entity references decoded; only cdc_discover_datasets truncates itcdc_list_catalog_vocabulary toolcdc_discover_datasets' category and tags filters are matched against — a value the catalog does not carry matches nothing, which is indistinguishable from a real value with no resultstag_limit (default 50, max 500) and tag_offsetresultSetSize: 100 beside them, so the under-count reads as completefilter narrows both vocabularies before the page is cut, matching on whole words in either direction ("vaccin" reaches Vaccinations and covid-19 vaccination). Not fuzzy — a misspelling returns nothing rather than a guessvocabularySize, the matched categoryCount/tagCount, and truncated/shown/cap/nextOffset; truncationCeiling bounds every omitted tag, since the list is ranked by the same countdata.cdc.gov and chronicdata.cdc.gov return the identical 55 categories and 1,583 tagscdc_query_dataset toolselect, where, group, having, order, plus full-text search across text columnsoffset capped at 1,000,000truncated is measured by an over-fetch probe (one row past the limit), never guessed from the row count; the whole response — structuredContent and content[] together — is bounded by a 200,000-character budget, so a wide page can end short of limit with a nextOffsetoffset > 0 is diagnosed with one probe at offset 0: the notice says whether the offset ran past the end or the query matches nothing, and names both causes if the probe failseffectiveQuery echoes the SoQL clauses sent in their original text, not URL-encoded, so a clause can be copied back into the parameter it came fromnot_queryable when every returned row carries no fields and the asset reports no columns — a chart or map ID, which Socrata answers 200 with a body of empty objects. When the asset does have columns, the same shape is a null-only projection and comes back as a success with a noticedataType from the schemacdc_query_wonder tooldatabase selects which of five mortality databases answers the query:
| Value | CDC database | Years | Race groups | mcd_icd10 |
|---|---|---|---|---|
underlying_1999_2020 (default) | D76 — Underlying Cause of Death | 1999–2020 | 4 bridged | — |
provisional | D176 — Provisional Mortality Statistics | 2018 → current year | 6 single-race | yes |
underlying_2018_2024 | D158 — Underlying Cause of Death, Single Race | 2018–2024 | 6 single-race | — |
multiple_1999_2020 | D77 — Multiple Cause of Death | 1999–2020 | 4 bridged | yes |
multiple_2018_2024 | D157 — Multiple Cause of Death, Single Race | 2018–2024 | 6 single-race | yes |
group_by: 1–4 of year, age_group, sex, race, each at most once; national totals only — no sub-national breakdown at any settingcause_icd10 and mcd_icd10 take one ICD-10 code or range, or a list of up to 50 matched as one union — e.g. the drug-overdose set ["X40","X41","X42","X43","X44","X60","X61","X62","X63","X64","X85","Y10","Y11","Y12","Y13","Y14"] returns one series with one set of rates. A range must be a chapter or block of WONDER's ICD-10 tree (X40-X49); any other span (X40-X44) is rejected, so list its codes insteadmcd_icd10 matches a cause recorded anywhere on the death certificate rather than only the underlying cause; accepted only by provisional, multiple_1999_2020, and multiple_2018_2024 — the others reject itage_groups must include "NS" (age not recorded) to match an unfiltered total; a year_range outside the selected database's span is rejected with that span namedSuppressed, Unreliable, Not Applicable) read null in rows and are named per cell in cellNotes; whole rows CDC hides (zero or suppressed deaths) are absent from rows with no gap marker — check messagesstructuredContent and content[] together, so a large grouping comes back a page at a time with an exact totalCount, truncated, and a nextOffset to continue from; limit (max 5,000) takes smaller pages and offset (max 10,000) resumescdc://datasets resourceassetType and columnCount for orientationcolumnCount: 0 marks a non-tabular entry (chart, map, story, file, or href); use cdc_discover_datasets for full catalog search with filtering and paginationcdc://datasets/{datasetId} resourceapplication/json; datasetId is a four-by-four identifier from cdc_discover_datasetscolumnCount and a truncated flag; wider schemas continue via cdc_get_dataset_schema with column_offset{?column_limit,column_offset} template would stop the bare cdc://datasets/{datasetId} form from matching at allanalyze_health_trend prompttopic required; timeRange and geography optionalBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CDC-specific:
cdc_query_wonder) — a separate XML-over-HTTP CDC system, spanning five mortality databases from 1999 through the current yeardomain input (data.cdc.gov, chronicdata.cdc.gov), allowlisted at the schema level — both front one tenant, so assets like PLACES and the Heart Disease & Stroke Atlas are reachable from eitherAgent-friendly output:
totalCount/truncated/nextOffset (or shown/cap) rather than a bare row count, so an agent can tell a complete result from a page of onerecovery hint on every declared reason (e.g. not_queryable, page_out_of_range) — actionable next steps, not just an error codeSuppressed, Unreliable, Not Applicable) are named per cell in cellNotes, and hidden-row notices surface in messageseffectiveQuery echoes the exact query sent (SoQL clauses or a WONDER summary), so a result is reproducible and a clause can be copied back into its parameterA public instance is available at https://cdc.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "streamable-http",
"url": "https://cdc.caseyjhand.com/mcp"
}
}
}
Add the following to your MCP client configuration file.
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cdc-health-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cdc-health-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cdc-health-mcp-server:latest"]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/cdc-health-mcp-server.git
cd cdc-health-mcp-server
bun install
cp .env.example .env
# edit .env and set optional overrides
| Variable | Description | Default |
|---|---|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http | stdio |
MCP_HTTP_PORT | HTTP server port | 3010 |
MCP_SESSION_MODE | HTTP session posture: stateful, stateless, or auto. src/index.ts declares stateless; this variable overrides it | stateless |
MCP_AUTH_MODE | Authentication: none, jwt, or oauth | none |
MCP_LOG_LEVEL | Log level (debug, info, warning, error, etc.) | info |
LOGS_DIR | Directory for log files (Node.js only) | <project-root>/logs |
STORAGE_PROVIDER_TYPE | Storage backend: in-memory, filesystem, supabase, cloudflare-kv/r2/d1 | in-memory |
CDC_APP_TOKEN | Socrata app token for higher rate limits | — |
CDC_BASE_URL | SODA host for requests that name no domain — in practice only the cdc://datasets/{datasetId} resource. The three Socrata tools always send a domain (data.cdc.gov by default), which overrides this | https://data.cdc.gov |
CDC_CATALOG_URL | Base URL for Socrata Discovery API | https://api.us.socrata.com/api/catalog/v1 |
OTEL_ENABLED | Enable OpenTelemetry instrumentation (spans, metrics, completion logs) | false |
See .env.example for the full list of optional overrides.
Build and run the production version:
# One-time build
bun run rebuild
# Run the built server
bun run start:http
# or
bun run start:stdio
Run checks and tests:
bun run devcheck # Lints, formats, type-checks, and more
bun run test # Runs the test suite
docker build -t cdc-health-mcp-server .
docker run --rm -p 3010:3010 cdc-health-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/cdc-health-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools/resources/prompts and inits services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). Four CDC data tools. |
src/mcp-server/resources | Resource definitions. Catalog overview and dataset detail. |
src/mcp-server/prompts | Prompt definitions. Health trend analysis workflow. |
src/services/socrata | Socrata SODA API service layer — HTTP client, catalog search, metadata, queries. |
src/services/wonder | CDC WONDER service layer — XML request builder and response parser. |
src/utils | Shared helpers, including escapeTableCell for format() output. |
tests/ | Unit and integration tests mirroring src/. |
See CLAUDE.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for logging, ctx.state for storagecreateApp() arraysIssues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @cyanheads/cdc-health-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-cdc-health-mcp-server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/cdc-health-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/cdc-health-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.