Search the Wayback Machine and IA library (40M+ items), fetch snapshots, item metadata, and text.
Search the Wayback Machine and IA library (40M+ items), fetch archived snapshots, retrieve item metadata and full text via MCP. STDIO or Streamable HTTP.
The Wayback Machine and Internet Archive library (40M+ items). Find and fetch archived snapshots of any URL, search the library by keyword and metadata, and retrieve item metadata, file manifests, and OCR text from any MCP client. Runs as a stdio process or a local Streamable HTTP server.
| Tool | Description |
|---|---|
ia_find_snapshots | Find Wayback Machine snapshots of a URL, by closest timestamp or full capture history |
ia_get_snapshot | Fetch archived page content at a specific Wayback timestamp |
ia_search_items | Search the IA library (40M+ items) by keyword and metadata filters |
ia_get_item | Retrieve full metadata and file manifest for an Archive item |
ia_get_text | Retrieve readable OCR text from a text item, with paging |
| Resource | Description |
|---|---|
ia://item/{identifier} | Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count |
All resource data is also reachable via ia_get_item.
ia_find_snapshots toolclosest mode: single lookup via the Availability API, returns the nearest capture to a given timestamphistory mode: full capture list via the CDX API; filter by date range (from/to), HTTP status (status_filter), and MIME typecollapse of timestamp:8 (one capture per day); adjustable to timestamp:N, N=1–14limit, default 100); resume_key pagination for large historiesno_snapshots (no matches), no_snapshot_available (closest mode, no capture near timestamp), cdx_unavailableia_get_snapshot tool200IA_MAX_SNAPSHOT_CHARS (default 50,000 characters)no_snapshot_available, content_fetch_failedia_search_items toolmediatype, collection, creator, language, and date range (date_from/date_to)sort, Solr syntax; default downloads desc)rows, default 50), 1-indexed pagetotal_found, page, rows for pagination; empty results return a notice with guidance rather than an erroria_get_item tooltitle, creator, description, subject, collection, licenseurl, rights, and language when present in upstream metadatafiles[] includes every manifest file — format, size, md5, and a direct download_urlitem_not_found for unknown identifiersia_get_text toolmax_chars (defaults to IA_MAX_SNAPSHOT_CHARS) and char_offset page through long documents; has_more signals additional text remainssource_file names the file fetcheditem_not_found, no_text_file, download_forbidden (restricted collections)ia://item/{identifier} resourceapplication/json — title, creator, mediatype, description, subject, collection, date, licenseurl, rights, language, and file_countidentifier comes from ia_search_items resultsitem_not_found for unknown identifiersBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Internet Archive-specific:
WaybackService (Availability + CDX), ArchiveSearchService (Solr), ArchiveMetadataService (Metadata + downloads)limit keep responses tractable for high-capture URLsIA_USER_AGENTAgent-friendly output:
total_found, page, rows (search) and resume_key (CDX history) so agents never have to guess whether results are completeno_snapshots, no_snapshot_available, item_not_found, no_text_file, download_forbidden) with recovery hints so callers can retry or explain to users without parsing textia_get_item response includes file-level metadata (format, size, URL) enabling agents to select the right file without a follow-up callNo API key required — the Internet Archive's APIs are fully public.
Add the following to your MCP client configuration file:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/internet-archive-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp
git clone https://github.com/cyanheads/internet-archive-mcp-server.git
cd internet-archive-mcp-server
bun install
cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
| Variable | Description | Default |
|---|---|---|
MCP_TRANSPORT_TYPE | Transport: stdio or http | stdio |
MCP_HTTP_PORT | HTTP server port | 3010 |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth | none |
MCP_LOG_LEVEL | Log level (debug, info, notice, warning, error) | info |
LOGS_DIR | Directory for log files (Node.js only) | <project-root>/logs |
STORAGE_PROVIDER_TYPE | Storage backend | in-memory |
OTEL_ENABLED | Enable OpenTelemetry instrumentation | false |
IA_USER_AGENT | Custom User-Agent for IA API requests | internet-archive-mcp-server/{version} (github.com/cyanheads/internet-archive-mcp-server) |
IA_REQUEST_TIMEOUT_MS | HTTP request timeout in milliseconds | 30000 |
IA_MAX_SNAPSHOT_CHARS | Default character cap for ia_get_text responses | 50000 |
See .env.example for the full list of optional overrides.
Build and run:
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
Run checks and tests:
bun run devcheck # Lint, format, typecheck, security
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-server
The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools, resource, and inits services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). Five tools across Wayback and IA library. |
src/mcp-server/resources | Resource definitions. ia://item/{identifier} item metadata resource. |
src/services/wayback | WaybackService — Availability API + CDX API client. |
src/services/archive-search | ArchiveSearchService — Solr Advanced Search client. |
src/services/archive-metadata | ArchiveMetadataService — Metadata API + file download client. |
tests/ | Unit and integration tests mirroring src/. |
See CLAUDE.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped logging, ctx.state for tenant-scoped storagesrc/mcp-server/*/definitions/index.tsIssues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @cyanheads/internet-archive-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-internet-archive-mcp-server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/internet-archive-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/internet-archive-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.