Back to Directory/Developer Tools

io.github.cyanheads/internet-archive-mcp-server

Search the Wayback Machine and IA library (40M+ items), fetch snapshots, item metadata, and text.

Developer ToolsTypeScriptv0.3.0

@cyanheads/internet-archive-mcp-server

Search the Wayback Machine and IA library (40M+ items), fetch archived snapshots, retrieve item metadata and full text via MCP. STDIO or Streamable HTTP.

5 Tools • 1 Resource

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework


Overview

The Wayback Machine and Internet Archive library (40M+ items). Find and fetch archived snapshots of any URL, search the library by keyword and metadata, and retrieve item metadata, file manifests, and OCR text from any MCP client. Runs as a stdio process or a local Streamable HTTP server.

Tools

ToolDescription
ia_find_snapshotsFind Wayback Machine snapshots of a URL, by closest timestamp or full capture history
ia_get_snapshotFetch archived page content at a specific Wayback timestamp
ia_search_itemsSearch the IA library (40M+ items) by keyword and metadata filters
ia_get_itemRetrieve full metadata and file manifest for an Archive item
ia_get_textRetrieve readable OCR text from a text item, with paging

Resources

ResourceDescription
ia://item/{identifier}Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count

All resource data is also reachable via ia_get_item.

Capability reference

ia_find_snapshots tool

  • closest mode: single lookup via the Availability API, returns the nearest capture to a given timestamp
  • history mode: full capture list via the CDX API; filter by date range (from/to), HTTP status (status_filter), and MIME type
  • Default collapse of timestamp:8 (one capture per day); adjustable to timestamp:N, N=1–14
  • Up to 10,000 records per call (limit, default 100); resume_key pagination for large histories
  • Typed errors: no_snapshots (no matches), no_snapshot_available (closest mode, no capture near timestamp), cdx_unavailable

ia_get_snapshot tool

  • Resolves to the nearest available capture when the exact timestamp has no snapshot; exact 14-digit timestamps skip resolution and assume status 200
  • Strips scripts, styles, and nav from the archived HTML, returning readable plain text alongside the canonical replay URL
  • Output capped at IA_MAX_SNAPSHOT_CHARS (default 50,000 characters)
  • Typed errors: no_snapshot_available, content_fetch_failed

ia_search_items tool

  • Solr query syntax plus structured filters: mediatype, collection, creator, language, and date range (date_from/date_to)
  • Sort by relevance, date, or downloads (sort, Solr syntax; default downloads desc)
  • Up to 200 results per page (rows, default 50), 1-indexed page
  • Output carries total_found, page, rows for pagination; empty results return a notice with guidance rather than an error

ia_get_item tool

  • Returns title, creator, description, subject, collection, licenseurl, rights, and language when present in upstream metadata
  • files[] includes every manifest file — format, size, md5, and a direct download_url
  • Typed error item_not_found for unknown identifiers

ia_get_text tool

  • max_chars (defaults to IA_MAX_SNAPSHOT_CHARS) and char_offset page through long documents; has_more signals additional text remains
  • Locates the best available text file — DjVuTXT preferred, falls back to plain text; source_file names the file fetched
  • Typed errors: item_not_found, no_text_file, download_forbidden (restricted collections)

ia://item/{identifier} resource

  • Returns application/json — title, creator, mediatype, description, subject, collection, date, licenseurl, rights, language, and file_count
  • identifier comes from ia_search_items results
  • Typed error item_not_found for unknown identifiers

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Internet Archive-specific:

  • No credentials required — all four APIs are public
  • Three service layers: WaybackService (Availability + CDX), ArchiveSearchService (Solr), ArchiveMetadataService (Metadata + downloads)
  • CDX collapse-by-day default and configurable limit keep responses tractable for high-capture URLs
  • Identifies via a custom User-Agent on every request as required by IA's terms of use; configurable via IA_USER_AGENT

Agent-friendly output:

  • Pagination context on every list response — total_found, page, rows (search) and resume_key (CDX history) so agents never have to guess whether results are complete
  • Typed error reasons (no_snapshots, no_snapshot_available, item_not_found, no_text_file, download_forbidden) with recovery hints so callers can retry or explain to users without parsing text
  • Structured file manifests — every ia_get_item response includes file-level metadata (format, size, URL) enabling agents to select the right file without a follow-up call

Getting started

No API key required — the Internet Archive's APIs are fully public.

Add the following to your MCP client configuration file:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/internet-archive-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • No external accounts or API keys required.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/internet-archive-mcp-server.git
  1. Navigate into the directory:
cd internet-archive-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment:
cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts.

VariableDescriptionDefault
MCP_TRANSPORT_TYPETransport: stdio or httpstdio
MCP_HTTP_PORTHTTP server port3010
MCP_AUTH_MODEAuth mode: none, jwt, or oauthnone
MCP_LOG_LEVELLog level (debug, info, notice, warning, error)info
LOGS_DIRDirectory for log files (Node.js only)<project-root>/logs
STORAGE_PROVIDER_TYPEStorage backendin-memory
OTEL_ENABLEDEnable OpenTelemetry instrumentationfalse
IA_USER_AGENTCustom User-Agent for IA API requestsinternet-archive-mcp-server/{version} (github.com/cyanheads/internet-archive-mcp-server)
IA_REQUEST_TIMEOUT_MSHTTP request timeout in milliseconds30000
IA_MAX_SNAPSHOT_CHARSDefault character cap for ia_get_text responses50000

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec
    

Docker

docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point — registers tools, resource, and inits services.
src/configServer-specific environment variable parsing and validation with Zod.
src/mcp-server/toolsTool definitions (*.tool.ts). Five tools across Wayback and IA library.
src/mcp-server/resourcesResource definitions. ia://item/{identifier} item metadata resource.
src/services/waybackWaybackService — Availability API + CDX API client.
src/services/archive-searchArchiveSearchService — Solr Advanced Search client.
src/services/archive-metadataArchiveMetadataService — Metadata API + file download client.
tests/Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
  • Register new tools and resources via the barrels in src/mcp-server/*/definitions/index.ts
  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
npx -y @cyanheads/internet-archive-mcp-server

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-cyanheads-internet-archive-mcp-server": {
      "command": "npx",
      "args": [
        "-y",
        "@cyanheads/internet-archive-mcp-server"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

@cyanheads/internet-archive-mcp-servernpm

Compatible MCP Clients

io.github.cyanheads/internet-archive-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More