Back to Directory/Developer Tools

io.github.cyanheads/evals-mcp-server

Author verifiable eval records through a draft→review→revise→submit loop with enforced graders.

Developer ToolsTypeScriptv0.1.5

@cyanheads/evals-mcp-server

Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.

9 Tools • 1 Resource

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework


Overview

Verifiable eval records, authored through a draft → review → surgical-revise → submit loop with server-enforced graders. Create a draft carrying its own executable grader, patch it surgically field by field, and submit through a committability gate that requires the gold to pass, a declared negative case to fail, and an independent verification to agree — then compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness. Runs as a stdio process or a local Streamable HTTP server.

Tools

ToolDescription
evals_describe_schemaReturn the required and optional fields plus grader options for a task type. Call before drafting.
evals_create_draftCreate a draft eval record carrying its own grader; returns the parsed record, a review protocol, and a verification subagent prompt.
evals_get_recordRead a draft or submitted record by id; the id is stable across submit.
evals_revise_draftApply a surgical set / append / unset patch to a draft by dotted path; re-runs the self-consistency check.
evals_discard_draftDelete a draft record by id. Draft-only.
evals_run_checkRun a grader spec against candidate answers and get PASS/REJECT per candidate, decoupled from any saved record.
evals_submit_draftFinalize a draft through the committability gate, then freeze it.
evals_list_recordsBrowse and filter records by status, domain, task type, or tag. Returns a compact summary per record.
evals_export_recordsCompile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness and write the artifact under exports/.

Resources

ResourceDescription
eval://record/{id}A single draft or submitted record by id — the same payload evals_get_record returns, for resource-capable clients.

All record data is also reachable through the tool surface — evals_get_record for a single record, evals_list_records to browse. The resource is a convenience mirror for clients that support resources, not the access path.

Capability reference

evals_describe_schema tool

  • Static — derived from the record and grader Zod schemas, no disk or runtime state
  • task_type is one of numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_response
  • Returns the gold shape, applicable grader kind(s), required/optional fields, and per-type authoring notes (e.g. mcq needs choices, free_response needs an llm_rubric grader)

evals_create_draft tool

  • Validates against the task_type discriminated union and persists the draft; mcq requires choices, free_response requires an llm_rubric grader
  • Runs a self-consistency check — the grader must PASS against gold and each discrimination.positive, and REJECT each discrimination.negative
  • Returns the normalized record, a per-field review protocol, a ready-to-paste verification-subagent prompt, and what's still required before submit
  • Accepts optional draft-time verification evidence and captures (EvalsIDs) when provenance is already in hand
  • Typed errors: grader_unexecutable, task_type_constraint, mcq_choice_mismatch
  • Stays draft — passing self-consistency proves the grader discriminates, not that the gold is correct

evals_get_record tool

  • Reads by id, stable across submit — resolves whether the record is still a draft or already submitted
  • Returns the full record, including its grader, discrimination cases, and verification evidence
  • not_found when no record matches; recovery points to evals_list_records

evals_revise_draft tool

  • Explicit set (dotted-path → value), append (dotted-path → array items), and unset (dotted paths) operations — never a full-record rewrite
  • Cannot target task_type or server-owned fields — start a new draft to change the discriminant
  • Re-validates the full record shape and per-task-type constraints after the patch, and re-runs self-consistency since the grader may have moved
  • Returns the updated record and an itemized changed list (op, path, before, after)
  • Draft-only — record_frozen on a submitted id
  • Typed errors: not_found, record_frozen, invalid_patch_path, task_type_constraint, mcq_choice_mismatch

evals_discard_draft tool

  • Deletes a draft record by draft_id
  • Draft-only — record_frozen when the id refers to a submitted record
  • A missing id reports not_found rather than a distinct "already discarded" error — effectively idempotent

evals_run_check tool

  • Runs a grader spec against one or more candidates (strings, numbers, objects, or arrays) without touching a saved record
  • Returns PASS/REJECT and a detail per candidate, plus the resolved comparison value (e.g. the math.js-evaluated numeric target)
  • gold applies only to gold-relative kinds (exact_match); it's a no-op for target-embedding kinds like numeric and mcq
  • llm_rubric cannot run here — submission relies on recorded independent verification instead
  • Typed errors: grader_unexecutable, mcq_choice_mismatch

evals_submit_draft tool

  • The committability gate: the gold must PASS its grader, ≥1 declared negative must be REJECTED, and a recorded, decorrelated independent verification must agree with the gold
  • Resolves and embeds any captures from EVALS_CAPTURE_DIR, cross-checking the gold against the authoritative captured value
  • Rejects duplicates by content_hash; confirm (or EVALS_REQUIRE_CONFIRMATION) can require human confirmation through multi-round input before finalizing
  • On pass, flips the record to submitted, stamps submitted_at and a checksum, and freezes it; otherwise refuses and the record stays a draft
  • free_response is admitted on recorded independent verification alone and flagged server_verified: false
  • Typed errors: not_found, record_frozen, verification_incomplete, grader_failed_on_gold, verification_disagrees_with_gold, missing_negative_case, negative_case_passed, duplicate, decorrelation_violation, capture_unresolved, submit_declined

evals_list_records tool

  • Filters by status (draft/submitted), domain, task_type, or tag; up to 500 per call (default 50)
  • Returns a compact summary per record (id, status, task_type, domain, tags, timestamps), newest-first — not full records
  • Discloses truncation (shown, cap, total count) when the limit is hit, so a partial set is never mistaken for the whole corpus

evals_export_records tool

  • Formats: jsonl (lossless), csv (flattened, lossy summary), inspect (UK AISI Inspect AI), lm-eval (EleutherAI lm-evaluation-harness)
  • Optional domain / task_type / tag filter
  • Only submitted records are exported — drafts are skipped
  • Writes the artifact under exports/ and returns its path, record count, byte size, and a short preview instead of dumping it inline

eval://record/{id} resource

  • Returns the same payload as evals_get_record, as application/json
  • id comes from evals_list_records or a draft/submit response
  • not_found when no record matches

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Eval authoring:

  • A draft → review → surgical-revise → submit loop, with the server acting as both scribe (normalize, persist, compile) and adversarial checker (runs the record's own grader, rejects what doesn't hold up)
  • Records are a Zod discriminatedUnion on task_type — numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_response
  • A typed grader DSL serialized with each record — deterministic kinds (numeric via math.js, exact_match, set_match, regex, mcq, json_match) run server-side; llm_rubric relies on recorded independent verification
  • An enforced committability gate at submit: the gold must pass its own grader, ≥1 negative must be rejected, and a recorded decorrelated verification must agree with the gold
  • Plain JSON files under EVALS_DATA_DIR — inspectable, diffable, version-controllable records, with drafts, submitted records, and exports kept separate

Agent-friendly output:

  • Instructional responses — evals_create_draft and evals_revise_draft return the parsed record parroted back, a per-field review protocol, and a ready-to-paste verification-subagent prompt
  • Self-consistency verdicts — every draft/revise response reports per-positive and per-negative pass/reject results, not just a boolean
  • Truncation disclosure — evals_list_records reports shown / cap / total count when the limit is hit, so a partial set is never mistaken for the whole corpus
  • Typed refusal — the submit gate fails with a typed reason plus a recovery hint, so a rejected record tells the agent exactly what to fix

Getting started

Add the following to your MCP client configuration file. Set EVALS_DATA_DIR to a writable folder — the server manages drafts/, submitted/, and exports/ under it.

{
  "mcpServers": {
    "evals-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/evals-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "EVALS_DATA_DIR": "/absolute/path/to/evals-data"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "evals-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/evals-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "EVALS_DATA_DIR": "/absolute/path/to/evals-data"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "evals-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "-e", "EVALS_DATA_DIR=/data",
        "-v", "evals-data:/data",
        "ghcr.io/cyanheads/evals-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 EVALS_DATA_DIR=./evals-data bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • A writable directory for EVALS_DATA_DIR. No external API key is required.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/evals-mcp-server.git
  1. Navigate into the directory:
cd evals-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment:
cp .env.example .env
# edit .env and set EVALS_DATA_DIR

Configuration

All server configuration is validated at startup via Zod schemas in src/config/server-config.ts.

VariableDescriptionDefault
EVALS_DATA_DIRRoot folder for record JSON; the store manages drafts/, submitted/, and exports/ under it../evals-data
EVALS_REQUIRE_CONFIRMATIONWhen true, evals_submit_draft requests human confirmation through multi-round input before finalizing.false
EVALS_DEFAULT_LICENSEDefault metadata.license applied when a draft omits one (e.g. CC-BY-4.0).—
EVALS_CAPTURE_DIRDirectory of framework-written tool-call captures; when set, captures EvalsIDs resolve to full dumps.—
MCP_TRANSPORT_TYPETransport: stdio or http.stdio
MCP_HTTP_PORTPort for the HTTP server.3010
MCP_SESSION_MODEHTTP session mode: auto, stateful, or stateless (auto resolves to stateful). A stateless HTTP start is refused, since a 2025-era client can answer the evals_submit_draft confirmation only over a live session. No effect on stdio.stateful
MCP_AUTH_MODEAuth mode: none, jwt, or oauth.none
MCP_LOG_LEVELLog level (RFC 5424).info
OTEL_ENABLEDEnable OpenTelemetry instrumentation.false

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
    
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec
    

Docker

docker build -t evals-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=stdio -e EVALS_DATA_DIR=/data -v evals-data:/data evals-mcp-server

The Dockerfile defaults to HTTP transport, stateful session mode, and logs to /var/log/evals-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

DirectoryPurpose
src/index.tscreateApp() entry point — registers tools and the resource, inits the record-store and exporter services.
src/configServer-specific environment variable parsing and validation with Zod.
src/mcp-server/toolsTool definitions (*.tool.ts).
src/mcp-server/resourcesResource definitions (*.resource.ts).
src/services/eval-recordThe record schema, draft builder, and submit gate.
src/services/graderDeterministic grader DSL execution and the committability check.
src/services/record-storeOn-disk JSON record CRUD, the draft→submitted move, and export writes.
src/services/exporterCompiling submitted records to JSONL/CSV/Inspect/lm-eval.
tests/Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging; records persist to disk via the record-store service, not ctx.state
  • Register new tools and resources in the createApp() arrays in src/index.ts
  • The server is the source of truth — validate inputs, run the grader as a hard gate, and never admit a record on assertion alone

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
npx -y @cyanheads/evals-mcp-server

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-cyanheads-evals-mcp-server": {
      "command": "npx",
      "args": [
        "-y",
        "@cyanheads/evals-mcp-server"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

@cyanheads/evals-mcp-servernpm

Compatible MCP Clients

io.github.cyanheads/evals-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More