Author verifiable eval records through a draft→review→revise→submit loop with enforced graders.
Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
Verifiable eval records, authored through a draft → review → surgical-revise → submit loop with server-enforced graders. Create a draft carrying its own executable grader, patch it surgically field by field, and submit through a committability gate that requires the gold to pass, a declared negative case to fail, and an independent verification to agree — then compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness. Runs as a stdio process or a local Streamable HTTP server.
| Tool | Description |
|---|---|
evals_describe_schema | Return the required and optional fields plus grader options for a task type. Call before drafting. |
evals_create_draft | Create a draft eval record carrying its own grader; returns the parsed record, a review protocol, and a verification subagent prompt. |
evals_get_record | Read a draft or submitted record by id; the id is stable across submit. |
evals_revise_draft | Apply a surgical set / append / unset patch to a draft by dotted path; re-runs the self-consistency check. |
evals_discard_draft | Delete a draft record by id. Draft-only. |
evals_run_check | Run a grader spec against candidate answers and get PASS/REJECT per candidate, decoupled from any saved record. |
evals_submit_draft | Finalize a draft through the committability gate, then freeze it. |
evals_list_records | Browse and filter records by status, domain, task type, or tag. Returns a compact summary per record. |
evals_export_records | Compile submitted records to JSONL, CSV, Inspect AI, or lm-evaluation-harness and write the artifact under exports/. |
| Resource | Description |
|---|---|
eval://record/{id} | A single draft or submitted record by id — the same payload evals_get_record returns, for resource-capable clients. |
All record data is also reachable through the tool surface — evals_get_record for a single record, evals_list_records to browse. The resource is a convenience mirror for clients that support resources, not the access path.
evals_describe_schema tooltask_type is one of numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_responsemcq needs choices, free_response needs an llm_rubric grader)evals_create_draft tooltask_type discriminated union and persists the draft; mcq requires choices, free_response requires an llm_rubric gradergold and each discrimination.positive, and REJECT each discrimination.negativeverification evidence and captures (EvalsIDs) when provenance is already in handgrader_unexecutable, task_type_constraint, mcq_choice_mismatchdraft — passing self-consistency proves the grader discriminates, not that the gold is correctevals_get_record toolid, stable across submit — resolves whether the record is still a draft or already submittednot_found when no record matches; recovery points to evals_list_recordsevals_revise_draft toolset (dotted-path → value), append (dotted-path → array items), and unset (dotted paths) operations — never a full-record rewritetask_type or server-owned fields — start a new draft to change the discriminantchanged list (op, path, before, after)record_frozen on a submitted idnot_found, record_frozen, invalid_patch_path, task_type_constraint, mcq_choice_mismatchevals_discard_draft tooldraft_idrecord_frozen when the id refers to a submitted recordnot_found rather than a distinct "already discarded" error — effectively idempotentevals_run_check toolcandidates (strings, numbers, objects, or arrays) without touching a saved recorddetail per candidate, plus the resolved comparison value (e.g. the math.js-evaluated numeric target)gold applies only to gold-relative kinds (exact_match); it's a no-op for target-embedding kinds like numeric and mcqllm_rubric cannot run here — submission relies on recorded independent verification insteadgrader_unexecutable, mcq_choice_mismatchevals_submit_draft toolcaptures from EVALS_CAPTURE_DIR, cross-checking the gold against the authoritative captured valuecontent_hash; confirm (or EVALS_REQUIRE_CONFIRMATION) can require human confirmation through multi-round input before finalizingsubmitted, stamps submitted_at and a checksum, and freezes it; otherwise refuses and the record stays a draftfree_response is admitted on recorded independent verification alone and flagged server_verified: falsenot_found, record_frozen, verification_incomplete, grader_failed_on_gold, verification_disagrees_with_gold, missing_negative_case, negative_case_passed, duplicate, decorrelation_violation, capture_unresolved, submit_declinedevals_list_records toolstatus (draft/submitted), domain, task_type, or tag; up to 500 per call (default 50)shown, cap, total count) when the limit is hit, so a partial set is never mistaken for the whole corpusevals_export_records tooljsonl (lossless), csv (flattened, lossy summary), inspect (UK AISI Inspect AI), lm-eval (EleutherAI lm-evaluation-harness)domain / task_type / tag filtersubmitted records are exported — drafts are skippedexports/ and returns its path, record count, byte size, and a short preview instead of dumping it inlineeval://record/{id} resourceevals_get_record, as application/jsonid comes from evals_list_records or a draft/submit responsenot_found when no record matchesBuilt on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Eval authoring:
draft → review → surgical-revise → submit loop, with the server acting as both scribe (normalize, persist, compile) and adversarial checker (runs the record's own grader, rejects what doesn't hold up)discriminatedUnion on task_type — numeric, exact_answer, set_answer, mcq, regex_answer, json_answer, free_responsenumeric via math.js, exact_match, set_match, regex, mcq, json_match) run server-side; llm_rubric relies on recorded independent verificationEVALS_DATA_DIR — inspectable, diffable, version-controllable records, with drafts, submitted records, and exports kept separateAgent-friendly output:
evals_create_draft and evals_revise_draft return the parsed record parroted back, a per-field review protocol, and a ready-to-paste verification-subagent promptevals_list_records reports shown / cap / total count when the limit is hit, so a partial set is never mistaken for the whole corpusreason plus a recovery hint, so a rejected record tells the agent exactly what to fixAdd the following to your MCP client configuration file. Set EVALS_DATA_DIR to a writable folder — the server manages drafts/, submitted/, and exports/ under it.
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/evals-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"EVALS_DATA_DIR": "/absolute/path/to/evals-data"
}
}
}
}
Or with npx (no Bun required):
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/evals-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info",
"EVALS_DATA_DIR": "/absolute/path/to/evals-data"
}
}
}
}
Or with Docker:
{
"mcpServers": {
"evals-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"-e", "EVALS_DATA_DIR=/data",
"-v", "evals-data:/data",
"ghcr.io/cyanheads/evals-mcp-server:latest"
]
}
}
}
For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 EVALS_DATA_DIR=./evals-data bun run start:http
# Server listens at http://localhost:3010/mcp
EVALS_DATA_DIR. No external API key is required.git clone https://github.com/cyanheads/evals-mcp-server.git
cd evals-mcp-server
bun install
cp .env.example .env
# edit .env and set EVALS_DATA_DIR
All server configuration is validated at startup via Zod schemas in src/config/server-config.ts.
| Variable | Description | Default |
|---|---|---|
EVALS_DATA_DIR | Root folder for record JSON; the store manages drafts/, submitted/, and exports/ under it. | ./evals-data |
EVALS_REQUIRE_CONFIRMATION | When true, evals_submit_draft requests human confirmation through multi-round input before finalizing. | false |
EVALS_DEFAULT_LICENSE | Default metadata.license applied when a draft omits one (e.g. CC-BY-4.0). | — |
EVALS_CAPTURE_DIR | Directory of framework-written tool-call captures; when set, captures EvalsIDs resolve to full dumps. | — |
MCP_TRANSPORT_TYPE | Transport: stdio or http. | stdio |
MCP_HTTP_PORT | Port for the HTTP server. | 3010 |
MCP_SESSION_MODE | HTTP session mode: auto, stateful, or stateless (auto resolves to stateful). A stateless HTTP start is refused, since a 2025-era client can answer the evals_submit_draft confirmation only over a live session. No effect on stdio. | stateful |
MCP_AUTH_MODE | Auth mode: none, jwt, or oauth. | none |
MCP_LOG_LEVEL | Log level (RFC 5424). | info |
OTEL_ENABLED | Enable OpenTelemetry instrumentation. | false |
See .env.example for the full list of optional overrides.
Build and run:
# One-time build
bun run rebuild
# Run the built server
bun run start:stdio
# or
bun run start:http
Run checks and tests:
bun run devcheck # Lint, format, typecheck, security
bun run test # Vitest test suite
bun run lint:mcp # Validate MCP definitions against spec
docker build -t evals-mcp-server .
docker run --rm -e MCP_TRANSPORT_TYPE=stdio -e EVALS_DATA_DIR=/data -v evals-data:/data evals-mcp-server
The Dockerfile defaults to HTTP transport, stateful session mode, and logs to /var/log/evals-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
| Directory | Purpose |
|---|---|
src/index.ts | createApp() entry point — registers tools and the resource, inits the record-store and exporter services. |
src/config | Server-specific environment variable parsing and validation with Zod. |
src/mcp-server/tools | Tool definitions (*.tool.ts). |
src/mcp-server/resources | Resource definitions (*.resource.ts). |
src/services/eval-record | The record schema, draft builder, and submit gate. |
src/services/grader | Deterministic grader DSL execution and the committability check. |
src/services/record-store | On-disk JSON record CRUD, the draft→submitted move, and export writes. |
src/services/exporter | Compiling submitted records to JSONL/CSV/Inspect/lm-eval. |
tests/ | Unit and integration tests mirroring src/. |
See CLAUDE.md/AGENTS.md for development guidelines and architectural rules. The short version:
try/catch in tool logicctx.log for request-scoped logging; records persist to disk via the record-store service, not ctx.statecreateApp() arrays in src/index.tsIssues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run test
Apache-2.0 — see LICENSE for details.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @cyanheads/evals-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-cyanheads-evals-mcp-server": {
"command": "npx",
"args": [
"-y",
"@cyanheads/evals-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.cyanheads/evals-mcp-server works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.