Goose-first MCP server for Workbench-owned acceptance evidence, validation gates, and analytics.
AI Workbench supervises AI coding agents, captures evidence, validates work, applies acceptance policy, and produces auditable PR-ready reports.
The PyPI package remains ai-workbench-mcp for this public alpha because the
ai-workbench package name is already occupied. The product and CLI are
AI Workbench:
pip install ai-workbench-mcp
ai-workbench --help
Current source metadata targets unpublished ai-workbench-mcp==0.8.0a0.
This public alpha consolidates local supervision, evidence capture, validation,
acceptance policy, and PR reporting into one product surface.
The supervisor is the preferred automated evidence path, but daemon, Codex hook, and OpenCode adapter coverage are alpha mechanisms. AI Workbench checks evidence quality and acceptance readiness; it does not prove the work is absolutely correct. High-risk work still requires human review.
validation_report.json.revision_decision.json.accept, needs_review, or block.Agent output is a proposal. Workbench accepts evidence.
MCP is the connection protocol. AI Workbench MCP is the tool server. Acceptance is decided by the selected validation profile and quality gate. The agent performs. Workbench accepts. MCP connects them.
Register a project once and start the local supervisor:
pip install ai-workbench-mcp
ai-workbench supervisor setup --project-dir . --task-type code_change
ai-workbench supervisor start
Run Codex, OpenCode, Goose, or another supported local workflow in the project. Then inspect the latest report:
ai-workbench supervisor status
ai-workbench reports show latest --project-dir .
Render PR-ready artifacts from a finalized run:
ai-workbench pr-gate --run-dir runs/<run_id>
The canonical local run ledger is:
runs/<run_id>/
task_metadata.json
final_prompt.md
model_selection.json
model_output.md
validation_report.json
revision_decision.json
run_log.jsonl
metadata.json
transcript.jsonl
commands.jsonl
workspace/
validation/
artifacts/
validation_report.json and revision_decision.json are the final acceptance
authority. Supporting supervisor reports are local evidence, not a substitute
for those Workbench artifacts.
Install project-local Codex hooks:
ai-workbench setup codex --project-dir . --task-type code_change
Restart Codex or start a new session, open /hooks, review the project hook,
and trust it once. Until a hook event is observed, supervisor status reports
Codex coverage as configured but unverified.
AI Workbench still exposes the same MCP tool lifecycle. Register the server with Goose or another MCP host using:
ai-workbench mcp serve
The seven MCP tools remain:
workbench_open_run
workbench_select_policy_pack
workbench_select_model
workbench_record_execution
workbench_validate_run
workbench_quality_gate
workbench_analyze_runs
Workbench PR acceptance consumes real Workbench run evidence:
ai-workbench pr-gate \
--run-dir runs/<run_id> \
--out runs/pr_gate/pr_comment.md \
--json-out runs/pr_gate/pr_decision.json
Outcomes are exactly:
acceptneeds_reviewblockMissing, unreadable, or scaffold-only evidence blocks. A green CI run, uploaded artifact, sticky PR comment, or model self-claim is not acceptance evidence.
To add starter configs, prompts, recipes, docs, and the GitHub PR-gate workflow to a repository:
ai-workbench bootstrap --target .
The bootstrap keeps runs/ ignored.
For a package-only synthetic demo:
ai-workbench demo --target ./workbench-first-run
This shows accept, needs_review, and block PR-gate outcomes with fixture
evidence. It is not a real target-repository acceptance run.
python -m pip install -e ".[dev,publish]"
python -m pytest -q -p no:cacheprovider
python -m ruff check . --no-cache
python -m mypy --no-sqlite-cache --no-incremental
ai-workbench demo --target runs/package_demo_smoke
ai-workbench validate --project ai_workbench_mcp --profile scaffold --run-dir runs/scaffold-smoke
Do not commit runs/. Committed sample evidence must be sanitized and live
under examples/.
Recipes:
Sample evidence:
Apache-2.0. See LICENSE. MIT-origin attribution for the consolidated Prove It code is retained in NOTICE.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx ai-workbench-mcpMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-hrishikesh-thakre-ai-workbench-mcp": {
"command": "uvx",
"args": [
"ai-workbench-mcp"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceai-workbench-mcppypiAI Workbench MCP works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.