Let one coding-agent CLI consult another: Codex, Claude Code, OpenCode, Copilot or Antigravity.
Orchestrator MCP is a local Model Context Protocol server that lets one coding agent consult another. It launches the Codex, Claude Code, OpenCode, or experimental Antigravity CLI already installed and authenticated on your machine, routes the request, and returns a structured answer.
It does not ask for a provider key, proxy provider traffic, or silently switch models. Authentication remains inside each vendor's CLI.
Quick start from Claude Code, with Codex already logged in (other hosts: Install):
uv tool install orchestrator-mcp-server # or pipx, pip, or brew
orchestrator-mcp-server init --host claude # writes ~/.orchestrator-mcp/config.yaml
claude mcp add orchestrator \
--env ORCHESTRATOR_CONFIG=$HOME/.orchestrator-mcp/config.yaml \
--env ORCHESTRATOR_HOST_RUNTIME=claude -- orchestrator-mcp-server
Then ask Claude Code for a second opinion from Codex.
| Without Orchestrator | With Orchestrator |
|---|---|
|
The conversation stays connected from the same client. |
Same subscriptions. Less context shuffling.
Claude Code host ──► Orchestrator MCP ──► Codex CLI
Codex host ──► Orchestrator MCP ──► Claude Code CLI
Any host ──► Orchestrator MCP ──► OpenCode CLI (DeepSeek, Qwen, Kimi…)
Any host ──► Orchestrator MCP ──► Antigravity CLI (experimental)
local routing
no provider API keys
no same-execution-identity loops
Ordinary orchestrator_consult calls exclude the host's entire runtime: Claude Code
cannot use that tool to consult Claude Code, and Codex cannot use it to consult Codex.
Reviews and workflows use the narrower (runtime, model) execution identity; when
consult.host.model names the host precisely, they may route to a provably different,
versioned model on the same runtime. The host can never route work back to its own
execution identity.
If Claude Code asking Codex is all you need, OpenAI's codex-plugin-cc does that, with background jobs. PAL MCP reaches many models through provider API keys. Orchestrator is for the rest:
| Orchestrator MCP | codex-plugin-cc | PAL MCP | |
|---|---|---|---|
| Direction | Either way: Claude Code asks Codex, Codex asks Claude Code, and either asks OpenCode or Antigravity | Claude Code asks Codex | Any MCP host asks API models; clink launches agent CLIs |
| Credentials | Each CLI's own login, no keys | The Codex CLI's login | Provider API keys (OpenRouter, Gemini, OpenAI, …) |
| Review | A panel of reviewers from different vendors, answered in parallel, then one synthesis | Codex review and adversarial review | codereview and multi-model consensus |
| Delegated edits | In a disposable worktree under an OS sandbox; the host applies the diff | /codex:rescue hands the task to Codex | — |
| Record | Every turn in a local database, with cost estimates and a dashboard | Job status and result | Conversation threads across models |
Calls block the turn that makes them, often for minutes. To keep working meanwhile in Claude Code, have a background subagent make the call.
Google's Gemini CLI stopped serving free and Google AI Pro/Ultra accounts on June 18, 2026; Google models are reached here through its successor, Antigravity CLI.
Claude Code: install as a plugin. Needs uv on your PATH, and at least one of Codex, Antigravity or OpenCode installed and logged in (step 1 below). What you consult them about — prompts, diffs, file contents — goes to those providers under your own logins, and the usage is billed to your accounts with them; see the privacy policy and the Security model. In Claude Code:
/plugin marketplace add crAK1644/orchestrator-mcp
/plugin install orchestrator-mcp@orchestrator-mcp
/orchestrator-mcp:setup
/orchestrator-mcp:setup writes ~/.orchestrator-mcp/config.yaml from the agent CLIs it finds, then runs doctor on it. It finds Codex and Antigravity; OpenCode goes into the config by hand, from its agent in config.example.yaml. Then reconnect plugin:orchestrator-mcp:orchestrator in /mcp, or restart Claude Code: /reload-plugins keeps the server that started without a config.
If you added the server earlier with claude mcp add orchestrator, remove that entry (claude mcp remove orchestrator -s <scope>, with the scope claude mcp get orchestrator shows), or two copies of the server run side by side.
Every other host, and Claude Code without the plugin, installs the server itself:
Homebrew:
brew tap crAK1644/tap
brew install orchestrator-mcp-server
Apple Silicon uses a prebuilt package. Intel macOS and Linux build dependencies from source; use the uvx option if you want a faster, temporary install.
uv, pipx or pip: uv tool install orchestrator-mcp-server, pipx install orchestrator-mcp-server, or pip install orchestrator-mcp-server into a virtualenv. Python 3.11 or newer.
Sign in to each agent you want Orchestrator to use:
codex login
claude auth login
These are the normal Codex and Claude Code login flows. Orchestrator checks readiness, but never reads or stores their credentials.
For OpenCode, sign in once with opencode auth login for whichever provider you plan to consult. Hosted providers only — this server does not run a model on your machine. See the OpenCode runtime section below.
orchestrator-mcp-server init --host claude # or codex: the client you run it under
init finds the Codex, Claude Code and Antigravity CLIs installed here, writes
~/.orchestrator-mcp/config.yaml (mode 0600, never over an existing file; --path
picks another), and prints the exact line for step 3. It picks the best reviewer
that is not the host. It writes no workflow: block, because only you can choose the
directories a workflow may work in, and no OpenCode agent, because its free models
rotate too often for a template; add both by hand from
config.example.yaml.
consult:
database_path: ~/.orchestrator-mcp/consultations.sqlite3
timeout_s: 180
agents:
codex:
runtime: codex
command: codex
model: gpt-5.6-sol
priority: 10
web_search: true
scores: { coding: 95, research: 90, reasoning: 95, review: 90 }
claude:
runtime: claude
command: claude
model: claude-opus-5-5
priority: 10
web_search: true
scores: { coding: 90, research: 95, writing: 95, review: 95 }
The key each agent is filed under is its id -- the name callers pass as target_agent,
and the name the dashboard writes into URLs. It has to be lowercase, start with a letter
or digit, and use only letters, digits, dots, dashes and underscores, up to 64
characters. Anything else is refused at startup with a message naming the key.
See config.example.yaml for a broader annotated configuration,
an OpenCode agent, and an experimental Antigravity example.
init printed this with your paths filled in.
claude mcp add orchestrator \
--env ORCHESTRATOR_CONFIG=$HOME/.orchestrator-mcp/config.yaml \
--env ORCHESTRATOR_HOST_RUNTIME=claude \
-- orchestrator-mcp-server
Add this to ~/.codex/config.toml:
[mcp_servers.orchestrator]
command = "orchestrator-mcp-server"
env = { ORCHESTRATOR_CONFIG = "~/.orchestrator-mcp/config.yaml", ORCHESTRATOR_HOST_RUNTIME = "codex" }
Add this to claude_desktop_config.json (Settings → Developer → Edit Config), with the
command that which orchestrator-mcp-server prints: Desktop does not read your shell's
PATH.
{
"mcpServers": {
"orchestrator": {
"command": "/opt/homebrew/bin/orchestrator-mcp-server",
"env": {
"ORCHESTRATOR_CONFIG": "~/.orchestrator-mcp/config.yaml",
"ORCHESTRATOR_HOST_RUNTIME": "claude"
}
}
}
}
VS Code reads .vscode/mcp.json in the workspace; Cursor reads ~/.cursor/mcp.json
with the same entry under "mcpServers" instead of "servers".
{
"servers": {
"orchestrator": {
"type": "stdio",
"command": "orchestrator-mcp-server",
"env": {
"ORCHESTRATOR_CONFIG": "~/.orchestrator-mcp/config.yaml",
"ORCHESTRATOR_HOST_RUNTIME": "claude"
}
}
}
}
Neither editor is one of the four runtimes, so name the one whose models its chat runs:
claude for Claude, codex for GPT. That agent is then left out of the routing, and a
question is never handed back to the model that asked it.
In opencode.json:
{
"mcp": {
"orchestrator": {
"type": "local",
"command": ["orchestrator-mcp-server"],
"environment": {
"ORCHESTRATOR_CONFIG": "~/.orchestrator-mcp/config.yaml",
"ORCHESTRATOR_HOST_RUNTIME": "opencode"
}
}
}
}
Restart the MCP client after changing its configuration.
[!TIP] Give
ORCHESTRATOR_CONFIGan absolute path or one under~. GUI-launched clients often start in a different working directory and inherit a smallerPATHthan your terminal.
ORCHESTRATOR_CONFIG=~/.orchestrator-mcp/config.yaml ORCHESTRATOR_HOST_RUNTIME=claude \
orchestrator-mcp-server doctor
One ok or FAIL line per check: the config loads, the database opens, and each
agent's CLI is installed and logged in. It exits 1 if anything failed. The only
things it runs are the CLIs' own login checks, so no project material leaves the
machine.
The client spawns the server and talks to it over stdin, so orchestrator-mcp-server
is not a command you start yourself. (The dashboard is the other
half of this distribution and is started by hand.) Two flags answer questions from
outside a client, besides init and doctor above:
orchestrator-mcp-server --version # which build the client will spawn
orchestrator-mcp-server --help # what the environment variables have to say
Both answer and exit without reading your configuration. Anything else on the command line, besides the reports below, is refused rather than ignored. To make the server read the configuration, run it with no arguments: a file it cannot accept leaves as a message naming the key -- one line for most mistakes, several for a schema violation, never a traceback.
Four read-only commands print what the database already holds. They run instead of the
server, read the same config for its database_path, and open the file read-only: they
never migrate it, so a database from an older version gets a sentence telling you to
start the server once.
orchestrator-mcp-server usage [--days 30] [--json] # turns, tokens and cost per agent and model
orchestrator-mcp-server history [--limit 20] [--kind review] [--json] # recent records, with their ids
orchestrator-mcp-server scorecard [--days 30] [--json] # how each reviewer answered, and what became of its findings
orchestrator-mcp-server export ID # one record and everything it owns, as JSON
A price is shown only when every turn in the group reported one; otherwise it reads
unknown beside the sum of the prices that were reported, never as zero. The scorecard
counts a finding as kept when the host marked it fixed or accepted_risk and
rejected when it marked it rejected; those are dispositions the host reported when it
finalized a review, which this server does not check, and a later recheck does not
update them. A reviewer's hit rate appears once ten findings are decided. Recheck
reviews are left out so the same finding is not counted twice, and a review with no
synthesis on record (not finalized yet, or store_full_content: false, under which
finalizing is refused) is counted as asked but not judged.
export takes a full id or a unique prefix of at least eight characters from history.
It prints the record as stored, run through the same masking as everything else, without
session ids or confirm-token hashes; a review or workflow names the consultations it owns
by id, and you export those to read them. What it prints is your prompts and the
answers, so treat it as you would the database. Opening a WAL database read-only can
create -wal and -shm files beside it; the database file itself is never written.
uvx insteadNo permanent server install is required:
claude mcp add orchestrator \
--env ORCHESTRATOR_CONFIG=$HOME/.orchestrator-mcp/config.yaml \
--env ORCHESTRATOR_HOST_RUNTIME=claude \
-- uvx orchestrator-mcp-server
For Codex:
[mcp_servers.orchestrator]
command = "uvx"
args = ["orchestrator-mcp-server"]
env = { ORCHESTRATOR_CONFIG = "~/.orchestrator-mcp/config.yaml", ORCHESTRATOR_HOST_RUNTIME = "codex" }
The PyPI distribution is named orchestrator-mcp-server; the shorter PyPI name belongs to another project.
| Capability | What it does |
|---|---|
| Second opinion | Ask another vendor's coding agent about code, research, writing, reasoning, or review. |
| Connected follow-ups | Continue the native CLI session by returning its consultation_id. |
| Predictable routing | Rank configured agents by capability score, priority, then agent ID. |
| Explicit model choice | Verify the responding model when the CLI exposes that information; fail on a detected substitution. |
| Review panel | Ask one reviewer, or up to five in deep mode, over the same approved material. |
| Three-phase workflow | Run a whole job — research and planning, implementation and testing, review and fixing — with eligible models bound to steps their runtime and configured execution mode permit. |
| Slash commands | Drive consultations, reviews and workflows by name, with their checkpoints written down rather than hoped for. |
| Local history | Store consultations, reviews, and workflows in SQLite, with an optional loopback dashboard. |
| Answer-only isolation | Codex, Claude Code, and OpenCode are prevented from using action tools; explicit web mode enables only the target runtime's web-search facility. Experimental Antigravity detects and fails reported tool use but cannot yet prevent it. |
| Tool | Purpose |
|---|---|
orchestrator_consult | Start or continue a structured consultation. |
orchestrator_consult_many | Ask 2 to 5 agents the same question at once — named, or the top count by score — and get one envelope each for the host to merge. Costs one consultation per agent; each is resumable by its own consultation_id, and all share the label group <group_id>. |
orchestrator_list_consult_agents | Show configured agents, routing scores, installation, and login readiness. |
orchestrator_get_consultation | Retrieve a stored consultation, its turns, usage, and routing decision. |
orchestrator_list_consultations | Recent ordinary consultations, newest first. Metadata only. |
orchestrator_delete_consultation | Delete one ordinary consultation and its local turns. |
orchestrator_request_delete_all_consultations / orchestrator_delete_all_consultations | Preview and confirm deletion of an exact ordinary-history snapshot. |
These deletion tools remove local SQLite records only. They cannot erase a consulted runtime's own CLI or provider history.
Three independent opt-ins: the consult tools are always advertised, the review tools only with a consult.review block, the workflow tools only with a consult.workflow block. Reviewers are not a workflow, and a workflow is not reviewers.
The server also serves MCP prompts, which a client that speaks prompts/list renders as slash commands. In Claude Code they appear as /mcp__<server-name>__<command>, where the server name is whatever you called it in your MCP client config — /mcp__orchestrator__review for the orchestrator entry shown above. Installed as the plugin, the server is named plugin_orchestrator-mcp_orchestrator, so the same command is /mcp__plugin_orchestrator-mcp_orchestrator__review.
| Command | Arguments | What it expands to |
|---|---|---|
consult | question, agent | Ask another agent, keep the consultation_id, and report the disagreements rather than smoothing them out. |
review | goal, deep | Plan the review, show the plan and secret_hits, stop for the user, then run and finalize. |
workflow | goal, workdir | Start the workflow, then plan-step, stop, run-step, check status, one step at a time. |
status | workflow_id | Report which reviews and workflows are unfinished and what each is waiting on. |
Every argument is optional; a command with none expands into an instruction to ask you for the missing part. A client that speaks completion/complete offers the configured agents for agent and recent ids for workflow_id. review and workflow are advertised only when their tools are, on the same two answers — a command that could only reply "no reviewers are configured" costs a round trip and reads like a bug.
Two things worth being clear about. Nothing is installed. These arrive over the same stdio connection as the tools: no command directory, no generated markdown, nothing written to your machine, and a client that does not speak prompts/list is unaffected. A prompt is text, not an action. Expanding one consults nobody, sends nothing, and starts no workflow — it reaches the host's conversation as if you had typed it, and the host then calls the tools, checkpoints and all. They exist because the flows worth having here are handshakes, and a host driving them from tool descriptions alone tends to skip the checkpoint that makes them worth having.
Each record is also a resource template, for a client that attaches resources rather than calling tools: orchestrator://consultation/{consultation_id}, orchestrator://review/{review_id} and orchestrator://workflow/{workflow_id} read as the JSON the matching get tool returns, and the id completes from recent history. Templates list no instances; the list tools are how you find an id. Each is advertised on the same opt-in as its tools.
In a host that renders MCP Apps, such as Claude Desktop, orchestrator_get_review, orchestrator_finalize_review and orchestrator_workflow_status show their result as a table: findings by severity, reviewer and location, or a workflow's steps and what can run next. The page ships in the package as ui://orchestrator/view.html, loads nothing from the network, and puts reviewer text on the page as text, never as markup. Other hosts get the same text result as before.
orchestrator_consult selects the eligible agent with the highest capability score. Lower priority wins a score tie; agent ID breaks the final tie. Set consult.score_margin to let priority decide among agents within that many points of the top score: give a cheap or free agent a low priority and a margin of 5, and its 92 beats a paid agent's 95. The default 0 keeps the plain score order. A missing capability or a score of 0 makes an agent ineligible.
The selected CLI runs under its existing login and returns one response envelope:
| Field | Meaning |
|---|---|
ok | False exactly when error is set. Check this before reading the answer. |
consultation_id | Handle for continuing the same native conversation. |
content | Answer, assumptions, uncertainties, follow-up questions, and sources. |
route | Agent, runtime, model, score, priority, and whether it was selected explicitly. |
usage | Token counts when the CLI reports them. |
latency_ms | End-to-end elapsed time. |
error | Stable error code, message, agent, and sometimes a command the user must run. |
If the chosen agent fails, Orchestrator returns that failure. It does not quietly fall through to a different model.
source_mode | What the consulted agent receives |
|---|---|
auto | document when context is present; otherwise model. |
document | Only the supplied context, with action tools disabled. |
web | The target CLI's own web search. web_turn_limit bounds it on Claude; Codex is bounded by timeout_s alone. |
model | No context and no web search; answer from model knowledge. |
Request fields:
| Field | Required | Meaning |
|---|---|---|
capability | yes | coding, research, writing, reasoning, review, planning, prompt_authoring, testing, or synthesis. |
prompt | yes | Task or question, up to 100,000 characters. |
context | no | Evidence, up to 1,000,000 characters. |
source_mode | no | auto, document, web, or model. |
consultation_id | no | Return the previous ID to continue the conversation. |
target_agent | no | Choose one configured agent instead of automatic routing. |
conversation_label | no | Label stored with the consultation, up to 200 characters. |
Agent configuration:
| Option | Default | Meaning |
|---|---|---|
runtime | required | codex, claude, opencode, or antigravity. |
command | required | Executable name or absolute path. |
model | required | Requested model and, where possible, verified responding model. |
priority | 100 | Lower wins a score tie, or any pick within consult.score_margin. |
enabled | true | Keep the agent configured but out of routing when false. |
scores | none | 0–100 per capability; missing means ineligible. |
web_search | false | Permit source_mode: web for this agent. |
reasoning_effort | unset | low, medium, high, xhigh, or max; Codex only. |
timeout_s | unset | Limit for one turn with this agent, overriding consult.timeout_s. |
A consultation asks one agent. A review asks one or more configured reviewers the same question over the same material.
plan review approve + run synthesize
sends nothing ──► reviewers answer ──► host records conclusion
│ in parallel │
└─ scope one-time token └─ every Critical and Important kept
reviewers
secret hits
request count
Enable reviews in config.yaml:
consult:
review:
reviewers: [codex] # standard: exactly one
deep_reviewers: [codex, claude] # deep: one to five
roots: [~/src] # context_paths is restricted to these trees
context_paths is a convenience for material too large to paste into a tool argument.
Orchestrator reads each named file beneath those roots and sends its contents to the
reviewers. The path string supplied by the caller is also included as the file heading
and manifest label, so an absolute path can disclose a username or directory layout. If
that is sensitive, pass the bytes through context with neutral labels instead.
Each root has to be an absolute path to a directory that exists. The server checks that
at startup and refuses to start naming the one it could not find, rather than starting
and failing the first context_paths call -- which is a reviewer's turn later.
The workflow is deliberately split:
orchestrator_review creates a plan and sends nothing. The plan shows reviewers, material size, web access, request count, and locations of credential-shaped text. A plan nobody runs within a day is dropped the next time a review is planned.orchestrator_review_run spends its one-time token and asks reviewers in parallel.orchestrator_finalize_review. Reviewer replies alone leave the review at awaiting_synthesis.The checkpoint binds the scope and makes the token single-use, but it is advisory: MCP gives the server no separate human channel, so it cannot prove who saw the plan. Human approval depends on the calling client's tool-confirmation experience or an external gate.
Finalization must preserve every machine-readable Critical and Important finding, even when other reviewers disagree with it. Deep mode also requires the host agent to record its own findings before seeing the reviewers' answers.
[!IMPORTANT] Material sent to a reviewer may remain in that vendor CLI's own history. Orchestrator cannot erase Codex, Claude Code, OpenCode, or Antigravity session logs.
| Tool | What it does |
|---|---|
orchestrator_review | Plan a review and show what would be sent. Sends nothing. |
orchestrator_review_run | Spend the token and ask reviewers in parallel. |
orchestrator_retry_review | Re-run failed reviewers without discarding successful answers. |
orchestrator_finalize_review | Record the host's synthesis; the only path to complete. |
orchestrator_cancel_review | Cancel a review while retaining answers already received. |
orchestrator_apply_fixes | Return selected findings and fix steps, and list the Critical and Important findings the selection leaves out. Changes no files. |
orchestrator_record_fix_round | Record the host's claim about a fix round. |
orchestrator_test_reviewers | Check installation and login readiness without sending project material. |
orchestrator_get_review / orchestrator_list_reviews | Read one review or recent review metadata. |
orchestrator_delete_review | Delete a review, its rechecks, and linked consultations. |
orchestrator_request_delete_all / orchestrator_delete_all_reviews | Preview and confirm deletion of an exact history snapshot. |
Reviews default to web: false. Reviewers cannot change files or run commands. orchestrator_apply_fixes is a plan for work the host agent performs; it never applies a patch itself.
A recheck is a review planned with parent_review_id. The server appends the parent's open findings and a recheck brief, so the host sends only the diff instead of the whole tree again. Set review.recheck_reviewers: raised to ask again only the reviewers behind an open finding (at least one); the plan lists the others in reviewers_skipped. The default all asks everyone. A recheck starts a fresh reviewer session rather than resuming the old one: a resumed session re-bills its whole transcript once the provider's prompt cache has expired.
A reviewer's prose comes back once, with the call that ran it, and its findings -- parsed out of that prose -- come back every time. orchestrator_get_review is the call that returns the prose again. That keeps a review's later calls from re-sending the same reviewer answers into your agent's context, where they would be charged for on every turn that follows.
Credential-shaped values are masked before storage. secrets="send_as_is" is an explicit escape hatch for a false positive: it requires the exact original goal and context again, sends those originals to the reviewers, and still stores only the redacted copy.
store_full_content: false does not apply here in full. A review's goal and context are stored either way — the second half of the approval handshake reads them back to send what was approved — and reviewer answers and findings are not. That leaves nothing to prove every Critical and Important survived synthesis, so orchestrator_finalize_review refuses, and the review stays at awaiting_synthesis. Finalization is refused on the same grounds when a reviewer answered only in unparseable prose, or when its findings were truncated.
A consultation is one question. A review is one body of material. A workflow is a
whole job held together: research and planning, implementation and testing, review and
fixing, with failed tests and open serious findings feeding a capped fix loop. Every
consultation and review it produces hangs off one workflow_id.
research? → plan → author_execution_prompt → implement
└→ [apply_patch if delegated] → test
test
├─ failed by default → fix → [apply_patch if delegated] → test
└─ passed / advance_on_failed_test → review → synthesize
synthesize
├─ clean → completed
└─ open serious findings → fix
`?` means research may be skipped. Fixing is capped at `max_fix_rounds`.
Enable it in config.yaml. Without a workflow: block the workflow tools are not
advertised at all, the same rule the review tools follow:
consult:
host:
runtime: claude # asserted against ORCHESTRATOR_HOST_RUNTIME
model: claude-opus-5 # optional; see "Host identity" below
workflow:
max_fix_rounds: 5
roots: [~/src]
advance_on_failed_test: false
review_policy:
different_from_implementer: true
different_from_planner: false
bindings:
research: {agent: codex-sol}
plan: {agent: codex-sol}
author_execution_prompt: {executor: host}
implement: {agent: nemotron-ultra, execution: patch}
apply_patch: {executor: host}
test: {executor: host}
review: {agents: [codex-sol, claude-opus]}
synthesize: {executor: host}
fix: {agent: codex-sol, execution: patch}
workflow: requires store_full_content: true and refuses at startup otherwise. A
workflow is its stored plans, briefs, patches and reports, and the review step cannot
finalize without them — the failure belongs at boot, not several paid steps in.
roots: follows the same rule as review.roots: — absolute after ~/$VAR expansion,
and a directory that exists, checked at startup. A relative root is refused rather than
resolved against wherever the client spawned the server, which would make the allowlist
something other than what the file says.
A binding is one of three shapes, and mixing them is a startup error:
| Binding | Meaning |
|---|---|
{executor: host} | The calling agent does this step itself and records the result. |
{agent: x, execution: patch} | That agent, in that execution mode. |
{agents: [x, y]} | Several agents. review is the only step that takes more than one. |
Leaving agent: out means auto: routed by capability score for the step, ties broken
by priority then agent id, the same rule orchestrator_consult uses. GPT, Claude,
DeepSeek, Gemini, Qwen — any model reachable through a supported runtime can take any
step it scores for and is permitted the mode for. A step with no binding falls to the
host, which is the conservative default — with one exception. review cannot be the
host's: its product is a review row that only the review service writes, and a
host-recorded outcome would be a synthesis written straight into a step. Because unbound
steps default to the host, a config that simply never mentions review is refused at
workflow_start, before anything has been spent, rather than at the review step after
research, planning and implementation have all been paid for.
Steps route on the four capabilities added for this: planning, prompt_authoring,
testing and synthesis, alongside coding, research and review.
Bindings are resolved and snapshotted at workflow creation, and so is the policy
they run under — the round cap, advance_on_failed_test and the review policy. Editing
config.yaml does not reroute a running workflow or move its cap; that takes
orchestrator_workflow_plan_replan and its own approval. A replan re-decides the steps
you name and leaves every other one on the routing the workflow already had.
execution_modes on an agent is operator trust, not capability. What actually
happens is the intersection of that list with what the runtime can be made to do, and a
refusal names which side said no.
| Mode | The agent gets | Repository access |
|---|---|---|
consultation | The read-only consult path, unchanged. | context_only |
patch | The same read-only path; it returns a unified diff. The host applies it. | context_only |
isolated_write | A disposable git worktree outside your repository, checked out at the workflow's baseline. The agent edits and runs commands there; Orchestrator reads the diff back out of git and the host applies it. Codex, and OpenCode where the sandbox holds. | worktree |
executor: host | Not an agent at all: the host edits its own checkout. | active_tree |
| Runtime | isolated_write | Why |
|---|---|---|
codex | supported | sandbox_mode: workspace-write with approval_policy: never is enforced by the CLI at OS level: a command aimed outside the worktree comes back Operation not permitted from the kernel, not from the model declining. Network is off, /tmp and $TMPDIR are excluded from the writable set. |
opencode | supported where the sandbox holds | Its own permission set isolates configuration, not filesystem effects, so the bound is Orchestrator's OS-level sandbox (seatbelt on macOS): writes are held to the worktree, and its runtime state is redirected into it and removed before the diff is read. The network is open, because the model is hosted: weaker than Codex, whose network is off. Stored opencode auth credentials are not carried in, so only providers that need none (such as OpenCode's free catalogue) work. Bubblewrap cannot grant that network yet, so Linux still refuses, and the refusal at startup says why. |
claude | refused | Same bar as OpenCode: its permission modes are requests, not kernel bounds. |
antigravity | refused | Writing needs --dangerously-skip-permissions, the one flag the adapter refuses by construction. |
A root allowlist and a prompt instruction are not containment. An agent that declares
isolated_write on a runtime that cannot be contained is refused at startup, not at
routing time.
No delegated agent writes to your working tree. In patch mode the agent never sees
your checkout — only the context the host sent it, exactly as a reviewer does — and the
diff comes back for the host to apply. In isolated_write it sees a copy: a worktree
under ~/.orchestrator-mcp/worktrees/<workflow_id>/<step_id>/, checked out at the latest
applied result (the workflow's baseline until the host has applied anything). It is
normally deleted after the diff is captured and its private recovery copy is written.
It is retained when capture fails or recovery cannot safely preserve the patch, because
the worktree may then hold the only usable copy. Either way the successful step ends in
awaiting_host_apply and the host owns the branch. ConsultAdapter gained nothing for
any of this: it is still three verbs with no way to ask for anything agentic, and write
capability lives in a separate package behind a separate protocol.
Four things about a contained run are worth knowing before you use one:
git add -A in the
worktree and takes the staged diff against the baseline, so files the agent created
are captured too — a plain git diff <baseline>.. would miss them. The model's summary
is stored beside the patch as a description of its work, never as the account of it.
Its list of commands is stored the same way, and is knowingly incomplete: codex omits
sandbox-denied commands from its event stream entirely.git add, commit, stash and checkout all fail from inside. It leaves the work in
the tree and this server records it. The step's timeout is
consult.workflow.execution_timeout_s (900s by default), not consult.timeout_s,
which is sized for a question.git init — makes git add -A refuse the whole tree. The step fails and the worktree
is kept, with its path in the error, because at that moment it holds the only copy
of the work. Ignored files are skipped by git add -A by design and can never appear
in a patch; they come back listed on the step's ignored field rather than vanishing
with the worktree..git is a one-line file naming its git directory, and it sits in the
directory the agent spent the whole step writing to. That gitdir is read at worktree
creation and passed explicitly to every later git call, so what capture diffs is not
decided by a file the step could rewrite.(runtime, model) is an execution identity, not an agent id, and it comes from trusted
startup configuration only — never a tool argument. runtime still comes from
ORCHESTRATOR_HOST_RUNTIME, which stays the authority; naming it under host: is an
assertion, and a mismatch refuses the boot.
model is the part the environment cannot supply. With it, a different model on the
host's own runtime becomes routable. Without it, every agent on that runtime is excluded
— today's consult behaviour. Write the versioned name: identity must be provably
different, so opus and claude-opus are treated as claude-opus-5 and refused, and
a name with no version at all is refused for being unprovable rather than assumed
distinct.
review_policy is checked against this resolved identity too, so two agent ids pointing
at one model are one reviewer however they are spelled.
plan_step run_step | record_host_step
returns a preview ──► spends that step's token
and a one-time token and runs or records it
There is no one-call form and no workflow-level token. One approval covering research, implementation, every fix round and every review would be an approval of nothing in particular. What a token proves is snapshot integrity — that the step being run is the step that was previewed, with the same agent, mode and inputs. It cannot prove a human saw anything, because MCP gives this server no channel to one; that depends on your client's tool-confirmation experience or an external gate.
A workdir must resolve beneath a configured root. / is never accepted and nothing is
inferred from the working directory. A dirty tree is refused without allow_dirty.
Nothing here reads your repository on an agent's behalf, so a delegated step starts with
no code in front of it. orchestrator_workflow_plan_step(workflow_id, step, context)
takes that material — the source of the files being changed, most of the time — and the
host decides what goes in it. Without it, an implementation step has the plan and the
brief and nothing to patch, and the honest models say exactly that instead of inventing
a file they were never shown.
The material is redacted with the same scrubber as everything else before it is stored and before it is sent, and it is covered by the step's prompt hash, so the text the preview described is the text that goes out. A review step gets the same material its coding steps did.
A coding agent's statement that it ran the tests is retained as reported information; it
is never the test result. A TestReport carries the exact command, working directory,
exit code, bounded output, duration and the commit tested, plus reported_by, which the
service assigns from which tool wrote the row and never reads out of a caller's
payload. A host-written report is host-attested, and its provenance says so.
reported_by: orchestrator has exactly one source: a test step bound to
isolated_write, where the exit codes come out of the CLI's own event stream rather
than out of anything the model wrote. Read what that does and does not claim. It means
every command the run reported returned zero — codex omits sandbox-denied commands
from that stream, so it is not a claim that the project's suite ran. A contained test
step that edited files while testing lists them on the report's changed_files; the
worktree is deleted either way, so nothing there is applicable, but "the tests pass" and
"the code was edited until they did" no longer look identical.
A failed test returns to fixing rather than advancing, unless advance_on_failed_test
is set.
Two fields are cross-checked rather than stored side by side. A report cannot be
passed with a non-zero exit code or with none at all — a command whose exit code was
never read is skipped, which is what a denied command or a killed process produces —
and it cannot be failed with a zero. The commit is stamped by the service from what
the workflow currently holds, not taken from the payload: loop_done compares them, so
a pass from an earlier round cannot carry a later one.
apply_patch is the one step whose whole job is a side effect on your tree, and the
only evidence it happened is a commit that was not there before. Recorded without one,
or with the commit it started from, the step is refused and marked failed rather than
advancing — there is no applied: false to write, because a step that did not do its
work is a failed step. Plan it again once the patch is applied and committed.
The review step goes through the review service — it does not write a synthesis
straight into workflow storage, which would route around the guarantee that every
serious finding survives. Findings gained a disposition (open, fixed, rejected,
accepted_risk), because "unresolved finding" was previously not representable;
rejecting or accepting the risk of a critical or important finding without a reason is
refused.
loop_done is computed here, never asked of a reviewer. It is true only when the
authoritative tests passed for the current commit, every reviewer's findings parsed and
were retained, no critical or important finding is still open, missing_serious
passed, and the workflow is in no exceptional state. Reaching max_fix_rounds with
serious findings still open ends the workflow needs_attention — not completed.
A fix round after review carries the findings that are still open, read back from the
review row rather than from workflow storage, so the round is an answer to the review
and not a second pass at the goal. A fix triggered by a failed test before review instead
carries that failed TestReport. A re-review is a recheck of the previous round's review
(parent_review_id): the server appends that review's open findings and tells the
reviewers to confirm or drop each one and to look for new problems only in the change.
Send the fix's diff as the step's context, not the whole tree again.
The coding prompt is assembled by this code: our EXECUTION_CONTRACT first, then the
scope, accepted plan, authored brief and prior findings as a JSON payload. An authored
brief is data inside that payload, and no field turns it into contract text.
That is code ownership, not a transport-level enforcement boundary. Claude Code has a real system-prompt channel; Codex and OpenCode receive one compiled text, so there the ordering is a prompt convention a determined model could argue with. Saying so is more useful than overclaiming.
| Tool | What it does |
|---|---|
orchestrator_workflow_start | Create a workflow: validate the workdir and root, resolve and snapshot bindings, record the git baseline. Sends nothing and returns no execution token. |
orchestrator_workflow_plan_step | Preview one step and mint its one-time token, with the optional context the step is shown. A review step returns the review plan's own token rather than an unrelated second approval. |
orchestrator_workflow_run_step | Spend the token and run the step through its bound agent. |
orchestrator_workflow_record_host_step | Record work the host did itself. The token is consumed as the host's attestation. |
orchestrator_workflow_status | State, artifacts, selected agents, round count, spend, and what may happen next. |
orchestrator_list_workflows | Recent workflows, newest first: id, goal, state. |
orchestrator_workflow_plan_replan / orchestrator_workflow_replan | Change the binding snapshot under the same preview-and-approve handshake. |
orchestrator_workflow_cancel | Cancel pending work and terminate a child this process owns, with the same caveat orchestrator_cancel_review carries about another process's children. |
orchestrator_delete_workflow | Delete one workflow with its steps, consultations and reviews. Refused while the workflow is open or a step's lease is live. |
orchestrator_request_delete_all_workflows / orchestrator_delete_all_workflows | Preview and confirm deletion of an exact workflow snapshot. |
A workflow deletes whole or not at all. Its consultations are excluded from every
consultation delete path and orchestrator_delete_review refuses a review that is a
workflow step, because a step pointing at a row that is gone still reads as intact —
and the next fix round would answer from the goal instead of from the review. So these
three tools are the only way any of those rows leave the database.
orchestrator_workflow_status is also the call that returns every step's output.
The tools that advance a workflow return the body of the step they touched and the
shape of the rest -- a plan, a patch and a test log do not change because a later
step ran, and re-sending them on every call is the same bytes accumulating in your
agent's context. Ask status when you want an earlier step's body back.
orchestrator_workflow_status reports what the workflow spent: per step, and totalled
over the workflow with its reviewers included. The numbers are rebuilt from the
consultations' turn ledgers at read time, so a step that took two turns counts both and
a re-read counts neither twice. A host step reports no usage — nothing was spent on it
here, which is not the same as it having cost zero. cost_usd is set only when every
turn behind it was priced: an agent on a free tier reports no price, and one unpriced
turn makes the total a floor rather than a sum, so it comes back unknown instead. The
dashboard shows the same numbers o
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx orchestrator-mcp-serverMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-crak1644-orchestrator-mcp": {
"command": "uvx",
"args": [
"orchestrator-mcp-server"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceorchestrator-mcp-serverpypiOrchestrator works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.