Work-queue kernel for AI coding agents: dependencies, worktree isolation, quality gates, recovery
A portable, agent-agnostic work-queue kernel for AI coding agents.
You keep a queue of phases and tasks with declared dependencies. You say "implement phase P2". Independent tasks fan out to parallel agents in isolated git worktrees; dependent ones wait. Every task passes a quality pipeline whose gates cannot be passed by assertion. If an agent crashes, its work is found rather than lost. If everything except the log is destroyed, the project's decision history rebuilds from the log alone.
One dependency beyond python3 and git (Jinja2, for the prompt templates; see
Extending it by writing text, not code). Works with Claude Code, Gemini CLI,
Codex, Copilot, Cursor, Kimi, opencode, Aider, a CI job, a Makefile, or a human at a
terminal — over a CLI
and an MCP server that are the same implementation.
If you are an AI agent and were asked to read this README, this section is your shorthand. Everything below it is detail.
What ddflow is. A work queue with rules the tools enforce, for projects that agents
build. Work is phases and tasks with dependencies and declared file globs. Each task gets
a lease and its own git worktree, then passes a gate pipeline: a check that could not run
is unavailable, never passed; one reviewer must be from another model family; a bug
closes only with a regression test that failed first. Everything is recorded in an
append-only log in .ddflow/events/, committed with the code. ddflow writes none of your
project's code — you do; it adds only its own config, a rules block and git hooks.
Every operation is both a shell command (ddflow …) and an MCP tool (ddflow_…).
If the project already uses it (.ddflow/ exists): call ddflow_brief (or run
ddflow brief) first and follow it. It replaces reading the project's rule and lesson
files. Then: ddflow_next → ddflow_claim → work only in the worktree it returns →
ddflow_gate_status, satisfy each gate → ddflow_merge → ddflow_complete. The full
driver is docs/ddflow/drivers/implement-phase.md in that project.
To onboard a project, ask the user first: it installs a tool, registers an MCP server and writes files in their repository.
uv tool install git+https://github.com/delian/ddflow-mcp (not on PyPI yet).ddflow adopt --agents claude (your agent's
name; see Wiring it into your agent). It writes
.ddflow/, the MCP registration, the commit hook and a rules block in AGENTS.md.
Restart the agent session so it loads the MCP server.ddflow import lists the plans, lessons, decisions and unmerged branches the
repository already has, and writes nothing. Run ddflow import --apply only once the
user agrees with that list.ddflow config --set gate.unit_tests.command "<their test command>". Then
ddflow companions lists the tools the gates expect that are missing: propose them,
never install one unasked.ddflow phase add …, ddflow task add … --globs …)
and finish with ddflow doctor.Exit codes, everywhere: 0 ok · 1 failure · 2 could not run / nothing to do · 3
refused. Never treat 2 as 0.
An AI coding agent is good at a task and weak at a project. One agent in one session mostly works. Run it for weeks, or run three at once, and the same failures come back:
ddflow is the layer between you and your agents that makes those failures structurally hard rather than a matter of discipline. It does not write code and it is not an agent. It is a queue, a set of rules the tools enforce, and a log of everything that happened.
| Without it | With ddflow |
|---|---|
| You decide what each agent does next, and keep the plan in your head or a chat. | The plan is a queue of phases and tasks with dependencies. ddflow next says what can start now and why everything else is blocked. |
| Parallel agents step on each other. | ddflow claim gives each task a lease and its own git worktree; tasks that declare overlapping files are refused, not merged over. |
| "Done" means the agent said so. | Every task passes a gate pipeline you configure. A gate that could not run is recorded unavailable, never passed; at least one reviewer must come from a different model family than the author; a bug cannot be closed without a regression test that failed first. |
| A crash loses work. | ddflow recover finds orphaned worktrees and reports what each holds. It never deletes work. |
| Every session starts from zero. | Lessons, decisions, research verdicts, bugs and your own prompts are recorded as you go. ddflow brief hands the agent the ones relevant to this task in a bounded amount of context, and ddflow recall searches all of it. |
| History is a transcript. | An append-only event log, committed in git. The board, the index and the reports are rebuilt from it; ddflow replay reconstructs the project's decisions from the log alone. |
--json form
and every command returns the same four exit codes, so a Makefile or a CI job can
drive it exactly as an agent does.It works with the agent you already use, because everything it does is reachable both ways: as a shell command and as an MCP tool. Use whichever your agent, script or CI job has. Your workflow is text, not code — the gate pipeline, the reviewer instructions and the agent-facing prompts are files in your repository that you can edit.
# ddflow-mcp is not on PyPI yet; until the first release, install from the repository:
uv tool install git+https://github.com/delian/ddflow-mcp # or: pipx install git+https://github.com/delian/ddflow-mcp
cd /path/to/your/project
ddflow adopt # registers the MCP server with your agents, writes .ddflow/ and a block in AGENTS.md
ddflow phase add P1 --title "Password reset"
ddflow task add P1.T1 --phase P1 --title "Reset-token endpoint" --globs 'src/auth/**'
ddflow next # what can start now, and why the rest is blocked
adopt registers the server you just installed, by its full path: an install that did
not come from a package index (from git, a local directory or an archive) carries a
direct_url.json in its metadata (PEP 610), and
for one of those adopt writes the ddflow-mcp installed beside its interpreter rather
than uvx ddflow-mcp, which would fetch from PyPI. Only an install from an index gets
uvx. --launch python still forces the interpreter-plus-PYTHONPATH form.
Then tell your agent "implement phase P1". The driver adopt installed tells it to
start with ddflow_brief, claim the task, work in its own worktree, satisfy each gate
and land the change. ddflow cannot make an agent follow instructions, but it makes
skipping them visible: the commit hook adopt installs flags a commit made without a
lease (or refuses it, if you set [enforce].commit_without_lease = "block"), and
ddflow complete refuses an item whose gates carry no outcome. Watch it with
ddflow board, and ask ddflow doctor at any point whether the project is healthy.
Every row is a command you can run in a terminal and a tool an agent can call over MCP — the same implementation, so neither drifts from the other.
| I want to… | CLI | MCP tool |
|---|---|---|
| see what the workflow is | ddflow workflow | ddflow_workflow |
| change the workflow | ddflow workflow pipeline task … · workflow gate <id> … · workflow drop <id> | ddflow_workflow_pipeline · _gate · _drop |
| change any setting | ddflow config --explain · --set <key> <value> | ddflow_configure |
| add a phase / a task | ddflow phase add P1 --title … · ddflow task add P1.T1 --phase P1 --globs 'src/**' | ddflow_phase_add · ddflow_task_add |
| get a plan into the queue | see From plan mode to the queue | same |
| know what to work on | ddflow next | ddflow_next |
| start a task | ddflow claim <id> → work → ddflow gate … → ddflow merge → ddflow complete | ddflow_claim, ddflow_gate_*, ddflow_merge, ddflow_complete |
| see everything about one item or bug | ddflow show <id> (phase, task or bug id) | ddflow_show |
| see progress / effort | ddflow progress · ddflow status · ddflow board | ddflow_progress · ddflow_status · ddflow_board |
| find out if we're going in circles | ddflow loops | ddflow_loops |
| record a lesson / decision / research / bug | ddflow lesson add · decision add · research · bug found|fixed | ddflow_lesson_add · ddflow_decision_add · ddflow_research_add · ddflow_bug_* |
| search everything the project remembers | ddflow recall '<regex>' | ddflow_recall |
| check a text against what is already filed (read-only) | ddflow similar '<text>' [--kind bug,task,...] [--json] -- exit 0 with candidates, 2 with none | ddflow_similar |
| record what happened this session | ddflow session start|prompt|note|end | ddflow_session_* |
| read the engineering log | ddflow history | ddflow_history |
| check the tooling around the gates | ddflow companions | ddflow_companions |
| find work a crashed agent left | ddflow recover | ddflow_recover |
| check the project's integrity | ddflow doctor | ddflow_doctor |
| rebuild everything from the log | ddflow replay --verify | ddflow_replay |
| invoke a workflow / a mode of your own | ddflow prompts list · prompts show <name> | prompts/list · prompts/get |
| see what this project left undone | ddflow doctor · ddflow status | the footer on tool results |
| ask the tool to explain itself | ddflow help [topic] | ddflow_help |
Every read command takes --json. Every exit code means the same thing everywhere:
0 healthy · 1 real failure · 2 could not run / nothing to do · 3 coordination
refused. 2 is never collapsed into 0 — "nothing is ready" and "everything is
fine" are different facts, and an agent that cannot tell them apart invents work.
$ ddflow help # what this is, the loop, every capability grouped
$ ddflow help workflow # workflow · import · gates · parallel · memory · recovery · config
Reachable as ddflow_help over MCP, and that is the point: an agent connecting had 59
tool descriptions and a state-aware handshake, neither of which answers "what is this,
and how am I meant to work here". A tool description explains one tool to someone who
already picked it; the handshake describes this repository right now.
Two halves, deliberately:
ddflow/templates/prompts/help/, so
ddflow prompts eject-style overriding applies — put your own
.ddflow/prompts/help/workflow.md in place and the tool teaches your workflow.Three ratchets keep the prose honest, because a page recommending a flag that was renamed is worse than no page — whoever finds nothing reads the code, and whoever finds a wrong answer trusts it. Every command a page names must exist as a CLI leaf or an MCP tool; every topic the index offers must resolve; and every tool must fall into a group, so a new capability has to be classified rather than quietly dropped from an inventory that claims to be complete.
Three pictures of what ships by default. Every step is configurable (see
The workflow, and changing it); ddflow workflow
prints what this project actually runs.
The agent's loop. One item at a time: pick it, lease it, clear its gates, land it.
Each arrow labelled with an exit code is what the tool says, not a convention —
2 means nothing to do, 3 means refused, and neither is ever read as success.
flowchart TD
S["Session start<br/>ddflow brief · recover"] --> R{"Crashed agent's<br/>work left over?"}
R -- yes --> SV["Inspect the worktree, salvage,<br/>ddflow release"] --> N
R -- no --> N["ddflow next"]
N -- "exit 2: nothing ready" --> W["ddflow wait<br/>sleeps until a holder lets go"] --> N
N -- "exit 0: items ready" --> C["ddflow claim<br/>lease + isolated git worktree"]
C -- "exit 3: refused<br/>(lease or file-glob conflict)" --> N
C --> P["Task pipeline<br/>satisfy every gate in order"]
P --> M["ddflow merge<br/>lands the branch from the primary checkout"]
M --> D["ddflow complete<br/>checks every gate has an outcome<br/>+ a different-family review"]
D -- "exit 3: refused" --> P
D -- done --> N
With [flow].integration = "pr", merge opens a pull request and parks the item in
review instead; next completes it when the request merges.
The task pipeline. Ten gates, in order. Heavy borders are the gates
gates.required names by default (implement, unit_tests, merge); the rest still
need some outcome — passed, failed, unavailable, partial, or an explicit skip with a
reason — because silence is not a pass.
flowchart LR
subgraph A["You, the agent"]
direction TB
g1["1 · research<br/>falsifiable claim + probe"] --> g2["2 · rules<br/>ddflow brief --item"] --> g3["3 · implement<br/>in the item's worktree"]
end
subgraph X["A different model family"]
direction TB
g4["4 · rubber_duck<br/>try to refute it"] --> g5["5 · critic<br/>diff vs. stated intent"]
end
subgraph T["Tooling"]
direction TB
g6["6 · standards<br/>linters, architecture"] --> g7["7 · unit_tests<br/>the suite, executed"]
end
subgraph B["You, again"]
direction TB
g8["8 · bug_hunt<br/>recurring bug classes"] --> g9["9 · dedupe<br/>already exists?"]
end
g10(["10 · merge<br/>ddflow lands it"])
A --> X --> T --> B --> g10
classDef req stroke-width:3px
class g3,g7,g10 req
The phase pipeline. A phase wraps its tasks and checks the whole before it lands:
flowchart LR
p1["research<br/>the phase's open questions"] --> p2["tasks<br/>each runs its own<br/>ten-gate pipeline,<br/>in parallel where<br/>globs allow"]
p2 --> p3["unit_tests"] --> p4["bug_hunt"] --> p5["dedupe"]
p5 --> p6["live_test<br/>run the real thing"] --> p7["corrections"] --> p8["docs<br/>README and docs<br/>match the change"] --> p9(["merge"])
Details: The task pipeline · The phase pipeline · Parallelism and coordination · Crash recovery.
The CLI is the whole product. The MCP server is a second surface over the same
commands, and tests/test_mcp_parity.py fails if the two diverge — every subcommand has
a tool, every flag is reachable, and each exemption carries a written reason.
$ ddflow init
$ ddflow config --set gate.unit_tests.command "python -m pytest -q"
$ ddflow phase add P1 --title "Billing" --globs "src/billing/**"
$ ddflow task add P1.T1 --phase P1 --title "Tax rules" --globs "src/billing/tax.py"
$ ddflow next # exit 2 = nothing actionable; exit 1 = unknown --phase
$ ddflow claim P1.T1 # exit 3 = refused, with the reason
leased P1.T1 · worktree .ddflow-worktrees/P1.T1 · branch ddflow/P1.T1
$ cd .ddflow-worktrees/P1.T1 && ... # do the work
$ ddflow gate status P1.T1 # what the pipeline wants next
$ ddflow gate run P1.T1 unit_tests # runs it; the exit code IS the evidence
$ ddflow gate record P1.T1 implement --outcome passed --evidence "added tax.py"
# failed/unavailable/partial/skipped also need --reason
$ ddflow complete P1.T1 # exit 3 lists whatever is unsatisfied
$ ddflow merge P1.T1
You get everything except the judgement. Command gates run themselves; agent gates wait
for a human to record an outcome, and ddflow gate skip <id> <gate> --reason "..." is
the escape hatch — recorded as a skip, never as a pass.
In CI, the exit codes are the interface:
check:
ddflow doctor # 1 = integrity problems, each named
ddflow workflow # 1 = the pipeline does not hang together
ddflow cadence # 2 = no periodic pass is due
2 is never "no problem". A job that treats it as success reports a green build for a
suite that never ran.
ddflow mcp speaks newline-delimited JSON-RPC over stdio. You rarely run it by hand —
ddflow adopt writes the launch entry into each agent's own config and leaves existing
servers alone:
22 agents are supported. The full table, with what each one gets, is in Wiring it into your agent.
It also copies the driver to docs/ddflow/drivers/, and installs the pre-commit hook
that enforces claim-before-you-edit.
What an agent sees the moment it connects, with no call to make:
initialize result itself — and state-aware:
what is ready, what is in flight, which setup is missing, whether this project has
history worth importing, whether an import was left unfinished.ddflow://board, ddflow://brief, ddflow://lessons,
ddflow://research.A refused call says so first. Over MCP a tool that did not do what was asked — a claim
refused for an overlap (exit 3), or any call whose result would otherwise be its success
shape in nulls — returns a JSON body whose first key is
"refusal": {"reason": ..., "outcome": ..., "exit": ...}, followed by whatever the
operation actually said (a refused claim's alternatives). Exit 2 ("nothing") keeps its
declared keys after that lead; a result that fills its declared shape is left as the CLI's
--json prints it, with the reason in the second content block.
Two tools exist so an agent can orient itself without being told: ddflow_help (what
is this, what is the loop) and ddflow_workflow (what are the rules here).
ddflow adopt writes it as a managed block between and. Your own prose around it is preserved; re-running updates only
what is inside. If you write it by hand, four things have to be in it:
ddflow_brief (or ddflow brief in a shell).ddflow_next → ddflow_claim → work in the worktree
it creates.ddflow_gate_status → satisfy each gate → ddflow_complete →
ddflow_merge.2 is not success.Without that block an agent sees the tools and has no reason to reach for them before editing. The block is what makes the queue authoritative rather than optional — and it is 232 words, because an instruction file nobody finishes reading is one nobody follows.
$ ddflow workflow
# The workflow this project runs
1. research agent
2. rules agent
3. implement agent (required)
4. lint command (required, NOT proven able to fail)
$ ruff check .
...
## The rules, and where each came from
gates.require_outcome True [default]
gates.enforce_order block [file]
schedule.max_parallel_tasks 4 [default]
One answer to "what are the rules here": every gate in order, which are commands and which you perform, which are required, which need evidence, which need a different-family reviewer, which have been proven able to fail — plus the completion rules, the caps, the reviewers, and where each value came from, so a deliberate choice is distinguishable from a default nobody touched.
$ ddflow workflow gate lint --command "ruff check ." --into task --after implement --required
$ ddflow workflow pipeline task research,implement,lint,unit_tests,merge
$ ddflow workflow drop dedupe
All four reach MCP — ddflow_workflow, ddflow_workflow_pipeline,
ddflow_workflow_gate, ddflow_workflow_drop — so an agent can change the workflow
with the operator's agreement. Their descriptions say to ask first and offer
dry_run, because a pipeline governs every future item, not the one in hand.
Nothing is written until it is checked, and the order is the point: compose the change, validate the result, then replace the file atomically.
gate record rejects the id as unknown — so the item can never be
completed at all, and nothing says why.[gatez] is valid TOML
and used to be written happily, breaking every later command — the write path
validated the merged text for syntax and then validated the config already on
disk, which is a writer checking the state it is replacing.required too, or it becomes a requirement that
quietly requires nothing.ddflow workflow and ddflow doctor both re-run those checks against what is on
disk. Everything is a file you can also edit by hand: gates in [gate.<id>], reviewers
in [[reviewer]], companions in .ddflow/companions.toml, and every prompt —
including the instructions your agent receives at connect — under .ddflow/prompts/.
One caveat with MCP: the connection instructions are computed once, when the server
starts. A workflow changed mid-session is live for every tool call immediately, but the
text the agent was handed is stale. Tell it to call ddflow_workflow, or restart.
The append-only event log is the source of truth; everything else is a projection that
can be deleted and re-derived. The SQLite index, the markdown boards, the search index,
the recovery bundle — all disposable, all rebuilt by ddflow rebuild.
That single inversion is what makes the four hard properties fall out for free rather than needing to be engineered:
| You get | Because |
|---|---|
| Two agents on two branches never conflict | Each appends to its own file. Measured: a real two-branch merge resolves clean. |
| A crashed agent loses nothing | State is folded, never written. Nothing is half-updated. |
| The project rebuilds from the log | Operator prompts are events. |
| An edited history is detectable | Event ids are content addresses. |
The design decisions, with the probes that settled each, are in docs/RESEARCH.md. The two that most shaped it:
One line in your agent's MCP config. Nothing else.
Until the first release is on PyPI,
uvx ddflow-mcphas nothing to fetch: install from the repository and runddflow adopt, which registers that installation instead, as in A first run.
{ "mcpServers": { "ddflow": { "command": "uvx", "args": ["ddflow-mcp"] } } }
uvx fetches and runs the published package in an ephemeral environment on first use —
no clone, no virtualenv, no PYTHONPATH, no install step for an operator to forget, and
no vendored copy to drift from upstream. ddflow needs one runtime dependency
beyond python3 and git (Jinja2), which is what lets it install inside
sandboxes, CI images and other tools' ephemeral containers.
Then, from the agent, with no shell at all:
| Call | What it does |
|---|---|
ddflow_setup | creates .ddflow/, writes the driver and the AGENTS.md section |
ddflow_configure with toml: '[gate.unit_tests]\ncommand = "pytest -q"' | sets your test command |
ddflow_reviewers_detect with write: true | finds a local model server and registers it as a cross-family reviewer |
ddflow_phase_add, ddflow_task_add | fill the queue |
ddflow_brief | start every session here |
That is the whole adoption. The per-project instruction text is 232 words — a
managed block in AGENTS.md, because the MCP tool descriptions already carry the
how, and a second copy of that would drift from the one the model actually reads.
uv tool install ddflow-mcp # or: pipx install ddflow-mcp
cd /path/to/your/project
ddflow adopt # every supported agent
ddflow adopt --agents claude,cursor,vscode,kimi # or name the ones you use
After upgrading ddflow, ddflow doctor notes any driver doc (implement-phase.md, an
adopted agent's delta) that differs byte-for-byte from the template the running ddflow
ships. ddflow adopt --refresh-docs (MCP: ddflow_setup with refresh_docs) rewrites
only those docs, the AGENTS.md/CLAUDE.md blocks and the agents' native rules -- never the
MCP launch, hooks or command files -- and refuses a project that was never adopted.
adopt is idempotent and writes managed blocks, so re-running after an upgrade updates
them and leaves your own prose alone. It writes the MCP registration into each agent's
own config location, merged with whatever servers are already there. From a source
checkout it points the config at that checkout instead of the published package, so
developing ddflow does not silently configure your project against the released
version.
{ "mcpServers": { "ddflow": { "command": "docker", "args": [
"run", "-i", "--rm",
"-v", "${workspaceFolder}:/repo",
"--add-host=host.docker.internal:host-gateway",
"ghcr.io/delian/ddflow-mcp:latest" ] } } }
ddflow adopt --launch docker writes exactly that. The image is 107 MB (Alpine;
ddflow is pure standard library, so there is no compiled dependency to worry musl
about) and behaves identically on Linux, macOS and Windows.
Four things go wrong when a containerised tool touches a bind-mounted git repo. All four are silent, one of them loses work, and all four are handled:
| Trap | What it looks like | Handled by |
|---|---|---|
| Worktrees land outside the mount | worktree.root defaults to ../.ddflow-worktrees, a sibling of the repo. In a container only the repo is mounted, so worktrees go to the ephemeral layer and are destroyed on exit with the agent's uncommitted work inside them. | container.default_worktree_root relocates a sibling root to .ddflow-worktrees inside the repo, and adopt gitignores it |
| Root-owned files | On a Linux bind mount the operator needs sudo to edit their own project afterwards | the entrypoint reads the mount's uid/gid and su-execs down to it |
| git refuses the mount | "detected dubious ownership", surfacing as an unexplained ddflow failure | safe.directory set in the entrypoint |
| No git identity | git commit fails with "Please tell me who you are" | entrypoint prefers GIT_AUTHOR_*, then the repo's own config, then a clearly-marked placeholder |
And one that cannot be fully handled, so it is reported: 127.0.0.1 inside a
container is the container. A model server on your own machine is not reachable from
there. ddflow rewrites loopback reviewer URLs to host.docker.internal, and ddflow doctor tells you that on Linux you must also pass
--add-host=host.docker.internal:host-gateway, because unlike Docker Desktop the Linux
engine does not provide that name.
The related portability fix: worktree paths are stored in the event log relative to the repo root. The log is committed and shared, so an absolute path is true only on the machine that wrote it — false for a teammate who cloned elsewhere, for CI, and for a container where the repo is
/repo. Pinned bytest_the_event_log_carries_no_absolute_paths.
Every prompt is an external template, resolved config → project → shipped:
ddflow prompts list # where each template currently comes from
ddflow prompts eject # copy the shipped ones into .ddflow/prompts/
$EDITOR .ddflow/prompts/review_system.md
Adding a mode of your own: [[macro]]. Overriding a shipped workflow needs no code,
and neither does adding one. A macro is a named, parameterised prompt — "enter debugger
mode" — that appears everywhere the shipped workflows do: prompts/list and prompts/get
over MCP, which is what a client turns into a slash command, and ddflow prompts list|show
in a terminal.
# .ddflow/config.toml (or .ddflow/macros.toml, if you prefer to split it out)
[[macro]]
name = "debugger"
title = "Enter debugger mode"
description = "Reproduce first, then bisect. No fix without a failing probe."
params = ["symptom"] # required, not optional
tools = ["ddflow_bug_found", "ddflow_gate_run", "ddflow_bug_fixed"]
prompt = """
You are debugging: {{ symptom }}
Reproduce it before you theorise. Paste the command and its output.
"""
Use prompt_file = "docs/modes/debugger.md" instead for anything long enough that TOML
quoting gets in the way.
When to reach for a macro rather than a gate. A gate is a step every item passes through, recorded against that item and blocking its completion. A macro is a MODE an operator enters, belonging to no item and recorded nowhere — "audit this release", "handle this incident". If the thing should hold up a task until it is done, it is a gate; if it is a way of working you want to name and re-enter, it is a macro. Putting a mode in the pipeline makes every task wait for something that was never about that task.
tools is declarative, not a sandbox. It is rendered into the prompt as the ordered
set the mode expects, so the agent is told what the mode is for and the next reader can
tell what it was supposed to do. It does not restrict what the agent may call — MCP has
no mechanism for that, and claiming a security property this cannot honour would be worse
than not having it. This is the deliberate departure from
dx-zero/mcpn, whose toolMode: situational lets the
model pick freely from a bound set with no recorded ordering: a session you cannot replay
is a session you cannot review, which is the property the event log exists to give you.
Three things a macro refuses, because each alternative fails quietly: a missing
parameter (a prompt with a hole in it reads as a complete instruction), a name that
belongs to a shipped command (silent shadowing leaves you editing a block that does
nothing), and both prompt and prompt_file (two sources for one body means one is
dead and looks live). A refused macro is refused alone and by name — the others still
load — and ddflow prompts list, prompts show, doctor and the MCP prompts/list say
which one and why; an undecodable prompt_file is a named problem too.
Including the one the agent actually reads first. mcp_instructions.md is the block
an MCP client injects into the model's context on connect — the workflow, the reporting
duties, and which companion tools to reach for. It is the file to edit when you want
this project to work differently:
ddflow prompts eject mcp_instructions
$EDITOR .ddflow/prompts/mcp_instructions.md # or [prompts] mcp_instructions = "..."
It renders against the live state — adopted, task_pipeline, setup_todo,
companions, missing_companions, gate_gaps, recoverable, loops — so the
instruction is the next concrete action rather than a fixed blurb the model learns to
skip. A broken override says so in the instruction block itself instead of falling
back to the default: this is the one surface where nobody would ever notice their edit
was not live.
Templates render with Jinja2, which is ddflow's one runtime dependency, and with a
strict standard-library renderer when it is absent — a stripped deployment with no
reachable package index still starts. The shipped templates use the subset both engines
agree on, and tests/test_template_engines.py walks the template REGISTRY, rendering
every entry through both engines and asserting the outputs are byte-identical.
That test is iterated rather than hand-listed for a reason. Its predecessor named three
templates in a dict, mcp_instructions.md was never added, and in 0.1.1 the largest and
most important template rendered correctly under Jinja2 and failed under the fallback —
so the entire MCP handshake for an unadopted repository, the first thing a new user
ever sees, degraded to ddflow's instruction template could not be loaded. Jinja2 was
not a declared dependency at the time, so developers had it and the project venv did
not: python -m pytest was green and uv run pytest was red on the same commit.
The fallback now raises on any construct it does not implement rather than copying
it through. The old regex engine emitted what it could not parse, so a condition as
ordinary as {% if a or b %} — which its single-name pattern never matched — reached
the client as literal template source.
Both renderers are strict about undefined variables: a prompt silently missing the diff it was supposed to carry is the vacuous review in template form — the model dutifully reviews nothing and reports no findings.
The rest is TOML: gates and their pipelines ([gate.*], gates.task_pipeline),
reviewers ([[reviewer]]), companions ([[companion]]), enforcement ([enforce]),
cadences, and the rest of the 141 knobs.
ddflow config --set <key> <value> edits one key in place, preserving comments.
ddflow recommends services; it never ships one person's configuration. Two layers:
| Layer | Files | Holds |
|---|---|---|
| committed | .ddflow/config.toml, .ddflow/gates.toml | generic project policy: the test command, the pipelines, the gates (human checkpoints included) — what every clone must agree on |
| machine-local, git-ignored | .ddflow/local/config.toml, .ddflow/local/gates.toml, .ddflow/local/reviewers.toml | your services: reviewer endpoints, model names of a private deployment, API-key variable names, LAN hosts, a worker count sized to this machine |
The local files are read last, so they win; ddflow config --explain reports such a
value's source as local. Every writer of an operator-specific value targets the local
layer by default:
ddflow reviewers add --preset ollama --model qwen3:8b # -> .ddflow/local/reviewers.toml
ddflow reviewers detect --write # -> .ddflow/local/reviewers.toml
ddflow config --local --set gate.unit_tests.command "pytest -q -n 48"
ddflow config --local --append-toml "$(cat my-reviewer.toml)"
--shared on reviewers add / reviewers detect --write commits the block to
.ddflow/config.toml instead — only for a service every clone reaches at the same
address. config --set and --append-toml stay committed unless you pass --local
(over MCP: ddflow_configure with local=true, ddflow_reviewers_detect with
shared=true), because a test command or a pipeline is project policy. The human-gate
guards hold on both layers: no writer sets gate.<id>.human, and none can drop a human
gate from a pipeline. .ddflow/local/ carries its own * .gitignore, so it stays
uncommitted even in a project whose .ddflow/.gitignore predates it. API keys are never
written anywhere — only the name of the variable that holds one.
Nobody should have to paste JSON into an IDE to use this. server.json is the
MCP registry manifest
(io.github.delian/ddflow-mcp), and publishing it is what makes ddflow findable in the
VS Code and Cursor marketplaces rather than something you configure by hand. It offers
three ways to run the same server, so a client picks whichever it supports:
| Package | Identifier | For |
|---|---|---|
pypi | ddflow-mcp, runtimeHint: uvx | Anything with uv — no clone, no install step |
oci | docker.io/delian/ddflow-mcp:<version> | Operators with no Python toolchain |
oci | ghcr.io/delian/ddflow-mcp:<version> | The same image, no Docker Hub account needed |
The image is built for amd64 and arm64, because an Apple-silicon operator running it
under emulation pays that cost on every tool call, and tool calls are all this server
does.
How CI authenticates — four mechanisms, one stored secret:
| Target | Mechanism | Stored secret? | Setup |
|---|---|---|---|
| PyPI | OIDC trusted publishing (id-token: write) | No | Add a trusted publisher on PyPI, once |
| ghcr.io | GITHUB_TOKEN, injected per run, expires with the job | No | none |
| Docker Hub | DOCKERHUB_USERNAME + DOCKERHUB_TOKEN | Yes | Create an access token, add both secrets |
| MCP registry | GitHub OIDC — proves control of the account that owns the io.github.delian/* namespace | No | none |
| tag + release | GITHUB_TOKEN (contents: write) | No | none |
Docker Hub is the only one that needs a long-lived credential, because it has no OIDC
equivalent. Use an access token scoped to this repository, never an account password. If
that is one secret too many, delete the Docker Hub login and its two tags — ghcr.io alone
satisfies the OCI entries a marketplace needs, and server.json lists both so a client
picks whichever resolves.
environment: release on the publishing jobs is a control worth knowing about: point it
at a GitHub environment with required reviewers and every release waits for a human,
with no change to the workflow.
Order matters and the workflow encodes it. mcp-publisher validates that every
package named in the manifest exists, so the registry step runs after both PyPI and
Docker — publishing the manifest first would advertise a version nobody can fetch.
The registry verifies ownership, and the workflow checks it first. The registry accepts
a package only when the artefact itself names the server: the PyPI package's README
carries `` (the first lines of this file),
and each image carries the label io.modelcontextprotocol.server.name with the same name.
The verify job runs tests/test_registry_ownership.py first, so a release the registry
would reject — a missing marker, or a label that does not match — is refused before
anything is uploaded; a server.json description over the registry's 100-character limit
fails the suite the same way. So do the registry's per-package rules, which the JSON schema
does not express and the registry enforces only at publish: an oci package must not
carry a version field (or registryBaseUrl/fileSha256) — the tag in its identifier is
the version — while the pypi package must carry one. publish #40 failed on exactly that, at
the last job, after PyPI and both images had shipped; tests/test_registry_manifest_rules.py
now encodes the rules offline and verify runs it first. The publish step also retries a transient registry failure
(a 504, a 408/429, a network error) with backoff, checking the registry for the exact
version after every attempt, because the publish behind a 504 may have committed; a 4xx
that retrying cannot fix fails at once, and only after the budget does it fail with a
"Re-run failed jobs" hint.
Four things gate a release, and each exists because the failure it catches is public and irreversible:
ddflow/__init__.py (the one declared version), server.json's version and
every OCI identifier's tag must agree — a :0.1.0 left behind while version moved on publishes a manifest
pointing at the previous image, installable and wrong;-m 'not slow' otherwise
excludes from every ordinary run;uv build succeeding proves the metadata parses, not that ddflow help works;initialize over stdio. A built image that cannot is a broken
release every marketplace will happily offer.$ git push origin main # that is the whole release
Every push to main that changes shipped code releases, with the PATCH version bumped.
Major and minor move only when you move them.
"Shipped" means ddflow/, pyproject.toml, uv.lock, Dockerfile,
docker-entrypoint.sh, .dockerignore, server.json, server.template.json or README.md (the PyPI page, and
the line the registry verifies): a push of other docs, tests or the ddflow event log
releases nothing, because PyPI keeps every version forever and one identical to the last
is noise nobody can withdraw. CI runs scripts/bump.sh patch, commits release 0.1.2 to
main, publishes PyPI, Docker Hub, ghcr.io and the MCP registry, then creates v0.1.2 and
a GitHub release — last, and only once every publish succeeded, because a tag pointing
at a half-release is worse than no tag: it looks authoritative.
Pull after a release. The release commit is CI's, so your main is one commit behind
it and the next push is refused until you git pull.
How the gate picks the version: it publishes the declared version if PyPI does not have it yet, and bumps patch only if it does. So moving major or minor is yours to do —
$ scripts/bump.sh minor # 0.1.4 -> 0.2.0 (or: major, or an exact 1.0.0)
$ git commit -am 'release 0.2.0' && git push origin main
— and CI publishes exactly 0.2.0; the next push that bumps nothing releases 0.2.1.
The same rule makes a release that failed before its PyPI upload retry its number with
the next push. One that failed after it — Docker Hub, ghcr.io, the MCP registry — does
not: PyPI has the number, so the next push moves past it; re-run that run's failed jobs
from the Actions page instead. A bump never lands on a number PyPI already holds (a v*
tag can publish one out of band): CI skips to the next free patch before it commits
anything. The bump is pushed to main before anything publishes: a push that
loses a race with another commit fails the run and publishes nothing, where pushed last it
would leave PyPI holding a version main does not declare. Runs are serialized, and only
main or a v* tag releases.
The version is declared in one place: __version__ in ddflow/__init__.py, the only
line a bump edits. pyproject.toml reads it (a hatch dynamic version, so uv.lock records no
version for the project and a bump never makes the lock stale), SERVER_INFO (what the
server tells every client
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx ddflow-mcpMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-delian-ddflow-mcp": {
"command": "uvx",
"args": [
"ddflow-mcp"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceddflow works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.