Back to Directory/Cloud Providers

io.github.HighlyLoadedEgo/clef-mcp

Local MCP server for the Cloudflare Clef-Flash decision model: structured decisions, fully offline

Cloud ProvidersTypeScriptv0.3.0
clef_decide — probability distributions for a production incident

clef-mcp

npm version npm downloads CI Official MCP Registry License: Apache-2.0

A "reflex" for AI coding agents: structured decisions with probabilities — not prose.

Local MCP server that gives agents (Claude Code, Codex, Cursor, ZCode, …) access to the Clef-Flash decision model (9B, Apache-2.0 by Cloudflare) through a single tool: clef_decide. Pass a state and typed questions, get a probability distribution over your options in one forward pass. Fully local, offline, no tokens burned.


Quick start

# 1. Detect hardware, download the model (~6 GB) + llama.cpp runtime, verify checksum + inference
npx clef-mcp install

# 2. Register the MCP server + agent skill in your clients (zcode, claude-code, codex, cursor)
npx clef-mcp setup

# 3. Run the MCP server (stdio)
npx clef-mcp

clef-mcp (step 3) never downloads anything. If the model is missing, tool calls return a structured MODEL_NOT_INSTALLED error with a hint. Steps 1 and 2 combine: npx clef-mcp install --setup.

Why not just ask the LLM?

Chat LLMclef_decide
Outputprose, you parse itstrict JSON: probability per option
Determinismvaries per runsingle forward pass, no sampling
Latency (1 decision)seconds of generation~0.5 s local
Context costgrows with every decisionfixed, small schema
Privacydepends on provider100% on-device, works offline
Calibrationvibessoftmax over trained option scores

Sweet spot: decision points inside agent loops — next action, routing, classification, severity, yes/no judgment — asked dozens of times per task.

The clef_decide tool

{
  "state": {
    "task": "Fix failing tests",
    "error": "TypeError: Cannot read properties of undefined"
  },
  "questions": {
    "next_action": {
      "type": "choice",
      "instructions": "What should the coding agent do next?",
      "criteria": {
        "inspect": "Inspect the code and gather more information",
        "modify": "Modify the code",
        "test": "Run additional tests",
        "ask_user": "Ask the user for clarification"
      }
    },
    "confidence": {
      "type": "score",
      "instructions": "How confident are you in this decision?",
      "criteria": ["very_low", "low", "medium", "high", "very_high"]
    },
    "is_outage": { "type": "noul", "instructions": "Is a service down?" }
  }
}
TypeCriteriaAnswer
choicemap option id → description, or a plain listprobability per option
scoreordered list (index = score)probability per level
nouloptional {"true": "...", "false": "..."}{"true": p, "false": 1-p}

Response — strictly structured, never prose. Each decision carries the model-reported confidence (when the runtime sends it), plus token usage for the call:

{
  "model": "clef-flash",
  "decisions": {
    "next_action": { "answer": { "inspect": 0.72, "modify": 0.12, "test": 0.14, "ask_user": 0.02 }, "confidence": 0.83 },
    "confidence":  { "answer": { "very_low": 0.01, "low": 0.04, "medium": 0.18, "high": 0.61, "very_high": 0.16 }, "confidence": 0.61 },
    "is_outage":   { "answer": { "true": 0.9, "false": 0.1 } }
  },
  "usage": { "input_tokens": 228, "output_tokens": 0, "latency_ms": 512 }
}

Act on the argmax only when the distribution is decisive — top p ≥ 0.8 and high confidence for destructive or security-adjacent calls.

Batch up to 64 questions per call — they are scored in one forward pass. state is treated strictly as data: never executed, never interpreted as instructions for the server.

Prompts & resources

The server ships four MCP prompts (canned, decision-shaped asks — your client lists them via prompts/list):

PromptPurpose
incident-triageaction + severity + user-impact questions for a production incident
next-actionwhat the coding agent should do next + confidence
ticket-routingclassify a message into a team + urgency
security-reviewvulnerability yes/no, risk scale, first mitigation

And three resources (read-only, no model needed):

URIContents
clef-mcp://capabilitieslive JSON: model, runtime, limits, error codes
clef-mcp://evals/schemahow to write eval cases
clef-mcp://evals/datasetthe bundled 30-case dataset

Measured, not marketed

Apple M4 Pro, Clef-Flash Q4_K_M (6 GB), single request through the full MCP stdio path:

ScenarioLatency
Cold start (incl. model load, once per session)~4.4 s
1 question~0.5 s
10 questions, one call~2.9 s
64 questions, one call~18.6 s

Quality gate: a 30-case evaluation dataset (coding / security / classification / routing / yes-no) — 86.7% pass on the live model. Run it yourself: clef-mcp evals.

Register with your MCP client

Claude Code
claude mcp add clef-mcp -- clef-mcp
# or, without a global install:
claude mcp add clef-mcp -- npx -y clef-mcp
Codex — ~/.codex/config.toml
[mcp_servers.clef-mcp]
command = "clef-mcp"
args = []
Cursor — .cursor/mcp.json
{
  "mcpServers": {
    "clef-mcp": { "command": "clef-mcp", "args": [] }
  }
}
ZCode — ~/.zcode/cli/config.json (user scope, auto-connect)
{
  "mcp": {
    "servers": {
      "clef-mcp": { "command": "clef-mcp", "args": [], "type": "stdio" }
    }
  }
}

Ready-made snippets: examples/.

Scripting & hooks

No MCP client required — hooks, CI jobs and shell scripts call the same model one-shot:

# Full document on stdin
echo '{"state": "checkout 500s after deploy", "questions": {"is_outage": {"type": "noul", "instructions": "Is a service down?"}}}' \
  | clef-mcp decide

# Or split across files
clef-mcp decide --questions questions.json --state state.json

stdout carries the strict JSON result (same shape as the MCP tool, including confidence and usage); errors go to stderr as structured JSON with exit codes: 2 invalid input, 3 model not installed, 4 runtime missing. decide never downloads anything.

Each plain decide invocation is a cold start (model load included, a few seconds) — fine for gates and triage. For repeated calls, start clef-mcp daemon once: it keeps the model warm on a permission-scoped unix socket in CLEF_HOME (no TCP port, unloads after CLEF_DAEMON_IDLE seconds, default 600), and clef-mcp decide --daemon answers in well under a second, falling back to a cold run when no daemon is running. See examples/hooks/ for a PreToolUse guard and a GitHub Action recipe.

The PreToolUse guard blocks a command when the model judges it destructive (p ≥ 0.9) and tells the agent to ask the user — a block is a pause plus escalation, not a wall; the gate is advisory by design and says so. First real firing on day one: caught a history-rewrite force push (p=0.94) and surfaced its own bypass vector, which is now fixed and documented in the recipe.

Teach your agent (skill)

The schema tells the client what clef_decide accepts; agents also need to know when to reach for it and how to frame decisions. The bundled clef-decisions skill covers decision patterns, batching, criteria writing, distribution interpretation and error recovery:

clef-mcp setup                                                    # automatic
cp -r skills/clef-decisions ~/.agents/skills/                     # manual, from repo
cp -r "$(npm root -g)/clef-mcp/skills/clef-decisions" ~/.agents/skills/  # from npm package

Architecture

flowchart LR
    subgraph clients [MCP clients]
        CC[Claude Code]
        CX[Codex]
        CU[Cursor]
        ZC[ZCode]
    end
    clients -- MCP stdio --> S[clef-mcp<br/>validation · limits · structured errors]
    S -- SystemOne adapter --> R[ClefRuntime<br/>llama.cpp subprocess<br/>127.0.0.1]
    R -- single forward pass --> M[("Clef-Flash<br/>9B · GGUF · local")]
    M -. probabilities .-> S -. strict JSON .-> clients

The ClefRuntime interface (load / decide / unload / health) isolates the engine: MLX or remote runtimes plug in without changing the MCP API. The wire format is POST /v1/systemone — the same contract across llama.cpp and other Clef runtimes.

CLI

clef-mcp              # run the MCP server on stdio (default command)
clef-mcp decide       # one-shot decision (no MCP session): JSON in, JSON out — for hooks, CI, scripts
clef-mcp daemon       # keep the model warm on a local unix socket; `decide --daemon` uses it
clef-mcp install      # detect hardware → download model + runtime → verify checksum → verify inference
clef-mcp setup        # register the MCP server + agent skill in zcode / claude-code / codex / cursor
clef-mcp models       # list models/quantizations and install status
clef-mcp status       # runtime, model, memory summary
clef-mcp doctor       # full diagnosis (platform, RAM, GPU, binary, model, checksum*, inference, MCP config)
clef-mcp uninstall    # remove the model (and optionally the managed runtime)
clef-mcp evals        # run the evaluation dataset against the installed model

Flags: install --quant Q8_0 --yes --skip-probe, install --setup, setup --clients zcode,cursor --no-skill, doctor --deep (re-hash the model file), uninstall --runtime --yes.

Runtimes: llama.cpp and MLX

Two local runtimes behind the same ClefRuntime interface:

llama-cpp (default)mlx
PlatformsmacOS, Linux, WindowsmacOS / Apple Silicon only
ModelGGUF from ggml-org/Clef-Flash-GGUFMLX 4-bit from mlx-community/clef-flash-4bit
Extrasnoneuv on PATH (managed Python env)
Installclef-mcp installclef-mcp install --runtime mlx

Switch at runtime with CLEF_RUNTIME=mlx (must be set for the MCP server process — e.g. in the client's env block). Both speak the same POST /v1/systemone contract. The MLX snapshot is fetched into CLEF_HOME via a uv-managed huggingface_hub (no global Python state) at a pinned revision.

Configuration

VariableDefaultMeaning
CLEF_MODELclef-flashModel id (per-call model also accepted)
CLEF_HOME~/.cache/clef-mcpCache/model home
CLEF_RUNTIMEllama-cppllama-cpp | mlx
CLEF_LOG_LEVELerrorerror | warn | info | debug (stderr only)
CLEF_LLAMA_BIN–Explicit llama-server binary path (llama-cpp runtime)
CLEF_LLAMA_RELEASE_TAGlatest nightlyPin the managed llama.cpp build
CLEF_LLAMA_BATCH8192llama.cpp physical batch (multi-question requests)
CLEF_MLX_UVuv on PATHExplicit uv binary (MLX runtime)
CLEF_DAEMON_IDLE600Seconds of idle before the daemon unloads the model (0 = never)
CLEF_MAX_QUESTIONS64Max questions per call
CLEF_MAX_STATE_BYTES1048576Max serialized state size
CLEF_MAX_INSTRUCTION_CHARS10000Max chars per question instructions

Runtime resolution: CLEF_LLAMA_BIN → managed binary in CLEF_HOME/runtime → llama-server on PATH.

Storage: CLEF_HOME/models/<model>/<quant>/ (model + manifest.json with repo/revision/sha256/license) and CLEF_HOME/runtime/llama.cpp/.

The model is downloaded from the pinned official GGUF conversion (ggml-org/Clef-Flash-GGUF) and sha256-verified against Hugging Face's content hash. It is never repackaged by clef-mcp. Also listed in the official MCP Registry as io.github.HighlyLoadedEgo/clef-mcp.

Error handling

{
  "error": {
    "code": "MODEL_NOT_INSTALLED",
    "message": "Clef model \"clef-flash\" is not installed.",
    "hint": "Run `clef-mcp install`."
  }
}

Codes: MODEL_NOT_INSTALLED, MODEL_LOAD_FAILED, RUNTIME_NOT_FOUND, RUNTIME_INIT_FAILED, RUNTIME_NOT_SUPPORTED, INVALID_INPUT, CLEF_INFERENCE_FAILED, UNSUPPORTED_PLATFORM, OUT_OF_MEMORY, CHECKSUM_MISMATCH, DOWNLOAD_FAILED. Input exceeding the 16k-token model context is rejected with a hint to reduce the state.

Security & data handling

  • No network servers, no telemetry, no accounts; everything runs locally.
  • The model downloads only on an explicit install, over HTTPS, checksum-verified.
  • state content is passed to the model as data; the server never executes or instruction-interprets it.
  • Filesystem access is limited to CLEF_HOME (plus reading standard MCP client config paths in doctor).
  • The managed runtime is the official llama.cpp build; pin it with CLEF_LLAMA_RELEASE_TAG.

See SECURITY.md for the full policy.

Development

npm install
npm run build
npm test          # unit + integration (fake llama-server, no model needed)
npm run evals     # needs an installed model; exit code reflects pass rate

See CONTRIBUTING.md and tests/evals/dataset.jsonl.

License

  • Code: Apache-2.0.
  • Clef / Clef-Flash model: © Cloudflare, Apache-2.0 — see NOTICE.
  • llama.cpp runtime: © its authors, MIT-licensed; downloaded as an official prebuilt binary.

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
npx -y clef-mcp

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-highlyloadedego-clef-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "clef-mcp"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

clef-mcpnpm

Compatible MCP Clients

io.github.HighlyLoadedEgo/clef-mcp works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More