ask-fable

MCP server for Anthropic's Claude Fable, Opus 5, and multi-model reasoning councils.

AI & MLPythonv0.13.0

ask-fable: Multi-Model Reasoning MCP Server

ask-fable is a portable, installable MCP (Model Context Protocol) server for AI coding agents. It works in Claude Code, OpenCode, Kimi Code, Grok, Cursor, Codex, and any other harness that can spawn a local MCP server.

It gives those agents guarded code and architecture reasoning from Anthropic's Claude Fable (the newest claude-fable-*), Claude Opus 5 (claude-opus-5), MiniMax (MiniMax-M3), Gemini, Codex, GLM, DeepSeek, Grok, Kimi, and Ollama Cloud models. It can query one backend, synthesize a parallel council, run an ordered refinement chain, or stage a structured adversarial debate.

Fable and Opus 5 use Claude Code's existing OAuth session (through the Agent SDK, with the claude CLI as a fallback). MiniMax, Gemini, Codex, Grok, and local Ollama similarly reuse authenticated local CLIs. GLM, DeepSeek, and Atlas Cloud are optional HTTP backends that need server-side API keys.

Start here

If you need to…Use
Ask one trusted coding model, with follow-up memoryask (Fable) / ask_opus5 (Opus 5)
Compare independent answers in parallelask_council
Draft, critique, then decide in orderask_chain
Stress-test a high-impact decisionask_debate
Grind a claim down to what survives evidenceask_falsify
Brainstorm an open question, models arguing to divergenceask_conference
Select a task-matched Atlas Cloud modellist_atlas_models → ask_atlas
Atlas council with GPT-5.6 Sol adjudicatingask_atlas_council
Reuse large code context without pasting it againcontext_write + context_ref
Investigate a request after it rantrace_list + trace_get

Start with ask for one hard question. Escalate to a council, chain, or debate only when the decision warrants the extra latency and cost.

What it gives you

ask-fable gives an MCP client six ways to reason:

ModeWhat happensBest for
AskOne model answers directly; Fable can remember a sessionEveryday debugging and design questions
CouncilSeveral models answer in parallel; Fable reconciles themComparing independent opinions
ChainModels work in order: draft → critique → decideDeliberate refinement and cost-tiered escalation
DebateA proposer and opponent test claims; Fable adjudicatesContentious, hard-to-reverse decisions
FalsifyClaims are asserted, attacked, and resolved by a code clerk; the ledger persists across callsGrinding a checkable claim down to what receipts actually support
ConferenceModels argue together over rounds; a rapporteur maps the disagreementOpen-ended ideation

The same guard, context bus, cache, audit trail, and tracing layer wrap every mode. Backends are optional: use Fable alone, call a specific provider, or mix Fable, Opus 5, MiniMax, Gemini, Codex, Grok, GLM, DeepSeek, Kimi, Ollama, Atlas Cloud, and OpenRouter. Unavailable council members are reported and skipped instead of failing the whole request.

A real example — ask_debate, lazy token bucket vs. background refill task for a per-user rate limiter (resolution: adjudicated):

Use the lazy token bucket. Do not build the background refill task — the timer only approximates at tick granularity what the lazy design computes exactly.

The debate surfaced traps neither side opened with (a 100 req/min bucket permits ~199 requests in a worst-case rolling minute; TTL eviction alone doesn't bound memory) and closed with four ship-it fixes. More real calls, one per mode: docs/EXAMPLES.md.

The cheapest real second opinion is the twin token — the twin flames. It expands to both Anthropic reasoners at once, Fable + Claude Opus 5, and both ride the same OAuth session as ask, so a two-model cross-check costs you no provider keys and no extra setup:

ask_council(models=["twin"])        # or tier="twin" — the pair, in parallel
ask_chain(pipeline="m3 > twin")     # cheap draft, then fable → opus in turn

Five features make the result useful to an agent, not just readable by a human:

  • Structured sidecar — every answer carries a machine-readable sidecar ({recommendation: apply|investigate|reject|needs_more_context, confidence, needs_context}) next to the prose, so an agent acts on it directly. When the model needs more, a followup tells it exactly what to paste, and a per-session terminator stops an unbounded re-ask loop (status:"context_exhausted").

  • Context bus — context_write a big codebase context ONCE under a key, then pass context_ref on any ask tool (or council) to pull it in instead of re-pasting. Shared by every agent on the server; context_read / context_list / context_delete round it out.

  • Council consensus — councils return a consensus signal (strong | partial | divergent | unknown) + material_disagreement computed from the panel's recommendations, each sources entry shows that model's recommendation, and the synthesis is anonymized (Expert A/B, Fable last) to blunt self-preference bias.

  • Correlated traces — every call includes a trace_id; inspect the ordered request timeline without storing raw prompts in the default safe mode.

  • Session hub — successful turns from local MCP instances are mirrored into a shared, visibility-only dashboard. Agents can use the same label to coordinate work without that shared history ever becoming model context.

How it works

A request enters through MCP, resolves any reusable context_ref, passes the guard, and is routed to the chosen reasoning mode. The result is normalized into an answer plus a machine-readable sidecar, persisted to the configured observability stores, and returned with a trace ID.

The project ships its own two-layer request gate: a size/sanity floor followed by a prohibited-use denylist. Fable's model prompt adds the final semantic scope contract. See The guard for the exact behavior.

The guard

Every question is checked before any model call:

  1. Sanity floor — rejects only empty / too-short (<3 chars) / too-long (>65536 chars) questions. Context is unbounded by default (any cap you set is floored to 512,000 chars). Breadth is allowed.
  2. Prohibited-use denylist — ask-fable's bundled offensive-security and biology dual-use patterns. Extend it via ASK_FABLE_DENYLIST_FILE (one term per line). Benign multi-word phrases (e.g. request payload) are neutralized before matching so an ambiguous word like payload used in an ordinary engineering sense doesn't false-trip; add your own via ASK_FABLE_ALLOWLIST_FILE (one phrase per line). This only rescues the exact benign phrase — a bare prohibited term still rejects.
  3. Model scope contract — Fable answers engineering questions, including conceptual/brainstorming ones with no code context (breadth is fine), and replies REFUSED: <reason> only when the question itself directly asks for offensive-security work (exploit development, attack tooling) or non-software domain knowledge (e.g. biology). Questions about security-related code are normal engineering.

Every decision is appended to an owner-only JSONL audit log (question hashed by default; ASK_FABLE_AUDIT_RAW=1 to store raw).

Quick start

1. Install

New here? The setup & usage guide walks through install, registering in Claude Code (OpenCode, Kimi Code, Grok, and other MCP clients use the same server — see below), setting up every backend (API keys, Ollama Cloud, MiniMax/Gemini CLIs), /mcp verification, and how to use every tool.

Want the big picture? The visual architecture map charts the whole server end to end — the request pipeline, the oracle bridges, council/chain orchestration, and on-disk state.

# not on PyPI yet — install from source:
pip install -e .
# or with pipx:
pipx install .

Requires the Claude Code CLI to be installed and logged in (that's the OAuth session Fable is reached through).

2. Register in your coding harness

ask-fable is a local stdio MCP server (ask-fable on PATH). Point any MCP-capable coding harness at it; only the config-file shape changes. Restart the harness after editing — most load MCP servers once at startup. All 40 ask_fable tools then become available. They are grouped into reasoning modes, direct provider calls, context management, configuration, and observability; see the tool guide for the short chooser or CLAUDE.md for the complete one-line inventory.

Claude Code — ~/.claude/.claude.json

Add to ~/.claude/.claude.json (root-owned — edit as the owner, e.g. via sudo):

{
  "mcpServers": {
    "ask_fable": { "command": "ask-fable" }
  }
}

(or "command": "python3", "args": ["-m", "ask_fable"]).

OpenCode — ~/.config/opencode/opencode.json

The docs/OPENCODE.md guide covers the full setup — the exact schema-valid MCP block, optional API keys, the restart-to-load behavior, and troubleshooting. Minimal registration:

{
  "mcp": {
    "ask_fable": {
      "type": "local",
      "command": ["ask-fable"],
      "enabled": true
    }
  }
}
Kimi Code — ~/.kimi-code/mcp.json
{
  "mcpServers": {
    "ask_fable": {
      "transport": "stdio",
      "command": "ask-fable",
      "toolTimeoutMs": 600000
    }
  }
}

Kimi Code's default MCP request timeout is ~60s; oracle calls often run longer. toolTimeoutMs keeps the host from aborting a still-running call. A Request timed out error from the client is that transport timeout, not a refusal — check trace_list before re-asking.

Grok — ~/.grok/config.toml
[mcp_servers.ask_fable]
command = "ask-fable"
enabled = true

Cursor, Codex, and other MCP clients take the same ask-fable command; only the config file shape differs.

3. Ask a question

In your MCP client, call ask with a focused question and the relevant code or error. Reuse the same session key for follow-ups:

{
  "question": "Why does this cache invalidate too early?",
  "context": "<relevant code and failing test output>",
  "session": "cache-investigation"
}

Tool guide

The server exposes 40 MCP tools, but you only need six entry points — ask, ask_council, ask_chain, ask_debate, ask_falsify, and ask_conference. Everything else selects a specific backend, manages reusable context, or inspects what ran.

GoalStart withEscalate when
Solve or debug one problemask (Fable) or ask_opus5 (Claude Opus 5 — ~half the price, faster)use context_ref for large reusable context
Get one alternate opinionask_m3, ask_deepseek, ask_glm (cheap direct APIs first), ask_gemini, ask_codex, ask_grok, ask_kimi, ask_ollama, ask_atlas, or ask_openrouter (~400 models, one key)use a council when you need comparison
Pick an Atlas model for a tasklist_atlas_models(task="…")call ask_atlas with the accepted selection or rendered picker
Cross-check with a second strong modelask_council(models=["twin"]) — Fable + Opus 5 on one OAuth session, no keysadd a third voice with models=["twin","m3"]
Compare several viewsask_counciluse ask_chain when order matters
Cross-check Atlas models, GPT adjudicatingask_atlas_councilpin the panel with configure_atlas_council
Make a contentious decisionask_debatekeep the scope narrow; it is the most expensive mode
Prove a claim before acting on itask_falsifypack the corpus it must cite; reuse the same session to compound evidence
Brainstorm an open questionask_conferenceraise rounds (default 3, up to 10) when a dilemma needs more back-and-forth
Inspect what happenedtrace_list then trace_getenable full mode only when redacted content is needed

Full reference: docs/TOOLS.md — every tool, its arguments, and when to reach for it. A one-line inventory of all 40 lives in CLAUDE.md.

Observability & response shape

Every answer carries a machine-readable sidecar ({recommendation, confidence, needs_context}) and a trace_id; councils add a consensus signal and debates a deterministic resolution. Results are cached, progress streams to the console, and all persisted state lives under a per-user state dir with owner-only permissions.

Details: docs/OBSERVABILITY.md — the full response contract, caching, console progress, and backend setup.

Configuration

Everything is optional environment variables set in the server's env block, with sensible defaults.

Reference: docs/CONFIGURATION.md — every setting grouped by backend, guard, storage, and observability.

Recommended agent instructions

The server injects a short standing instruction so agents reach for these tools unprompted. But weak local models under-attend to system prompts, so for the best results also drop a decision ladder into your project's CLAUDE.md / AGENTS.md / opencode.md (agents re-read those). Copy this block:

## Using ask_fable (external reasoning)
Reach for the ask_fable MCP tools on the hard 5% — cheapest option first:

1. **Answer it yourself** for trivial, low-blast-radius, or already-in-context work.
2. **Double-strike rule:** the moment you've failed the SAME bug/error twice, STOP
   and call `ask` before a third guess. Include what you tried and the exact error.
3. **`ask`** (single Fable, multi-turn) for a real design trade-off, a subtle bug
   hypothesis, "am I reasoning about X right?", or a change spanning >2–3 files.
   Reuse the `session` key for follow-ups on the same problem.
4. **`ask_council`** only for a contentious or hard-to-reverse decision
   (architecture, concurrency, data model, public API, migration). One council
   call per problem, max. Check `quorum`/`degraded` and `consensus` in the result —
   a `1/N` answer (or a `divergent` panel) is not agreement. Reach for **`ask_chain`**
   instead when you want *ordered* refinement rather than a parallel vote — e.g. a
   cheap model drafts and Fable finalizes, or draft → red-team → decide.
5. **Reuse context:** for a big codebase context you'll ask about repeatedly,
   `context_write` it once and pass `context_ref=<key>` — don't re-paste each time.
6. **Recommend it, don't just skip it:** if one of these tools would clearly help
   but you're not calling it, say so in one line — which tool and why — so the
   operator can opt in.

Frame questions tightly: paste the real code + real error (don't paraphrase), state
ONE specific decision (ideally A-vs-B), and the constraints. Act on the result's
`sidecar.recommendation`; if you get a `followup`, paste exactly what it names (but
check `likely_already_pasted` and re-read your own paste first) and re-ask on the same
`session`. If tests or a linter can verify the answer, run them instead of asking again.

Companion skills

skills/ ships four skills that drive these tools from Claude Code, OpenCode, Grok, Kimi Code, and other skill-capable harnesses (copy or symlink into ~/.claude/skills/, ~/.agents/skills/, or the harness equivalent):

  • ubercode — treat Fable (and, via ask_council, MiniMax-M3) as a smarter reasoning partner for the hard 5%: oracle escalation when you're stuck, and cross-checked adversarial review before a high-consequence diff.
  • uberplan — fan out N diverse candidate plans locally, use Fable as a comparative judge (optionally cross-checked with ask_council), then synthesize one final plan.
  • uberarch — open-ended architectural ideation: fan abstract ideas out to the oracles (ask_council / ask_chain) for multi-model trade-off analysis before any code exists.
  • uberbrainstorm — design-first, approval-gated brainstorming for the fuzzy front end ("what should we build and why"), with the council red-teaming the chosen design; hands off to uberplan.

Development

uv sync --extra dev           # or: uv pip install -e '.[dev]'
uv run pytest -q              # 841 tests, no network needed
uv run ruff check src tests

salient-core (a richer prohibited-use denylist) is unpublished and therefore not declared as an extra; the guard picks it up automatically at runtime if it is installed in the environment.

Documentation

License

MIT

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
uvx ask-fable

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-baggybin-ask-fable": {
      "command": "uvx",
      "args": [
        "ask-fable"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

ask-fablepypi

Compatible MCP Clients

ask-fable works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More