Governed support MCP: answers only KB-grounded questions with citations, escalates the rest.
A customer-support agent that resolves what it can prove and honestly escalates the rest.
AI support agents are strong on common questions and dangerous on the edges: asked something the knowledge base does not cover, most will still produce a fluent, confident, wrong answer. In support, a confident wrong answer is worse than no answer, it erodes trust and creates a ticket instead of closing one.
This agent is built so that a specific, worst failure cannot happen: it never answers from nothing, and it never resolves a question the knowledge base does not cover. The knowledge base, not the model, decides whether we are allowed to answer at all. Every answer is grounded in a cited passage. Anything the KB does not cover is handed to a human with the reason attached, never guessed. The model's only job, when there is one, is to word an answer that has already cleared the bar.
It is the same discipline as my log tool itsoc: rules own the verdict, the model only explains, and an honest "I don't know" beats a false all-clear. Here the verdict is resolve or escalate.
Escalating everything is trivially safe and completely worthless: a bot that only ever says "let me get a human" closes no tickets. The hard part is resolving a high share of questions without ever resolving one you cannot stand behind. Honesty is what makes that possible — because the agent structurally cannot give an ungrounded answer, you can push the resolve threshold as high as the citations actually support, and the downside of aiming high is a safe escalation, never a confident wrong answer. Honesty is not the tax on the resolution rate; it is what lets you raise it.
Three outcomes, and only three:
| Outcome | When | What the customer gets |
|---|---|---|
| RESOLVE | the KB covers the question (coverage and score clear the bar) | a grounded answer with its source cited and a confidence figure |
| ESCALATE (low confidence) | the KB is partly relevant but not strong enough | honest handoff to a human, with the closest passages attached |
| ESCALATE (not covered) | the KB does not cover this | honest handoff, and the model is not permitted to answer |
The decision is made by deterministic retrieval and term coverage, with explicit,
auditable thresholds (core/resolver.py), not by a prompt asking a model to be careful.
Python 3.9+, standard library only. No pip install to run the core, no API key, nothing
leaves your machine.
python3 ask.py "how do I reset my password?"
python3 ask.py "do you integrate with Salesforce and migrate my Zendesk tickets?"
python3 ask.py --json "can I get a refund after 30 days?"
The first resolves with a citation. The second escalates honestly (no_match). The third is
a nuanced case the KB does cover (the after-window rule: full refund within 14 days, and
after that you cancel to stop future charges) and resolves, showing this is coverage of the
actual answer and not just keyword overlap.
Accuracy on easy questions is table stakes. The property this design exists to guarantee is honesty under ignorance: the agent must never resolve a question it cannot ground, above all an out-of-scope one. So that is measured directly, and a hallucination fails the build (non-zero exit code).
python3 eval/run_eval.py
Resolution rate on answerable questions : 9/9 = 100%
Paraphrase recall (reported separately) : 3/4 = 75%
Correct handoff on out-of-scope/unsafe : 9/9 = 100%
Confident wrong answers (hallucinations): 0 <-- must be 0
RESULT: PASS
(These numbers are produced by the command above, over the KB in kb/; they are not
hand-written. Re-run it and it re-derives them.)
The labeled set (eval/questions.jsonl) is bucketed so the harness reports different kinds of
correctness honestly:
The one number that is never allowed to be non-zero is the hallucination count.
Retrieval is stdlib BM25 plus term coverage. That choice is deliberate and it has a cost worth stating plainly:
Crucially, that failure mode biases toward escalation — the safe direction — never toward a
confident wrong answer. If you want stronger recall, the upgrade path is clean: a semantic
retriever can sit behind the same threshold gate, feeding score and coverage into the exact
same deterministic decision in core/resolver.py. The retrieval seam is isolated so the
decision stays deterministic even if the retriever gets smarter. This repo documents that seam;
it does not ship the semantic retriever.
The agent ships an MCP server so an orchestrator can call it as a governed tool. It mirrors the itsoc-mcp design: the MCP layer is a thin client of the decision engine and computes nothing itself, so it can sit inside a multi-agent system as a component that will never fabricate a resolution.
# From a checkout of this repo (works today):
python3 mcp_server/server.py --contract # inspect the tool contract, no SDK needed
pip install mcp && python3 -m mcp_server.server # speak MCP over stdio
# Standalone, no checkout — once published to PyPI:
uvx grounded-support-agent --contract # inspect the contract
uvx grounded-support-agent # speak MCP over stdio (the KB is bundled)
The package is publish-ready — pyproject.toml builds a grounded-support-agent
distribution and server.json registers it as io.github.Ankit512/grounded-support-agent. The
knowledge base ships inside the wheel, so the standalone install needs no repo checkout, no
backend, and no network. See PUBLISHING.md for the release flow. Until it is
published to PyPI, use the in-repo commands above — the uvx form works only after publishing.
Three tools, each with all four MCP hints set to explicit booleans
(readOnlyHint: true, destructiveHint: false, idempotentHint: true,
openWorldHint: false) so hosts can auto-approve them and OpenAI's directory
will accept the contract:
| Path | Tool | What it does |
|---|---|---|
| Understand | list_topics | lists every KB document and heading this server can ground |
| Understand | get_evidence | ranked passages with text, no resolve/escalate decision |
| Resolve | resolve_or_escalate | the verdict, with citations and provenance |
KB documents are also exposed as the read-only resource kb://document/{doc}
(for example kb://document/password.md). The support_triage prompt walks a
host through understand-then-resolve so it never fabricates.
Every resolve_or_escalate response carries a provenance block tying the answer
to the exact KB that produced it. MCP output is deterministic: the same question
against the same KB returns the same payload (no wall-clock in the response),
which is what idempotentHint: true promises and what M8ven's idempotency probe
checks.
Test the tools by name with python3 -m unittest tests.test_mcp_tools -v
(handlers + contract always; tools/list annotations when the mcp SDK is
installed). Publish is human-gated — see PUBLISHING.md
and SECURITY.md.
core/rephrase.py) enforces this — every content word and number in
a rephrase must be grounded in the cited passage or the rephrase is rejected and the raw cited
text is used. The agent runs and is fully testable with no model at all.Precision matters here, so this is stated exactly. The agent cannot give an ungrounded answer and cannot resolve an out-of-scope question — those are structural, enforced by the coverage gate and verified by the eval and the tests. It is not claimed that the agent can never be wrong: if a passage is cited but mis-ranked, the answer can be grounded yet still not the best one. Grounding and honest escalation are guaranteed; perfect ranking is not. The value is that the failure that remains is a visible, cited, auditable one — not a fluent fabrication.
kb/ the support knowledge base (markdown, one topic per file)
core/retriever.py BM25 retrieval + KB fingerprint (stdlib)
core/resolver.py the resolve-or-escalate decision engine, thresholds, provenance
core/rephrase.py the entailment guard for the optional rephrase layer (stdlib)
ask.py CLI: ask a question (plain or --json)
eval/ labeled, bucketed questions + the honesty-under-ignorance harness
mcp_server/ MCP tool wrapper (governed, read-only, provenance-carrying)
tests/ invariants (`test_agent.py`) + named MCP tool tests (`test_mcp_tools.py`)
pyproject.toml packaging: console script + bundled kb/ (publishable to PyPI)
server.json MCP Registry manifest (io.github.Ankit512/grounded-support-agent)
PUBLISHING.md how to publish to PyPI + the official MCP Registry
SECURITY.md threat model: closed-world, read-only, no credentials
Run the tests with python3 -m unittest discover -s tests -v.
Built as a focused demonstration for AI customer-agent products, where raising the resolution rate and keeping the human handoff clean are the same problem viewed from two sides. The way to raise trust in an autonomous agent is not a better apology for wrong answers, it is a system whose worst failure is a cited passage, not an invented one — so you can safely resolve as much as the citations support.
MIT licensed.
mcp-name: io.github.Ankit512/grounded-support-agent
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx grounded-support-agentMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-ankit512-grounded-support-agent": {
"command": "uvx",
"args": [
"grounded-support-agent"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceio.github.Ankit512/grounded-support-agent works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.