Typed, calibrated and certified decisions over a context, with abstention and evidence.
Decisions your agents can act on. MIMIR is a non-generative decision model: give it a context, a question and the options, and get back a typed answer with calibrated probabilities, the evidence behind it, and a certified verdict on whether to act or escalate. No text generated. Nothing to parse. Nothing to hallucinate.
It beats Laya and GLiNER2.5-Decide head-to-head on six of ten tasks — by 54.7 points on Banking77, 42.3 on MASSIVE, 36.3 on typed decisions — and where it cannot back an answer, it abstains instead of guessing.
pip install "mimir-decisions[local]" # CPU engine
pip install "mimir-decisions[local-gpu]" # CUDA engine
pip install mimir-decisions # data models and HTTP client only
Python 3.11+. Documentation: https://abderahmane-ai.github.io/mimir/
The shipped policy certifies the fp32 CPU configuration; the CUDA (fp16) graph ships without a policy in this release.
Most agents route, classify, and verify using a general-purpose language model: slow, expensive, and impossible to audit. MIMIR is built for structured decisions. It runs on ONNX Runtime in milliseconds, returns calibrated probabilities with every answer, and issues a mathematical certificate — a formal guarantee that its realised error rate stays at or below the risk level you ask for, measured on held-out data.
from mimir import Mimir
model = Mimir.from_pretrained("Mythologic/MIMIR-1")
result = model.choose(
"My card was charged twice for the same order.",
"Which team should handle this ticket?",
options={"billing": "Billing: payments, refunds", "security": "Security: account access"},
)
result.status # Status.DECIDED, Status.ABSTAINED or Status.DEFERRED
result.answer # an option id, or None when no option applies
result.probabilities # calibrated probability of each option id
result.certificate # the certified threshold the decision was checked against
answer is the model's prediction. status is the policy's verdict:
DECIDED — the answer is an option and is certified at the requested risk level.ABSTAINED — no listed option applies, and that is certified.DEFERRED — not certified; result.deferral.reason is below_threshold, out_of_distribution, or no_certified_threshold.The first call downloads the model from the Hugging Face Hub at the revision this package version pins, verifies its Sigstore signature, checks every file against the manifest's SHA-256, and loads it.
| Spec | Answer |
|---|---|
Choice(question, options) | an option id, or None |
MultiChoice(question, options) | the option ids that apply |
YesNo(question) | True or False |
Verify(claim) | supported, contradicted, or not_enough_information |
Rank(question, candidates) | candidate ids, best first |
Rate(question, levels) | a level id; levels given lowest first |
Estimate(question, low, high, unit) | a number in [low, high], with a confidence interval |
from mimir import Context, Field, Passage, Rate, Table
context = Context(
passages=[Passage(title="Ticket #4412", text="The export has failed every night this week.")],
tables=[Table.from_rows([["2026-03-02", "failed"]], header=["date", "status"])],
fields=Field.from_json({"customer": {"plan": "enterprise", "seats": 240}}),
)
result = model.decide(context, Rate("How urgent is this?", ["low", "medium", "high"]), risk=0.01)
A context can be a string, a list of strings, a dict read as a JSON state, or a Context of typed passages, tables, and fields. Table.from_dataframe(frame) reads a pandas or polars DataFrame. decide_many batches multiple decisions, and every method has an async counterpart (adecide, adecide_many, …).
decide takes a risk level certified by the loaded policy (model.info().risk_levels). A decision is taken only when its calibrated confidence clears a threshold certified on held-out data to keep the realised error rate at or below that risk with 95% confidence, and when the context passes the out-of-distribution gate. decide_uncertified returns the raw model answer with no policy applied.
A certificate covers one exact configuration: model files, variant, ONNX Runtime version, execution provider, and options. On hardware not listed in the certificate, the first load runs the release's equivalence set and requires every decision to match. To certify thresholds on your own labelled data:
mimir calibrate labelled.jsonl --risk 0.01 --confidence 0.95 --out policy.json
model = Mimir.from_pretrained("Mythologic/MIMIR-1", policy="policy.json")
from mimir import MimirClient
remote = MimirClient("https://mimir.internal", api_key="...")
remote.choose("...", "Which team?", options=["billing", "security"])
MimirClient has the same interface as Mimir, so all code, decision tools, and framework adapters accept either. It requires only the base install. Connection errors, timeouts, and 429/502/503/504/529 responses are retried with exponential backoff that honours Retry-After.
from mimir import Choice
route_ticket = model.tool(
"route_ticket",
Choice("Which team should handle this ticket?", options=["billing", "security"]),
description="Route a support ticket to the team that owns it.",
)
route_ticket("My card was charged twice")
route_ticket.input_schema, route_ticket.output_schema
Tools can also be declared in a YAML file, which the HTTP and MCP servers load:
tools:
- name: route_ticket
description: Route a support ticket to the team that owns it.
decision:
type: choice
question: Which team should handle this ticket?
options: [billing, security]
A tool-call check decides, against rules you write, whether an agent's pending tool call may run. A certified yes allows it, a certified no denies it, and anything else escalates to a person.
check = model.tool_call_check(
["Refunds above 500 dollars need a manager's approval."], tools=["issue_refund"]
)
outcome = check("issue_refund", {"order": "4412", "amount": 900})
outcome.permission # Permission.ALLOW, Permission.DENY or Permission.ESCALATE
outcome.reason # one sentence for the agent or the approver
Each adapter turns decision tools into the framework's native tool type and wires a tool-call check into that framework's own approval hook.
| Framework | Install | Tools | Tool-call check |
|---|---|---|---|
| OpenAI Agents SDK | mimir-decisions[openai-agents] | as_function_tool | guard: escalations pause the run for approval |
| LangChain / LangGraph | mimir-decisions[langchain] | as_structured_tool | ToolCallCheckMiddleware: escalations interrupt with the human-in-the-loop request |
| PydanticAI | mimir-decisions[pydantic-ai] | as_toolset | guard: escalations end the run with DeferredToolRequests |
| CrewAI | mimir-decisions[crewai] | as_crewai_tool | tool_call_hook: escalations go to your approver |
| Google ADK | mimir-decisions[adk] | as_adk_tool | tool_call_callback: escalations ask for ADK confirmation |
| Microsoft Agent Framework | mimir-decisions[agent-framework] | as_function_tool | ToolCallCheckMiddleware: only certified calls run |
| LlamaIndex | mimir-decisions[llamaindex] | as_llamaindex_tool | none |
| smolagents | mimir-decisions[smolagents] | as_smolagents_tool | none |
from agents import Agent
from mimir.integrations.openai_agents import as_function_tool
agent = Agent(name="support", tools=[as_function_tool(route_ticket)])
Every framework also reaches MIMIR through its own MCP client. examples/ has a native, an MCP, and a checked agent for each framework, plus a Vercel AI SDK agent in TypeScript.
pip install "mimir-decisions[local,server]"
MIMIR_API_KEYS=key-one,key-two mimir serve --host 0.0.0.0 --tools tools.yaml
| Route | Does |
|---|---|
POST /v1/decide | one certified decision: {context, decision, risk, alpha} |
POST /v1/decide/uncertified | the model's raw answer: {context, decision} |
POST /v1/decide/batch | up to 64 decisions in one call |
POST /v1/tools/{name} | a tool from --tools, given only {context} |
POST /v1/systemone | Jev's request and response format |
GET /v1/models | model, revision, runtime and certified risk levels |
GET /healthz, GET /readyz | liveness, and readiness once the model is loaded |
GET /metrics | Prometheus metrics |
Concurrent requests are batched. With keys in MIMIR_API_KEYS, every route except the probes requires Authorization: Bearer <key>. A server with no keys listens only on loopback unless started with --allow-no-auth. The OpenAPI 3.1 document is openapi.json.
Each configured tool becomes an MCP tool that takes only a context; --generic-tools adds mimir_choose, mimir_verify, mimir_rank, and mimir_rate. A deferred decision is a normal result telling the agent to escalate.
uvx --from "mimir-decisions[local,mcp]" mimir-decisions mcp --tools tools.yaml # stdio
MIMIR_API_KEYS=... mimir mcp --http --host 0.0.0.0 --tools tools.yaml # Streamable HTTP at /mcp
mimir mcp --tools tools.yaml --remote https://mimir.internal # forward to a server
mimir serve --mcp --tools tools.yaml # HTTP API and /mcp together
In Claude Code:
claude mcp add mimir -- uvx --from "mimir-decisions[local,mcp]" mimir-decisions mcp --tools /path/to/tools.yaml
claude mcp add --transport http mimir https://mimir.internal/mcp --header "Authorization: Bearer ..."
Claude Desktop, Cursor, and VS Code take the same command or the same URL and header in their MCP configuration. The server is registered in the MCP Registry as io.github.abderahmane-ai/mimir.
docker run -p 8000:8000 -e MIMIR_API_KEYS=... -v mimir-models:/models ghcr.io/abderahmane-ai/mimir:1.0.0-cpu
docker run --gpus all -p 8000:8000 -e MIMIR_API_KEYS=... -v mimir-models:/models ghcr.io/abderahmane-ai/mimir:1.0.0-cuda
Images carry the runtime, never the model weights. On first start, the model is downloaded at the revision the package version pins, verified, and cached in /models. To run from that cache with no network access, append serve --host 0.0.0.0 --model-cache /models --offline.
Images are signed with Sigstore by the release workflow:
cosign verify ghcr.io/abderahmane-ai/mimir:1.0.0-cpu \
--certificate-identity https://github.com/abderahmane-ai/mimir/.github/workflows/release.yml@refs/heads/main \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
| Command | Description |
|---|---|
mimir serve | the HTTP server; --mcp also serves MCP at /mcp |
mimir mcp | the MCP server, over stdio or --http |
mimir decide | one decision from flags, or a JSON request on stdin |
mimir bench FILE | accuracy, coverage and realised risk on labelled decisions |
mimir calibrate FILE | certify thresholds on labelled decisions |
mimir schema | JSON Schemas of every spec, result and request |
mimir download | download and verify a release for offline use |
mimir doctor | report the environment; --verify loads the model and runs the equivalence check |
Releases are loaded from a pinned Hugging Face revision. Before any model file is read, the manifest's Sigstore signature is verified against the abderahmane-ai/mimir release workflow, every file is checked against the manifest's SHA-256, and the ONNX graph is checked against its operator allowlist and signature. No pickle is used anywhere.
mimir.compat.systemone.v1 converts Jev /v1/systemone requests and responses, and mimir.compat.laya.v1 exposes load(...).predict(state, questions) in Laya 0.3.20's shape. See the migration guides for step-by-step instructions.
The MIMIR SDK is licensed under Apache-2.0; that license covers the software only, not the MIMIR model weights. The MIMIR-1 model weights are licensed separately under the MIMIR Model License.
Eligible community users may use MIMIR-1 commercially without royalties, subject to the MIMIR Model License. Anyone may download, benchmark, evaluate and prototype with MIMIR without registering.
Organizations exceeding the Revenue Threshold (US$1,000,000 annual gross revenue), or requiring enterprise, OEM, redistribution, hosting, or other additional rights, may obtain a commercial agreement from Mythologic. See COMMERCIAL-LICENSING.md.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx mimir-decisionsMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-abderahmane-ai-mimir": {
"command": "uvx",
"args": [
"mimir-decisions"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceMIMIR works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.