Local-first RAG engine with MCP server for AI agent integration.
A local-first RAG engine that ingests documents, indexes them with BM25 + dense embeddings, and exposes search via an MCP server for AI agent integration.
Built by BrightDotDev.
pip install darwin-rag
# Interactive — detects hardware, pick your models
darwin-admin setup interactive
# Or one-shot (embedding-only, no prompts)
darwin-admin setup --preset required
# stdio mode — for AI agent subprocess (Claude Desktop, Cursor, etc.)
darwin mcp
# Or HTTP mode — for remote clients
darwin mcp --http --port 8765
Configure your MCP client:
{
"mcpServers": {
"darwin": {
"command": "darwin",
"args": ["mcp"],
"env": {
"OPENAI_API_KEY": "sk-..." // At least one LLM provider key
}
}
}
}
Or generate config automatically:
darwin config claude # Claude Desktop config
darwin config cursor # Cursor config
darwin config all --copy # All clients + copy to clipboard
| Tool | Description |
|---|---|
search_darwin | Query the knowledge base with hybrid/semantic/keyword search |
search_lists | Search structured data (CSVs, JSON arrays) by field values |
get_search_results | List saved search results |
get_search_result_by_id | Load a saved search result by filename |
get_schema | Inspect schemas for structured files (keys, types, record counts) |
run_pipeline | Ingest + index documents from a path or URL |
purge_artifacts | Delete pipeline artifacts for specific files |
create_store | Create a new isolated data store |
list_files | List all tracked files with pipeline status |
file_status | Detailed status for a single file across all stages |
get_logs | Query session logs (oldest first, INFO excluded) |
Full documentation: docs/mcp.md
Start the server on a network-accessible endpoint:
# SSE transport (legacy)
darwin mcp --sse --host 0.0.0.0 --port 8765
# Streamable HTTP transport (recommended for production)
darwin mcp --http --host 0.0.0.0 --port 8765
Configure your MCP client with the URL:
{
"mcpServers": {
"darwin": {
"url": "http://your-host:8765/mcp" // or /sse for SSE mode
}
}
}
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY | No* | OpenAI provider key |
ANTHROPIC_API_KEY | No* | Anthropic provider key |
GEMINI_API_KEY | No* | Google Gemini provider key |
MISTRAL_API_KEY | No* | Mistral AI provider key |
GROQ_API_KEY | No* | Groq provider key |
COHERE_API_KEY | No* | Cohere provider key |
TOGETHER_API_KEY | No* | Together AI provider key |
OPENROUTER_API_KEY | No* | OpenRouter provider key |
DEEPSEEK_API_KEY | No* | DeepSeek provider key |
DARWIN_BASE_DIR | No | Override the base data directory |
NO_COLOR | No | Set to any value to disable ANSI color output |
* At least one LLM provider key is required for answer generation. Search/indexing works without any.
| Command | What it does |
|---|---|
darwin-admin setup interactive | Guided setup — detect hardware, choose models |
darwin-admin setup --preset required | Download embedding model only (fastest) |
darwin-admin setup --preset recommended | Embedding + reranker + OCR models |
darwin-admin setup logging | Reconfigure logging only |
darwin-admin setup validate | Validate current setup |
See docs/setup.md for the full walkthrough including Docker, from-source install, and API key configuration.
darwin — User CLI
| Command | Description |
|---|---|
darwin mcp | Start MCP server (stdio, --sse or --http for network) |
darwin config [client] | Generate MCP client config |
darwin-admin — Power-user CLI
| Command | Description |
|---|---|
darwin-admin setup | Setup models, logging, and configuration |
darwin-admin status | System status overview |
darwin-admin models | Model registry: list, install, switch, keys |
darwin-admin store | Data store: status, files, audit, health, repair |
darwin-admin pipeline | Ingestion pipeline: run, ingest, index, purge |
darwin-admin search | Interactive search |
darwin-admin logs | Structured log viewer and management |
darwin-admin system | System information |
darwin-admin uninstall | Remove Darwin data and configuration |
See docs/admin.md for the full command reference.
For embedding darwin-rag as a library in your own app:
from core import Darwin
d = Darwin()
d.ingest("./papers", recursive=True)
results = d.search("what is this paper about")
records = d.search_records(filters={"status": "active"})
Full reference: docs/api.md
| Doc | What |
|---|---|
| setup.md | Full setup walkthrough |
| mcp.md | MCP server, tools, resources, transports |
| admin.md | Admin CLI reference |
| architecture.md | For developers and contributors |
| pipeline.md | Ingestion & indexing |
| retrieval.md | Search engine |
| storage.md | DarwinStore |
| models.md | Model registry & inference |
| logger.md | Structured logging |
| orchestrators.md | High-level business logic |
| api.md | Python API (Darwin class) |
Found a bug? Want to add something? You're welcome here.
Read CONTRIBUTING.md for the full guidelines.
MIT with Attribution — see LICENSE.
Core architecture and implementation by BrightDotDev.
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx darwin-ragMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-brightdotdev-darwin-rag": {
"command": "uvx",
"args": [
"darwin-rag"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceDarwin RAG works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.