Gmail MCP server for email triage: search, labels, archive, trash, unsubscribe, snooze.
A reliable, native Gmail MCP server — full mailbox triage for AI assistants, with the feature no other Gmail MCP server ships: mailbox-side snooze.
threads.list — the call any thread search goes through — can
answer is:unread from a stale thread-level read state: measured in one real mailbox, 86% of
the threads it returned held no unread message at all; in a second mailbox, no drift whatsoever.
You cannot tell which mailbox you are in without looking, so search re-verifies every hit against
its live labels. Paginated via pageToken/nextPageToken. (the measurements)get_thread reports the SPF/DKIM/DMARC results
the receiving server recorded, so "is this really from my bank?" is answered from the message's
own headers instead of from how the domain looks. It reads the receiving server's report only —
a message can carry forged ones of its own — and says unchecked when nobody checked, because a
missing check is not a passing one.bulk_modify archives/labels everything matching a query at
1000 messages per API request — with per-chunk partial-success reporting instead of
all-or-nothing. The snooze sweep uses the same batch path.outputSchema and returns validated
structuredContent alongside fenced JSON text — no parsing guesswork for clients. Failures are
structured as well: a code and a retryable flag, so a client can tell "try again later" from
"re-authorize" without reading prose.messages.send, every draft endpoint, permanent deletion and forwarding
settings, whatever a compromised or careless caller asks for. One deliberate exception:
unsubscribe / bulk_unsubscribe (manage tier) contact the opt-out endpoint named in a message's
own header — the only non-Google host mailwarden ever reaches, and a read-tier deployment makes
no outbound request at all. Details under
Security & privacy and
Unsubscribing.=?UTF-8?B?…?= → readable text),
bodies decoded in their declared charset (no mojibake for ISO-8859-1/Shift_JIS mail),
429/5xx retried with exponential backoff.Connectors that sync or cache your mailbox can lag behind it — and even Gmail's own search index is sometimes loose (see below). mailwarden talks straight to the live Gmail API (no cached snapshot) and re-verifies what the index returns, so what you see is what's actually there. It's a generic Gmail capability layer — keep your own rules/logic in your AI client, not in the server.
search goes one step further than the raw API: Gmail's threads.list index can answer read-state operators from a stale copy of that state, so is:unread returns threads you finished reading weeks ago — in one measured mailbox, the large majority of what came back. Since every hit is fetched live anyway, search re-checks the unambiguous predicates (is:unread/is:read/is:starred/in:inbox/category:…, with negation) against each thread's true labels and drops the index's false positives.
Most Gmail MCP servers cover the same read/label/send surface. Three capabilities are still unique to mailwarden (mailbox-side snooze, search re-verification, sender authentication), and one deliberate omission is a security feature, not a gap. Google's own server is also narrower than it looks: draft-only, and no trash, filters or unsubscribe.
| Capability | mailwarden | Google official | taylorwilsdon | a-bonus | aaronsb | klodr |
|---|---|---|---|---|---|---|
| Mailbox-side snooze — archive now, resurface in the inbox on a date/time or preset | ✅ | — | — | — | — | — |
| Search-result re-verification — drops the thread index's false positives against live labels | ✅ | — | — | — | — | — |
| Sender authentication — SPF/DKIM/DMARC as the receiving server recorded them, on every message | ✅ first header only, values token-validated | — | — | — | — | — |
| Sweep / bulk over a query — one action across every thread a search returns | ✅ 1000/req, partial-success | — | ⚠️ batch by explicit ids | — | ⚠️ batch by explicit ids | ⚠️ batch by explicit ids |
| Unsubscribe — per-sender overview + RFC 8058 one-click opt-out, no send scope needed | ✅ | — | ⚠️ header shown, no action | — | — | — |
| Inbox triage overview — one call that buckets what is waiting | ✅ sender/label/age + header signals | — | — | ✅ heuristic flags + stats | ⚠️ unread list, no buckets | — |
| Server-side filters — rules that keep triaging with no assistant in the loop | ✅ never forwarding | — | ✅ | — | — | ✅ |
| No send tools — by design — a prompt-injected mail has no exfiltration path | ✅ no compose at all | ⚠️ draft-only | ❌ sends | ❌ sends | ❌ sends; ⚠️ opt-in draft-only policy | ❌ sends |
| Least-privilege tool tiers — OAuth scopes derived from the tools you enable | ✅ | ⚠️ scope split | ⚠️ --permissions narrows scopes per service; tiers narrow tools only | — | ⚠️ per-account read/readwrite, at call time | ⚠️ inverse: tools gated by granted scopes |
| Token encryption at rest | ✅ AES-256-GCM, opt-in (MAILWARDEN_TOKEN_PASSPHRASE) | n/a (hosted) | ⚠️ file mode 0600; bucket CMEK on GCS | — | — | — |
| No vendor cloud — you operate the server | ✅ | ❌ Google-hosted | ✅ | ✅ | ✅ | ✅ |
Structured outputs — every tool declares an outputSchema | ✅ | — | — | — | — | ⚠️ one tool (download_email), more planned |
Snapshot as of 3 September 2026 — the oldest of the five column checks, and so the only date the table as a whole can claim. Three columns were read that day against that day's state of each project: the two repositories by diff against the revision recorded in docs/comparison-sources.json, Google's by re-reading the tool reference. The taylorwilsdon column is newer — its least-privilege cell was corrected on 7 September 2026 at 54b1c56, and the column was brought forward again on 18 September 2026 to e844d805, 73 commits on, by reading the diffs: the Gmail changes are all on the sending path (bare newlines in an HTML body, a signature wrapper) and then stop altogether — gmail/ does not appear in the sixteen most recent commits at all, and auth/permissions.py, auth/scopes.py and core/tool_registry.py appear in neither stretch, so the least-privilege cell holds on identical code. — = not offered / not documented. Columns are the servers a reader is most likely to reach for — Google's first-party one, plus the three largest community servers still under maintenance — and klodr, which comes closest to mailwarden's own least-privilege design. aaronsb was added on 8 September 2026, and the reason it was missing is worth stating plainly: the selection had been made by capability overlap, so a server an order of magnitude larger than the smallest column sat outside the table while that column stayed in it. It earns its place on the send row rather than on size. It is the only server here besides mailwarden that will refuse to send at all — GWS_SAFETY_POLICY=draft-only-email blocks send, reply, replyAll and forward — and the difference that keeps the row honest is where the promise lives: there, sending is the default and the refusal is an operator's environment variable; here, no tool to send was ever written. Its no-delete policy stands in the same relation to this project's trash-only rule. The filter row asks the same question from the other end, and a third project shows why it is one: c0webster/hardened-google-workspace-mcp, a fork of the taylorwilsdon column made for no reason but this one, removes create_gmail_filter and delete_gmail_filter outright because a filter can forward — which is exactly the risk, answered by dropping the capability. mailwarden keeps the capability and removes the risk instead: create_filter has no way to express a forward, and list_filters surfaces the ones that already exist so they can be audited. That fork also shows where a promise of this kind usually leaks, which is why the row above is about the tool surface and the egress guard rather than about scopes: it drops gmail.send but keeps gmail.compose — commented "for draft creation/editing only (NOT sending)" — and gmail.modify. Google documents the first as "Manage drafts and send emails" and the second as "Read, compose, and send emails", so the promise does not rest on one scope too many but on two, and gmail.modify alone would be enough — the same scope this project holds in its own manage tier, which is exactly why the row points at the tool surface. Its token can send; only its tool list cannot. The most installed Gmail server is absent for that reason and not by oversight: GongRzhe/Gmail-MCP-Server is archived, its last commit dating to August 2025, and it still drew 112,163 npm downloads in the month to 29 August 2026. Reach and currency are different questions, and a comparison of what a server does today can only answer the second. Send capability is listed as a security property: mailwarden's lack of it is intentional (see Security & privacy). The encryption row asks who holds the key: mailwarden encrypts the token itself from a passphrase you set — and does nothing without one, which is why the cell says opt-in rather than showing a bare tick; taylorwilsdon relies on file permissions locally and on the storage bucket's own CMEK when hosted on GCS — protection against a stolen file in the first case, against a stolen disk in the second. The last row asks who operates the server, not where it happens to run: self-hosting is common ground here, and every community server on this table offers some remote deployment except klodr (stdio only) — mailwarden via --http, taylorwilsdon over streamable HTTP with OAuth 2.1, a-bonus on Cloud Run. Running one of them on your own host is not a cloud copy; running it on the vendor's is.
The moat isn't any single row — it's snooze + live re-verification together: an actual inbox-workflow layer that acts on the mailbox's current state, not a cached snapshot. Where others have caught up it's noted honestly above: at-rest encryption (taylorwilsdon), scope-driven tool gating (klodr inversely; taylorwilsdon in our direction, and further than this row said until it was corrected — besides --read-only, which switches the OAuth flow to the read-only scope map, he has had per-service permission levels since 25 February 2026: --permissions gmail:organize builds the requested set from cumulative levels — readonly, organize, drafts, send, full — and his tool registry then disables every tool whose declared scopes fall outside them, deliberately skipping scope-hierarchy expansion so that organize really does exclude gmail.send. What separates that from this row's claim is the direction, not the strength: the level is picked per service rather than derived from the tools you enable, so --tool-tier core --tools gmail still asks for the full Gmail scopes unless a level is named, and neither mode has a guard at the request itself. Our August round read only the tier path and understated him here from 26 August on; re-read in his auth/permissions.py, auth/scopes.py and core/tool_registry.py at 54b1c56 on 7 September 2026 — the day we read, at the revision his repository stood at. The tier half was checked on 26 August 2026 and corrected the same day by csitte.at, who verified it against their own clone rather than taking our word for it), a richer per-message triage heuristic (a-bonus), and bulk organize over a mailbox (the hosted mcpemails.com, which has no snooze either). What none of them do is act on a query and check the mailbox's answer before acting on it.
mailwarden is a Gmail server, not a Workspace suite — if you want Calendar, Drive, Docs and Sheets
from one place, a broad server like taylorwilsdon/google_workspace_mcp covers ground this one never
will, and the two are not mutually exclusive. Adding both is a reasonable setup, and the reason to is
the token, not the tool count: a suite server that can send mail holds a credential that can send
mail, for every mailbox it is pointed at. Giving Gmail to mailwarden instead means the mail half of
your setup has no compose, reply, forward or send tool at all. Where that promise rests differs by
tier, and the distinction matters: on read Google enforces it at the token (gmail.readonly,
which the send endpoints reject), while on manage it rests on the tool surface — Gmail does
accept gmail.modify for sending, so the scope alone is no guarantee. In both cases the egress
guard refuses messages.send and every draft endpoint in the server itself,
so an injected message in your inbox has no tool to reach for and no endpoint to reach.
Practical shape: point the suite server at the services you want and disable its Gmail tools
(--disabled-tools, or a tier that omits them), and run mailwarden alongside for mail. Keep the
tier rule in mind — at most one mailbox per client config should carry
writing tiers.
Ask an assistant to "archive the unread promotional mail that's already skipped my inbox" and it will reach for the obvious query, category:updates is:unread -in:inbox. A server that trusts Gmail's index now archives threads you had already read — mail you never meant to touch, gone in a bulk action you can't easily reverse.
Measured, not asserted. In one real mailbox (~70,000 messages), category:updates is:unread returned 131 threads through threads.list, and only 17 of them held an unread message — 87% stale. The same query, same mailbox, same minute, asked through messages.list instead: 19 messages, none stale. So this is not "Gmail search is unreliable" — the thread view of read state lags while the per-message view does not, and search goes through threads.list. A second mailbox, measured identically on the same day, drifted not at all.
Method, all three queries, the controls, and what the finding is not (it is not the index dropping the predicate, and not a quirk of exotic operator combinations): Gmail's thread index can answer is:unread from a stale read state — a standalone report, every figure traced to a recorded measurement.
Which is the whole point: a server cannot know which kind of mailbox it is in. Re-verification costs nothing where nothing drifts, and saves you where it does — in the measurement above, every thread search dropped was genuinely read, and it discarded no genuinely unread mail.
Where it is not free: the bulk tools. search re-verifies because it fetches every hit anyway; bulk_modify (and create_filter's applyToExisting sweep) is sized in thousands of messages, where one fetch per hit is a different order of cost. Both act on what the index returns — and bulk_modify reports unverifiedPredicates, the conditions from your query that were taken on the index's word (+UNREAD, -INBOX, …). Empty means there was nothing to distrust. Non-empty and the result has to be read-state-precise? Resolve the set with search first and act on those thread ids. A dryRun does not close this gap: it re-reads the same index, so it confirms how big the set is, never whether it is right.
The filter sweep does not report that field, deliberately. Its query is not yours — create_filter builds one from the criteria, parenthesising a query criterion and double-quoting every from/to/subject value, and both of those switch the predicate derivation off. The list would come back empty for every filter you can define, and an empty list here would read as "nothing to distrust" when it means "could not tell". So the caveat stays in prose, where it holds without exception: treat a sweep as acting on the index's answer, whatever it matched. Use verify: true to learn what the sweep actually changed, and search when the set has to be right before anything is written.
A cheaper half-measure, honestly labelled. bulk_modify also takes crossCheck: true, which asks Gmail the same question a second way before writing: every derived predicate is re-run as a label filter (labelIds) rather than as a query operator, and any message the two routes disagree about is left untouched and reported. The cost is one extra list call per predicate — flat, independent of how many messages match — where re-verification costs one fetch per hit. What it buys is bounded and worth stating plainly: a disagreement is real evidence, agreement is none at all, because both routes read the same index and an index can be consistently wrong. So unverifiedPredicates still reports what it always did, cross-check or not, and only search re-checks against the mailbox itself. Whether the two routes ever diverge in practice is unmeasured — node scripts/probe-crosscheck.mjs measures exactly that in your own mailbox, read-only and ids only.
mailwarden fetches every hit live anyway, so search re-checks the unambiguous predicates (is:unread, is:read, in:inbox, category:…, with negation) against each thread's true labels and drops the index's false positives before any tool sees them. The bulk action then runs on exactly the set you asked for. This is the difference between acting on what Gmail indexed and acting on what's actually in the mailbox right now — and it's why snooze/sweep are safe to hand to an assistant: the sweep resurfaces only threads whose snooze is genuinely due, verified against live labels at run time.
See it yourself — no Gmail account needed. From a clone of the repo (the demo is a repo-only verification script, not part of the npm package):
git clone https://github.com/csitte/mailwarden && cd mailwarden
npm install && npm run build
node scripts/demo-reverify.mjs
There is a second script next to it, node scripts/probe-reverify.mjs, which measures the same thing in your mailbox instead of a fake one — read-only, metadata only (no subject, sender or body is fetched), printing counts and label names. It is how the numbers above were produced, and how you can check whether your mailbox drifts at all.
The demo drives the real search() against a fake Gmail API whose index is deliberately stale (returns a read thread for an is:unread query, exactly as Gmail does) and shows mailwarden dropping the false positive. It asserts the outcome, so it exits non-zero if the behavior ever regresses. The same case is locked by unit tests in test/gmail.test.ts ("drops index false positives via live-label re-verify").
A recurring check — what came in since I last looked — is the expensive shape for a live server: the
obvious way to answer it is to search the whole slice again and compare. what_changed (read tier)
answers it from Gmail's own event log instead. Hand it the historyId a previous call or get_profile
returned, and it comes back with what arrived, what left, and which labels went on or came off, plus the
next id to keep.
This is not a cache, and the distinction is the whole design. The only thing that persists between calls is one number, and it persists in the caller. mailwarden still stores nothing about the mailbox, keeps no mirror and no index, and every call remains live against the Gmail API — the same rule as everywhere else here.
Two properties worth knowing before relying on it. It reports events, not state: a message marked
unread and then read appears under both, and both are true — for how the mailbox looks now, ask
search. And Gmail keeps roughly a week of history, after which an id is refused; mailwarden turns that
refusal into an error rather than an empty result, because nothing changed and I can no longer tell
you what changed call for opposite reactions and only one of them is safe to act on.
authentication answers, and what it doesn'tEvery message from get_thread carries an authentication object: the SPF, DKIM and DMARC
results the receiving server recorded, plus the domains each check actually validated.
{
"spf": "pass", "mailedBy": "forwarder.example", // envelope sender SPF checked
"dkim": "pass", "signedBy": "routing.example", // domain whose key signed it
"dmarc": "pass", "headerFrom": "authority.example", // the From domain DMARC evaluated
"authservId": "mx.google.com", // who asserts all of the above
"returnPath": "srs0=…=authority.example=…@forwarder.example"
}
Read dmarc first. It is the only one of the three that ties a passing check to the From
address a human sees. spf: "pass" on its own says an envelope sender was authorised to send —
something a lookalike domain gets in minutes.
The three domains do not have to match, and a mismatch is not a finding. The object above is a
real message from a public authority, forwarded through a custom domain on a mail-routing service
before it reached the mailbox. Every domain differs from the others, and the mail is genuine:
forwarding rewrites the envelope sender (mailedBy becomes the forwarder), the forwarder signs
with its own key (signedBy), and only headerFrom still names the original sender — which is
exactly why DMARC, not SPF, is the check that carries meaning here. Treat the domains as the
explanation of a result, not as a test of their own.
What a pass does not mean. That the mail really came from that domain — not that the domain deserves anything. A phisher holds perfect SPF, DKIM and DMARC on the lookalike domain he registered this morning; authentication tells you who sent it, and the answer can be "exactly who it claims to be, and that is the problem".
What unchecked: true means. The message carried no Authentication-Results header at all —
nobody looked. It is not a failure, and it is not a pass. The header is written by a server that
receives a message, so anything that never arrived from outside — your own sent mail, for
instance — should be expected to have none.
Forged reports. A message can carry Authentication-Results headers of its own — an attacker
writes whatever he likes into the mail he sends. Only the first such header is read, because
each hop prepends its own and the first one is therefore the receiving server's; authservId names
who is asserting the result (for Gmail, mx.google.com) and otherReports counts the ones that
were not read. Values are validated as tokens rather than passed through, so a field that reads
like a verdict cannot carry a sentence. If two results for the same method disagree — a second DKIM
signature that failed — the disagreement shows up in alsoReported instead of being swallowed.
| Tool | What it does |
|---|---|
search | Gmail query syntax → thread summaries (from/subject/date/labels/snippet); read-state/category predicates are re-verified against each hit's live labels; paginated via pageToken/nextPageToken. Each hit carries signals — newsletter (List-Id / List-Unsubscribe / Precedence bulk or list), automated (Auto-Submitted, auto-reply/suppress headers, no-reply-style senders), calendar (text/calendar or .ics part), replyToMismatch (Reply-To on another domain than From; a subdomain of the same domain counts as the same) — read off the first message's headers/MIME, no extra call. Spam and trash are excluded unless the query says in:spam / in:trash — see Looking in spam |
get_thread | Full thread: headers, plaintext + HTML bodies, attachment metadata. Every message also carries authentication — SPF/DKIM/DMARC as the receiving server reported them, plus the domains each check validated (signedBy/mailedBy/headerFrom), who asserts it (authservId), the returnPath, alsoReported for results that contradict each other, otherReports for reports that were not read, and unchecked: true when the message carried no report at all — see Judging a sender. full: false fetches headers and labels only — it then omits plaintextBody/htmlBody/attachments and sets metadataOnly: true, rather than reporting them empty for a request that never looked |
list_labels | All labels (system + user) |
get_profile | Connected account's address, total message/thread counts and the mailbox's current historyId — confirm which mailbox is wired up before acting |
what_changed | Mailbox events since a historyId you hold — arrivals, removals, labels on and off, in one call |
triage_digest | Structured overview of a mailbox slice for decisions: top senders (each with the signals its threads carry), label and age buckets, unread + attachment counts, and how many threads are newsletters / automated / calendar invites / reply-to mismatches — instead of a raw thread list |
list_unsubscribe | What opt-out options a thread advertises (List-Unsubscribe), plus body links when it advertises none — contacts nobody |
list_subscriptions | A mailbox slice grouped by sender: thread/unread counts, the date span each was seen over, and each one's opt-out options — one header fetch per sender, contacts nobody. sendersFound reports how many senders there were before topN truncated the list |
create_label | Create a user label (idempotent; nested via Parent/Child) and return its id; an optional backgroundColor/textColor pair colours it, including one that already exists |
modify_labels | Add/remove labels by name or id — an unknown name in add is auto-created (archive = remove INBOX, read = remove UNREAD) |
bulk_modify | Batch label changes for every message matching a query — 1000 messages per API request, partial success reported per chunk (thread-id list capped at 500, submittedThreadCount has the total). Counts say submitted, because messages.batchModify answers 204 with no body and ignores unknown ids silently; verify: true reads the labels back and returns verified {applied, notApplied, unverifiable} — the only observed outcome on offer. Acts on the raw index, so unverifiedPredicates names the conditions it could not vouch for (see below). dryRun: true resolves the query and reports the matched threads and the labels it would create, touching nothing |
archive / mark_read / mark_unread | Convenience wrappers |
trash / untrash | Move to / restore from Trash |
download_attachment | Save an attachment to a local path (never overwrites — collisions get a numeric suffix) |
unsubscribe | One-click opt-out (RFC 8058) using the endpoint from the message's own header — the only tool that contacts a non-Google host (details) |
bulk_unsubscribe | The same for several threads, sequentially and at most one request per sender (remembered across calls for as long as the server runs, so a retry contacts nobody twice); partial success reported per thread. dryRun: true runs the same header reads and dedupe and reports the endpoint each thread wouldCall — contacting nobody |
snooze | Archive now, resurface on/after a date (YYYY-MM-DD), a date+time (2026-06-20 9am), or a preset (tomorrow, tomorrow 9am, weekend, next week, a weekday name, in N days, in N hours) |
unsnooze | Cancel a snooze, return to inbox now |
list_snoozed | All snoozed threads + due dates |
sweep_snoozed | Resurface threads whose snooze is due (run on demand, via cron, or the daemon); batched, with partial-failure reporting. dryRun: true answers "what is due right now?" (dueLabels/dueThreads) without waking anything |
list_filters | All Gmail filters (criteria + label actions); surfaces any forward address on existing filters for auditing |
create_filter | Create a server-side auto-triage rule (criteria → label actions only; no forwarding — see below). Optionally applyToExisting to also sweep matching mail already in the mailbox |
delete_filter | Delete a filter by id |
All tools declare an outputSchema and return structured content (validated, machine-readable)
alongside the same JSON as fenced text — clients never have to parse prose.
A failure is structured too: isError plus a fenced JSON body with a code
(not_authorized, needs_reauth, insufficient_scope, forbidden_operation, not_found,
rate_limited, upstream_unavailable, network_error, invalid_input, internal_error) and a
retryable flag, alongside the sentence a human reads. So "wait and try again" versus "re-run
mailwarden --auth" is something a client can decide, not something it has to infer from wording
that may be reworded next release. (No structuredContent on errors: that is validated against the
tool's outputSchema, which describes a success.) A rate_limited failure also carries
retryAfterSeconds, because there the wait is the remedy. An insufficient_scope failure names
the gap rather than the possibilities — which scope is missing, what it covers, and what the saved
authorization grants instead — by reading the token's own scopes when it lands, and the server
prints the same line on stderr at startup. It explains the refusal rather than preventing it:
the call still goes to Google, and Google is what turns it down. That matters most with an encrypted token: the tier gate
at registration cannot decrypt, so the surface is advertised in full and the gap would otherwise
surface as a bare 403 mid-task.
Gmail bills each call against a per-user budget of 6,000 units per minute, and the units are not uniform — fetching a thread is the expensive part, in whatever format it is fetched:
| Tool | Units per call |
|---|---|
search | 10 per list page, plus 40 for every candidate thread it fetches — about 1,000 at the default size, up to about 4,000 when the query carries a read-state or category condition |
get_thread (full and metadata alike) | 40 |
trash, download_attachment | 20 |
archive, mark_read, untrash | 10 |
modify_labels, snooze | 10, plus 1 for each label looked up by name and 5 for each label it has to create |
bulk_modify | 5 per 500 matching messages, plus 50 per 1,000 modified; verify: true adds 40 per thread, crossCheck 5 per condition and 500 matches |
list_labels, get_profile | 1 |
mark_unread costs what mark_read does.
Two consequences worth knowing before you build on this. Parallel agents share one budget —
they reach Gmail through a single server process as a single user, so four of them reading threads
at once spend four times as fast. And bulk_modify is cheaper than looping from six threads
on: up to 500 matching messages it costs 55 units, the price of five and a half archive calls,
and it grows by the thousand messages, not by the thread. verify: true reverses that — reading
the labels back costs 40 per thread, more than the loop it replaced.
If you do exceed it, mailwarden fails the call with rate_limited and retryAfterSeconds: 60.
That failure is temporary and waiting genuinely fixes it — the budget is counted per minute.
mailwarden absorbs one such wait (about 20 seconds) itself before reporting; past that it hands the
decision back rather than holding your tool call open. That holds for the tools that fetch many
threads in one call too: search fails with rate_limited rather than returning a list the quota
cut short, and bulk_modify with verify and list_subscriptions stop fetching and report what
they did not reach as unverifiable and unknown (for bulk_modify the change is already made,
and failing would hide it). Note that Gmail reports this as a 403 as readily as a 429, so a
client keying on the status alone will mistake it for a permission problem — key on the code.
snooze removes INBOX and applies a dated label MCP/Snoozed/<key>, where the key is either YYYY-MM-DD (due all day) or YYYY-MM-DDTHHMM (due at that local minute). The until argument takes an explicit date, a date+time (2026-06-20 9am, …T17:00), or a preset resolved server-side — today, tomorrow, weekend (next Saturday), next week (next Monday), a weekday name (monday–sunday, next occurrence), in N days, or in N hours — and a date preset may carry a trailing time (tomorrow 9am, monday 8:30), so the caller never has to compute the moment itself. sweep_snoozed finds due labels and returns those threads to the inbox (marked unread); a timed snooze wakes at the first sweep on/after its minute, so wake latency equals your sweep interval. Run the sweep:
sweep_snoozed tool),mailwarden --sweep,MAILWARDEN_AUTO_SWEEP=1 (hourly sweep while the server runs).create_filter sets up a Gmail server-side rule: mail matching the criteria automatically gets the
given label actions — the mailbox keeps triaging itself with no assistant in the loop.
from, to, subject, query (full Gmail search syntax), negatedQuery,
hasAttachment, excludeChats, and size + sizeComparison (smaller/larger, given together).
At least one is required.addLabels / removeLabels, by name or id (an unknown name in
addLabels is auto-created, nested via /). Common recipes: skip the inbox → removeLabels: ["INBOX"];
auto-mark-read → removeLabels: ["UNREAD"]; auto-trash → addLabels: ["TRASH"];
star → addLabels: ["STARRED"]; never-spam → removeLabels: ["SPAM"]; file under a label → addLabels: ["Receipts"].applyToExisting: true to also apply the same actions once to mail already in the mailbox —
mailwarden builds a Gmail search from the criteria and runs a bulk modify (up to maxMessages,
default 1000; same unverified-index caveat as bulk_modify, and the one-off pass excludes Spam/Trash).
This requires at least one positive criterion (from/to/subject/query/hasAttachment:true/size):
an exclusion-only rule (negatedQuery or hasAttachment:false) is refused for applyToExisting
because it would match almost the whole mailbox — create such a filter without the flag.
The outcome comes back under applied (the query used, matchedMessages/submittedMessages/submittedThreadCount
counts, capped when the match set hit maxMessages, per-chunk failed, and an error string if the whole
pass failed); it's null when applyToExisting was not set. submittedMessages is what was handed to the API,
not what changed — pass verify: true alongside applyToExisting to read the labels back and get
applied.verified {applied, notApplied, unverifiable}, the same check as bulk_modify's verify
(one extra read per affected thread, so off by default). The filter is created first, so a partial or
failed backlog pass is reported in applied, never raised — the rule still stands.gmail.settings.basic scope; re-run --auth once if you authorized an older version.
Not available in read-only mode.list_unsubscribe (read tier) reports what the sender offers, without contacting anyone. It reads the
newest message that actually carries a List-Unsubscribe header — a reply threaded onto a newsletter
sits at the end and advertises nothing, which would otherwise read as "this list has no opt-out".
When a thread advertises no opt-out header at all, list_unsubscribe looks in the message body and
reports the unsubscribe links it finds there as bodyCandidates — plenty of senders put the link only
in the footer, and answering "no opt-out options" for them is true about the headers and wrong about
the mail. Those links are reported, never fetched, they cannot be handed to unsubscribe, and
hasUnsubscribe stays false for them, because that flag has always described the headers. The search
costs one full thread fetch and happens only in that case.
list_subscriptions (read tier) does the same across a whole slice, grouped by sender, so you can see
who keeps writing and which of them can actually be left — one header fetch per sender rather than
per thread. unsubscribe and bulk_unsubscribe (manage tier) act on it — and that is the only
place mailwarden ever talks to a host that isn't Google, so the rules are tight:
navbuildz/gmail-mcp-server, does fetch them, following redirects, when the
header is missing.List-Unsubscribe-Post.
A plain https: link is meant for a human in a browser and is handed back, not fetched.mailto: opt-outs are never performed. They would require sending mail, which mailwarden cannot
do. The address is reported so you can act on it yourself.List-Unsubscribe=One-Click and is
never derived from anything; the response body is cancelled unread. What returns to the model is the
status code and the URL actually called — no content from the endpoint, so it cannot answer with
instructions. (A 301/302/303 redirect is followed as a GET, i.e. with no body at all.)bulk_unsubscribe takes thread ids
(never a query — a query-driven bulk would fire off a request per matched sender before anyone had
looked). Threads from a sender whose request already went out are reported with duplicateOf and
cost no second request: two threads from one list share an opt-out, and calling it twice only
confirms your address twice. A sender is only recorded once a request actually reached an
endpoint, so a refusal or a dropped connection still leaves the next thread its own try — and if
the skipped thread advertises a different endpoint, the reason says so, since one sender can run
several lists. That memory spans calls for as long as the server runs, and unsubscribe shares
it: a call that times out is safe to repeat, and asking twice for the same newsletter contacts the
sender once. Pass force: true to unsubscribe for a deliberate second attempt — after an endpoint
answered 500, say. It is kept in memory only: persisting it would mean a second kind of local state
beside the token, which this server deliberately does not keep, so a restart forgets.
Capped at 25 threads and 60 seconds per call; whatever the budget doesn't cover comes
back as skippedOutOfTime rather than silently undone. None of it can be reversed, which is why all
three limits exist.::1 and 0:0:0:0:0:0:0:1 alike); an
address that does not parse is refused. DNS resolution and all hops share one 10-second budget. Not
rebinding-proof (fetch resolves again when it connects) — see SECURITY.md; what
survives that gap is a blind POST whose response is never read.Check it against your own mail before you trust it. From a repo clone (repo-only, not in the
npm package), after npm run build and mailwarden --auth:
node scripts/probe-unsubscribe.mjs --vet # category:promotions, 25 threads
node scripts/probe-unsubscribe.mjs "from:substack.com" --max 50 --vet
It prints each real List-Unsubscribe header next to what the parser made of it, and --vet also
runs the endpoint through the URL vetting and the address guard — so you see both whether the parser
understood the header and whether the guards would have let that opt-out through. Strictly
read-only: no request is ever made to a sender, and nothing in the mailbox changes.
What it can't undo: the request tells the sender your address is live. A sender that ignores its own
opt-out is beyond any client's reach — pair unsubscribe with create_filter or trash for those.
Not offering an automatable option is reported as unsubscribed:false with the alternatives, not as an
error. A read-only deployment gets list_unsubscribe and list_subscriptions, and never makes the
request at all.
A query that does not name a place never sees spam or trash. Gmail excludes both from any
search that does not say in:spam / in:trash, so from:someone returns nothing for a mail that
is sitting in the spam folder — and nothing in the answer says so. Measured against a live mailbox:
the same from: query returned 0 hits by default and 1 with spam included.
This matters because of why mail gets misfiled. A spam filter judges a message on its own; it cannot know that you signed up for something a minute ago, requested a password reset, or placed an order — so the confirmation you are waiting for is exactly the kind of mail that lands there. You know what you just did. The filter does not.
So when mail someone expects is missing, ask again with the place named:
search("in:spam newer_than:2d") # what got filed as spam recently
search("in:spam from:example.com") # the confirmation that never arrived
A thread returns to the inbox with modify_labels (remove SPAM, add INBOX), and a sender that
keeps being misjudged is best fixed for good with a never-spam rule — create_filter with
removeLabels: ["SPAM"] (see Filters).
Two things this server deliberately does not do. It does not scan the spam folder and judge what belongs there: measured over one real spam folder, 89% of it carries no mailing-list machinery at all, so "looks unlike bulk mail" flags nearly the whole folder and filters nothing. And it does not act on that judgement by itself — releasing mail from spam is a decision, and the context that makes it obvious ("I just registered there") lives in the conversation, not in the mailbox.
For the full threat model — trust boundary, per-threat mitigations, explicit non-goals, and how to report a vulnerability — see SECURITY.md. The highlights:
--http listener binds to 127.0.0.1
(not the LAN) and refuses to start without a MAILWARDEN_TOKEN bearer token — set
MAILWARDEN_ALLOW_NO_TOKEN=1 to override on a trusted, isolated network. On a loopback bind it
also validates the Host header (DNS-rebinding defense). For remote hosting, set MAILWARDEN_HOST
and front it with TLS.create_filter follows
the same rule: it can label, archive, trash, star or mark mail, but never creates a forwarding
filter (which would be an exfiltration path). list_filters still surfaces any forwarding filter
already on the account, so you can spot one. This holds because no such tool exists and none can be
registered at runtime; for the stronger variant, where Google refuses to send rather than
mailwarden declining to, see Read-only mode below. Read the promise as narrow, because it
is: it closes the route out through this server, not the assistant's other routes — a client
that can fetch a URL, write to a synced folder or run code still has everything an injected
message needs, and this server is what puts that message in front of it. See
non-goals.unsubscribe tool is the only code path that
contacts a non-Google host. Its endpoint is read from the message's List-Unsubscribe header —
never from a tool argument — the request body is fixed and the response body is discarded, so it
cannot become a data channel. https/default-port only, redirects re-validated, and any hop resolving
to a private, loopback, link-local or metadata address is refused. See
Unsubscribing.MAILWARDEN_TOOLS advertises only the tiers
you name — read (the read tools), manage (mailbox mutations, snooze, downloads), filters
(server-side filter CRUD, the only tier whose tools need gmail.settings.basic). Default is all
three; e.g. read,manage gives a full triage surface without filter management. The OAuth scopes
requested at --auth are derived from the enabled tiers — a read deployment asks only for
gmail.readonly, and gmail.settings.basic is requested only when the filters tier is on. And the
filter tools are hidden automatically when the stored token doesn't carry gmail.settings.basic
(e.g. a token authorized before you enabled the tier) — re-run --auth to grant it. Older tokens
without a recorded scope are advertised as before, with the runtime insufficient-scope message as the
fallback.MAILWARDEN_READONLY=1 (shorthand for MAILWARDEN_TOOLS=read) and only the
read tools (search, get_thread, list_labels, list_snoozed, get_profile, what_changed, triage_digest,
list_unsubscribe, list_subscriptions)
are registered — nothing that can change the mailbox or write
files is even advertised to clients (the filter tools, which need the broader gmail.settings.basic
scope, are excluded too). Recommended for shared/HTTP deployments that only triage.
It is also the only tier whose no-send property Google enforces: it holds a gmail.readonly
token, which Gmail's send endpoints reject outright. manage needs gmail.modify, and Gmail
does accept that scope for sending — mailwarden simply exposes no tool that would. So a read
deployment could not send even if this binary were replaced; a manage one cannot send because
there is nothing to call. (There is no send-free write scope to switch to — see
SECURITY.md, threat 1.)Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y mailwardenMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-csitte-mailwarden": {
"command": "npx",
"args": [
"-y",
"mailwarden"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referencemailwardennpmio.github.csitte/mailwarden works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.