Back to Directory/Developer Tools

io.github.Davmunrey/trazum

Trazum as an MCP server: let an agent price and budget its own prompts before it sends them.

Developer ToolsTypeScriptv2.4.1

Trazum

Most of your LLM bill is not the prompt. Trazum finds where it is.

A deterministic cost analyser for prompts and usage logs. Offline, free, same answer every time. It reports sixteen findings priced in dollars per month: caching you are not getting, a model tier you may not need, a schema you pay to describe on every call. Shortening the prompt is one of them, and it is rarely the biggest.

CI CodeQL MIT licence Node Runtime dependencies

Start here: one line, nothing installed

npx @trazum/cli bill ~/.claude/projects

Point it at a file or a folder of whatever usage you have — Claude Code transcripts, OpenTelemetry spans, a LiteLLM, Helicone or LangSmith export, an Anthropic, OpenAI or OpenRouter usage report, or a plain usage log — and it tells each file's shape from its own text, prices it, and prints the receipt. A file it cannot read is named, never guessed at; a model it cannot price is a named gap, never a zero. Nothing you point it at leaves your machine.

Your agents spend money in a loop. This prices the call before it happens.

One agent costs what it costs. A fleet of them spends in a loop nobody is watching per iteration, and the bill arrives a month later as one number with no per-decision detail inside it.

Trazum installs into that loop. The MCP server's first tool is spend_guard, and it is the only one here whose trigger is not a sentence somebody types:

"May I spend this?"  ->  yes, no, or cannot-tell.

A refusal carries the cheaper ways to make the same call, each priced for that call and each naming what it assumes. The ceilings come from your trazum.config.json; the spend so far comes from the usage log your host already writes. Nothing is called and nothing is spent to answer.

claude plugin marketplace add Davmunrey/Trazum
claude plugin install trazum@trazum

That one line brings the skill and the MCP server. Every other MCP client reads the same mcpServers JSON, so Cursor (.cursor/mcp.json), Windsurf, Claude Desktop and the rest are one paste:

{ "mcpServers": { "trazum": { "command": "npx", "args": ["-y", "@trazum/mcp"] } } }

The server is listed on the official MCP registry as io.github.Davmunrey/trazum, and every release updates the listing.

Or run it right now, without installing anything: the Playground

That link opens the CLI's pure subset running in the page, against sample files already loaded, through the same @trazum/core functions the terminal runs. Nothing you paste leaves your browser.

The argument, in one screenshot

trazum optimize on a wordy support prompt: 238 tokens down to 142 (-40.3%), $24.00/month saved by the rules, and an advisory pointing at $528.40/month, 22 times more

Real output, transcribed. Read the last two lines: the rules recovered $24.00 a month, and the advisory above them is worth $528.40, 22 times more. That gap is the entire argument for this tool.

What goes in: a prompt or a usage log. What is computed on your machine, offline: @trazum/core with zero dependencies, plus the CLI, the MCP server, the web app and the Action. What leaves: a receipt, a report, a gate verdict. What never crosses the line: the prompt text, the model's answer, file paths, branch names and credentials.

Generated by npm run draw:architecture, not drawn. architecture-image.test.js fails the build if a package exists that the picture does not show, if the core takes a dependency while the picture says it has none, or if a third module is allowed to reach a network without the picture saying so.

The prompt is the part everyone looks at, and usually the cheap part. In the run above, forty percent of the text came out and it moved 3.5% of the bill. What moved the rest was a question nobody was asking: does this task need the model it is running on?

Every figure has a receipt. Sixteen advisories, each priced per month and reproducible on a single file: caching you are not getting, work that could go through the Batch API, a schema costing tokens on every call to describe a shape the request could carry as a parameter. Underneath them, twelve deterministic rules that shorten the text itself: same input, same output, free, offline, and never touching code, URLs, email addresses or placeholders. On top, an optional LLM pass for the compression rules cannot do, through whichever provider you configure, which never runs unless you ask.

                      ┌──────────────┐
                      │ @trazum/core │   the library: rules, tokens, pricing
                      └──────┬───────┘   zero dependencies, browser-safe
    ┌────────────┬─────────────┼─────────────┬─────────────┐
 @trazum/cli  @trazum/mcp  trazum-vscode  @trazum/web    action/
 51 commands   MCP server    the editor      Next.js    comments on
              for your agents  status bar               pull requests

              @trazum/tokenizer-openai   optional: exact counts for OpenAI
                    install it or do not, nothing else changes

Contents


What it actually does

1. Tells you where the money actually is. This is the part worth reading first, because it is where the numbers are. Every advisory is priced per month against your own call volume, and none of them is about making the text shorter:

AdvisoryWhy it matters
Prompt cachingReading from cache costs 10% of input. The saving is computed over the real stable prefix: in a template with {{placeholders}}, only what precedes the first one is cached — not the whole prompt.
Reorder the templateStable instructions sitting after the first variable placeholder never cache today. Trazum prices moving them in front — and with --reorder, does it.
Batch API50% off input and output when the work tolerates latency.
Cheaper modelComplexity heuristic: if the task looks simple, what dropping a tier would save.
Output-dominated costIf you pay more for the answer than for the prompt, shortening the prompt has a ceiling.
Promotional pricingWarns when you are budgeting with an introductory price that expires.
Context windowIf the prompt does not fit, the call is going to fail.
Contradictory instructions"Answer in English" three paragraphs above "reply in the customer's own language". The model has to pick one, and which one can change between calls — a correctness problem that also costs tokens twice.
Redundant examplesFew-shot examples that are near-copies of an earlier one, and what they cost per month.
Output format stated twiceA schema shown in a code block and then walked again in prose. The block is the version worth keeping.
Schema the request could carryA schema block introduced by "Output format:" is paid for in input tokens on every call. Every major API now takes a response schema as a parameter — and moving it there is both cheaper and stricter. See below.

The last four are advisory only. A contradiction has a right answer that only the author knows, and an example that looks redundant may be demonstrating a boundary case on purpose. Trazum points; it does not cut.

The one finding that is not a trade-off

Most of what Trazum reports is a choice: shorter against clearer, cheaper against more capable. Moving an output schema out of the prompt is neither.

→ The output schema could travel in the request instead of the prompt
  A schema block introduced by "output format" defines `category`, `reply`,
  `escalate_to_human`, `confidence`, costing about 62 tokens on every call.

Those tokens are paid on every call to have the model read a shape and be asked, politely, to match it. output_config.format, response_format, responseSchema — whatever your provider calls it — takes the same shape as a request parameter, where the decoder is constrained rather than persuaded. Cheaper and stricter.

Trazum reports it and never does it, because it is not a change to the prompt: it is a change to the code that sends the prompt. A rule that deleted the schema would leave a prompt asking for a shape it no longer describes, sent by a client nobody updated — strictly worse than what it started from.

The one way this could do harm, and what stops it. Output format: {...} is a contract and moving it is free; Input: {...} inside a few-shot example is data the prompt needs, and moving it breaks the prompt. So nothing is guessed: a block counts only when a phrase from the output-cue dictionary appears immediately before it, in one of the seven languages the rules cover. No phrase, no finding — a false negative, which is the right direction to be wrong in.

The example detector finds near-copies — the way few-shot blocks actually grow. It deliberately does not flag paraphrases: that case needs a model, and is on the roadmap for the LLM pass.

2. Then it trims the prompt itself. Twelve deterministic rules: courtesy, filler, verbose phrasing, duplicated paragraphs, decorative separators, shouting in capitals. Two levels — safe (no semantic risk) and aggressive (read the diff). This is the smallest number on the page more often than not, and it is reported that way rather than dressed up.

3. And never touches what would break the prompt. Code fences, indented code blocks, inline code, URLs, email addresses, template placeholders ({{x}}, ${x}, {x}, {% %}) and XML/HTML tags are isolated before any rule runs. If a rule ever did make one of those disappear, that rule is discarded and the rest carry on.

Reviewing an aggressive run. Every rule reports what it actually changed, so the level that saves the most is judged rule by rule rather than as one wall of diff — and a single rule you disagree with comes off with --disable:

  [aggressive] Intensifiers (3×, ~6 tokens)
      VERY → —
      extremely → —
      quite → —
  [aggressive] Self-verification instructions (1×, ~17 tokens)
      You should double-check your answer before re… → —

4. Optionally, runs it past an LLM. The result is only accepted if it is shorter and leaves protected content byte-identical. Otherwise the deterministic version stands. It never returns something worse than where it started. With --suggest it proposes phrases one at a time — You should always make sure to → Always — each checked against your prompt before you see it, so eight surviving out of ten is a useful morning rather than a rewrite to read end to end.

5. Answers the questions that come before "shorten this". Trimming one file is the smallest thing here. optimize is one of 51 commands — the table above names what each answers — because knowing a prompt is wasteful is not the same as knowing which prompt, whose change made it so, or whether the shorter version still works.

check, diff, rank and blame all take --markdown-out, so the answer can land in a pull request comment rather than a terminal nobody is looking at.

It gates in whatever CI you already run. One binary, two exit codes, and worked recipes for GitLab CI, Jenkins, CircleCI and a pre-commit hook — no vendor plugin, because each one would be a second code path that drifts from the exit codes it is supposed to relay.


The 51 commands

CommandWhat it answers
trazum initWhat is in this repository, and what is the one thing worth fixing? The first command to run.
trazum optimizeWhat can come out of this prompt, and what is that worth a month?
trazum checkDoes this prompt fit its token budget, and has the repository drifted past its recorded baseline? Exits 1 when either fails — this is the CI gate.
trazum baselineWhat does this repository's prompts cost right now? Records it, to commit.
trazum diffWhat did this edit cost?
trazum rankOf these forty prompts, which is worth an afternoon?
trazum doctorWhat is wrong across the whole workspace?
trazum pruneWhich few-shot examples earn their tokens? Measured, and it asks before spending.
trazum blameWho made this prompt expensive, and when?
trazum evalDoes the shorter prompt still do the job?
trazum whereWhich prompts are hiding inside my source files?
trazum modelsWhat does each model cost, and what is its cache minimum?
trazum profileWhere did the money actually go? Reads a usage log, not a prompt.
trazum routeIs the cheaper model good enough? Measured, and it asks before spending.
trazum planOf everything the log shows, what do I do first, and what is each move worth?
trazum verifyDid the plan's savings actually arrive? Three outcomes, never two.
trazum historyWhat have twenty reports been saying that no two of them could? Shapes, never forecasts.
trazum connectWhat did the provider actually bill me? Read from their API, nothing exported by hand.
trazum storeWhat have I measured and kept? Aggregates only — no prompt text, ever.
trazum watchHas anything crossed a budget? Measured crossings only — never a forecast.
trazum serveWhat will this call cost, and is there budget? Answered in milliseconds, halves kept apart.
trazum gatewayCan it stop the call instead of advising against it? Refuses; never substitutes.
trazum ladderIs cheap-first-escalate-on-failure saving money, or costing it? Break-even rate, stated.
trazum experimentWhich of two arms is better on real traffic? A winner only when there is one.
trazum qualityDid that prompt change quietly make the product worse? Refuses to blame what it cannot attribute.
trazum semanticDoes this prompt say the same thing twice, or contradict itself? The model proposes; the checker disposes.
trazum ownersWhose budget does this land on? The unallocated is never spread.
trazum commitmentWhat would that committed-use deal have been worth? On measured months, both directions priced.
trazum reportWhat did the year actually look like? No new data, and it lists its own blind spots.
trazum schemaWhich fields must a document of this format carry? A JSON Schema, for validators that are not Trazum.
trazum conformDoes the document my tool emits conform, and what will it not be able to answer?
trazum rollupFour of us measured four things — what is the total, and what did merging lose? A format and a merge, not a service.
trazum pulseDid the things that are supposed to run, run? Runs nothing itself — your CI is the thing that notices.
trazum positionWhere does the month stand against every ceiling? Measured, denominators attached, no forecast anywhere.
trazum receiptWhat did this cost, in a form that still answers when read somewhere else? Counts, and the money split so it can be added up rather than believed. No prompt text, no answers, no paths: there is no field for them.
trazum from-claude-codeWhat did my Claude Code sessions cost? Reads the transcripts already on disk — the numbers only, never the words.
trazum from-otelWhat did the LLM calls in my OpenTelemetry export cost? Reads the GenAI spans any exporter already emits — the counts only, never the prompts.
trazum from-litellmWhat did the calls my LiteLLM proxy logged cost? Reads the spend log the gateway already writes — the counts only, never the prompts, keys or addresses on the same row.
trazum reconcileDoes what Trazum computed match what the provider charged? Sets the two figures beside each other, never merges them, and leaves the unexplained remainder standing on its own.
trazum billWhat did all of this cost? One door: reads a file or a directory of anything the converters read, tells each file's shape from its own text, converts, prices, and ends on the receipt. A file no shape claims is named, never guessed.
trazum from-openrouterWhat does OpenRouter say I used? Reads the activity report your own management key fetched, keyed by the same slugs the live pricing overlay uses; what OpenRouter charged is printed beside Trazum's figure and never merged, and reasoning tokens are counted, not added.
trazum from-openaiWhat does OpenAI say my organisation used? Reads the completions usage report your own admin key fetched; the record keeps OpenAI's own cached-inside-prompt shape, batch rows and non-default tiers are left out and named, and audio or image tokens are never priced at a text rate.
trazum from-anthropicWhat does the provider itself say my organisation used? Reads the usage report your own admin key fetched; Trazum never holds the credential, refuses to price a batch row at a standard rate, and labels by workspace only from a mapping you write.
trazum from-heliconeWhat did the requests my Helicone proxy kept cost? Prices the model that answered, not the one that was asked for, and counts the substitutions.
trazum from-langsmithWhat did the model calls in my LangSmith traces cost? Only the llm runs, because a trace is a tree and summing it bills the same tokens twice — and it refuses to price a call by the client class that made it.
trazum switchShould we move this traffic, and when does moving pay? Measured delta, declared migration cost, break-even as division on the past — and the required evaluation itself priced.
trazum ownrateWhat does my self-hosted model cost per million tokens? Your GPU rate over your measured throughput — derived from your declaration, never guessed.
trazum benchHow fast is Trazum here, and on what? One shot per workload, no judgement — run it before and after a change.
trazum writeWhat should this prompt say, and what will it cost before I ever send it? Asks; nothing is generated.
trazum rulesWhich rules exist, and what does each one do?
trazum feedbackWhere do I report this, and what will you ask me for? Sends nothing.

Getting started

npx @trazum/cli init

No install, no key, no network. It reads what is already here — your prompts, which provider your code calls, a usage log if one is lying around — writes a config out of what it can actually justify, and prints the single most valuable thing it found. See the first five minutes.

Or start from one file:

npx @trazum/cli optimize your-prompt.txt --cost

Either way, keep it around:

npm install -g @trazum/cli     # the terminal
npm install @trazum/core       # the library
npm install @trazum/mcp        # the MCP server, for an agent
npm install @trazum/tokenizer-openai  # optional: exact counts for OpenAI models

Or hand the whole thing to Claude Code as a plugin — the trazum skill plus the MCP server, installed together, nothing else to configure:

claude plugin marketplace add Davmunrey/Trazum
claude plugin install trazum@trazum

The plugin's skill is the same document this repository's own agents work from, derived by scripts/build-plugin-skill.mjs with only the invocation changed — a test fails the build if the two drift apart in any other way.

From source, if you are working on Trazum itself
npm install
npm run build      # core + cli
npm test           # every suite: core, CLI, web, Action
npm run verify     # the above plus typecheck and the web build

The test count used to be written here as a number. It said 580 while the real figure had reached 798, because nothing checked it — so it now says what the command covers instead. A number nobody maintains is worse than no number.

The first five minutes: trazum init

npx @trazum/cli init
What is here

  Running inside a terminal.
  1 prompt file found.
  Usage log found: usage.jsonl.

What the config would say
  + usage.model  100% of the measured bill went to claude-opus-5
  + usage.callsPerMonth  240 calls over 30 days, stated as 240 a month
  + usage.avgOutputTokens  96000 output tokens over 240 calls averages 400
  · usage.cacheHitRate  this log has no cache columns at all, which is not the same as a hit rate of zero
  · usage.batchEligible  whether the work can wait for a batch window is a product decision, and no log records it
  · labels  1 label in the log, and nothing here proves which prompt file sends which
  · spend.maxUsd  a budget is a policy, so it is yours to set — the measured figure is $38.40 over 30 days

The most valuable thing found
  240 calls labelled "classify" went to Claude Opus 5 over 30 days.
  They cost $38.40.
  The same work fits Claude Sonnet 5, which is cheaper per token.
  The Batch API halves both halves of the bill, for work that can wait.
  Together: $30.72 over the same 30 days.

It is a detection, not a wizard. Nothing is asked. Each line above is something that was found — a prompt, a provider named in your code, a log — or a key it declined with what would settle it. --dry-run prints the config and writes nothing; --yes replaces one that is already there; without it an existing config is left alone.

Every key it writes carries the arithmetic that justified it. A generated config full of guessed thresholds is one nobody trusts and everybody deletes, and it is worse than an empty one, because six weeks later it reads as a decision somebody made.

Four things it refuses to write, and they are the interesting four:

  • A budget. A log says what your traffic was; a budget says what it may cost, which no log can answer. "The measured month plus twenty per cent" would be this tool inventing a threshold and then grading you against it. So the measured figure is handed over and the limit stays yours.
  • A monthly rate from a short window. Twenty-eight days minimum, so every weekday appears the same number of times. Four days multiplied by seven is a forecast wearing a measurement's clothes.
  • A cache hit rate from a log with no cache columns. Not recorded is not not-happened. Writing 0 there would tell every later caching advisory that caching is doing nothing — a finding invented out of a missing field.
  • batchEligible, in either direction. Whether the work tolerates a batch window is a product decision, and no log records it. false would quietly delete the batch lever from every report; true would sell a saving on latency nobody agreed to give up.

It also declines a model when your code names a provider and no model. where prints a provider's default because a reader can see it is a guess; a config file cannot.

No usage anywhere? It says so, and points at docs/usage-logs.md — Anthropic, OpenAI, the Vercel AI SDK and an OTel collector, with records you can copy.

trazum init --json is the same proposal as data, including every declined key and its reason — contracted in docs/json-output.md. It writes nothing.

CLI

node packages/cli/dist/index.js optimize prompt.txt --calls 50000 --diff
Input tokens
  190 → 137   -27.9% (estimated, ±6%)

Rules applied
  [safe] Repeated paragraphs (1×, ~19 tokens)
  [safe] Wordy phrasing (1×, ~3 tokens)
  [safe] Politeness formulas (4×, ~19 tokens)
  [safe] Filler and throat-clearing (2×, ~11 tokens)

Cost with Claude Opus 5
  50,000 calls/month · 300 output tokens per call
  $422.50 → $409.25   saving $13.25/month (3.1%)

Beyond shortening the prompt
  → This task may not need Claude Opus 5 ~$327.40/month
  → If the work tolerates latency, use the Batch API ~$204.62/month

Every other command, each with its own chapter in the command reference:

trazum doctor                        # survey the whole workspace
trazum plan usage.jsonl              # the findings as a ranked plan
trazum verify plan.json --against new.jsonl   # did it work?
trazum history reports/              # the long run, from stored reports
trazum connect anthropic             # your bill, read from the provider
trazum store                         # what is kept, and what a prune takes
trazum watch --once                  # did anything cross, this afternoon
trazum serve                         # answer before the call is sent
trazum rollup a.json b.json          # several people's bills, one roll-up
trazum profile usage.jsonl --html-out report.html   # the report somebody forwards
trazum pulse --max-stale-hours 36    # did anything stop running?
trazum bench                         # how fast is Trazum on this machine
trazum rank prompts/                 # which one to fix first
trazum blame prompts/system.txt      # who made it expensive, and when
trazum diff old.txt new.txt          # what this edit cost
trazum check prompts/ --max-tokens 2000
trazum eval prompts/system.txt --cases cases.json
trazum where src/agent.ts            # which provider this actually calls
trazum models                        # pricing table and cache minimums
trazum rules                         # what each rule does, and its id
trazum --help

When redirected it writes only the optimised prompt, so it pipes cleanly:

cat prompt.md | node packages/cli/dist/index.js optimize - > prompt.optimised.md

To install it as a trazum command:

npm link -w @trazum/cli

Token budgets in CI. trazum check exits 1 when the prompt busts its budget, so a template that grows unchecked breaks the build instead of the bill:

trazum check prompts/system.txt --max-tokens 2000
# FAILED 2,481 tokens busts the budget of 2,000.
#   Optimised with "trazum optimize --level safe" it would land at ~1,913 tokens and fit.

Before it reaches CI: a pre-commit hook.

ln -s ../../scripts/pre-commit .git/hooks/pre-commit
trazum: these prompts are over their token budget:
  prompts/system.txt

  trazum doctor .          shows how far over, and what it costs
  Shorten them, raise the budget in trazum.config.json, or commit with --no-verify.

It blocks only on prompts your commit actually touches. A hook that refuses a commit over a different prompt somebody else committed last month is one people learn to pass --no-verify to — and then it is worse than no hook at all. TRAZUM_HOOK=0 disables it; nothing staged, no Trazum installed, no prompts or an unreadable config each say so once and exit 0. One real limitation: it reads the working tree, not the staged blobs, so it judges a prompt's newest edit even when an older version is what is staged.

In GitHub Actions, use the packaged action — nothing to install:

- uses: actions/checkout@v7
- uses: Davmunrey/Trazum@b6dd9349b7c1502688b609a7c1adbe20aac82c79  # 2.4.1
  with:
    target: prompts/system.txt
    max-tokens: 2000

One-click fixes, as suggestions. suggest-fixes: true posts the optimised prompt as a GitHub suggested change, which a reviewer applies with one button:

permissions:
  contents: read
  pull-requests: write
with:
  target: prompts/
  suggest-fixes: true
  github-token: ${{ secrets.GITHUB_TOKEN }}

A suggestion, not a commit, and that is deliberate. Committing the fix would need contents: write; a suggestion lands in the same place with the same one click on the pull-requests: write the comment mode already uses, and you stay the one who commits. Two limits, both real: it uses the safe level only — a one-click apply is not the moment for a diff that wants reading — and a suggestion can only anchor to lines in the pull request's diff, so a PR that edits three lines of a forty-line prompt gets a notice explaining why there is no suggestion rather than a partial rewrite.

Pinned to a commit SHA, not a tag — the same rule SECURITY.md states and security.test.js enforces on every third-party action in this repository. A tag is a mutable pointer: whoever can move v1 can change what runs in your workflow with your token. The # 1.0.0 comment names the version at that commit, and is what Dependabot reads to offer you the bump.

The report lands in the run summary automatically — every run, pass or fail, with no token and no permissions. To also post it as a pull request comment that replaces its own previous one:

permissions:
  contents: read
  pull-requests: write     # the action cannot grant itself this

steps:
  - uses: actions/checkout@v7
  - uses: Davmunrey/Trazum@b6dd9349b7c1502688b609a7c1adbe20aac82c79  # 2.4.1
    with:
      target: prompts/            # a directory uses trazum.config.json budgets
      comment: true
      github-token: ${{ secrets.GITHUB_TOKEN }}

Commenting can never fail your build. No pull request, comments disabled, or a read-only token — each prints a notice and carries on, because the report has already reached the run summary. That matters on pull requests from forks, where GITHUB_TOKEN is read-only by design and the comment simply will not post.

If you go looking for a way around that, the answer you will find is pull_request_target. Don't. It runs with a writable token against the base repository while checking out code the contributor controls, which turns "we wanted to comment on a PR" into arbitrary code execution with your secrets. The run summary is there precisely so you do not need it. Trazum asserts in CI that it uses pull_request_target nowhere.

A passing report is collapsed; a failing one is not. A green table that stays green on every push is the thing you learn to skip — and then you skip the red one too.

The spend gate, packaged. The same action gates the bill itself when handed a usage log instead of prompts — mutually exclusive with target, because one run gates tokens before the money is spent or the spend itself, and saying which is the caller's job:

- uses: Davmunrey/Trazum@b6dd9349b7c1502688b609a7c1adbe20aac82c79  # 2.4.1
  with:
    usage-log: logs/yesterday.jsonl
    max-usd: '50'            # exit 1 over budget — no period assumed
    # against: logs/day-before.jsonl
    # max-growth-usd: '10'
    # label: chat            # one workload's budget
    # since: '2026-08-11'    # one period's — until includes its whole day
    # until: '2026-08-17'

The profile report lands in the run summary either way, and a failing gate still writes it — a red build with no report is a mystery, and mysteries get deleted from pipelines.

The report leaves the terminal in three shapes. --markdown-out for a CI summary or a PR comment, --csv-out for whoever signs off the bill (one row per workload and model, no total row, empty cells where dollars are unknown), and --json for anything built on top — documented field by field in docs/json-output.md, with a schemaVersion and a test that fails if the two ever disagree. Point profile at a directory and a month of rotated logs is read in name order as one bill.

Or by hand, if you already have the repo checked out:

- run: npm ci && npm run build
- run: node packages/cli/dist/index.js check prompts/system.txt --max-tokens 2000

The rest of the commands, in their own book

optimize, check and init above are the front door. Every other command has its own chapter — same prose, same worked examples, one page — in the command reference: the measured multiplication (--from-log), the cache reorder, the CI baseline, the fleet, the plan and its verification, the provider pull, the gateway, the evaluations that spend money and say so first, and everything else the table above links to.

trazum --version prints the version on its own, and works when your config is broken — which is exactly when somebody is asking.

Web

npm run build:web
npm run dev:web        # http://localhost:3000

An interface for pasting a prompt, tuning the usage scenario, and reading the word-by-word diff, the saving and the advisories. Includes optimisation history stored only in the browser — nothing leaves your machine.

Reordering for the cache is available here too, behind a checkbox rather than a level, with the same warning the CLI prints and the same refusals reported. It is the largest saving Trazum can make, and it should not need a terminal to find.

And there is a Compare tab. Two versions of a prompt, and what the edit did: the token delta, what it costs per month, and which advisories and rules it introduced or resolved. Every figure is after - before, so positive means worse — the opposite of the rest of Trazum — and the page says so above the numbers rather than beside them, because a reader arriving from Optimise has the opposite convention already loaded.

Compare what the rules would leave is off by default and the default is the interesting half: your edit changed the text as written, so the text as written is what you are being asked about. Trimming both sides first hides a prompt that doubled in length and happened to double in courtesy.

The usage scenario is shared between the two tabs. Setting 50,000 calls on one and reading 10,000 on the other would make their answers incomparable while looking like they were about the same workload.

So are phrase-level rewrites. Two switches: one asks the model for suggestions, the second takes them. They are listed above the saving, one line each — You should always make sure to → Always ~4 ×2 — with a count of how many the checks threw out, because "four did not survive" is the useful fact and which four is noise unless you are debugging the model. Nothing is applied unless the second switch is on, and turning the first one off clears it.

And a "Your bill" tab, which is trazum profile in the browser: drop or paste a usage log and read where the money went — the spend split, whether caching paid for itself, the levers that would actually move the bill, conversation growth, and the answers that were cut off mid-generation. The log is parsed entirely in the page against the bundled pricing catalogue. Nothing is uploaded: there is no fetch in that component, a test fails if one appears, and the only analytics event carries two booleans. A usage log names your workloads, spend and conversation counts — exactly the file nobody should have to hand to a server to see a report on it.

The drop zone reads more than logs. A Claude Code project folder — ~/.claude/projects as it sits on disk — prices every transcript in the page, labelled by project, with the counts crossing and never the words. An OpenTelemetry export prices its GenAI spans the same way. And a price card — an OpenRouter /models response, or the same overlay JSON a --pricing file holds — widens the catalogue every figure in the tab prices with, so a model the bundled snapshot has never met (your Qwen, your self-hosted rate from trazum ownrate) gets the same exact arithmetic, still without a single request leaving the page.

Under the report, the rest of the loop. The ranked plan — each action with its money as a projection or a measured stake and never both, the typed assumption it rests on, and the command that would check that assumption — and below it, Did it work?. Save plan.json writes byte-for-byte what trazum plan -o writes, so a plan made in a tab can be committed, gated on in CI, and opened back here later. Opening a saved plan turns the log in the tab into the check on it: three outcomes, never two, with the three cannot-tell reasons kept distinct. Saved as a file rather than offered as a link, because a link would mean this page storing somebody's bill somewhere — an access-control question nobody has designed. The plan format is documented.

A guided tour walks the public tabs — Optimise, Write, Compare, Bill and the Playground — ringing each panel in place with a sentence on what it answers. It never auto-plays: a first visit is offered it once, and the compass in the rail starts it any time after. The Playground tab is the CLI itself in the page — 13 commands that spend nothing and touch no network, over sample files already loaded, through the same @trazum/core functions the terminal runs, so trazum profile usage.jsonl can be tried before anything is installed, against data that never existed outside the browser.

The HTTP API behind it is public and small:

# Metadata: models, and whether an LLM is configured on the server
curl https://your-deployment/api/optimize

# Optimise
curl -X POST https://your-deployment/api/optimize \
  -H 'content-type: application/json' \
  -d '{
    "prompt": "Please, in order to help me, analyse {{x}}. Thanks!",
    "level": "safe",
    "locale": "en",
    "reorder": false,
    "suggest": false,
    "applySuggestions": false,
    "usage": { "model": "claude-opus-5", "callsPerMonth": 20000, "avgOutputTokens": 300 }
  }'

reorder, suggest and applySuggestions are honoured only on a literal true — the body is untrusted, and a truthy check would let "false" rearrange somebody's prompt. With reorder, the response carries what moved and what was declined, and original stays the text you sent so a diff shows the move. With suggest, it carries every proposal that survived the checks and everything rejected and why — present even when the model proposed nothing, so "nothing was found" is distinguishable from "you did not ask". applySuggestions without suggest is a 400, not a no-op, refused before any call to the model.

# Compare two versions: what did this edit cost?
curl -X POST https://your-deployment/api/compare \
  -H 'content-type: application/json' \
  -d '{
    "before": "Classify {{x}}. Answer with the category only.",
    "after": "Please kindly classify {{x}}. Thank you!",
    "optimizeBoth": false,
    "usage": { "model": "claude-opus-5", "callsPerMonth": 50000 }
  }'

POST /api/compare returns every figure as after - before, so positive means worse. Both endpoints are rate limited (30/min per IP), with a bucket each. And /api/optimize will not fetch an LLM endpoint a caller names: a request may only select one from TRAZUM_ALLOWED_LLM_ENDPOINTS, empty by default — stricter than filtering the URL, because a hostname an attacker registered can resolve wherever they like. See SECURITY.md.

Signing in (optional)

Off by default, and a deployment that leaves it off is the tool this README has been describing all along: paste a prompt, get an answer, nothing remembered.

Set three variables and the sidebar grows a Sign in button at its foot:

TRAZUM_GITHUB_CLIENT_ID=Iv1.xxxx
TRAZUM_GITHUB_CLIENT_SECRET=xxxx
TRAZUM_PUBLIC_URL=https://trazum.example

A fourth, TRAZUM_DATABASE_URL, points it at any Postgres so sign-in survives a restart; without it sessions live in memory and the account menu says "temporary session". Trazum asks GitHub for read:user and nothing else, never stores the access token, and stores session cookies only as their SHA-256. Misconfigure any of it and sign-in simply stays off, with /api/auth/* answering 503 naming the variable to set.

Signed in, a Library tab appears: prompts you saved and every version of each, append-only, token counts recomputed on read rather than stored. On the Compare tab, Create share link publishes a comparison at /c/<token> for anyone holding the URL — expiring after thirty days by default, revocable, kept out of search engines, and saying what it publishes before the button. Every share link doubles as a README badge at /badge/<token>.svg, recomputed on every load, with no script and no prompt text. Set TRAZUM_ADMINS and /admin totals what every prompt on the deployment adds up to — names and token counts, never anybody's prompt text, and deliberately not a spend report.

docs/accounts.md has the setup, the schema, every security decision and why, the limits, and an explicit list of what is not covered.

Deploying to Vercel

The repo is an npm workspaces monorepo; Vercel handles it with no special configuration:

  1. Import the repository in Vercel.
  2. Root Directory: apps/web. The rest — installing from the workspace root, building @trazum/core via prebuild — is automatic.
  3. Optional variables: TRAZUM_LLM_* to offer the LLM pass without users supplying keys, NEXT_PUBLIC_POSTHOG_KEY for analytics, TRAZUM_GITHUB_* and TRAZUM_PUBLIC_URL for sign-in.

Vercel runs more than one instance, so if you enable sign-in there, set TRAZUM_DATABASE_URL as well. Without it each instance keeps its own sessions in memory and a browser is signed in against one and signed out against the next.

Library

import { optimize, refineWithLlm, openAiCompatible } from '@trazum/core';

const result = optimize(prompt, {
  level: 'safe',
  locale: 'en',
  usage: {
    model: 'claude-opus-5',
    callsPerMonth: 50_000,
    avgOutputTokens: 500,
    cacheHitRate: 0.9,
    batchEligible: false,
  },
});

console.log(result.optimized);
console.log(result.savings.monthlySavingsUsd);

reorderForCache is the API behind --reorder. It returns the original text unchanged when nothing can safely move, and always reports what it declined and why — a saving Trazum chose not to take is one the caller cannot evaluate:

import { reorderForCache } from '@trazum/core';

const r = reorderForCache(prompt, { minPrefixTokens: 1024 });  // the model's minimum

r.text;                 // the rearrangement, or `prompt` byte-for-byte
r.tokensMoved;          // moved out of paid-every-call into the prefix
r.prefixTokensBefore;   // 14
r.prefixTokensAfter;    // 1174
r.declined;             // [{ reason: 'backward-reference', phrase: 'above', text }]

minPrefixTokens is a bar on the resulting prefix, not on the amount moved. A prefix below the model's minimum caches nothing at all, so a rearrangement that does not clear it buys nothing — but a head that already clears it gains from any block that joins it, however small.

comparePrompts is the API behind trazum diff. Note the sign: everything it returns is after - before, so positive means worse — the opposite of result.savings, and the reason it lives in its own module.

import { comparePrompts, formatSignedUsd } from '@trazum/core';

const change = comparePrompts(oldPrompt, newPrompt, { usage });

change.tokenDelta;                      //  +37   (grew)
formatSignedUsd(change.monthlyDeltaUsd) //  "+$9.25"
change.advisories.appeared;             //  ['contradictory-instructions']
change.rules.noLongerFiring;            //  what the edit cleaned up

Two entry points. @trazum/core is browser-safe and imports no Node builtins — that is enforced by a test that walks the import graph, not by convention, because the web app bundles it and one node:fs import anywhere in that graph fails the build. Anything that reads the filesystem lives on @trazum/core/node:

import { loadConfig, walkPrompts } from '@trazum/core/node';

const { config, path } = await loadConfig();   // null path = none found
const { files, truncated } = await walkPrompts('prompts/');

parseConfig and budgetFor are pure functions of their arguments, so they sit o

Installation

Source-derived launch command. Check the maintainer’s required arguments and credentials before running:

bash
npx -y @trazum/mcp

Set up in your AI client

Merge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.

json
{
  "mcpServers": {
    "io-github-davmunrey-trazum": {
      "command": "npx",
      "args": [
        "-y",
        "@trazum/mcp"
      ]
    }
  }
}

Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.

Claude Desktop setup reference

Package

@trazum/mcpnpm

Compatible MCP Clients

io.github.Davmunrey/trazum works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.

  • Claude Desktop~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.
  • Cursor~/.cursor/mcp.jsonRestart Cursor for changes to take effect.
  • VS Code.vscode/mcp.jsonReload VS Code window for changes to take effect.
  • Windsurf~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect.
  • Claude Code.mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.

Learn More