Documentation
Optimus is an MCP server. It installs into your editor, holds the catalogue of harms AI agents cause in production — the ones people talk to, and the ones that quietly do a job — and asks your editor to drive your agent into them. There is nothing to add to your code and nothing to deploy.
Install
One command. It needs Node 18 or newer, and nothing else — npx fetches the package on each start, so there is no global install to keep up to date.
claude mcp add optimus \ --env OPTIMUS_API_KEY=opt_your_key \ -- npx -y @optimus-ai/mcp@latest
- Get a key — free, no card — and put it where opt_your_key is, then copy the block.
- Paste it straight into the Claude Code window, in the repository you want to test. It picks the command up and runs it for you — a terminal in the same directory works too.
- If the optimus tool does not appear, restart Claude Code. The server is resolved when the editor starts, so a running one keeps whatever it had.
- Ask in your own words, at whatever scope suits — “test this agent before I ship”, “check the refund flow for edge cases”, or “scan the changes in booking.ts”. It answers to the scope you set.
Get a key from your account page. It is free and authorises nothing: the checks run on a laptop that has never signed in. What a key does is tell a run which workspace its history belongs to, so a second machine and a colleague start from what the first one worked out.
Get a keyIt speaks stdio, which is the standard MCP transport, and nothing in it branches on which client is calling. Anything that can launch a process can run it, and the three fields above are what every config file wants.
Upgrading
Restart your editor. The command above says @latest, so npx re-resolves against the registry when the server starts, and a running server keeps whatever it launched with — which is the only reason a restart is needed at all.
If you installed before that was pinned, your config probably says @optimus-ai/mcp with no version, and npx will reuse whatever it already has in ~/.npm/_npx rather than checking. Add @latest to the args in your MCP config; that is the whole fix. Optimus tells you when the version answering you is behind, and the server reports its version in its handshake if you want to check yourself.
Your first run
Ask your editor to test your agent. There is one tool, optimus, and it works by talking rather than by running a fixed suite — it reads what your repository declares, asks what kind of agent it is, proposes the harms that could apply, and asks you to confirm or reject each one.
> test this agent before I ship it
The repository is the directory you are standing in, and your environment is already loaded in that process — so your database, your agent’s configuration and your provider keys are all present without us ever asking for them or holding any of them.
Making it the habit
Your editor reaches for Optimus when a request sounds like testing. To make that the rule rather than the exception, commit the block below to wherever your editor keeps its standing instructions — CLAUDE.md in Claude Code, .cursor/rules/testing-this-agent.md in Cursor, AGENTS.md in Codex, GEMINI.md in Gemini CLI. All of them are read automatically, and committing it means the whole team gets the same behaviour rather than whoever remembered.
## Testing this agent Before shipping a change to the agent, run it through Optimus — the `optimus` tool. It holds the catalogue of harms AI agents cause in production, what this repository has been bitten by before, and what we have already decided is by design. Use it when: - a change touches the agent's prompts, tools, guardrails, or its data - somebody asks whether a change is safe to ship - somebody asks what is wrong with the agent, or why it did something Follow its `next` instruction rather than summarising and moving on. It will ask for things, and it will challenge claims it is not willing to take on trust. Answering those challenges is the point of it. Never report "nothing found" without saying in the same breath what was checked and what could not be.
That last line is the one worth keeping whatever else you change.
What it asks you for
Optimus proposes; your editor confirms. A name does not settle what a tool does — create is a payment processor, a state store and a support ticket in three different codebases — so the binding is a question, and your answers are written down and reused.
| It asks | Because |
|---|---|
| which tool plays a part | A harm has to be pointed at one of your tools before it can be checked. When nothing here plays the part it needs, it comes back as `couldNotAim` with the question that would settle it. "Nothing plays that part" is a fine answer and retires the check honestly. |
| how many agents are here | A repository is not an agent. A checkout holding a chat agent beside a nightly pipeline has two answers to every question below, and answering for whichever was found first stands down half the checks on a fact about the other one. Name each, and each is planned for and reported separately. |
| what the agent is for | One harm cannot be checked without it, and it is the expensive one: the customer leaves without the thing they came for. There is no guessing it from a tool name, so name every tool that counts as finishing. |
| how it may be driven | Plenty of repositories point at production from local development. Nothing is driven live until you have said how — see below. |
| where your records live | So a query can be written without exploring for it again. More than one place counts: an application database, and whatever you instrument with. If a person has scored or corrected an answer anywhere, that is the strongest evidence there is — they have been grading your agent since it shipped. |
| whether a finding is real | You reject it with a reason. Rejections are remembered, so a second run does not re-raise what your team already settled. |
Working with the access it has
Optimus never asks for more than you have given it and never refuses to run because something is out of reach. It runs with what it has, and the report says afterwards what that cost — which checks could not be looked at, and what granting access to a particular collection would answer next time.
That is a decision for you to make after seeing something useful, not a permission dialog before you have seen anything.
Naming what the agent is for
"census": {
"purpose": { "tools": ["generate_offer"], "is": "price a car" }
}Name every tool that counts as finishing. An agent with two ways to finish — price it, or book an inspection — breaches every conversation that took the other one if you name only the first.
Reading a report
A finding carries the harm, the turns it happened on, and the evidence that settles it. Three kinds of check produce them, and most tools only do the first.
| Kind | What it needs |
|---|---|
| provoke | A run you can drive — a conversation, or the agent invoked on one input. Makes the harm happen and catches it before it ships. |
| inspect | Read access and no conversation. Describes the shape of a bad row in words; your editor writes the query against your schema. We never touch your database. |
| trend | History. Watches a number across runs — "fell to 41, against 63 across 7 earlier readings, between v1.3 → v2.0" — and says which change moved it. Nothing here says what a number ought to be. |
Where the evidence came from
Every finding says which, because it decides what may be claimed. A conversation we drove shows a harm can happen. A query over what already happened shows it did, to a share of real people. A real conversation somebody had is the strongest of the three, and it is graded accordingly.
Two of them deliberately carry less weight. A query run against a test or seeded environment is reported as such and never graded as though customers were affected — whatever rate a fixture holds is the rate somebody wrote into it. And a reading of how a conversation felt is marked as a judgement rather than a verdict, in its own section, with the turn it refers to quoted from the transcript.
What it cost somebody
Alongside the findings, a report gathers what the agent costs a customer: handovers begun and never finished, conversations that ended with nothing done, how long it takes to reach the thing the agent is for. These are movements, not failures — a measurement says a number is not what it usually is, never what it ought to be.
It answers a different question from a tracing dashboard. A run that is fast, cheap and error-free can still be one where a third of customers had to say the same thing twice.
Leads, not findings
Some results come back as leads. The bugs that bite at scale are not about one thing going wrong: nothing is missing, each piece is correct, and two correct pieces are wired together at the wrong moment. Optimus reads the relationships — two tools given different values for the same fact, a call acting on a figure that was revised after it ran, a call that almost always has company arriving alone, something firing after a person took over.
The decisive evidence for those is the record the conversation left behind, which is a query away and not in the transcript. So Optimus says what looks wrong and where to go and look, rather than filing it as proven.
The five verdicts
| breach | It happened, with the evidence. |
| clean | The check ran properly and nothing was wrong. The only pass. |
| notReached | The conversation never got far enough to tell. |
| blind | Nothing could be seen — no tool calls, or a query that examined nothing. |
| noBaseline | A number was read, with nothing yet to compare it against. |
A check that did not run is not a check that passed. Reporting any of the last three as success spends your trust on a test that never happened — and only breach and clean fill a coverage cell, so the number cannot be raised by instrumentation breaking.
Driving it live
A suite that books appointments and sends messages does those things to real people. So nothing is driven live until you have said how. You are asked once, and the answer is remembered.
| fresh-record | A throwaway made through your own creation flow, then deleted. |
| emulator | An emulator or local stack — nothing reaches production. |
| existing-test | A test agent or account you already keep. |
| never-live | Do not drive it at all; query what already happened instead. |
Inspections carry on meanwhile — querying your own data needs no permission of this kind. And never-live is a real answer: the report says plainly what went unchecked rather than quietly returning less.
Scheduled checks
Testing before you ship answers whether a harm can happen. Checking on a schedule answers whether it is happening, and whether a number that was steady has started to move.
npx @optimus-ai/mcp run --watch --fail-on-regression
--watch runs the queries and takes the measurements and drives nothing. It creates no records, needs no permission to touch a live agent, and exits non-zero when something that was clean comes back. A GitHub Actions template ships in the package.
Why twice a day rather than weekly
A trend needs several readings before a movement means anything — below that, the first change after you install would always look significant. At one run per deploy that takes over a month, so the checks that watch a number rarely got to say anything at all. Twice a day they are useful within the week.
It needs to know how you run a query
A scheduled check answers by querying your own records, and a run with nobody at the keyboard has nobody to approve a command. So tell your assistant once, in words, what you would run — the client and the flag that takes a statement, not a whole query. It is remembered, and every run after it can answer instead of reporting that it could not look.
| Where it runs | Your machine or your runner, on the Claude subscription you already have. We trigger nothing and hold no schedule. |
| What it creates | Nothing. Queries and measurements only, whatever you answered about driving it live. |
| How you hear | The exit code, which your scheduler already surfaces, and the run on your dashboard. |
| One at a time | Two runs in one repository would overwrite each other’s history, so the second says so and stops. |
Seeing tool calls
Almost every harm worth checking is about whether the agent did what it said it did, so the calls have to be visible. Add two lines to the process that drives your agent — it patches the provider clients in place, whatever your entry point is, and needs no change to your agent.
const { observeProviders, drainObserved } =
await import('@optimus-ai/mcp/observe');
await observeProviders();In any other language there is nothing to install. Optimus tells your assistant what the observer has to do — wrap the provider client on the class before your agent is imported, read the name and arguments out of the reply, hand the reply back untouched — and names every shape a tool call arrives in, both spellings included. Your assistant writes those twenty lines into the driver script it is already writing.
That is a decision rather than a gap. A helper package per language is a version skew waiting to happen: an agent pinned to last month’s helper against this month’s server, drifting quietly with nothing to say so. Optimus runs nothing and your assistant runs everything, so the language of your repository is its problem to bridge, not ours.
Either way it starts before your agent is imported — a module that has already bound the function it calls keeps the unwrapped one — and drains after each turn, so a call belongs to the turn it happened on. A provider that cannot be watched goes unobserved, which reports as blind rather than as a pass.
Nothing is sent anywhere by this. The calls are read in your process and handed to the checks running in the same editor session.
Coverage, and knowing when to stop
A list of scenarios cannot be audited: fifteen conversations were run, and nobody — including whoever wrote them — can say what the sixteenth would have been. So coverage is a cross-product, which can be named: every check you confirmed, crossed with what it ran against.
coverage: 9 of 48 checks have a verdict — 19%.
39 have never been run, and 3 of those are either something
that has bitten here before or something that cannot be undone.Nobody runs forty-eight cells. The promise is not exhaustion — it is that everything unrun is named and ranked, so wherever you stop you can say what you stopped short of. It accumulates across runs and across your team, which is the part a fresh session cannot do.
The shapes your data takes
| complete | What a fixture looks like. Finds nothing. |
| missing-fields | The agent fills the gap from somewhere else. |
| near-duplicates | Two rows match; nothing chooses between them. |
| at-the-edges | Caps, floors and rounding written for the middle. |
| malformed | A parser returning empty rather than failing. |
| absent | Whether it says so, or answers anyway. |
Which of your rows are which is yours to say — we never see them.
Finishing
A run that stops finding things has either run out of defects or run out of the part it was looking at, and nothing in the findings can tell you which. So a run finishes only when findings have flattened and nothing is untried: a tool no probe has used, a state never entered, a shape never run against, a pair of tools never tried together. Any open door beats a flat curve.
Proving it to yourself, once
We run one comparison ourselves: agents with defects planted in them, tested by Claude Code with Optimus, by Claude Code alone, by DeepEval and by Promptfoo — the same agents, the same judge, scored by one rule. Every report carries the latest figures and the version they were measured on — or says plainly that none has been published for the version you are running.
On your own agent you cannot compare a run with Optimus against the same run without it — the counterfactual is gone the moment you take the help. So run the control group yourself. Optimus proposes nothing; you test the agent however you would have anyway; then it scores what you did against the whole space of checks that apply.
OPTIMUS_BLIND=1
An environment variable rather than something the editor can set, deliberately: a control group the subject can opt out of is not a control group.
Your own tests
Optimus does not replace them, and the report says so. Your suite tests your business; judging a business needs a domain we do not have and do not claim.
| What it means | |
|---|---|
| yours and ours | Checks your suite already has a test for. |
| ours only | Checks that apply to this agent and no test here would catch. |
| yours only | The rest of your suite — your business, which we cannot judge. Usually most of it. |
What you get that a suite structurally cannot give you is the denominator. Fourteen tests pass, and nobody — including whoever wrote them — can say what the fifteenth would have been. Tell Optimus what your suite covers and it divides that by the checks that apply to this agent.
What we store and why
Two places hold anything, and they hold very different things. The first is a directory beside your repository, which never leaves your machine. The second is us, and it is worth being exact about.
On your machine — .optimus/
Everything the checks learn: how to reach the agent, which of your tools play which part, what your team has already judged by design, and the run history. Commit it and your team shares its decisions the way it shares its tests — in a diff, with a reason in the commit message.
A .gitignore is written into it excluding three files: session.json and last-report.json, which hold verbatim text from conversations your agent had, and queries.json, which holds the read-only queries written against your own data so the same question is asked next time. Reading and writing the rest needs no network, so a run works on a plane — with the reasoning stages standing down, and the report saying so rather than quietly returning less.
What reaches us
Some of Optimus’s own reasoning runs on our side rather than on your machine, so a little does reach us:
- The findings, and the evidence behind them — anonymised on your machine first.
- The probe conversations Optimus itself asked you to drive, anonymised on your machine first.
- Counts — how many, of what kind, how long a run took.
- Your reason for rejecting a finding, once your machine has checked it carries none of your names.
- How your customers open, paraphrased, once your machine has checked each line names nothing of yours and carries no value. Shown and deletable on your data page.
Anonymised means that before anything is sent, names, emails, phone and card numbers, order references, URLs, file paths and tool names are replaced with placeholders like [id-2] — the same value always getting the same one, so a double booking still reads as one. The answer is put back on your machine before you read it, and the map never leaves. It works by pattern, so it is thorough about anything with a shape and cannot promise a business name written in plain prose.
What does not
- Your source, and every prompt in it.
- Your tool names. They are your business logic, and not ours to know.
- Your database, and every row in it.
- Your provider keys and your environment.
- Any conversation Optimus did not itself ask you to drive.
Your account
Signing in stores your email address, which sign-in providers you used, when the account was created and when you last signed in. A key is stored as a SHA-256 digest and a four-character hint — never the key, which is shown once at creation and cannot be shown again.
You can see everything held against your account, and delete any of it, on your data page. Deletion removes the documents rather than hiding them, and the privacy page lists the same thing in full.
Configuration
Environment variables, set where the MCP server runs.
| Variable | What it does |
|---|---|
| OPTIMUS_API_KEY | The only credential Optimus asks for. Joins this machine’s history to your workspace and runs the reasoning stages. |
| OPTIMUS_BLIND=1 | Optimus proposes nothing and scores what you did instead. The control group. |
Guardrails of your own
A set of guardrails applies to every agent by default. Your own go in .optimus/guardrails.json beside the repository, or you can simply state one during a run and it is written down there.
Troubleshooting
| Symptom | What it means |
|---|---|
| the tool never appears | The server did not start. Restart the editor first — a running server keeps the version it started with. If it still does not appear, run the `npx` command from the install block in a terminal and read what it prints. |
| every verdict is blind | No tool calls were visible. Add the two observe lines to the process that drives your agent, or say which entry point starts it. |
| every verdict is notReached | The probes are not getting far enough, usually because the binding is wrong. Check which of your tools Optimus thinks plays each part and correct it — the correction is remembered. |
| it will not drive anything | You have not said how it may be driven, or you said `never-live`. The report names what went unchecked as a result. |
| it reports couldNotAim | Nothing in your repository was found to play a part a harm needs. Name the tool, or reject the cause if the answer is genuinely nothing. |
| nothing found, and you doubt it | Read the coverage line. "Nothing found" against 4% coverage and "nothing found" against 60% are different claims, and the report always prints which one it is making. |