DocsWhat can it doThe runtime and its seams

The runtime

The execution trust layer for AI agents. An agent tells the runtime what it intends to do, the action, the target and the environment, and the runtime makes a deterministic decision before anything executes.

It is Apache-2.0 licensed, runs on your own machines, and needs no account, no API key and no network call to do its job.

Every action walks the same five stages, on your own machine, whichever agent asked. The verdict is settled by the first stage with something to say about it, and a withheld call never reaches the ones after it.

What it is not

It is a gate, not a worker. It answers "does this violate a rule, and who authorizes it if nobody has?" and never does the work itself.

Memnox does not read your code. It has no import graph, no diff scanner and no editor integration. It never generates, edits or commits anything and never reviews a pull request. The one wall it draws is around your own agent, when memnox run --untrusted starts it inside a kernel sandbox, and nothing runs in there on Memnox's behalf.

Running it

bash
npx memnox setup              # see what your agents reach, answer once, and it is wired
memnox status                 # where the machine stands, and what happened today

Setup puts the interceptors in ~/.memnox/bin, a hook in front of every tool call in Claude Code, Codex, Cursor, Gemini CLI and Windsurf where each is installed, every MCP server through the proxy, a baseline of rules, and the daemon under launchd or systemd as a user service. The baseline is split in two: the denies on secret reads are about this machine rather than a repository, so they go in ~/.memnox/machine.policies.toml and apply in every repository, and the rest go in memnox.policies.toml. It starts in observe. Each step is still its own command for somebody who wants one at a time:

bash
memnox protect --yes          # write a baseline of rules
memnox protect --interceptors # put the wrappers on PATH
memnox mcp wrap               # route the MCP servers through the proxy
memnox daemon --install       # hand the daemon to the machine
memnox run -- claude          # start an agent behind all of it

The daemon then keeps what setup drew: an agent installed later is hooked, a hook something removed is put back, and a new MCP server goes through the proxy, each with a desktop notice. It never puts back what a person took out on purpose. The daemon keeps the boundary has the whole of it, and Untrusted repositories and new agents has the layer under the rules that nobody has to write: writes kept in the repository, the egress proxy, and probation.

There is no server to run and no port to open. The runtime is a CLI, a local socket and files under ~/.memnox/, which is why "no account required" is architecture rather than a free tier.

An agent's own tools meet the same rules

The hook setup installs rules on each tool call before it runs, so a read of ~/.ssh, a web fetch or an MCP call is decided even where no wrapper stands in the way. A file read, search or write becomes filesystem.read or filesystem.write on the absolute path, a fetch becomes http.request on its host, a command line goes through the shell classifier, and an MCP tool becomes mcp.<server>.<tool>. An allow says nothing, so the agent's own permission prompt still runs, and a tool the table does not know is left to the agent.

Each agent's hook system allows a different amount, and this is what each one lets Memnox enforce:

Agent

Claude Code

Every tool. Allow is silence, deny refuses, and ask asks when a person is at it. In bypassPermissions there is nobody to ask, so an ask is refused instead.

Codex

Every tool its hook reports. Deny only: no prompt is documented, so an ask is refused instead.

Gemini CLI

Every tool. Deny only, so an ask is refused instead.

Cursor

Commands, MCP calls, file reads and writes, with ask for commands and MCP. Cursor does not send the MCP server's name, so its calls match as mcp.*.<tool>, and web fetches have no hook at all.

Windsurf

Reads, writes, commands and MCP calls. Deny only, so an ask is refused instead, and web fetches have no hook.

When the rules cannot be read, the mode decides. In enforce a hook that could not rule refuses the call and says the refusal is Memnox's and why. In observe it lets the call through, as every other seam does while watching.

Three positions, and each one is honest

Memnox does not have to sit in the middle of every request to be useful. It can stand in three places relative to an agent, and most teams use all three at once, on different agents.

Position

Context

Somebody asks what governs a task before an agent starts it, with memnox check. Nothing is recorded and nothing is refused. It exists because a refusal after the work is done costs the work, and a briefing before it starts costs a question.

Governance

The agent, or the platform running it, asks whether an action may happen and honours the answer. Memnox is not in the execution path, so what this adds over the first position is authority and the record.

Enforcement

The runtime sits between the agent and its tools. A denied action does not run, and there is nothing for the agent to ignore.

Only the third is unconditional. In the first two the agent could fail to ask, so put the proxy in the path wherever the consequences are real: production, payments, customer data. Nothing forces the march, and the position is chosen per agent rather than all at once.

The seams, one product at a time

Each row is a seam that ships, and the column that makes the table worth reading is the last one. A governed agent with an unwatched side channel is worse than an ungoverned one, because somebody has stopped worrying about it.

Agent

Claude Code

The MCP proxy, the governed shell, the PATH wrappers, the git hooks, and its own permission file through protect --apply-native. Blind to the model's own reasoning.

Cursor, Cline, Roo

The MCP proxy, the PATH wrappers and the kernel guard. Blind to edits made inside the editor rather than through a tool or a command.

Codex CLI

The governed shell, the PATH wrappers, git, and the kernel guard over the credentials it was handed. Blind to execution that happens on the provider's side.

Copilot coding agent

Issue assignment, the pull request gate, required checks and the Actions boundary. Blind to everything before the pull request appears.

Agents in their own VM

Network egress, brokered credentials, and the systems it reaches. Blind to the whole interior.

Agents inside products you buy

MCP if the vendor speaks it, and otherwise only the credential and its scope. Blind to everything the vendor does.

Any MCP client

The proxy, on every tool call and every tool result. Blind to very little, which is why this is the seam to reach for first.

Pipelines

The OIDC exchange, environment protection and the deploy call. Blind to steps that need no credential.

Why MCP first

It is the one seam that is provider neutral, already in the developer's config, and carries the result as well as the request. Every client speaks it, so one proxy governs Claude Desktop, Cursor and Zed at once rather than three integrations that drift apart.

The result direction is the half that is easy to skip and expensive to skip. A tool result is content the agent read, and content an agent read must never become an instruction it follows. See Need to know for how that is enforced by type rather than by a detector.

The seam that always exists

An agent that cannot be wrapped can still be starved. Hand it a capability scoped to one operation on one resource for a few minutes instead of a key that lives a year, and every system it reaches becomes a place a verdict can stand.

  1. 1

    It asks by operation, not by secret

    postgres.write on the orders database, not "the database password". The request is a decision request like any other, and it is evaluated the same way.

  2. 2

    The broker mints a lease, not a credential

    Scoped, counted, and carrying the id of the decision that allowed it, so the record says why an agent held a credential and for how long.

  3. 3

    Expiry belongs to the issuer

    Never to the agent's good behaviour. A lease that has run out is refused by the broker whatever the agent believes.

Turn one on at a time

Enforcing everything on the first day is how a team turns all of it off on the second. Each seam has its own mode, so a machine can run the proxy in enforcement while its shell wrapper is still only observing.

Where it runs

One personA teamAn organization
Installnpx memnox setupthe same, on each machinememnox login, then memnox setup, plus the hosted console
Account requirednonenoneSSO via your identity provider
StorageSQLite and files under ~/.memnox/the same, per machineabove, plus retention
Approvalsthe terminal, or a second terminalthe samerouted to Slack, with identity
Recorda local ledger, exportable signedthe sameacross every machine

One person genuinely means zero infrastructure, and everything above it is additive: nothing about the local path changes when a team adopts it. What the hosted half adds is the things that only mean something across more than one person, which is a different product rather than a bigger runtime.

Retention is a setting rather than a flag:

bash
memnox config set retentionDays 365
memnox purge --dry-run

Design principles

  • Deterministic core. No LLM, no network call, no clock and no randomness in the decision path.
  • Fail closed. Unknown identity, unreadable state or ambiguous input denies rather than guesses. An unknown verb is reported as unknown and counted, never silently treated as safe.
  • Everything auditable. A decision that cannot be proven afterwards did not happen.
  • Small, inspectable pieces. The policy engine has zero dependencies, and every store sits behind an interface defined in @memnox/core.
  • Ports over lock-in. Files are the only implementation shipped; any backend can implement the same interfaces.

You do not need to know any of that to run it. It matters because it is the reason "no account required" is true rather than a marketing line.