Documentation · Operating
Setting up auditing
This page is an outline. What is here is accurate, but it is
not yet the whole story — each section ends with a note on what is still to be
written. For anything it does not answer, DEPLOYMENT.md and
USING-MCP.md in the repository are the complete references.
There are two halves, and the first one is on by default.
A local trail, kept on the machine the servers run on and downloadable from the console. Every installation has it, free tier included. How long it is kept is what a licence raises.
Streaming, which sends the same records off the box to Kafka, Splunk, S3 or syslog. That is a licensed feature, and it is what you want when the record has to outlive the container or live somewhere the audited process cannot reach.
| Variable | Meaning |
|---|---|
MCP_AUDIT_LEVEL | names (default) or values. |
MCP_AUDIT_MAX_BYTES | Cap on the recorded response, default 8192. 0 for no cap. |
The local trail
| Variable | Meaning |
|---|---|
MCP_AUDIT_FILE_DIR | Where the trail lives. The console defaults it under the config directory. |
MCP_AUDIT_RETENTION_DAYS | How long to keep it. Default 1. 0 keeps nothing. |
MCP_AUDIT_FILE_MAX_BYTES | Hard cap so a busy server cannot fill the volume. Default 512 MB. |
Records are JSONL, one file per writer, rolled into segments named for the moment they were closed —
so ls shows how far back the trail goes. Admin → Audit shows how much is held and downloads
the whole thing as one file, every source merged into one timeline, oldest record first.
That merge is the point of it. The proxy records who called which tool and whether it was allowed; each generated server records what the tool did. Two files would leave you to do the join.
Records that age out are gone. That is the feature rather than a limitation: the window is what makes a store of production data defensible, and one that quietly kept everything would be a worse position to be in. The window is enforced per segment, so a record can outlive it by up to one segment.
MCP_AUDIT_RETENTION_DAYS=0 is supported and means what it says. Nothing is kept here; the only
audit bytes on disk are records the spool has not yet delivered. It is the setting for a deployment
that streams to a SIEM and wants no production data resting on this machine. It is never a default —
an unset window is one day, and a value that cannot be read stops start-up rather than quietly
becoming zero.
Streaming it off the box
MCP_AUDIT_SINK is the class name of a com.mcpdbwizard.pub.McpAuditSink. A mistyped one stops
start-up rather than leaving you believing records are leaving the machine when they are not. Four
ship:
MCP_AUDIT_SINK | Needs | Key settings |
|---|---|---|
com.mcpdbwizard.pub.KafkaAuditSink | kafka-clients on the classpath | MCP_AUDIT_KAFKA_BOOTSTRAP, _TOPIC |
com.mcpdbwizard.pub.SyslogAuditSink | nothing | MCP_AUDIT_SYSLOG_HOST, _PORT, _PROTOCOL |
com.mcpdbwizard.pub.SplunkAuditSink | nothing | MCP_AUDIT_SPLUNK_URL, _TOKEN_FILE, _INDEX |
com.mcpdbwizard.pub.S3AuditSink | the AWS SDK on the classpath | MCP_AUDIT_S3_BUCKET, _PREFIX, _REGION |
The local trail and a stream run together: setting a sink does not switch the trail off.
Syslog: use TCP. Over UDP a record goes onto a socket and nothing comes back, so a collector that is down is indistinguishable from one that recorded everything — tolerable for logs, a poor fit for evidence. Messages are RFC 5424, octet-framed so a record containing a newline cannot split in two at the collector.
Splunk: prefer MCP_AUDIT_SPLUNK_TOKEN_FILE to MCP_AUDIT_SPLUNK_TOKEN. A token in an
environment variable is visible in docker inspect and in whatever holds your task definition. The
console will not let you type the token itself for the same reason — it would land in a settings
file in plaintext.
S3: records become objects, not a stream. They are batched and rolled by size or age into
<prefix>/<config>/yyyy/MM/dd/<timestamp>-<uuid>.jsonl, so a lifecycle rule can expire a prefix and
Athena can partition on the date. No credentials are configured — the SDK’s default chain finds
the task role or service account, which is what a deployment in AWS actually uses.
names is the default on purpose
At values the record carries argument values and the response body — production data, chosen by a
model. That turns your sink into a store with retention, encryption and erasure obligations.
Switch it on deliberately, not by accident.
A truncated response still carries its full byte size and a SHA-256 of the whole payload, with
truncated: true. Truncation costs readability, not integrity.
At values, the trail is a personal-data store and you are the data controller for it. Three
things follow, and they are worth deciding before you switch it on rather than after:
MCP_AUDIT_RETENTION_DAYSis your storage-limitation control. Pick a window you can justify.- Erasure is by that window expiring. Records are append-only, so there is no way to remove one
and leave the rest — which is also what makes the trail worth citing.
0keeps nothing at all. - The console’s download is an export of production data. It is admin-only, and the page says so at the point you click it.
names stays the default and an upgrade will never change it for you.
Surviving an outage: the spool
MCP_AUDIT_SPOOL_DIR turns on a write-ahead spool. Every record goes to disk before any
delivery attempt and is removed only once the sink confirms it, so records survive a sink outage and
survive the server dying — the spool is read back and replayed on the next start.
| Variable | Meaning |
|---|---|
MCP_AUDIT_SPOOL_DIR | Spool directory. Unset means no spool. Put it on a persistent volume. |
MCP_AUDIT_SPOOL_MAX_BYTES | Cap, default 100 MB. |
MCP_AUDIT_SPOOL_ON_FULL | drop (default) or block. |
MCP_AUDIT_SPOOL_FSYNC | never (default) or always. |
MCP_AUDIT_SPOOL_KEY / _FILE | Encrypts each spooled record (AES-256-GCM). Unset means plaintext. |
Three sentences to read before relying on it.
Delivery is at-least-once. A crash between delivering a batch and deleting it replays that batch, so
a consumer can see a record twice. Every record carries an id — dedupe on it.
fsync=never survives the process, not the machine. Records live through a crash, an OOM kill or a
container restart. They do not survive power loss. always covers that too, at a disk round trip per
tool call.
A full spool refuses new records rather than discarding old ones. Watch for Audit record not spooled in the log — it means the trail has holes.
Expect a web/ subdirectory under the spool: the console records proxied requests while each
generated server records its own tool calls, and a spool tolerates exactly one writer. That is the
isolation working, not a stray directory.
Encrypting the spool
It protects the FILE, not the PROCESS. Anyone who can read the container’s environment or its memory has the key. It defends what outlives the process and travels: disk images, volume snapshots, backups.
Encrypting the volume is usually the better answer — no code, and it also covers the configs, the accounts and the runtime workspaces sitting in the same directory, which this does not.
Switching it on over an existing spool is safe. Switching it off, or changing the key, is not: those records can no longer be read, and an undecryptable segment is quarantined rather than retried or deleted. Drain the spool before rotating the key.
To write. The record’s field-by-field shape, and writing your own sink against the SPI.