MCPDBWizard

Documentation  ·  Operating

Setting up auditing

This page is an outline. What is here is accurate, but it is not yet the whole story — each section ends with a note on what is still to be written. For anything it does not answer, DEPLOYMENT.md and USING-MCP.md in the repository are the complete references.

There are two halves, and the first one is on by default.

A local trail, kept on the machine the servers run on and downloadable from the console. Every installation has it, free tier included. How long it is kept is what a licence raises.

Streaming, which sends the same records off the box to Kafka, Splunk, S3 or syslog. That is a licensed feature, and it is what you want when the record has to outlive the container or live somewhere the audited process cannot reach.

VariableMeaning
MCP_AUDIT_LEVELnames (default) or values.
MCP_AUDIT_MAX_BYTESCap on the recorded response, default 8192. 0 for no cap.

The local trail

VariableMeaning
MCP_AUDIT_FILE_DIRWhere the trail lives. The console defaults it under the config directory.
MCP_AUDIT_RETENTION_DAYSHow long to keep it. Default 1. 0 keeps nothing.
MCP_AUDIT_FILE_MAX_BYTESHard cap so a busy server cannot fill the volume. Default 512 MB.

Records are JSONL, one file per writer, rolled into segments named for the moment they were closed — so ls shows how far back the trail goes. Admin → Audit shows how much is held and downloads the whole thing as one file, every source merged into one timeline, oldest record first.

That merge is the point of it. The proxy records who called which tool and whether it was allowed; each generated server records what the tool did. Two files would leave you to do the join.

Records that age out are gone. That is the feature rather than a limitation: the window is what makes a store of production data defensible, and one that quietly kept everything would be a worse position to be in. The window is enforced per segment, so a record can outlive it by up to one segment.

MCP_AUDIT_RETENTION_DAYS=0 is supported and means what it says. Nothing is kept here; the only audit bytes on disk are records the spool has not yet delivered. It is the setting for a deployment that streams to a SIEM and wants no production data resting on this machine. It is never a default — an unset window is one day, and a value that cannot be read stops start-up rather than quietly becoming zero.

Streaming it off the box

MCP_AUDIT_SINK is the class name of a com.mcpdbwizard.pub.McpAuditSink. A mistyped one stops start-up rather than leaving you believing records are leaving the machine when they are not. Four ship:

MCP_AUDIT_SINKNeedsKey settings
com.mcpdbwizard.pub.KafkaAuditSinkkafka-clients on the classpathMCP_AUDIT_KAFKA_BOOTSTRAP, _TOPIC
com.mcpdbwizard.pub.SyslogAuditSinknothingMCP_AUDIT_SYSLOG_HOST, _PORT, _PROTOCOL
com.mcpdbwizard.pub.SplunkAuditSinknothingMCP_AUDIT_SPLUNK_URL, _TOKEN_FILE, _INDEX
com.mcpdbwizard.pub.S3AuditSinkthe AWS SDK on the classpathMCP_AUDIT_S3_BUCKET, _PREFIX, _REGION

The local trail and a stream run together: setting a sink does not switch the trail off.

Syslog: use TCP. Over UDP a record goes onto a socket and nothing comes back, so a collector that is down is indistinguishable from one that recorded everything — tolerable for logs, a poor fit for evidence. Messages are RFC 5424, octet-framed so a record containing a newline cannot split in two at the collector.

Splunk: prefer MCP_AUDIT_SPLUNK_TOKEN_FILE to MCP_AUDIT_SPLUNK_TOKEN. A token in an environment variable is visible in docker inspect and in whatever holds your task definition. The console will not let you type the token itself for the same reason — it would land in a settings file in plaintext.

S3: records become objects, not a stream. They are batched and rolled by size or age into <prefix>/<config>/yyyy/MM/dd/<timestamp>-<uuid>.jsonl, so a lifecycle rule can expire a prefix and Athena can partition on the date. No credentials are configured — the SDK’s default chain finds the task role or service account, which is what a deployment in AWS actually uses.

names is the default on purpose

At values the record carries argument values and the response body — production data, chosen by a model. That turns your sink into a store with retention, encryption and erasure obligations. Switch it on deliberately, not by accident.

A truncated response still carries its full byte size and a SHA-256 of the whole payload, with truncated: true. Truncation costs readability, not integrity.

At values, the trail is a personal-data store and you are the data controller for it. Three things follow, and they are worth deciding before you switch it on rather than after:

names stays the default and an upgrade will never change it for you.

Surviving an outage: the spool

MCP_AUDIT_SPOOL_DIR turns on a write-ahead spool. Every record goes to disk before any delivery attempt and is removed only once the sink confirms it, so records survive a sink outage and survive the server dying — the spool is read back and replayed on the next start.

VariableMeaning
MCP_AUDIT_SPOOL_DIRSpool directory. Unset means no spool. Put it on a persistent volume.
MCP_AUDIT_SPOOL_MAX_BYTESCap, default 100 MB.
MCP_AUDIT_SPOOL_ON_FULLdrop (default) or block.
MCP_AUDIT_SPOOL_FSYNCnever (default) or always.
MCP_AUDIT_SPOOL_KEY / _FILEEncrypts each spooled record (AES-256-GCM). Unset means plaintext.

Three sentences to read before relying on it.

Delivery is at-least-once. A crash between delivering a batch and deleting it replays that batch, so a consumer can see a record twice. Every record carries an id — dedupe on it.

fsync=never survives the process, not the machine. Records live through a crash, an OOM kill or a container restart. They do not survive power loss. always covers that too, at a disk round trip per tool call.

A full spool refuses new records rather than discarding old ones. Watch for Audit record not spooled in the log — it means the trail has holes.

Expect a web/ subdirectory under the spool: the console records proxied requests while each generated server records its own tool calls, and a spool tolerates exactly one writer. That is the isolation working, not a stray directory.

Encrypting the spool

It protects the FILE, not the PROCESS. Anyone who can read the container’s environment or its memory has the key. It defends what outlives the process and travels: disk images, volume snapshots, backups.

Encrypting the volume is usually the better answer — no code, and it also covers the configs, the accounts and the runtime workspaces sitting in the same directory, which this does not.

Switching it on over an existing spool is safe. Switching it off, or changing the key, is not: those records can no longer be read, and an undecryptable segment is quarantined rather than retried or deleted. Drain the spool before rotating the key.

To write. The record’s field-by-field shape, and writing your own sink against the SPI.