Agent audit logs often outgrow tool output because one visible action generates many control-plane, provenance, policy, retry, and verification events around it.
A home agent may return a two-line confirmation after renaming one file, yet its audit trail records the prompt, plan, identity, authorization, target resolution, tool call, result, verification, and memory update. Parallel retrieval, retries, streaming snapshots, and repeated payloads multiply that difference. Growth reflects event fan-out and representation choices rather than only the number of bytes returned by tools.
Control-Plane Fan-Out Multiplies One Visible Action
A single tool operation may create events for planning, policy evaluation, capability issuance, approval, queueing, execution, timeout, retry, verification, and final response. Each event carries identifiers, timestamps, status, and enough context to reconstruct the decision path.
A managed provenance metadata storage system automatically captures lineage as managed metadata and discusses both the new functionality and overhead this creates. The same principle explains why agent accountability records relationships that ordinary application output omits.
Multi-step agents amplify the ratio because each step can branch into retrieval or validation subcalls. Ten short checks around a 200-byte tool result may produce kilobytes of headers and relational metadata before any prompt text is retained.
Payload Duplication and Snapshots Dominate Byte Growth
Systems often log full prompts, retrieved chunks, tool arguments, tool results, and updated state at several layers. Streaming token events and before-and-after snapshots repeat mostly unchanged content, while base64 images or embeddings inflate records further.
whole-system provenance captures whole-system provenance by observing information flow at the operating-system layer. Its approach shows how comprehensive lineage creates dense event graphs even when application outputs remain small. This distinction remains visible during later household testing.
Content-addressed blobs can store one payload once while events reference its hash. Deltas can replace full snapshots, and schemas can separate required reconstruction fields from optional debug detail; compression helps repetition but does not justify collecting sensitive content without purpose.
Retries, Retention, and Integrity Add Records That Never Reach Users
A failed tool attempt, policy denial, rollback, or verifier disagreement still belongs in the audit trail even though the final response may hide it. Hash chaining, signatures, indexes, and replication add integrity and query overhead beyond the event payload.
Research on append-only request auditing keeps append-only request records and historical versions inside a storage security perimeter. It reports that auditing has measurable but bounded performance cost, demonstrating that accountability is a separate stored workload.
The failure boundary is indiscriminate completeness. Logging every token and retrieved document forever increases privacy exposure and can make investigation slower. Define the questions the log must answer, then tier retention, deduplicate payloads, and preserve immutable summaries instead of removing causal identifiers.
Build a Per-Stage Audit Byte Ledger
Run representative one-step, five-step, and twenty-step workflows with success, denial, retry, timeout, and rollback. Count events and compressed bytes by planner, retrieval, policy, approval, tool, verification, memory, payload blob, index, and integrity metadata. The intermediate result must remain inspectable before automation follows.
Use the reconstruction goal in decision reconstruction logs to mark which fields are required to explain each consequential action. Repeat with payload hashing, deltas, sampling of low-risk reads, and tiered retention while confirming that investigators can still rebuild the decision.
Set budgets by event class rather than comparing logs only with tool-output bytes. If growth comes from repeated payloads, deduplicate them; if it comes from necessary causal edges, retain the edges and shorten optional diagnostic detail.
Tech & AI HUB
More to Read

What Features Enable a Home AI Trust Boundary Around Sensitive Files?
See how classification, capability-scoped access, isolated parsing, retrieval filters, egress policy, approvals, and audits contain sensitive home files.

What Factors Determine Whether Merkle-Tree Backups Detect Silent Change Efficiently?
Learn how chunk size, fan-out, trusted roots, cached hashes, change locality, metadata scope, and scrubbing determine Merkle backup verification cost.

What Components Enable Verifiable Backups of AI Indexes and Model State?
See how coordinated snapshots, content manifests, checksums, version locks, restore drills, and query tests prove that AI state can actually recover.

