Define the minimum event contract
| Event field | Minimum record | Why it matters |
|---|---|---|
| Identity | Event ID, actor, account, owner and time | Connects actions to an accountable route. |
| Workflow | Approved job, matter or project reference and data class | Shows the boundary the run was expected to follow. |
| System | Product, provider, model ID, version and region | Preserves the technical candidate actually used. |
| Evidence | Instructions, source identifiers and retrieval results | Lets a reviewer reproduce material assertions. |
| Action | Tool, arguments, permission, result and destination | Shows what the agent attempted and changed. |
| Control | Approval, denial, interruption, escalation and stop reason | Records whether guardrails operated. |
| Outcome | Output, reviewer corrections, disposition and system of record | Separates a model response from completed legal work. |
Preserve the useful inventory layer
Begin with sanctioned and unsanctioned tools, browser extensions, embedded vendor features, API keys, service accounts and automations. For each, record the business owner, users, workflow, data classes, provider, contract, integrations and system of record. This retains the useful discovery purpose of the original firm-wide audit.
Inventory alone is insufficient for agentic systems because the same tool can run under different accounts, models, permissions and destinations. Link every production event to a current inventory record and reject events whose route is missing or expired.
Record identity, time and version without ambiguity
Use a unique event ID and reliable timestamps. Record the human initiator, service identity, workspace or project, environment and accountable workflow owner. Preserve the exact model or endpoint when exposed; a product family or moving alias is not enough for later reconstruction.
Keep the release or configuration identifier for the surrounding application. A model can remain constant while the system prompt, retrieval index, tool schema or permission set changes. The audit chain should reveal those changes without requiring a forensic search across unrelated systems.
Minimize and classify inputs and sources
Record the approved data class and source identifiers, not unnecessary duplicate content. For public authority, keep the URL, publisher, date and retrieval time. For internal or client material, use controlled references, hashes or repository identifiers where appropriate and permitted. Define when the full prompt or document must be retained and who may access it.
Source lineage should show which material influenced the answer and which authority the reviewer opened independently. If retrieval returns stale, conflicting or unverified material, the event should record the uncertainty and the escalation rather than silently presenting one source as controlling.
Log tool proposals, authority and results
For every tool call, record the tool name and version, proposed arguments, permission decision, execution identity, destination, result, retry and error. Separate read, draft and write authority. A model can propose an action; only the approved system and person can authorize it.
Capture denials and interruptions as first-class evidence. A blocked external send, rejected file access or human stop shows that a guardrail operated. Logs that preserve only successful calls hide the control behavior a security or compliance reviewer needs to inspect.
Make approvals specific and reproducible
An approval should identify the decision, approver, time, evidence shown, scope and expiry. “Human in the loop” is not an auditable field. Record whether the person approved a source set, tool action, external destination, final output or exception. Do not reuse approval for one stage as authority for another.
Define stop conditions before the run: unknown data class, missing source, unsupported authority, permission expansion, unexpected destination, material output defect, incident signal or unavailable reviewer. Preserve the stop reason and follow-up owner.
Separate model output from reviewed outcome
Keep the raw or addressable model output when policy permits, the reviewer’s material corrections, the final disposition and the repository that holds the authoritative work product. Measure unsupported assertions, missed issues, citation failures, tool errors, review minutes and rework. Do not call a generated draft a completed matter event.
The record should also show when the output was rejected or used only for orientation. A clean distinction between generated, reviewed, approved and sent states prevents an audit dashboard from overstating production or legal reliance.
Apply retention and access controls to the audit trail itself
An audit trail can contain prompts, client identifiers, source excerpts, tool arguments, credentials or sensitive outputs. Classify the log, minimize content, encrypt it, restrict roles and define deletion, preservation and incident rules. Vendor logs and the firm’s own observability stack may retain different copies.
OpenAI, Google and Anthropic publish different data and safeguard structures. Link the event to the provider-specific retention record rather than asserting that “logging” has one universal period. Test whether the firm can export and delete its own traces without losing evidence it is required to preserve.
Connect audit events to change and incident response
When a model, contract, data route, prompt, tool or permission changes, identify affected workflows and reopen the relevant approval. Use the model-change runbook to preserve baseline, regression, pilot and rollback evidence.
If an event indicates unexpected access, disclosure or action, preserve the minimum facts needed for the reviewed incident process and restrict further access. Notification, legal hold and reporting decisions depend on the event, agreement and jurisdiction. The audit guide does not prescribe a universal deadline.
Classify workflows by data, action and consequence
Preserve the original audit guide’s useful classification step, but avoid a single generic risk score. Record the sensitivity of inputs, whether the system can act outside its interface, the importance of the output, the number of users and matters affected, and how quickly a failure can be detected and reversed. A public-source summarizer with no write access differs materially from an agent that can open client repositories and send external messages.
Use the classification to choose logging depth, approval points, sample rate and review cadence. High-consequence actions may require pre-execution approval and complete event capture. Lower-consequence orientation work may use sampling and shorter retention. Counsel and security should review categories and exceptions rather than relying on a label alone.
Combine complete control events with sampled content review
Capture identity, route, version, tools, permissions, approvals, stops and destinations for every production event when feasible. Review content proportionately. A full copy of every prompt and document can create unnecessary confidentiality and retention exposure, while metadata alone may not explain a material failure. Define which events trigger content preservation and which can use controlled references or hashes.
Sample ordinary runs and always review defined exceptions: unsupported citations, permission denials, external writes, new data classes, user complaints, incidents and material model changes. Keep the sample method, reviewer and result so audit coverage can be evaluated over time.
Reconcile records across the complete chain
The model provider, application, identity system, tool gateway, source repository and destination may each hold part of the event. Use a stable correlation identifier and synchronized timestamps to join them. Reconciliation should reveal missing segments, duplicate actions, retries, changes in execution identity and outputs that reached a destination without the expected approval.
Test reconstruction with a non-sensitive canary. Starting from the final work product, trace back to the reviewer, model response, tool results, source set, permission decision and initiating actor. Then start from the event ID and trace forward to the final disposition. Record gaps as control defects rather than filling them with assumptions.
Turn findings into owned remediation
Classify a finding by affected workflow, data, action and evidence gap. Name an owner, containment step, target date, validation method and reopening trigger. Examples include removing a personal account, narrowing tool permissions, adding a missing source field, changing retention, pausing an external action or rerunning a model regression. A finding is not closed because a policy document was edited.
Preserve exceptions with approver, scope, compensating control, expiry and review date. Expired exceptions should return to the decision queue automatically. Repeated findings can indicate that the workflow design, training or product choice needs to change rather than another reminder.
Report control health without overstating legal results
Useful operating measures include inventory coverage, events linked to an approved workflow, current model and source records, tool calls with explicit permission, stop events handled on time, deletion tests completed, overdue reviews, unresolved findings and time to reconstruct a sampled event. Show Unknown when a source is unavailable.
These measures indicate whether the audit system is maintained. They do not establish that confidential information was never exposed, that every output was legally correct or that privilege was preserved. Executive reporting should keep observed events, supported inferences, hypotheses and unknowns visibly separate.
Validate the audit system, not only individual runs
Periodically test whether clocks align, identifiers join across systems, permissions match the approved route and deletion works at each copy. Use a canary event to confirm that an ordinary run, a denied tool call, a human stop and a completed review all appear correctly. Record expected and observed evidence.
Independent review can sample configurations and findings, but it should not be described as a legal or security certification unless a defined assurance engagement actually supports that statement. Failed reconstruction is itself a finding with an owner and remediation date.
Test continuity when a provider or model changes. The historical event should remain understandable even if a current catalog no longer lists the identifier, a source URL moves or a tool version is retired. Preserve source snapshots where permitted, correction notes and replacement references. The goal is a defensible chronology, not an immutable claim that the original configuration remains current.
What an audit trail does not prove
NIST describes the AI RMF as voluntary. The event contract can support governance, reconstruction and review, but it is not certification and does not establish privilege, work-product treatment, discoverability, regulatory compliance or the absence of harm. Those conclusions require the actual facts and legal analysis.
FAQ
What should a law-firm AI audit log contain?
At minimum: event and actor, account and time, workflow and data class, model and version, instructions and sources, tool actions and permissions, approvals and stops, output destination, review and final disposition.
Should a firm retain every prompt forever?
No. Define a proportionate retention and access policy. Use controlled references or hashes where appropriate, preserve full content only when required and avoid turning the audit trail into an uncontrolled duplicate.
Does an AI audit trail determine privilege or discoverability?
No. It preserves facts for review. Legal treatment depends on the workflow, content, purpose, agreement, actions and jurisdiction.
Are denied tool calls worth logging?
Yes. Denials, interruptions and stops show whether guardrails operated and help reconstruct attempted actions.
Sources checked
- AI Risk Management Framework, checked 2026-09-03.
- Data controls in the OpenAI platform, checked 2026-09-03.
- Gemini API Additional Terms, checked 2026-09-03.
- Fable and Mythos 5.1 system card, checked 2026-09-03.
Operational information, not legal advice. Verify current terms, account configuration and applicable professional duties before use.