METR's Viewer Disclosure: Can Your Team Trust the AI Audit Screen?
A founder's guide to checking stored agent events, independent exports, and review screens before using an AI activity log to approve customer-facing actions.
On October 6, METR disclosed a proof of concept in which content from an AI agent could change what a human saw in the Inspect transcript viewer. The underlying trajectory remained in METR's database. METR says it has not observed agents exploiting this in its evaluations, and the vulnerability was patched after reporting. This was an isolated staging investigation, not evidence that a production audit database had been erased. See METR's original disclosure.
For founders building support agents, approval workflows, research assistants, or products that show customers an activity history, the immediate lesson is concrete: retaining an event and presenting that event faithfully are separate product obligations. A polished history screen can become the place where a person makes an unsafe decision if it hides, reinterprets, or fails to load decisive evidence.
This guide offers a three-part evidence check, a reusable acceptance worksheet, and a supplier brief. It helps a small team decide whether its review interface supports a real release decision. We have not tested your product or reproduced the exploit. The scenarios and acceptance rules below are proposed defensive checks, not measured results or a claim that every AI dashboard has this vulnerability.
1. Read the disclosure narrowly before changing your roadmap
METR describes a client-side JavaScript injection in a transcript renderer. Agent-controlled text could influence the displayed page, including a download control. Meridian Labs patched the reported issue within one day; METR also points to an October 1 change introducing an untrusted rendering mode. Those facts support reviewing the software that displays evidence. They do not establish a general rate of concealed agent behavior.
A founder should resist two tempting interpretations. “The records still existed, so nothing important happened” ignores the person relying on the screen. “Agents can now erase all audit evidence” ignores the stated boundary of this demonstration. The useful middle position is that a review decision can be wrong even when a recoverable record exists elsewhere.
Ask your team where customer-impacting decisions depend on a transcript, generated summary, activity feed, or downloadable report. Examples include approving a refund, accepting a supplier comparison, investigating an unexpected email, or deciding whether an agent stayed inside its permissions. If the screen is merely a convenience, the urgency differs from a screen that authorizes the next action.
Your roadmap response should follow that dependency. First identify the decision and its evidence route. Then inspect the relevant rendering and export controls. Do not commission a complete monitoring rewrite because a different product disclosed a patched issue. Equally, do not accept a supplier's statement that it “has logs” as an answer to whether your operator can reliably inspect them.
2. Separate the event, the export, and the view
An event record is the stored account of an occurrence: which actor requested which action, against which object, with what result. An export is a representation retrieved from storage for examination. A view is the interface that formats, filters, summarizes, or groups those records. Provenance describes where a record came from and which transformations occurred. A trust boundary separates content supplied by a less trusted actor from capabilities belonging to a more trusted system.
These terms matter because the same screen can combine all three stages. A green card may be a generated interpretation rather than a stored tool result. A download button may export only filtered rows. A transcript may show an assistant's claim of completion while the external service rejected the operation. None of those differences requires malicious behavior to mislead an operator.
Inspect's log documentation describes programmatic access to evaluation logs as well as interactive viewing. That provides a useful example of separate examination routes. It does not promise that arbitrary products have immutable storage or independent exports.For your own app, draw a simple chain: action producer → event collector → retained record → export service → review interface → human decision. Mark every place that changes content. Compression, redaction, translation, sorting, and summaries belong on the drawing, even when they are legitimate features.
Then ask who can change each stage. If the agent can write its own “approved” label, that label is a claim. If the event service creates the label after receiving the approver's decision, it has a different origin. The interface should make that distinction visible without asking an operator to infer it from color or confident language.
3. Treat rendered agent content as a product boundary
An activity viewer often displays Markdown, equations, code, images, and links because plain text is cumbersome. Those conveniences expand the software involved in displaying a record. A formatter, media loader, or mathematical renderer can interpret content in ways that an ordinary text display does not.
OWASP's XSS prevention guidance distinguishes output encoding, HTML sanitization, and safe text destinations. It also warns that changing sanitized content afterward can undermine the protection. The engineering detail belongs with your implementer; the product requirement is that agent-supplied content must not acquire control over reviewer navigation, approval controls, or unrelated records.A practical purchasing question is: “Which parts of this review screen interpret agent output, and how are they isolated from the controls I use to make decisions?” An answer naming only the frontend framework is incomplete. Ask about the whole transformation path, including export previews and embedded attachments.
Current Inspect viewer documentation offers inspect view --no-trust-content to show log content as plain text. This is a documented option, not a command we ran for this article. A team using Inspect should check its installed version and actual launch configuration rather than assuming a recent documentation page describes an older installation.
For a customer product, the equivalent could be a plain-text inspection mode with clearly separate action controls. The tradeoff is readability: tables and equations may become harder to interpret. Preserve enough structure to find the relevant event, and keep a path to attachments without silently executing or loading everything they reference. A safer mode that operators cannot use is likely to be bypassed.
4. A support agent scenario: the refund review that looks complete
Consider a hypothetical product, HarborSupport. Its agent reads a ticket, proposes a refund, and asks an operator to approve it. The operator sees a timeline containing the customer's message, the proposed amount, a tool call, and a green completion summary. This scenario illustrates a design problem; it is not an observed customer incident.
Suppose the customer qualifies for a partial refund, but the proposal concerns the wrong order. The review screen collapses the earlier order lookup and highlights the final assistant message. The operator approves because the amount looks plausible. The underlying records could be accurate while the default view still makes the wrong decision easy.
Now suppose an integration times out. The agent says the refund completed, but the product has no confirmed result from the payment service. A generated timeline summary replaces “confirmation unavailable” with “processed.” This is a separate failure: interpretation has promoted uncertainty into success. It can occur without a browser injection.
The founder's acceptance question is whether an ordinary operator can identify the order, amount, requester, approver, external result, and unresolved uncertainty before choosing the next action. A full transcript is available for investigation, but the main screen should surface the decisive fields with their origin.
Have the builder prepare a synthetic record where those fields deliberately disagree. The wrong order appears in the assistant narrative; the structured tool event names the intended order; the external result remains unknown. Ask an operator to explain what happened using the shipped interface, then compare that explanation with the prepared record. Record the explanation, not just whether the page loaded.
This exercise measures the product's evidence presentation under a controlled case. It does not measure an agent's attack capability. If the operator misses the disagreement, improve the information hierarchy and uncertainty labels before expanding automated authority.
5. Use an independent route for decisive evidence
“Download the transcript” sounds like independent verification, but it may be another function of the same page. METR's demonstration makes this distinction relevant. For a consequential review, specify where the export is produced and how a reviewer can retrieve it without trusting the rendered transcript to supply the file.
Independence is relative to a failure. A server-produced export bypasses a client renderer but can still share a faulty event collector. A storage snapshot can preserve original bytes but omit an event the collector never received. A payment receipt can establish an external effect but says little about whether a user authorized it. Choose the second route for the question you need answered.
OWASP's logging guidance includes verification of logging systems, access controls, and failure behavior. Use that as a reason to place the evidence path inside acceptance testing, rather than treating it as incidental operational plumbing.For HarborSupport, the proposed review packet would contain a stable run identifier, the retained event sequence, the export timestamp, the viewer version, and a separately obtained external refund status. The operator would reconcile object identifiers and outcomes. Where access controls redact a field, the packet would say it is unavailable to that reviewer instead of implying it never existed.
A hash can help compare a file with a trusted reference copy. It cannot prove that the original record described reality, that collection was complete, or that the party supplying both file and hash was independent. Likewise, append-only storage protects a particular history from particular modifications; it does not automatically validate every statement in that history.
These limits should appear in the supplier brief. Buy a demonstrable route from a defined source to a defined decision, not a vague promise of “tamper-proof AI.”
6. Reuse this evidence acceptance worksheet
The following matrix is a proposed minimum for a bounded pilot. Use synthetic accounts and harmless records. Ask an engineer to design any security-specific tests in an isolated environment; do not paste exploit material into a live customer console. The matrix tests product behavior and operator understanding, not universal resistance to adversarial agents.
| Case | Prepared condition | What the operator must be able to establish | Evidence to retain |
|---|---|---|---|
| Ordinary completion | One requested action and one confirmed result | Correct object, actor, approval, and outcome | Run ID, event list, operator explanation |
| Conflicting narrative | Assistant says one object; tool event names another | The disagreement is visible before approval | Both original fields and review decision |
| Uncertain outcome | Tool request exists; final confirmation does not | Completion remains unknown | Missing-confirmation label and follow-up route |
| Long output | Decisive warning sits inside a large transcript | Warning is discoverable without reading everything | Search/filter settings and chosen event |
| Rendering fallback | Rich formatting cannot be displayed safely | Content remains inspectable; controls remain separate | Plain-text output and viewer configuration |
| Filtered export | Screen shows a subset of a run | Reviewer knows whether the export is full or filtered | Scope label, event counts, export origin |
| Access restriction | Reviewer lacks one protected field | Restricted is distinguished from absent | Role, redaction notice, escalation owner |
| Collector failure | A synthetic logging interruption occurs | Review does not claim a complete history | Gap indication and release decision |
| Out-of-order arrival | A late event arrives after a summary | Timeline and decision state acknowledge the update | Sequence identifiers and updated status |
| Export mismatch | View and independent export disagree | Approval pauses while the difference is reconciled | Both copies, discrepancy, responsible owner |
Define expected answers before the demonstration. Otherwise, a convincing presenter can turn an unexpected result into the new acceptance criterion. For the conflicting-narrative case, a pass might require the operator to name both objects, state that approval is blocked, and identify who resolves the mismatch. That is a proposed requirement for this workflow, not a published benchmark threshold.
Capture the exact configuration as well as the screenshot. A test of plain-text mode says little about the rich mode customers actually use. A test by an administrator may miss what a support agent with narrower permissions sees. Include the role, browser, export path, filtering state, and relevant software versions in the record.
For each failure, decide whether to fix, limit, or defer the feature. A broken convenience filter might justify shipping without that filter. A screen that mislabels an unknown external effect as completed should block reliance on that screen for consequential approvals. Make the decision about a specific use, not about whether the entire product feels trustworthy.
7. Buy an operator workflow, not just a dashboard feature
A small team can turn the worksheet into a short supplier acceptance brief. Name the review decision, its consequence, the people who make it, and the original records they need. Specify the ordinary mode and the fallback mode. Request a demonstration of a disagreement, a missing event, and an unavailable export.
Use this reusable brief as a starting point:
We will use the activity screen to decide whether [action] may proceed for [user/object]. The screen must distinguish agent claims, recorded tool requests, human approvals, confirmed external results, and unknown outcomes. Agent-controlled content must not control reviewer actions or change unrelated evidence. A reviewer with [role] must retrieve [defined scope] through [export route] and reconcile it with [independent source]. If evidence is incomplete or contradictory, the product must show [blocking state] and route resolution to [owner]. Acceptance includes the attached synthetic cases in the actual shipped configuration.
Add delivery obligations: who owns the renderer dependency list, who assesses relevant security updates, who can disable rich display, and who confirms that an export still matches the retained scope after an upgrade. Ask the supplier to show the installed configuration, not simply a link to upstream documentation.
The Inspect command reference clarifies that its plain-text override applies regardless of a log's own trust configuration. This kind of precise operational behavior is what a supplier should explain for your chosen product. It should not be generalized into a promise about every viewer or copied as an untested setting for a different stack.
Price the work honestly. Independent export, permission-aware redaction, operator training, and dependency maintenance can require more effort than adding a timeline component. If the budget cannot support consequential review, narrow the pilot's authority. A draft-only assistant with manual verification in the original service may still provide value without pretending to offer a complete audit system.
8. Keep privacy, accessibility, and evidence usefulness together
A response to viewer risk can create a different problem if it copies every conversation into a broadly accessible archive. OWASP’s logging guidance identifies sensitive fields that should generally be excluded or protected. Auditability is not a reason to retain secrets, customer documents, or personal information without a defined purpose. Establish who needs which fields, how long the packet is kept, and how access is revoked.
Separate the operator's view from the investigator's deeper access. An operator may need an order identifier and confirmed result while an investigator needs the relevant tool payload under tighter permissions. The interface should explain why evidence is restricted and offer an accountable escalation route. “No records” is an inaccurate substitute for “records exist but your role cannot read them.”
Usability is part of evidence quality. A status shown only through a changing color may be missed by a reviewer. An update that arrives without an accessible announcement may leave someone acting on stale information. WCAG 2.2's status-message criterion requires status information to be programmatically determinable for assistive technologies without moving focus. This is one relevant requirement, not a full accessibility assessment of your product.
Ask operators to find a missing confirmation with keyboard navigation, zoom, and the actual assistive technology they use. Show textual states such as “confirmation unavailable” and “evidence incomplete,” and explain the next action. Preserve accessible structure in fallback mode. Do not make a safer display option usable only by the engineer who built it.
Balance these requirements through scoped access and clear explanations. A smaller evidence packet that reliably supports a named decision is often more useful than an unrestricted transcript dump nobody can interpret or safely share.
9. Decide what blocks launch and what can wait
For a read-only research assistant whose activity log is supplementary, a review-screen defect may justify disabling a rendering feature while continuing the pilot. For an agent that sends messages, changes records, or initiates financial actions, the same defect may remove the team's basis for supervision. The consequence and the available independent route determine the release decision.
Proposed stop conditions include an agent-controlled field changing an approval control, an unexplained mismatch between a view and its defined export, a missing event being represented as a complete history, or an unconfirmed external result appearing as successful. These are acceptance conditions we recommend for the described workflows, not facts established about your app.
A content security policy can add browser protection, but it is not a complete answer. OWASP's CSP guidance describes it as a defense layer, rather than a cure for underlying vulnerabilities. Let the engineer choose and validate the policy for the application; do not treat the existence of a header as proof of trustworthy evidence.
A controlled rollout can start with synthetic review packets, move to read-only customer cases with permission, and expand authority only when operators consistently identify conflicts and uncertainty. Keep a switch to disable rich rendering or stop consequential actions. State who can use it and where evidence remains available after it is used.
This framework is insufficient for high-assurance security evaluations or systems whose adversary can compromise the collector and all independent stores. Those require deeper threat modeling and technical verification. It also does not prove that an agent's reasoning trace is a faithful account of its motives. It evaluates the route by which your team receives recorded evidence.
10. What a founder can do today
Begin with one consequential decision in the product. Ask the builder to trace the decisive fields from their producer through storage, export, and display. Identify agent-written narrative separately from service-confirmed results. Choose one synthetic disagreement and one missing-confirmation case, and have the actual operator review them.
Next, obtain the export through a route whose independence is explicit. Compare scope, identifiers, and outcomes, not merely page appearance. If the screen and export disagree, stop using that screen as the basis for approval until the cause is understood. Preserve both versions and the configuration so the discrepancy can be investigated.
Finally, put responsibility into the release record. Name the person who maintains rendering dependencies, the person who reconciles missing evidence, and the person who can narrow agent authority. Decide which feature can be disabled without losing access to the records needed for follow-up.
The value of METR's disclosure for a small product team is a better question at procurement and launch: can the reviewer reach the evidence needed for this decision, through a path the reviewed content cannot silently redefine? Answer that with a scoped demonstration and retained artifacts before making the audit screen part of your customer promise.
References
- METR: AI systems could cover up misbehavior, October 6, 2026. Original disclosure and demonstration limits.
- Inspect: Log Viewer. Interactive viewing and plain-text option.
- Inspect: Log Files. Retention, programmatic access, and logging configuration.
- Inspect: inspect view command reference. Trust configuration behavior.
- OWASP: Logging Cheat Sheet. Logging verification and access considerations.
- OWASP: Cross Site Scripting Prevention Cheat Sheet. Encoding, sanitization, and safe rendering.
- OWASP: Content Security Policy Cheat Sheet. Browser defense in depth and its limits.
- W3C: WCAG 2.2, Status Messages. Accessible status communication.