Google AQuA: Turn Silent AI Failures Into a Product Repair Queue
Google's new AQuA diagnoses production agent failures. A founder's guide to choosing repairs, checking coverage, and proving user tasks improved.
On October 8, Google published AQuA, an Ambient Quality Agent reference implementation for finding recurring problems in production agent conversations. It addresses a familiar product failure: the service responds successfully, but the customer’s task goes wrong. Google’s announcement describes a diagnostic assistant that operates beside the application. It does not change the agent or approve releases on its own.
For founders shipping an AI support assistant, booking product, or internal workflow app, the immediate question is practical: when automated review produces another list of issues, which deserves the next afternoon of your team’s time? A finding can be plausible, frequent, and still point toward the wrong repair.
This guide proposes a repair queue built around the user’s intended task, the strength of the evidence, and the cost of leaving the problem unresolved. You will get a reusable issue worksheet, a small-team review meeting, and a way to decide when a repair has enough support to expand. These are recommendations, not results from a YBuild experiment. No customer outcomes, model comparisons, or production savings were measured for this article.
1. Understand what changed without buying a promise of automatic quality
The noteworthy change is the availability of a public diagnostic workflow that connects production conversations with the application version behind them. AQuA’s documented sequence samples sessions, reviews them, groups related findings, checks candidate groups, and tracks surviving issues. Investigation can consult a deployment’s source snapshot. Its automatic resolved status means an issue has not appeared for fourteen days; it does not establish that a repair caused improvement.
The public recipe is a starting point with integration prerequisites, not a plug-in guarantee for any no-code application. A founder should ask their builder whether the necessary traces and deployment records exist before buying an implementation project. A downloadable repository is different from a supported service with a support commitment.
Three terms help make the purchasing conversation concrete. A trajectory is the sequence of messages and tool events during a task. A finding is a claim about a particular departure from expected behavior. A repair queue is your team’s list of decisions about those claims: investigate, contain, change, verify, defer, or reject. The queue is an operating practice proposed here, not an AQuA feature name.
This matters even if you never install AQuA. Ask an existing analytics supplier to show a user task that failed while infrastructure looked healthy. Then ask what evidence lets a product owner distinguish a bad recommendation, missing information, a tool problem, and an intentional refusal. If the supplier can show only a score and a generated explanation, the diagnosis still needs work before it becomes a development instruction.
2. Define the user promise before configuring the reviewer
An AI reviewer cannot decide what your product owes the customer unless somebody states it. “Be accurate and helpful” does not settle whether a booking assistant may reserve an alternative appointment, whether a refund assistant may promise payment, or whether a drafting tool should ask a follow-up question.
Start with one sentence describing an observable promise. For a scheduling product: “When a customer changes the location, the assistant must check availability for the new location before describing an appointment as available.” Follow it with the evidence that would satisfy the promise: a current availability result for the selected location and time, followed by an accurate status shown to the customer.
Google’s ADK evaluation guidance distinguishes the final answer from the steps taken to produce it and asks teams to define objectives before choosing metrics. That distinction is useful at the product level. A pleasant final answer can conceal a missing check; an awkwardly worded answer can accompany a correct reservation. Decide which matters for the task before grading tone.
Write exceptions beside the promise. A user browsing options may leave without booking. A customer may explicitly choose to wait for a human. Neither is necessarily a failure. Conversely, a refusal that protects an unavailable appointment is useful only if the customer understands what happened and can continue.
Choose a narrow first scope, such as appointment changes, rather than reviewing every interaction against every possible virtue. Ask the product owner and a person who handles customer escalations to agree on examples. Their disagreement is valuable: it exposes an unsettled product rule. Resolve that rule before asking an automated judge to enforce it at scale. Otherwise, the queue will become a machine for generating arguments your team has not decided how to answer.
3. Separate service health, review coverage, and user success
A healthy endpoint, an active reviewer, and a successful task are three different observations. Google’s agent monitoring documentation lists operational signals such as request counts and latencies. These help operators understand whether the service is running. They do not tell a founder whether the chosen appointment exists or whether a customer’s revised request was honored.
Your product report should show three separate rows. First, service health: did requests complete and were tools reachable? Second, review coverage: which eligible tasks supplied usable evidence and were actually reviewed? Third, task outcome: among the reviewed tasks, what did the evidence establish?
Keep the denominators explicit. Suppose a hypothetical app had 200 location-change tasks during a reporting window, 80 usable traces, and 40 reviewed cases. A statement about those forty cases describes those forty cases. It cannot be presented as the failure rate of all two hundred tasks without a justified sampling and estimation method. These counts illustrate reporting, not measured traffic.
Also distinguish a complete trace from one containing only the final response. The latter might establish what the assistant told the customer, but not whether an omitted availability event means no check occurred. Missing instrumentation and missing behavior can look alike. Put evidence gaps into the queue rather than resolving the ambiguity in the model’s favor.
The online monitor documentation exposes sampling settings and caps, reinforcing that reviewed traffic is selected traffic. For your first deployment, ask for a run report that includes excluded workflows, incomplete traces, evaluation errors, and time windows with no usable data. A blank chart should trigger “we do not know,” not “nothing went wrong.” Record changes to the sampling policy too, because a new filter can move a score without changing the product.
4. Walk one fictional failure from complaint to decision
Consider SlotDesk, a fictional scheduling assistant for a small service business. A customer first asks for a Thursday appointment at the central branch. Later they switch to the north branch. The assistant retains the original availability result and presents the north-branch slot as available. The conversation looks smooth, the request returns normally, and no tool reports an error.
The proposed queue item is specific: “After a branch change, the assistant can describe availability using the previous branch’s result.” It is not “the agent hallucinates,” which leaves the builder guessing where to start. Attach the user’s change request, the earlier result’s branch, the displayed claim, and the application version. Mark the example fictional; a real implementation must supply its own records.
Next, consider alternatives. Perhaps a fresh north-branch lookup happened but was not captured. Perhaps the user was shown a draft itinerary labeled unconfirmed. Perhaps the business allows interchangeable branches and the display is merely unclear. The reviewer should establish the product rule and the evidence before assigning a defect to a model or prompt.
If the promise was violated, containment can precede the deeper repair. SlotDesk could temporarily describe changed-location options as unconfirmed and offer a human availability check. That costs convenience and staff attention, but prevents an unsupported promise while the team investigates. The founder owns that tradeoff; a diagnostic score cannot decide it.
The eventual change could affect state handling, tool routing, instructions, or the interface. The important acceptance result is the same: after the branch changes, the product obtains appropriate current evidence or explains that confirmation is pending. A builder should return a demonstration of that behavior, with a neighboring case that still works, rather than only a patch summary. This scenario is a planning example, not a test of AQuA or evidence that any commercial scheduler has this defect.
5. Rank by consequence and evidence, then consider frequency
A frequently observed annoyance should not automatically outrank a rare unsupported commitment. Begin with the consequence: what happens to the person if the issue remains? Then assess whether the attached evidence supports that consequence. Only after those questions should frequency and repair effort influence ordering.
The NIST AI RMF Core recommends prioritizing documented risks using impact, likelihood, and available resources. The following table adapts that principle into a small-team product queue. Its categories are proposed operating choices, not a regulatory classification or vendor scoring system.
| Situation | Product decision | Evidence required before the next step |
|---|---|---|
| A supported finding shows an unauthorized or unsupported commitment | Contain the affected capability and assign an owner | Actual user-visible claim, relevant action result, applicable rule, affected version |
| Repeated task failure has a supported mechanism and a recoverable outcome | Schedule a bounded repair | Representative complete cases and a defined expected outcome |
| A severe claim relies on incomplete records | Investigate urgently; apply precaution proportionate to possible harm | Missing-event explanation and a safe way to confirm the claim |
| Frequent complaints describe unclear status rather than incorrect state | Review interface and communication | What users saw, underlying state, and how they interpreted it |
| A reviewer flags behavior the product intentionally permits | Reject or revise the reviewer rule | Documented exception and an example that fits it |
| Low-impact inconvenience has clear workarounds | Defer with a revisit condition | Named owner, reason for deferral, and escalation trigger |
Use frequency carefully. Count distinct eligible tasks, not repeated spans, repeated retries, or multiple findings in one conversation. A large cluster can contain several mechanisms requiring different owners. Split it when the expected outcome or proposed change differs.
Avoid multiplying arbitrary severity, confidence, and frequency numbers into a precise-looking score. If your team cannot explain why an item outranks another, the formula has not solved the decision. A short written rationale is often more useful: “Contain location-change confirmation because an unsupported slot creates an external promise; investigate greeting repetition later because customers can still complete the task.”
6. Use a worksheet that survives the handoff to a builder
The reusable artifact below is an issue record. Copy it into your tracker or spreadsheet and keep one record per failure mechanism. Link to restricted evidence rather than pasting entire customer transcripts into a broadly shared ticket. For SlotDesk, the record would name location-change availability as its scope and branch-specific confirmation as its expected result.
| Field | What the owner must record |
|---|---|
| User task and promise | The outcome the customer sought and the rule the product owes them |
| Observed departure | What happened, including the user-visible consequence |
| Evidence and access | Trace or receipt references, completeness, and who may read them |
| Version and conditions | Application revision, relevant model/tool settings, time, locale, and workflow |
| Coverage boundary | Eligible tasks, captured tasks, reviewed tasks, exclusions, and errors |
| Competing explanations | Missing telemetry, intentional exception, external change, or alternative mechanism |
| Consequence and containment | Who can be affected and the temporary operating decision |
| Repair owner and hypothesis | Named person, proposed change, and why it should help |
| Acceptance cases | Original case, meaningful variations, and neighboring behavior to preserve |
| Release and reversal | Bounded rollout, stop condition, and how to restore safe behavior |
| Closure evidence | Results, remaining failures, observation window, and reviewer |
| Revisit trigger | Rule or version changes, renewed complaints, or newly available evidence |
This worksheet deliberately separates diagnosis from acceptance. A finding can justify investigation while its proposed cause remains uncertain. A completed change can justify a test while the product remains unready for a broad rollout. Use plain status labels your team understands: investigating, contained, ready to test, limited release, accepted within scope, and deferred.
Ask your builder to return evidence in the same record. “Updated the prompt” is an implementation description. “The north-branch request now receives a north-branch result, and unavailable slots are labeled pending” describes behavior you can review. Both are useful, but only the latter answers the product promise.
Keep unresolved observations visible after acceptance. If one variation still fails, record the exclusion and the operating restriction. Closing a narrowly scoped repair must not erase adjacent risk or silently rename it success. The worksheet should let another team member understand that decision next month without asking the original author to reconstruct it from chat.
7. Verify the repair in an environment that matches the claim
Begin by reproducing the relevant departure safely, then compare the proposed change with the previous behavior. Fix the rule, conditions, and expected outcome before examining results. Preserve the original case as a regression case, but add variations that prevent a narrow patch from merely recognizing the wording.
For SlotDesk, variations might include changing branch twice, changing the date after the branch, an unavailable slot, and a tool that cannot return a result. Include a neighboring case where the location never changes. These are suggested cases; their number is not a validated minimum or a statistical guarantee.
A replay of conversation text is not a reconstruction of the world. Inventory can change, a customer record can be edited, and a tool can answer differently. Use disposable test data or controlled tool responses for the first comparison, and ask the builder to state which conditions were reconstructed and which were simplified. Never replay a real booking or payment into a live system just to obtain a screenshot.
Inspect meaningful outcomes rather than requiring every internal step to be identical. If the assistant correctly asks for missing information, that may be an acceptable alternative to a direct answer. Decide those alternatives in advance. Human review is particularly useful when the expected customer experience has several legitimate forms.
For production expansion, compare the changed version with a suitable control and watch the affected workflow separately. Google’s canary release guidance explains why aggregate metrics can conceal a problem in a small release population and why stateful tests need care. Your product-specific stop condition might be any observed unsupported confirmation, subject to an agreed evidence check. Zero such cases in a bounded pilot is a result within that pilot, not proof that the probability is zero everywhere.
8. Protect the evidence without destroying its usefulness
Production review creates another place where customer conversations can be read. Before enabling it, identify the data that the reviewer needs, the people allowed to access it, and the systems that receive it. A tool running within your cloud project still needs an appropriate collection purpose, permissions, retention, and a clear operational owner.
OpenTelemetry’s sensitive-data guidance recommends minimizing collection and describes removal and redaction mechanisms. It also warns about limits of hashing. Replacing a name with a stable hash does not automatically make a detailed conversation anonymous.Use a fictional version of the SlotDesk case for general planning. Restrict the actual customer record to the people investigating that issue. A developer can often receive a synthetic reproduction with branch identifiers and expected results, while a designated reviewer retains access to the original evidence. State when redaction removes a fact needed for diagnosis; a sanitized record is not necessarily a complete record.
Trace storage, exported issue details, screenshots, and test datasets may have different retention paths. Make a small inventory of those copies and assign responsibility for each. Do not promise that deleting a support ticket deletes every diagnostic copy unless your actual data flow supports that claim.
Finally, keep customer communication separate from internal speculation. “We are checking an appointment-confirmation issue and will confirm your slot with staff” is a usable operational message. Sending a generated root-cause theory to the customer as established fact adds confusion. Explain the current status, the safe next action, and what the customer should rely on. Privacy review should reduce unnecessary exposure while preserving enough accountable evidence to make and audit the repair decision.
9. Run a short review meeting that ends in decisions
A small team does not need a daily ceremony around every generated insight. Set a review cadence that matches task volume and potential harm, with an immediate route for serious incidents. The scheduled meeting should include the product owner, the builder responsible for changes, and somebody familiar with customer outcomes.
Bring a short queue with evidence ready to open. For each item, answer four questions: what promise was broken, what supports the finding, what should happen now, and what would change our decision? If the evidence is unavailable, assign retrieval rather than debating a generated summary. If the rule is disputed, assign a product decision rather than pretending the issue is purely technical.
Google’s SRE monitoring chapter argues for actionable signals and cautions against noisy alerting. Apply that principle to the meeting: a new cluster does not automatically deserve a page. Route urgent customer harm to incident handling; route recurring low-impact friction to the planned queue.
Track the time spent investigating rejected claims as well as accepted defects. That makes reviewer noise visible. Also track issues with no owner and repeated deferrals. More diagnostic output is useful only if it leads to better decisions without consuming the team’s capacity to deliver repairs.
End each item with a named owner, an action, and a review trigger. If the same issue returns, reopen the record with new evidence rather than creating an unrelated ticket. If you change the reviewer’s rule, annotate the trend so a lower finding count is not mistaken for a better product. A quieter dashboard can mean better behavior, narrower review, or a broken pipeline; the meeting should establish which explanation is supported.
10. Decide whether the outer loop is worth adding now
AQuA is most relevant when you already have a launched agent, accessible trajectories, identifiable application versions, recurring workflow failures, and somebody able to act on findings. An ADK-based team may find the reference implementation a useful starting point. Other stacks can borrow the operating practice while choosing different instrumentation and review tools.
If you have a handful of pilot conversations, manual review against a clear task promise may be enough. If you cannot identify the deployed version, first improve that record. If nobody owns customer escalation, appoint an owner before expanding diagnostic output. If critical actions need real-time authorization, build that control in the action path; a later review cannot stop a commitment already made.
Start with one workflow and a bounded pilot whose success condition is operational: can the team distinguish a supported defect from an evidence gap, assign a justified action, and review whether the change improved the intended task? Record software cost, review effort, repair effort, and remaining uncertainty. Compare that burden with a manual process using the same product rules. No universal saving follows from adopting a diagnostic agent.
The durable deliverable is the completed issue record, not the number of insights on a screen. For SlotDesk, the founder should be able to show what changed, which cases were checked, what remains excluded, and who will respond if unsupported confirmations recur. That is enough to make a bounded decision today while keeping the next decision open to evidence.
References
- Google Developers Blog: AQuA announcement, October 8, 2026. Original product account; its demonstrations are not YBuild tests.
- Google ADK recipes: Ambient Quality Agent. Implementation and prerequisites; related to the announcement, not independent validation.
- ADK: Why evaluate agents. Objectives, trajectories, and final-response evaluation.
- Google Cloud: Monitor an agent. Operational monitoring signals.
- Google Cloud: Continuous evaluation with online monitors. Sampling and evaluation monitoring.
- NIST: AI RMF Core. Risk prioritization, measurement, and post-deployment management.
- Google SRE Workbook: Canarying releases. Release comparison and stateful testing limitations.
- OpenTelemetry: Handling sensitive data. Data minimization and redaction limitations.
- Google SRE Book: Monitoring distributed systems. Actionable monitoring and alert noise.