Google's AI Expertise Study: Redesign Onboarding Around User Judgment
AI-assisted success does not prove user readiness. Turn Google's new research discussion into an onboarding storyboard for correction, stopping, and escalation.
On October 7, Google Research published an explanation of a three-month experiment with patent lawyers. AI-assisted work improved, but independent professional judgment did not improve uniformly. The underlying NBER working paper was issued in September; this week's development is the public research discussion, not a newly completed experiment. Google's research account makes a distinction founders should carry into product design: helping someone produce a result and helping them understand a decision are different achievements.
If you build an AI support tool, proposal assistant, research workspace, or app builder, your onboarding dashboard probably celebrates a first successful output. That is useful evidence of activation. It does not tell you whether the customer can notice a wrong assumption, change the scope, or stop an inappropriate action tomorrow.
This guide turns the research into a proposed onboarding storyboard: demonstrate a useful result, expose one consequential choice, let the user repair a safe example, and make help reachable. It includes a reusable design brief and a way to evaluate the flow without turning every customer into a test subject. These are product recommendations, not an intervention proven by the patent study. No onboarding, retention, or customer outcomes were measured for this article.
1. Read the result without turning it into a universal warning
The NBER study description reports a randomized trial involving 133 patent lawyers at eleven U.S. firms. Expert graders were blinded. Access improved assisted drafting; after three months, an unassisted redlining advantage was concentrated among senior lawyers. Junior lawyers showed no average improvement, with more low and more good scores. That is neither universal skill loss nor proof that juniors cannot learn with AI.
The sample, profession, tool, and observation window matter. Patent judgment is not identical to configuring a customer-support assistant. An experiment on one workflow cannot establish that your signup sequence causes dependency. Treat the working paper as evidence for separating outcomes, then investigate your own users.
For an AI product, the useful question is narrower: what judgment does a person need to use this feature responsibly? A scheduling assistant may require recognition of a timezone conflict. A proposal assistant may require recognition that a claimed capability is unsupported. An app builder may require recognition that a preview is showing sample data rather than live customer records.
Name that judgment before writing tutorial copy. Otherwise, onboarding will teach the easiest interaction to instrument: pressing Generate. A person can become fluent in your buttons without becoming able to make the decision your product asks of them.
Also preserve the distinction between a tool's quality and a user's capability. If your model becomes more accurate, customers may legitimately need fewer corrections. You should not manufacture mistakes or friction to make the product look educational. You should ensure that the remaining consequential choices have clear evidence, usable controls, and an available owner.
2. Separate activation, readiness, and learning in your product promise
Activation means reaching a useful first outcome. Readiness means being able to handle the decisions required for the next real task. Learning means a capability persists or transfers beyond the exact guided example. These definitions are the working vocabulary of this guide, not a standardized certification scheme.They answer different questions. Did the user receive a usable draft? Can they recognize an unsupported promise? Can they apply that understanding to a different request later? A single completion event cannot answer all three.
There is good reason to keep the immediate value. The Generative AI at Work study found productivity benefits in customer support, especially among less-experienced workers, and offered suggestive evidence of learning. Its setting and outcomes differ from the patent trial. Together, these studies resist a simple claim that assistance either always builds expertise or always destroys it.
Your marketing promise should identify which outcome you actually offer. “Create a first draft” is a delivery promise. “Become better at deciding what to send” is a capability promise requiring additional evidence. If you have only measured generation success, do not quietly convert that finding into an education claim.
Readiness is also role-specific. A frontline operator may need to recognize a questionable recommendation and escalate it; an administrator may need to set the policy and review escalations. The operator does not need to understand model training. They do need to know which source is authoritative and whether their button sends a message or merely saves a draft.
Write one sentence for each role: “Before using this feature on live work, this person must be able to ___.” Choose an observable action. “Understand AI” cannot guide a design review. “Identify which policy date applies before sending a refund explanation” can.
3. Find the decision hidden inside the happy-path demo
Most demos remove uncertainty. They use a tidy input, an available source, an obvious answer, and a forgiving example. That makes the value visible. It can also conceal the point where a real user must exercise judgment.
Walk through your demo and ask what would change the correct next action. A missing requirement? A stale source? A customer outside the permitted scope? A draft that sounds confident while omitting a qualification? Pick one issue that your intended users actually encounter. Do not start with a spectacular adversarial example if ordinary ambiguity explains most support requests.
Research on the jagged technological frontier emphasizes that AI usefulness can vary across tasks within a workflow. That supports examining the actual decision, rather than labeling an entire job “AI-ready.” It does not identify your product's boundary for you.
Map the chosen decision into four plain questions:
- What information changes the answer?
- Where can the user inspect that information?
- Which action can they safely take when it is missing?
- Who owns the case if they cannot resolve it?
Design around the real boundary, including permission and workflow constraints. Sometimes the right improvement is a source panel. Sometimes it is a disabled send button with a clear reason. Sometimes it is changing the promise so that the product delivers a draft instead of an autonomous action.
4. Use a safe correction example, not a surprise trap
Consider ReplyDesk, a hypothetical product that drafts answers for a small software company's support team. This is an invented design scenario, not a customer or tested deployment.
The existing onboarding asks a new operator to paste a question, generate an answer, and send it to a sandbox inbox. The proposed revision uses a fictional customer asking whether their plan includes a feature. The visible plan record and the draft disagree. No real customer data or messages are involved.
The flow first shows a correct example so the operator understands the benefit. It then announces a practice case with an intentional mismatch: “This draft includes a claim you should check before sending.” That disclosure matters. A training example should not trick someone into making a real mistake or create distrust through a hidden test.
The operator opens the plan record, identifies the unsupported feature, removes the promise, and saves a corrected draft. A final unfamiliar case omits the plan record entirely. Here the appropriate action is to request information or route the case, not guess. The source remains accessible where it exists; the exercise withholds an AI recommendation, not the evidence needed to judge.
A product designer can implement this as a static practice screen before building adaptive tutoring. It needs a fictional record, a draft, an editable field, and a visible help route. Keep the correction small enough that the user can see why it matters. Do not bury the lesson in a long document full of incidental errors.
For an app builder, the analogous example could show a polished directory populated with sample customers. The user identifies the sample-data notice and chooses a preview step instead of a production announcement. They are practicing a product decision, not being required to debug unfamiliar code.
5. Offer help in layers while keeping production work usable
A useful practice flow can offer progressively more help: reveal the relevant source, point to the conflicting claim, explain the rule, and finally show a corrected example. This is a proposed interaction pattern. It should be evaluated with your users, and it should not become a mandatory obstacle in urgent work.
The PNAS mathematics experiment compared a standard AI interface with a tutor designed to protect learning. The standard interface helped practice performance but harmed later unassisted performance; the tutoring design largely mitigated that harm without establishing a positive unassisted gain. This was a school setting, not SaaS onboarding. It motivates attention to help design, not a claim that hints will improve your activation or retention.
In ReplyDesk practice, the first hint might say, “Check the plan record.” The next could highlight the relevant entitlement. The final example explains why the draft must remove the unsupported promise. Each layer should connect the action to a verifiable fact. An AI-generated explanation can itself be wrong, so maintain the practice answer against an approved source.
In production, do not force someone to solve the practice puzzle every time. Make evidence and corrections fast, preserve the user's ability to request a full draft, and route cases beyond their role. A learning mode and a work mode can share controls while having different pacing.
Let experienced users bypass basic practice after seeing the relevant boundary. Let new users return to examples without losing work. Support people who need accessible formats or additional time. Ability to type a polished explanation is not the same as ability to notice a wrong entitlement.
6. Build the controls that make judgment actionable
A user who spots a problem but cannot correct or stop it is not meaningfully in control. This is where onboarding design becomes a product requirement.
Microsoft's human-AI interaction guidelines include setting expectations, supporting correction and dismissal, and communicating changes. The accompanying HAX Toolkit helps teams prioritize these practices across the user journey. Neither source guarantees that a specific interface teaches durable skill.
For ReplyDesk, useful controls are concrete: inspect the plan record, edit the draft, choose request-information, hand off the case, and distinguish save from send. Every control needs a clear consequence. If “approve” both stores an internal note and emails a customer, the label hides a consequential action.
Avoid making the tutorial depend on capabilities the user's account will not have. A sandbox administrator who can inspect all records is a poor stand-in for an operator with limited access. Test the practice under the intended role, and explain how missing permissions are resolved without exposing restricted information.
Give feedback on the decision rather than the person's character. “This response promises an entitlement absent from the current plan record” is useful. “You don't understand AI” is neither accurate nor actionable. If the source is ambiguous, accept a justified escalation instead of demanding a single confident answer.
A correction control also needs somewhere to save its result. Otherwise the user may learn to edit text locally while the original draft remains the version used downstream. Verify that the practice's visible action and its resulting state match. That check is a product behavior check, separate from whether the user learned anything.
7. Copy this onboarding design brief into your next product review
Use the following artifact for one feature and one user role. It is a proposed design brief, not a benchmark, and the examples are illustrative.
| Design field | What the team must decide | ReplyDesk example |
|---|---|---|
| First useful outcome | What value appears before practice? | A grounded draft for a fictional customer |
| Required user judgment | What consequential choice remains human? | Check entitlement before promising access |
| Authoritative evidence | Which source settles the choice? | Current approved plan record |
| Safe practice defect | What small, disclosed mismatch can be repaired? | Draft includes an unsupported feature |
| Corrective action | What can the user actually change? | Remove promise and save revised draft |
| Missing-information path | What happens when evidence is unavailable? | Request plan information or hand off |
| Help layers | What appears before the full solution? | Source cue, conflict cue, rule, example |
| Next unfamiliar case | What changes while the principle stays relevant? | Different feature, absent record |
| Live-work boundary | What requires extra authorization or expertise? | Exceptions remain administrator-owned |
| Maintenance owner | Who updates practice after product changes? | Support operations owner |
After completing the table, storyboard the screens in order: value example, disclosed practice case, evidence view, repair action, feedback, missing-information case, and return to work. State what users see, what they can do, and what happens next. A storyboard with no consequence after a click is incomplete.
Add the sentence you will use in research consent or participant instructions: “We are checking whether this flow explains the decision clearly; we are not assessing your job performance.” Adapt it honestly to the actual context. If the exercise will affect access to a consequential feature, disclose that separately and provide a support or review route.
Give your builder this brief instead of asking for “smart onboarding.” It defines observable behavior without prescribing a particular framework or requiring a custom model. The team can estimate the work and review the resulting screens against a shared problem.
8. Observe the next decision without overclaiming learning
Start with a small, consented usability study using fictional cases. Ask participants to work through the flow, then encounter a different case that requires the same judgment. Watch what they inspect and do. Record which help they use. A correct choice made only after revealing the full solution is different from a correct choice made after inspecting the source.
The Microsoft and Carnegie Mellon survey examined self-reported critical-thinking effort among knowledge workers. Confidence in AI was associated with less reported effort. It was not a causal test of your users' competence. That distinction suggests collecting observed actions alongside confidence ratings, rather than treating “I feel ready” as the outcome.
For each session, note the task version, role, source availability, final action, assistance used, and the participant's explanation or accessible alternative. Use fictional data, collect only what the study needs, explain access and retention, and let participants decline recording. Include ambiguous cases where requesting information is acceptable. Ask a domain owner to resolve disputed answers rather than using the same AI that generated the exercise as the sole judge.
Do not convert a handful of sessions into a population pass rate or a claim about long-term learning. Small studies can reveal confusing labels, unreachable evidence, and a tendency to follow a prominent recommendation. They cannot establish that your onboarding changes retention or prevents every error.
If you later compare versions, keep the underlying feature, cases, and evaluation criteria stable where possible. Distinguish changes in model quality from changes in the onboarding experience. For a durable-learning claim, a delayed unfamiliar task matters; a repeated immediate quiz mostly measures recall of your example. Decide in advance what evidence would justify the particular claim you intend to make.
9. Interpret friction as a design signal, not a character flaw
A slower first session can have several meanings. A user may be learning a needed boundary, reading inaccessible copy, waiting for a source, or struggling with a confusing action label. Completion time alone cannot tell you which.
Segment observations by relevant experience and role, with participants' consent, rather than assuming all beginners share the same problem. Someone experienced in support may be new to your interface. Someone fluent in your interface may be unfamiliar with the underlying policy. Design help for the missing capability, not for a simplistic “junior” label.
Review activation and readiness evidence side by side, but do not collapse them into one score. If users get value and handle the practice boundary, continue evaluating on real workflows. If they get value but repeatedly miss the boundary, inspect evidence placement, defaults, and escalation access before adding more tutorial prose. If they understand the decision but cannot complete the flow, investigate the interaction cost.
If neither value nor readiness is visible, the feature may be too broad, the target role may be wrong, or the first example may not match the user's job. Narrow the feature before building a larger lesson library. A customer who needs a simple draft should not have to complete an unrelated course.
Track how often users choose a justified handoff, and whether that handoff is actually resolved. A lower send rate can reflect appropriate restraint rather than failure to activate. Conversely, a higher send rate can hide unsupported commitments. The metric needs the meaning of the action, not just its frequency.
These interpretations are hypotheses for product investigation. None is a demonstrated result of the patent experiment, and none establishes a revenue effect. Their value is that they point the team toward a specific screen, workflow, or promise to change.
10. Keep the boundary current and know when this approach does not fit
Onboarding examples age when plans, permissions, sources, and model behavior change. Assign an owner to review the practice case when a relevant product rule changes. Version the source and the expected action together. A tutorial that teaches yesterday's entitlement is worse than an honest note that the policy is being updated.
The NIST AI RMF Core treats oversight roles, deployment context, and ongoing evaluation as explicit organizational concerns. It is voluntary guidance, not certification for this storyboard. The practical implication here is to name who maintains the lesson and who handles decisions outside the user's role.
This approach fits features where a person has a bounded, meaningful choice: verify a claim, revise a draft, identify missing information, or ask for help. It is less useful for low-consequence decoration or a feature where the user has no real discretion. Do not add compulsory tests to a color-palette generator merely because AI is involved.
It also cannot compensate for missing domain expertise, unsafe automation, or inaccessible evidence. When a decision requires a qualified professional, a tutorial cannot confer that qualification. Keep the action with the appropriate role, reduce the feature's scope, or provide a genuine expert service. Teaching “pause and escalate” may be the right user capability.
For this week's product review, choose one live feature, one role, and one consequential judgment. Fill the brief, walk the storyboard under that role's permissions, and observe an unfamiliar safe case. Then make a scoped decision about the next change. Google's research discussion gives founders a timely reason to ask whether customers can use the value they receive with understanding. Your product evidence must supply the answer.
References
- Google Research: Does better work always mean better workers? — October 7, 2026 public explanation; related to source 2, not an independent replication.
- NBER Working Paper 35720: Does AI Assistance Enhance or Erode Expertise? — September 2026 study description and abstract; working paper, not a universal product finding.
- NBER Working Paper 31161: Generative AI at Work — customer-support evidence from a different setting.
- PNAS: Generative AI without guardrails can harm learning — education experiment; transfer to product onboarding is untested.
- Microsoft Research: The Impact of Generative AI on Critical Thinking — self-report survey, not a causal skill-loss experiment.
- Harvard Business School AI Institute: Navigating the Jagged Technological Frontier — original research overview on task-dependent AI performance.
- Microsoft Research: Eighteen Human-AI Interaction Guidelines — design guidance, not proof of learning.
- Microsoft HAX Toolkit: Guidelines for Human-AI Interaction — related design resource; not an independent study from source 7.
- NIST: AI RMF Core — voluntary risk-management guidance, not a certification or prescribed onboarding checklist.