MiMo-V2.6's Repeated Tool Calls: A Founder Release Gate for AI Agents
Xiaomi's MiMo-V2.6 fix exposes an expensive agent failure: useful-looking activity without progress. Use a task-level replay, stop policy, and release receipt before shipping.
On September 27, Xiaomi published its diagnosis of repeated tool calls in MiMo-V2.6. In some agent environments, the model kept issuing identical or nearly identical calls while making little progress. Xiaomi says it has deployed updated API models and released MOPD checkpoints. The interesting product question is not whether one model has been fixed. It is whether your customer-facing agent can recognize that it is busy but stuck, stop safely, and account for any action it has already taken.
This matters to nontechnical founders and small teams building assistants that research, schedule, draft, edit, or operate connected apps. A conventional success dashboard can miss a loop: the API responds, tools return data, and the final text may even sound plausible. Meanwhile the user waits, pays, or receives duplicate side effects. This guide gives you a vocabulary for the failure, a replay protocol, an acceptance matrix, and a release receipt you can hand to a product and engineering team. Xiaomi's measurements are vendor-reported. The example and thresholds below are proposed tests, not results from YBuild or an independent MiMo evaluation.
The release teaches a broader lesson than “use the fixed model”
Xiaomi describes repeated calls across MiMo Desktop, MiMo Code, OpenCode, and other harnesses. Its technical post reports pre-fix, response-level rates that vary by model and harness; for example, OpenCode is shown at 1.02% for Flash-RL and 0.54% for Pro-RL. Those figures are not a forecast for your app. They show that orchestration changes the observed failure rate, so a model benchmark cannot stand in for a test of your product's tools, prompts, permission scopes, and user tasks.
The post says the updated API models have been available since September 25 at 06:00 UTC+8 under the same mimo-v2.6-pro and mimo-v2.6-flash names. Separately, the released Flash-MOPD checkpoint and Pro-MOPD checkpoint make a self-hosted comparison possible for teams equipped to run them. The API name remaining the same is a practical release concern: a model label alone may not tell you which behavior a customer encountered. Record the provider response ID, time, configuration, and any available model revision, then preserve a trace of the task.
Xiaomi says its specialized teacher reduced repetition in its replay set and that broader benchmark performance held steady after MOPD. That is meaningful engineering evidence about its own setup, but it does not establish your completion rate, latency, duplicate-send risk, or bill per accepted customer result. A small team should use the release as a reason to test those outcomes on its own workflow. You do not need to reproduce the training method to learn whether the app is safe to ship.
Define the failure before trying to count it
A tool call is a request by the model to an external capability: search a document, read a calendar, create a draft, send a message, or charge a card. A turn here is the group of calls emitted before the model receives tool feedback. A trace is the time-ordered record of model decisions, calls, results, policy checks, and user-visible state. OpenAI's agent evaluation guidance treats traces as the place to inspect workflow behavior rather than only final answers.
Three patterns require different responses. Parallel work can be useful: searching three distinct sources in one turn is not a loop. Flooding means making more calls than the task or infrastructure reasonably needs, perhaps filling a queue. Repetition means repeating the same effective action without new information or a changed state. A retry after a timeout can be justified; rereading the same unchanged file after a successful read usually is not. Repeating a payment request without checking the first request's effect is dangerous even if the original timed out.
Xiaomi's reproducible metric is deliberately narrow: within one turn, normalize JSON arguments, then count calls with the same tool and identical normalized arguments. If a turn has N calls and U unique tool-and-argument pairs, the exact duplicate share is (N - U) / N. Its post explains that this misses near-duplicates, repetition across turns, and calls hidden inside a code-execution wrapper. Treat it as a cheap detector, not as a complete definition of user harm. An agent can waste minutes by searching the same question with slightly different wording while the exact duplicate rate remains zero.
For a product, add a progress signal and an effect signal. Progress asks whether the call produced information or state that advances the customer's task. Effect asks whether a tool changed anything outside the model, such as sending an email or creating a record. These are product judgments, not properties that token counters can infer. Keep a human-reviewable sample of ambiguous cases so a guard does not stop legitimate verification.
A concrete customer scenario: the meeting brief that keeps looking
Consider a hypothetical meeting assistant, CedarMeet. A customer asks it to prepare a brief from a calendar event, three linked documents, and the latest account notes, then save a draft for approval. The intended outcome is one accurate draft in the right workspace. The agent reads the event, fetches the documents, and queries account notes. It then calls search_notes("renewal blockers") repeatedly. Some calls return the same page; others use slightly altered queries that surface the same entries. The activity panel shows “researching,” so the customer assumes useful work is happening.
A naive tool-call cap may stop the task but still leave a partial draft. A naive retry may create a second draft. A final fluent summary might omit the fact that one attachment was inaccessible. The product needs to distinguish four states: fresh evidence found, a justified retry after a transient error, a stalled search with no new evidence, and a write whose outcome is uncertain. The last state must be reconciled before any retry. The customer should see “Draft saved; two sources checked; one source unavailable” or “Paused before saving; review needed,” depending on the trace, rather than a generic success badge.
This scenario is illustrative, not a claim that MiMo or any particular provider behaved this way in CedarMeet. Its value is that a nontechnical founder can judge the result. Ask your team to replay a fixed CedarMeet-like task with the previous and updated model path. Have a reviewer label whether the brief is acceptable, whether the same source was fetched without reason, whether the draft was written once, and whether the screen accurately described the stopping point. That is more informative than asking whether the new model “feels less repetitive.”
Build a replay set from tasks your customers actually pay for
Start with twenty to fifty consented or synthetic tasks that mirror your product's high-value paths. Include straightforward success, missing data, permission denial, timeout, rate limit, ambiguous tool output, and a task where verification genuinely requires a second call. Do not put raw customer secrets into an external evaluation service without checking your data terms. Preserve each case's objective, allowed tools, initial state, expected output, and forbidden effect. OpenAI's trace grading guide describes structured labels on end-to-end traces; use the idea even if you grade in a spreadsheet.
Freeze what you can: the same input set, connector permissions, tool descriptions, evaluation rubric, and stop rules. Capture the actual model version or deployment window, because unchanged public names can mask changed behavior. Record both the old and new paths where feasible. If a live connector changes between runs, note it and avoid claiming a clean model comparison. Prefer a staging tenant with controlled fixtures for write actions. If your product cannot replay a task because every run depends on mutable external state, your first artifact should be a small test fixture, not a performance verdict.
A trace needs more than prompts and final text. Capture tool name, normalized arguments or a privacy-safe hash, call start and end, response status, result fingerprint, state version before and after, retry reason, permission decision, idempotency key for writes, and the UI state shown to the customer. Link these to a task ID. Anthropic's tool-use documentation distinguishes client-side tool execution and returned tool results, which is a useful reminder that the model's requested action and the application’s actual action are separate events. A failed tool result must not be silently converted into a completed task.
Use human labels for “necessary duplicate,” “avoidable duplicate,” “uncertain,” and “harmful side effect.” Ask reviewers to explain uncertain labels briefly. The aim is not a perfect universal detector. It is a reliable way to see whether a model or orchestration change makes your own customer's task worse. Keep the same rubric for each candidate release so a threshold is not relaxed after looking at the data.
Score accepted outcomes, not just tool-call volume
The release decision should use a small set of measures that tell a product story. Accepted task rate is the share of runs meeting a reviewer-defined output and effect contract. Avoidable repeat rate is the share of calls a reviewer says added no needed information after a successful or unchanged result. Time to useful state is the time until the user can act on an honest result, rather than the time until the first token appears. Cost per accepted task is total provider and tool spend divided by accepted runs, including failed attempts and retries. Duplicate effect count is the number of repeated writes, sends, charges, or other external changes.
A drop in exact duplicate calls can coexist with a worse product. A model could make fewer calls by skipping necessary checks. A hard cap could reduce spend while raising incomplete tasks. A retry policy could improve apparent completion while duplicating drafts or notifications. Conversely, two identical reads can be appropriate when a document changed between them. Compare the call trace with state versions and task acceptance. Anthropic's tool-design guidance notes that redundant calls can also signal badly sized pagination or unhelpful error responses; the model is only one possible cause.
Do not borrow Xiaomi's reported rates as your pass threshold. Choose a threshold based on your own harm. For a read-only research feature, an avoidable repeat may mainly waste time and money. For a billing or sending feature, one duplicate effect may be enough to block release. Report both numerator and denominator, plus the sample's task mix. A rate based on a handful of easy tasks is a weak reason to enable the feature for every customer.
Use a stop policy that preserves state and tells the truth
A stop policy should combine several signals. The first exact duplicate within a turn can trigger a warning and a trace marker. Repetition after the same successful result can trigger a pause or a request to use a different plan. A repeated transient error can be retried only within an explicit retry budget. An uncertain write cannot be retried until the downstream system is checked. The rule should be stricter for externally visible effects than for reversible reads. OWASP's AI Agent Security Cheat Sheet recommends least privilege, tool-call limits, monitoring, and human approval for sensitive actions.
The guard belongs in the application or tool layer, where the team can enforce it even if the model ignores a prompt. It should identify the task, tool, arguments, result fingerprint, and state version. It should also record why a call was allowed: new information sought, changed state, explicit user request, or a permitted retry. If the reason is unknown, the product may safely pause and ask the user to confirm or narrow the task. “Stopped due to repeated research” is a useful product state; an endless spinner is not.
Set a total budget as well as a repetition guard: maximum elapsed time, tool calls, provider spend, and write attempts. These are guardrails, not quality targets. Hitting a budget should yield a partial result with an honest list of completed and uncompleted steps. Do not label the task successful merely because the model produced a final paragraph. A customer must be able to resume from the saved state or hand the case to a person without rerunning uncertain writes.
Separate reads, retries, and external effects
A repeated read may be wasteful; a repeated write may be irreversible. Put each tool in one of three product categories: read-only, reversible write, or high-impact effect. A read-only search can often be cached against a source version for a short period. A reversible draft creation needs a stable task-to-draft mapping so the product can reopen the existing draft. A payment, public send, or destructive action needs a stronger permission and reconciliation path. OWASP's Excessive Agency guidance recommends narrowing functionality and permissions and independently approving high-impact actions.
An idempotency key is a stable identifier the receiving service uses to recognize retries of the same intended operation. Stripe's idempotent-request documentation explains that repeating a request with the same key returns the saved first outcome within its retention policy. This protects against a network retry creating a second object when the provider supports the guarantee. It does not prove the first action was desired, and a new key can still create a new effect. The app must bind the key to a customer intent and inspect the downstream outcome before deciding what to show.
For connectors without reliable idempotency, keep a write ledger in your own application. Before execution, record the intended effect and target. After execution, store the downstream object ID or a state check. If the call times out, move to “outcome unknown” and reconcile; do not immediately generate another write. This ledger is a design recommendation, not a universal guarantee: third-party APIs vary in consistency, retention, and searchability. Test the exact connector your customers use. A release gate should reject any path that converts an uncertain side effect into a blind retry.
A reusable acceptance matrix for the next model or agent update
The matrix below is a template. Replace the examples with your own task and source fixtures. For each row, save the trace and the user-visible screen. A “pass” means both the action and the explanation matched the expected contract; the table does not report a measured MiMo result.
| Replay case | Evidence the team needs | Expected product behavior | Block if |
|---|---|---|---|
| Independent parallel reads | Distinct source IDs and results | Finish without penalizing useful parallelism | Guard cancels valid sources |
| Exact duplicate before feedback | Same tool and normalized arguments | Mark duplicate; execute at most as policy allows | Same no-op request proceeds repeatedly |
| Near-duplicate search | Similar intent and unchanged source version | Pause or change strategy after no new evidence | Query wording changes hide a loop |
| Timeout on a read | Error, retry count, source version | Retry within budget or explain unavailable source | Unlimited retries or false success |
| Updated source | Before/after version markers | Permit a justified second read | Guard mistakes fresh evidence for repetition |
| Uncertain draft write | Intent ID, downstream object check | Reconcile before retrying | A second draft appears |
| High-impact send | Approval and downstream receipt | Send once after explicit approval | Duplicate or unauthorized send |
| Budget exhaustion | Time, calls, cost, saved state | Stop with honest partial status and resume path | Spinner, silent abandonment, or “done” |
Have product, engineering, and support agree on the reviewer answer for each row before running it. If one person says a repeated search is necessary and another says it is pointless, write down the missing context. Perhaps the tool result is truncated, the source changed, or the user asked for independent verification. That disagreement is useful. It tells you what state the app must expose before any automatic loop detector can work reliably.
Decide whether to ship, limit, or hold
A practical release decision has three outcomes. Ship when the candidate meets the accepted-task contract, avoids new harmful effects, stays within task budgets, and reports partial states honestly across the replay set. Limit when it improves routine reads but has unresolved behavior in a specific connector, task type, or customer segment; route only the safe subset and monitor it. Hold when a duplicate write, hidden stalled state, or unacceptable regression remains, even if model-level benchmark scores look good.
Stage the rollout. Start with read-only internal tasks, then a small consented cohort, then broader use after inspecting live trace samples. Separate alerts for abnormal call volume, repeated intent without new results, budget stops, and duplicate effects. Include a kill switch that disables an affected tool or routes to a safer manual path. Define who receives alerts and who may authorize a rollback. A monitoring panel without an owner is not a control.
Keep a one-page release receipt: date and deployment identity; task-set version; provider/model route; tool and permission versions; accepted-task rate and denominators; avoidable-repeat labels; cost and latency; duplicate-effect count; known gaps; ship/limit/hold decision; owner; rollback trigger. OpenAI's agent evals documentation supports trace-based regression work, but the receipt should be understandable without an eval platform. It is a decision record, not a marketing claim.
Where this approach applies, and where it does not
This gate is useful for AI app builders with external tools, especially research, support, scheduling, document processing, coding, and operations workflows. It is most valuable when a user can see an agent appear active while progress is unclear. It is also relevant when a provider swaps a model behind a familiar name or a team changes a prompt, tool schema, or retry policy. The same trace can reveal whether the culprit is the model, the orchestration layer, tool feedback, or a connector.
It is not a complete safety review. Exact duplicate detection misses semantically similar calls, loops hidden inside exec, and harmful single actions. A passing replay set does not prove broad reliability, especially for rare high-impact cases. Privacy, authorization, prompt injection, and domain-specific correctness need their own tests. Nor does this guide establish that Xiaomi's fix works on your workload: the vendor post reports its own evaluation and the public checkpoints enable, but do not replace, independent testing.
The founder's question for the next release meeting is simple: “When our agent repeats an action, what changed, what did it cost, what effect occurred, and what did the customer see?” If the team can answer from a trace and a release receipt, a model update becomes a controlled product decision. If it cannot, the most urgent improvement may be the application's stop and reconciliation design, not another model comparison.
References
- Xiaomi MiMo, “Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6”
- Xiaomi MiMo, Flash-MOPD model card
- Xiaomi MiMo, Pro-MOPD model card
- OpenAI, Evaluate agent workflows
- OpenAI, Trace grading
- Anthropic, Tool use with Claude
- Anthropic, Writing effective tools for AI agents
- Stripe, Idempotent requests
- OWASP, AI Agent Security Cheat Sheet
- OWASP, LLM06:2025 Excessive Agency