WeatherNext Raises the Bar: A Forecast-to-Action Contract for AI Products
A founder launch framework for turning probabilistic AI forecasts into timely, calibrated, reviewable product decisions after WeatherNext Cyclones.
Google DeepMind has open-sourced WeatherNext 2 and WeatherNext Cyclones alongside a Nature paper and evidence from a real hurricane season. The headline result is unusually strong: across historical cyclones from 2023 and 2024, the system's three-day forecasts for track, intensity, and wind structure reached roughly the error level that leading alternatives reached at two days. A specialized version also ran in real time as guidance for the US National Hurricane Center during 2025.
That does not mean every storm now comes with a guaranteed extra day. It means matched forecast error improved on average. It also does not mean an app should turn 1,000 model scenarios into an automatic evacuation, inventory order, insurance decision, or customer notification. The National Hurricane Center used the system as one input to expert judgment and official warning processes.
This distinction matters far beyond weather. Small teams are adding forecasts for demand, churn, fraud, delivery delay, equipment failure, cash flow, and user intent to AI products. The model can be impressively accurate while the product still fails because a forecast arrives late, its probability is misread, the action threshold is arbitrary, or nobody owns the final decision.
This guide is for nontechnical founders, AI app builder users, and small product teams turning uncertain predictions into real actions. You will get precise terms, a three-layer evidence model, a worked logistics scenario, failure tests, a reusable forecast-to-action contract, and a staged launch decision. The central judgment is simple: ship the forecast, its delivery promise, its action rule, and its fallback as one product release.
What WeatherNext changed, and what it did not
The Google DeepMind release says WeatherNext Cyclones predicts a storm's track, intensity, and wind structure with one system. It was evaluated on historical cyclones from 2023 to 2024. Google reports more than 24 hours of average lead-time advantage at matched error, and the 2026 system can generate a 1,000-member ensemble to expose possible trajectories and localized wind probabilities.
A useful cyclone system has to represent where a storm may go, how strong it may become, how wide damaging winds may extend, and how uncertain those outcomes remain. A single attractive path is not enough.
The evidence is also stronger than a vendor-only retrospective. The National Hurricane Center's 2025 verification report covers the season in which AI guidance entered its real-time workflow. The current NHC model summary lists the Google DeepMind ensemble mean under GDMN/GDMI for track, intensity, and wind radii. NOAA had already documented a formal collaboration in which Google would provide near-real-time forecasts and NOAA would conduct routine evaluations.
But three boundaries remain.
First, an average lead-time gain is not a guarantee for each event, region, forecast horizon, or type of error. Second, a model forecast is guidance; NHC explicitly tells users to rely on official forecasts and warnings rather than raw model output. Third, availability is part of accuracy in practice. A forecast that misses the decision deadline cannot help, even if its late result would have been correct.
For a founder, WeatherNext exposes the distance between prediction quality and product reliability.
Define the terms before you promise an outcome
Teams often use “accurate,” “confident,” and “early” as if they were interchangeable. They are not.
Forecast target is the event or quantity being predicted. “Bad weather” is not a target. “Tropical-storm-force winds at Depot A during the next 72 hours” is closer to one. In another product, the target might be “invoice remains unpaid 14 days after due date.” Lead time is the interval between issuing the forecast and the event or decision deadline. More lead time is valuable only if the forecast remains useful enough to change an action. Matched-error lead-time gain asks how far in advance one system reaches the error level another system reaches later. WeatherNext's “extra day” is this kind of average comparison. It is not a promise to be correct exactly 24 hours sooner on every cyclone. Deterministic forecast gives one outcome or best estimate. A probabilistic forecast gives a distribution over possible outcomes. An ensemble is a set of plausible model trajectories used to estimate that distribution. A thousand members are not a thousand independent observations of reality. Calibration asks whether stated probabilities match observed frequencies over a defined population. If events labeled 20% happen close to one time in five over a sufficiently large comparable sample, that probability band is calibrated. Calibration is always conditional on the population, horizon, version, and event definition. Resolution asks whether the system meaningfully separates higher-risk from lower-risk cases. A product that says 10% for everything may be calibrated in a population with a 10% base rate, yet useless for choosing which cases need attention. Decision threshold maps a forecast to an action. It is a product and business choice, not a universal property of the model. A threshold for showing a planning note can be lower than a threshold for canceling a shipment. Operational availability is the share of required forecast cycles that arrive complete, valid, and before the decision deadline. It includes upstream data, model execution, transformation, delivery, and user-visible rendering.These definitions prevent a team from promising a better decision without measuring the system around a better model.
Model skill is not the same as product value
A forecasting product has at least four linked layers:
- The model estimates possible future states.
- The product converts those states into a probability for a defined user event.
- A decision policy maps the probability, timing, and consequences to an action.
- A person or system performs the action and records the outcome.
Google's own WeatherNext use-case and limitations guide makes the translation problem concrete. It says the models target global operational analysis rather than exact ground measurements, may require bias correction for local applications, omit some variables, and carry limitations in precipitation targets and fine-scale artifacts. Fast global inference does not automatically create a reliable local product.
The same logic applies to a churn model. A good account-level score does not tell a customer-success manager whom to call today, what offer is permitted, or whether outreach changed retention. A fraud score does not define who can freeze a payment, what evidence the user sees, or how an appeal works. A demand distribution does not decide whether the cost of excess inventory outweighs the cost of a stockout.
The founder owns the missing layers, even when the model comes from a respected vendor.
Require three kinds of evidence, not one benchmark
A forecast should move through three evidence tiers. Skipping a tier creates false confidence.
Tier 1: historical skill
Backtesting asks what the current model and pipeline would have predicted from information genuinely available at past forecast times. The comparison needs a named baseline, fixed event definitions, fixed horizons, and a representative evaluation window. It must prevent future information from leaking into features, postprocessing, labels, or manual corrections.
For a probabilistic product, do not reduce evaluation to one accuracy percentage. Check calibration, resolution, false-alarm and miss rates at the thresholds the product may use, and performance across relevant segments. The foundational research on calibration and sharpness explains why a useful probabilistic forecast should be as concentrated as possible while remaining statistically consistent with outcomes. WMO's forecast-verification guidance similarly separates accuracy, bias, reliability, resolution, and sharpness.
Historical skill answers “Could this signal have helped?” It does not prove today's production chain will deliver it.
Tier 2: live operational reliability
Run the new forecast in shadow mode beside the current process. Record whether every expected cycle arrives on time, contains valid fields, maps to the right user and location, and remains stable through retries or partial upstream failures. Compare the exact production artifact, not an analyst's later reconstruction.
The NHC evidence is valuable partly because WeatherNext guidance encountered real schedules, real initial data, and human forecasters during a season. Its verification material also notes timeliness issues for some AI guidance. That kind of failure disappears from a clean benchmark unless delivery is measured explicitly.
Tier 3: decision value
Finally, test whether the forecast improves the user decision. Measure avoided loss, unnecessary action, decision time, review load, customer comprehension, and reversibility. Compare the new policy with the old policy, not merely the new model with the old model.
A slightly less skillful forecast with reliable delivery and a clear action protocol can create more value than a higher-scoring model that arrives late.
Turn an ensemble into a user event before showing a percentage
An ensemble is a collection of possible futures. A user needs a probability attached to a specific event. The product must define that conversion.
Suppose 250 of 1,000 cyclone scenarios show tropical-storm-force winds near a depot. “25% risk” is incomplete. Which exact coordinates define “near”? What wind threshold is used? Is the time window 24, 48, or 72 hours? Were all ensemble members valid and equally weighted? Was bias correction applied? Does the percentage refer to any moment in the window or a sustained condition? Which model revision and initialization time produced it?
Changing any field changes the displayed number's meaning. Store an immutable event definition with every forecast. Preserve the raw probability rather than showing only “low,” “medium,” or “high.” Labels must not erase the number, threshold, horizon, issuance time, or expiry.
The Met Office guidance on ensemble decision-making makes the action layer explicit: users choose probability thresholds that balance alerts against false alarms for their application. That choice belongs to the user's consequences. It cannot be copied from the model card.
In a customer-facing app, every probability should answer five questions without a hidden tooltip:
- What event is being predicted?
- Where and during what time window?
- What probability did the current system assign?
- When was the forecast issued, and when will it expire?
- What action, if any, does the product recommend or permit?
Use this forecast-to-action contract
The following artifact keeps the model, product promise, and operating response in one versioned release. The values are fields to fill, not defaults to copy.
forecast_release:
id: depot-wind-risk-v1
owner: operations-product
intended_user: regional-logistics-manager
user_job: decide whether to reroute tomorrow's inbound deliveries
forecast:
provider: named-model-and-version
event: sustained_wind_at_or_above_defined_threshold
geography: depot-service-polygon-v3
horizon: 72h
issuance_schedule: every_6h
probability_method: documented_ensemble_to_event_transform
valid_until: next_successful_cycle_or_expiry
evidence:
backtest_window: fixed-and-documented
baseline: current-operating-process
required_segments: [region, season, horizon, risk_band]
calibration_report: versioned-link
shadow_cycles_required: team-defined
decision_policy:
monitor: probability_below_review_threshold
review: probability_at_or_above_review_threshold
act: official_warning_plus_named_human_approval
prohibited_actions: [public_safety_warning, automatic_evacuation]
delivery_slo:
deadline_after_source_cycle: team-defined
completeness: required-fields-list
stale_after: team-defined
unavailable_behavior: show_stale_state_and_use_official_source
receipt:
store: [model_version, issue_time, event_definition, probability,
threshold_version, recommendation, approver, final_action, outcome]
rollback:
trigger: calibration_or_delivery_stop_condition
fallback: official_guidance_only
owner: named-person
The contract forces several decisions that a dashboard can hide. The product needs a named user job, not “insights.” The event definition and geography must be versioned. The threshold policy must distinguish monitoring from review and action. Prohibited actions make authority boundaries visible. The delivery service-level objective, or SLO, makes late and partial forecasts observable. The receipt connects what the system knew to what a person did.
Treat changes to the model, event transformation, threshold, delivery schedule, or fallback as release changes. A new model behind the same API can still create a new product.
Worked scenario: a logistics app that does not pretend to be a weather service
HarborLane is a hypothetical five-person team building an operations app for regional food distributors. Customers need to decide whether tomorrow's inbound trucks should use the coastal depot or an inland backup. HarborLane wants to add WeatherNext-based wind probabilities because a single deterministic route forecast hides meaningful uncertainty.
The weak implementation shows a red banner whenever any model member approaches the depot. It has no event definition, no issue time, and no distinction between internal planning and an official warning. A dramatic tail scenario can trigger unnecessary rerouting. A missing model cycle can silently clear the banner. Users may assume HarborLane is issuing safety advice.
The stronger implementation defines the user event as a specified wind threshold within the depot's documented service polygon during a fixed 72-hour window. It displays the raw probability, issue time, horizon, data status, and link to the official local weather authority. The model can move a case from routine monitoring into an operations review, but it cannot issue a public warning or order an evacuation.
Before release, HarborLane replays two historical seasons using only data available at each simulated issue time. It compares the new pipeline with the team's existing official-guidance workflow. It checks probability bands by region and horizon, not only average error. Then it runs live for six weeks without changing customer operations. Every six-hour cycle receives one status: on time, late, incomplete, invalid, or missing.
The team learns that the model signal is useful at 48 to 72 hours, but its location mapping overstates risk for one coastal polygon. It corrects the transformation and repeats the holdout evaluation. It also learns that managers do not want an automatic reroute. They want a review task that includes supplier deadlines, inventory at risk, reroute cost, and the latest official warning.
HarborLane therefore launches in assisted mode. A forecast can create a review task; a named manager chooses the action. The receipt records the probability, official guidance state, affected deliveries, approver, and outcome. After the season, the team can ask whether earlier reviews reduced emergency reroutes without producing unacceptable false alarms.
This hypothetical scenario demonstrates the work required to turn a high-quality forecast into an honest product feature; it does not prove WeatherNext fits HarborLane.
Choose thresholds from consequences, not confidence adjectives
“High confidence” is not an action rule. A threshold should reflect the cost of a missed event, the cost of a false alarm, the time needed to act, the reversibility of the action, and human review capacity.
Use consequence bands before choosing numbers:
| Product action | Cost if unnecessary | Cost if missed | Reversibility | Appropriate release mode |
|---|---|---|---|---|
| Show an informational planning note | Low | Low | Immediate | Automated after basic validation |
| Create an internal review task | Low to moderate | Moderate | Easy | Automated with receipt |
| Recommend a reroute or inventory change | Moderate | High | Time-limited | Human approval |
| Cancel service or restrict a user | High | High | Difficult | Named authority plus evidence |
| Issue life-safety instructions | Extreme | Extreme | Often irreversible | Outside ordinary app authority |
A low threshold may be rational for a cheap, reversible review task. The same threshold can be reckless for an expensive or rights-affecting action. If the action becomes more consequential, do not merely increase the probability number. Add stronger evidence, narrower authority, an independent source, and a better appeal or reversal path.
Test thresholds on a holdout period and report the number of actions they would create. A 15% threshold can look cautious until it generates 400 daily reviews for a team that can inspect 20. Capacity is part of the policy. Overflow must have an explicit behavior rather than quietly approving or discarding cases.
Thresholds should also vary only when there is a documented reason. Changing them repeatedly in response to the latest dramatic event overfits the policy to anecdotes. Version every change and re-evaluate all relevant horizons and segments.
Make delivery a first-class acceptance test
Forecast products expire. This makes operational checks unusually important.
For each expected cycle, record source availability, model start and completion time, transformation completion, product publication, notification delivery, and first user view. Validate location identifiers, units, horizons, ensemble-member counts, probability bounds, and event-definition versions before publication.
Use five user-visible states:
- Current: the expected cycle arrived complete and passed validation.
- Delayed: the expected cycle has not arrived by its normal time but is not yet beyond the decision deadline.
- Stale: the last valid forecast has passed its approved age.
- Partial: some fields, regions, or ensemble members are missing.
- Unavailable: the product cannot produce a valid forecast.
The product team should set a release stop condition for delivery, such as an unacceptable share of cycles missing the decision deadline. The exact number depends on the job, but it must be selected before the launch data arrives. Otherwise the team will rationalize outages after seeing favorable model results.
Preserve human authority without creating approval theater
Human review helps only when the reviewer has a real decision, enough context, and the authority to disagree.
An approval screen should show the event definition, current probability, previous forecast, direction of change, official or independent source, affected objects, action cost, deadline, and recovery option. It should not show a green “AI recommends” button with the evidence hidden below it.
Assign separate owners for four responsibilities:
- model and data quality;
- event transformation and product presentation;
- decision policy and business consequences;
- incident response and user communication.
For high-consequence decisions, require independent corroboration or an authoritative source. The NOAA–Google collaboration is instructive: Google supplied near-real-time model output, while NHC evaluated it within a broader technical and expert process. The operational relationship did not erase institutional responsibility.
Test the failure modes that a good benchmark will miss
Run these tests against the complete product, not only the model endpoint.
The matched-error headline test
Ask ten target users what “one extra day” means. Fail if they interpret it as a guarantee for every event. Replace the phrase with the exact population, metric, average comparison, and limitation wherever it appears.
The stale-zero test
Block one scheduled forecast. Fail if the UI shows zero risk, repeats an expired value without its issue time, or continues an automatic action. The correct response is an explicit delayed, stale, partial, or unavailable state.
The threshold swap test
Change the review threshold without changing the model. Confirm that the release system identifies a new product-policy version, reruns the action-volume analysis, and requires the named approval.
The geography and unit test
Shift a polygon, time zone, or unit conversion. Fail if a plausible probability is attached to the wrong location, period, or threshold without validation catching it. Semantic plausibility is dangerous because users may not notice.
The tail-risk test
Create an ensemble with a low-probability, high-consequence branch. Confirm that the product neither hides it inside an average nor presents it as the most likely outcome. The interface should connect it to the appropriate review rule.
The authority test
Try to use the feature for an action outside its contract, such as issuing a public safety instruction. The product should block or redirect the request, not merely display a warning that can be clicked through.
The rollback reconstruction test
Take one past decision and reconstruct the model version, probability, event definition, threshold, recommendation, approver, action, and observed outcome. If the team cannot reproduce the decision record, it cannot investigate harm or improve policy responsibly.
Avoid five seductive misreadings
“Open source means production ready.” Public code and weights improve inspection and experimentation. They do not provide your local calibration, delivery SLO, support process, or authority model. The WeatherNext repository is an important release artifact, not a substitute for product validation. “A thousand scenarios make rare probabilities exact.” More ensemble members can represent tails more richly, but members share data, architecture, and assumptions. Model dependence, event transformation, and limited observations still matter. “A successful named event validates the whole system.” Hurricane Melissa is a meaningful operational case. The NHC storm report helps establish what forecasters observed and decided. One memorable success cannot replace multi-event, multi-region verification. “A probability is objective, so the action is objective.” The event definition, threshold, error costs, and action authority are human choices. Hiding them behind a score does not remove judgment; it makes judgment harder to audit. “A human in the loop transfers responsibility.” A reviewer who lacks time, context, alternatives, or permission to disagree is not a control. Measure reversal rates, review time, skipped evidence, and outcomes rather than counting approval clicks.Know when this framework fits, and when it does not
Use a forecast-to-action contract when an AI feature estimates a future event or quantity and that estimate may change user behavior. Typical examples include demand planning, delivery risk, churn intervention, fraud review, maintenance, cash-flow alerts, appointment no-shows, capacity planning, and weather-sensitive operations.
Use a lighter version when the forecast is purely informational, cheap to ignore, clearly labeled, and has no automated consequence. You still need a target, issue time, horizon, and stale state, but the authority and rollback process can be simpler.
Do not use this framework to justify an AI product acting as an official weather, medical, legal, credit, employment, insurance, or public-safety authority. Those domains can require qualified professionals, specific regulation, validated instruments, procedural rights, and accountable institutions. A generic product checklist does not create that authority.
Do not use it when the product is not forecasting anything. A summarizer needs evidence fidelity and completeness tests. A coding agent needs change and execution controls. An image generator needs claim, accessibility, and visual-acceptance checks. Calling every model output a forecast makes the terms less useful.
A 48-hour founder rollout
In the first four hours, write one user job, one forecast target, one horizon, one decision deadline, and one prohibited action. Name the baseline process and the person accountable for the product promise.
By hour eight, complete the forecast-to-action contract. Inventory every transformation between model output and the displayed event probability. Define current, delayed, stale, partial, and unavailable states.
On day one, build a small historical evaluation that preserves past information boundaries. Report calibration and action volume for the exact thresholds under consideration. Split the report by the segments that could change user consequences.
On day two, start shadow delivery. Log every expected cycle and generate receipts without changing user actions. Exercise the stale-zero, geography, threshold, tail-risk, authority, and reconstruction tests. Set stop conditions before reviewing the first live results.
Do not promise a launch date based only on the 48-hour setup. The point is to create the evidence pipeline quickly. The required shadow period depends on event frequency, seasonality, consequence, and how much historical data genuinely represents the live environment.
Make the release decision explicit
Use four release states:
| Decision | Evidence required | Product behavior |
|---|---|---|
| Ship assisted | Historical skill, acceptable calibration, live delivery pass, trained reviewer | Forecast creates a recommendation or review task; human owns action |
| Ship limited | Promising signal but incomplete segment or season evidence | Restricted users, geography, horizon, or low-consequence actions |
| Hold | Unclear calibration, unstable thresholds, missed deadlines, ambiguous authority | Continue shadow mode and repair the named gap |
| Reject | No useful lift over baseline, unacceptable harm, no viable fallback, or prohibited use | Remove the feature or redesign the user job |
“The model is better” is not one of the decisions. A release record should name the accepted use, excluded uses, evidence window, unresolved risks, next review date, and rollback owner.
WeatherNext deserves attention because it combines a serious research result, released artifacts, and real operational evidence. Its most transferable lesson is not that every small team should run a weather model. It is that a forecast becomes valuable only when uncertainty reaches the right person before the decision, through a tested policy, with visible limits and accountable authority.
Build that chain before you build the confidence badge.
References
- Google DeepMind, WeatherNext Cyclones release
- Alet et al., Operational tropical cyclone forecasting with AI, Nature
- Google DeepMind, WeatherNext open-source repository
- Google for Developers, WeatherNext model guide
- Google for Developers, WeatherNext use cases and limitations
- National Hurricane Center, 2025 Forecast Verification Report
- National Hurricane Center, track and intensity model summary
- National Hurricane Center, Hurricane Melissa tropical cyclone report
- NOAA, research partnership with Google for tropical cyclone forecasts
- Gneiting, Balabdaoui, and Raftery, Probabilistic forecasts, calibration and sharpness
- World Meteorological Organization, Guidelines for Streamflow Forecast Verification
- Met Office, Using ensemble forecasts in decision-making