SciGroveBeta
Climate

Can AI Weather Models Predict Beyond Two Weeks? A Quantitative Benchmark and Analysis of Long Rollouts

Fanny Lehmann, Firat Ozdemir, Yun Cheng, Torsten Hoefler, Sebastian Schemm, Benedikt Soja, Siddhartha Mishra

Featured June 7, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

When AI weather models predict far into the future, they often break down; this paper identifies three ways they fail and shows that good models fix tiny errors, making them stable and realistic over long periods.

In depth
This study addresses the critical issue of instabilities in AI weather models during long-term autoregressive rollouts, which limit their applicability for climate prediction. The authors introduce a formal taxonomy to categorize these failures into blow-up, drift, and loss of seasonality. Their analysis reveals that stable models act as denoisers of small spatio-temporal scale energy, generating unique and realistic weather trajectories, demonstrating genuine generalization rather than memorization.

Key Takeaways

  • 1
    The paper establishes a formal taxonomy of instabilities in AI weather models for long rollouts, categorizing failures into blow-up, loss of seasonality, and small-scale artifacts, and provides quantitative metrics for each.
  • 2
    Stable AI weather models, such as Aurora and SFNO, demonstrate a crucial denoising capability at small spatio-temporal scales, preventing error amplification and enabling realistic long-term predictions from noisy inputs.
  • 3
    Ablation studies on the Aurora model reveal that long-term stability is largely architecturally invariant to many design choices, but the time embedding is essential for preserving seasonality, and a minimum model capacity is required.

Conceptual Flow

HIGH LEVEL
1
Methodology: How AI Weather Models Break Down

The study tests many AI weather models by making them predict far into the future, then carefully watches how they fail in different ways.

Many AI Weather Models
Long-Term Forecasts
Run & Observe
Different Ways Models Fail
2
Results: Stable Models Clean Up Noise

The best models don't just predict; they also clean up any small errors, making their long-term forecasts stay realistic and unique.

Initial Weather State
Small Errors Added
Clean Up & Predict
Realistic Future Weather
Unique Forecast Paths