§ 41 — Use-case analysis
May 2024 G5 Storm Tested AI Grid Models – Here's What Worked
The May 2024 G5 geomagnetic storm offered the first real-world test for AI-driven disruption models. This article reviews the verified outcomes and gaps — from satellite anomaly predictions to ground-level power grid forecasting — providing a benchmark for procurement evaluations.
- Function
- demand-forecasting
- AI technique
- mixture-of-experts
- Failure pattern
- validation gap
- Evidence source
- Nature Scientific Reports (October 2025)
For anyone buying or approving AI for power grid disruption planning after geomagnetic storms, the May 2024 G5 storm is the test case that vendor decks had been missing. It was not a lab replay, not a mild disturbance, and not a historical backtest dressed up as readiness. The storm reached a Dst below -400 nT, solar wind speeds above 900 km/s, and was the strongest geomagnetic storm since 2003.[1]
That timing matters. The previous G5 event arrived before today’s AI grid-planning market existed in any recognizable form. The 2024 storm arrived after years of claims about predictive operations, autonomous response, and infrastructure resilience. It gave buyers a cleaner question to ask: which model was tested, on which physical system, during what intensity of event, and against which operational decision?
The answer is useful, but narrower than many sales narratives would prefer. AI did produce a strong, peer-reviewed result tied directly to the G5 onset. That result was for satellite power-subsystem anomaly prediction, not transformer protection, utility dispatch, or end-to-end ground-grid response.

The strongest evidence came from satellite power systems
The cleanest benchmark is a Nature Scientific Reports study published in October 2025 that evaluated a hybrid statistical and machine-learning framework on MisrSat2 satellite power-subsystem data during the May 2024 storm onset. The model family was a Mixture-of-Experts framework, and the measured target was solar-panel current anomalies on a satellite, not a terrestrial utility network.[1]
That distinction should not make the result look small. Satellite power subsystems are physical infrastructure exposed to space-weather stress. A model that performs well under a rare G5 onset has done something more valuable than win a routine training benchmark. The study reported R²=0.921 and MAE=0.063 A for solar-panel current anomaly prediction, outperforming LSTM at 0.881, Random Forest at 0.854, and Linear Regression at 0.742.[1]
| Model or method | Reported result in the Nature study | What the result supports |
|---|---|---|
| Mixture-of-Experts framework | R²=0.921; MAE=0.063 A | Strong prediction performance for MisrSat2 satellite solar-panel current anomalies during the May 2024 storm onset |
| LSTM | R²=0.881 | A weaker but still relevant machine-learning comparison point |
| Random Forest | R²=0.854 | A conventional ML baseline beaten by the MoE framework |
| Linear Regression | R²=0.742 | A simpler baseline with materially lower fit |
The comparison is important because it prevents the usual fog around “AI performed well.” The study did not merely say a neural model found patterns. It placed a named architecture against other model classes and reported the gap. For procurement teams, that is the difference between a claim that sounds modern and a result that can be written into evaluation criteria.
The anomaly detection layer also matters. The study used CUSUM change-point detection and identified 13–17 statistically robust anomaly events per solar panel during the storm onset.[1] That gives the result an operational flavor: not just a curve fit, but a way to flag event timing in a power subsystem under stress.

The most disciplined sentence in the result may be the one that keeps it from being oversold: all observed deviations remained below 4% of design tolerances.[1] That means the model was detecting statistically meaningful disturbances, not documenting a satellite power failure. It supports confidence in sensitivity and measurement, not a claim that AI prevented damage.
Why this is not yet proof for ground-grid disruption planning
The category slide is where buyers need to slow down. A validated satellite power-subsystem anomaly model is adjacent to grid-resilience planning, but it is not the same asset class, not the same control problem, and not the same consequence chain.
A satellite solar panel current anomaly is a bounded measurement inside a known spacecraft subsystem. Ground-level power-grid disruption planning has to deal with geomagnetically induced currents, transformer conditions, topology, protection settings, operator procedures, restoration constraints, and the possibility that a forecast is technically correct but still not actionable at the asset where a decision has to be made.
That does not diminish the Nature result. It makes it more useful. It gives buyers a benchmark for how a serious AI claim should look: named system, event intensity, measured subsystem, baselines, anomaly method, and design-tolerance context. A vendor claiming broader geomagnetic-storm readiness for utilities should be able to produce evidence at least as specific.
What buyers should not accept is a chain of implication that runs from “AI detected satellite current anomalies during a G5 storm” to “AI can plan power-grid disruption response during a G5 storm.” The first statement has peer-reviewed support. The second still needs operational evidence from ground-grid deployment under comparable event conditions.
The storm also showed why this is an operations problem, not just a space-weather problem
The May 2024 storm did not have to collapse a power grid to become commercially relevant. One published study cited by Space.com estimated about $500 million in losses to the U.S. agricultural industry through GPS navigation disruption.[2] That is the kind of downstream impact that reaches supply-chain teams, equipment operators, insurers, and ERP planners even when the electric grid does not become the headline failure.
The procurement lesson is not that every storm-response tool should now become a GPS-risk platform. It is that geomagnetic events can move through infrastructure layers unevenly. A model may be strong for one layer, weak for another, and irrelevant to a third. Buying a broad “resilience” label without checking the tested layer is how organizations end up with dashboards that look persuasive until the first real event.
DAGGER is promising, but it is a different evidence lane
NASA’s DAGGER model deserves attention because its operating cadence is close to what planners actually need. NASA describes it as an AI model that can provide global geomagnetic disturbance predictions 30 minutes in advance, updated every minute, and released as open source.[3]
A 30-minute warning horizon is not cosmetic. It can change whether an alert is merely observed or whether someone has time to adjust procedures, notify field teams, stage staff, or prepare contingency workflows. In operational systems, minutes are not evenly valuable; the difference between ten and thirty can determine whether a response is planned or improvised.
But DAGGER sits in a different evidence lane from the Nature MoE result. NASA’s validation described performance against 2011 and 2015 storms, not a demonstrated May 2024 G5 deployment.[3] That makes it credible and worth watching, especially because of its cadence and open-source status. It does not make it proof that an AI system handled the 2024 G5 storm in live utility operations.
For a buyer, that distinction changes the requirement language. DAGGER-like capability can be valued as an earlier-warning input. It should not be treated as evidence that a vendor can forecast transformer-level damage, manage dispatch decisions, or automate an integrated storm-response plan unless the vendor can show those downstream links with sourced operational data.
The GIC forecasting work is closer to the grid, but still not a complete response system
The CIRES and University of Colorado Boulder work is more directly relevant to terrestrial power-grid planning because it addresses geomagnetically induced currents. CIRES reported a machine-learning method that predicts GICs one hour ahead, compared with a traditional warning time of about ten minutes, and said it was tested on the 50 largest U.S. GIC events since 2000.[4]
That is a meaningful procurement signal. One hour of lookahead can support staffing, situational awareness, coordination with reliability procedures, and preparation for possible current flows in vulnerable parts of the system. The testing set also matters: using the 50 largest U.S. GIC events since 2000 is more relevant to grid planning than a generic space-weather benchmark.[4]
The limitation is equally important. A GIC forecast is not the same as a full operational storm-response deployment. It does not, by itself, prove asset-granular transformer damage prediction, cascading-effects modeling, or automated response across utility control rooms and enterprise systems. Ground-level GIC magnitude forecasts remain uncertain enough that buyers should ask how the model output is converted into thresholds, reviews, and operator actions.
This is where IT and ERP leads have a legitimate role, even if they are not space-weather specialists. If a forecast changes crew scheduling, spare-equipment staging, customer communications, or supply-chain prioritization, then the model is no longer just a scientific instrument. It is part of an operational workflow, and the buyer needs to know where human review enters, what decision is delayed or accelerated, and who owns the consequence of a false alarm.
What the May 2024 benchmark allows buyers to credit
The 2024 storm does not leave buyers empty-handed. It gives them a practical way to separate three classes of evidence.
- Validated during the May 2024 G5 event: satellite power-subsystem anomaly prediction using the Nature-published MoE framework, with reported performance and anomaly counts tied to the storm onset.[1]
- Validated on earlier or historical storms: NASA DAGGER’s 30-minute global predictions updated every minute, validated against 2011 and 2015 storms, and the CIRES GIC method tested on major U.S. GIC events since 2000.[3][4]
- Still unproven at G5 operational level: comprehensive AI planning for ground-level utility disruption, including transformer-level impact prediction, cascading effects, and live response orchestration during a comparable G5 event.
That third category is the one most likely to be blurred in procurement conversations. It is reasonable for a vendor to say that space-weather AI is advancing. It is reasonable to cite satellite anomaly detection as proof that AI can extract meaningful signals under an extreme geomagnetic event. It is not reasonable to let that stand in for verified utility-grid operations unless the asset class, decision point, and event intensity match.
The useful procurement questions are therefore concrete rather than philosophical:
- Was the model tested during the May 2024 G5 storm, or only on earlier and weaker events?
- Did the model predict satellite subsystem anomalies, geomagnetic disturbance, GICs, transformer risk, or operational actions?
- What warning horizon did it provide, and how often were predictions updated?
- Which baselines did it beat, and by how much?
- Were observed effects inside normal design tolerances, or did the system face actual degradation or failure?
- Who used the output in operations, and what decision changed because of it?
The safest reading of the evidence
The May 2024 G5 storm validated a real AI capability, but not the broadest one being sold. The strongest sourced evidence shows that a Mixture-of-Experts framework improved prediction of satellite solar-panel current anomalies during the storm onset, with better reported performance than LSTM, Random Forest, and Linear Regression, and with statistically robust anomaly detection that stayed within design tolerance limits.[1]
Adjacent models expand the picture. DAGGER points toward useful global warning cadence, and the CIRES GIC work points toward longer ground-grid warning horizons.[3][4] Those are important capabilities. They are not the same as an audited, operational G5 deployment by a utility AI vendor across grid assets, enterprise workflows, and response decisions.
So the benchmark is mixed, but not vague. Buyers can credit vendors for capabilities similar to the verified satellite anomaly work. They can cautiously value earlier-warning GIC and geomagnetic-disturbance models when the validation history is clear. They should discount claims of comprehensive geomagnetic-storm planning for ground-level grids until a comparable G5 deployment produces sourced operational data.
References
- A hybrid statistical–machine learning framework for evaluating geomagnetic storm effects on MisrSat2 satellite power subsystems, Nature Scientific Reports, Oct 2025.
- A worst-case solar storm could knock out satellites, GPS and power grids, report warns, Space.com.
- NASA-enabled AI Predictions May Give Time to Prepare for Solar Storms, NASA.
- Forecasters can now predict space weather-induced disruptions in power grid an hour in advance, CIRES.
§ 42 — Cited evidence
Flag an inaccuracy or submit a comparable account — Contribute or read how claims are verified in Methodology.
