CRBC News
Technology

3 GW Left PJM — How Uncoordinated Data‑Center Protections Are Hitting the Grid (NERC Warned in 2024)

3 GW Left PJM — How Uncoordinated Data‑Center Protections Are Hitting the Grid (NERC Warned in 2024)

Summary: On July 22, 2026 a transmission fault in Ashburn caused more than 3 GW of data‑center load to leave PJM, echoing a 2024 NERC investigation that found customer‑side protection (counting relays) disconnected most of the load. Regulators (NERC, ERCOT, FERC) are moving toward mandatory ride‑through rules and registry thresholds. Data centers must document ride‑through curves, share protection settings, and implement controlled reconnection ramps — or face conservative default assumptions that penalize them.

On July 22, 2026, a single transmission fault in Ashburn, Virginia, triggered protection actions that removed more than 3 GW of data‑center load from the PJM Interconnection — roughly 3% of regional demand at the time. PJM recorded a measurable frequency deviation but no immediate loss of bulk‑system reliability. Still, the scale and mechanism of the disturbance reveal a recurring, avoidable failure mode NERC already documented in 2024.

What Happened (And Why It Matters)

Inside individual facilities, the response looked like success: static transfer switches acted in fractions of a cycle, UPS systems transferred to battery, and generators started — IT kept running. But when thousands of facilities execute the same customer‑side protections simultaneously, the aggregate effect is a multi‑gigawatt instantaneous load rejection that the generation fleet must absorb.

Lessons From NERC’s 2024 Investigation

On July 10, 2024, a lightning arrester failure on a 230‑kV line produced a permanent fault and a sequence of automatic recloses (three attempts at each end), producing six successive faults over 82 seconds. Voltages dipped to 0.25–0.40 per unit; the shortest fault lasted 42 ms, the longest 66 ms. About 1.5 GW of load disappeared, and NERC found the loss was exclusively computational/data‑center type load. Crucially, none of it was shed by utility protection devices — it was disconnected by customer‑side protection and controls.

Most of the sustained loss (≈1.26 GW) did not drop on the initial fault; it fell on the third voltage depression and remained offline for hours. NERC traced this to counting relays inside data centers that trip after several short voltage depressions (commonly set to three in one minute) and then hold on backup until manual reconnection.

Why This Is An Engineering Problem

On millisecond timescales, a data center looks less like a passive load and more like a distributed protection system: static transfer switches, UPS logic, breaker trip units, and generator controllers — thousands of devices per campus — each continuously monitoring voltage and each armed to separate the facility from the grid. Protection engineering requires coordination and grading (time and reach) so upstream devices delay long enough for local devices to clear a fault. The current, largely duplicated customer settings produce highly correlated behavior that protection engineers are trained to avoid.

Regulatory Response And The Near‑Term Compliance Calendar

Regulators are moving fast. Key recent actions include:

  • NERC Level 3 "Essential Actions" alert on computational load (May 4, 2026), responses due Aug. 3, 2026.
  • ERCOT NOGRR 282 ride‑through requirements, approved by the Texas PUC July 9, 2026, effective Aug. 1, 2026.
  • FERC directed NERC to develop new or revised reliability standards for computational loads by Dec. 31, 2026, with registry criteria and a Phase II plan due March 1, 2027.

NERC’s draft registry criteria would capture entities with an aggregate connected load of 20 MW or more at a single point of interconnection (60 kV or above) that host at least 1 MW of computational load — a threshold that reaches beyond hyperscale into many colocation and enterprise sites.

3 GW Left PJM — How Uncoordinated Data‑Center Protections Are Hitting the Grid (NERC Warned in 2024)

Modeling Defaults Versus Verified Capability

Where planners lack verified ride‑through data, they will apply conservative default assumptions. ERCOT’s large‑load stability guidance recommends modeling uncertain facilities as tripping below 0.75 per unit for ≥20 ms and not recovering for the rest of the simulation. NERC’s alert offers an illustrative relay of 0.85 per unit for 30 seconds. Those defaults penalize facilities that do not document and verify their true performance; conversely, verified ride‑through curves replace punitive assumptions.

What Data Centers Must Do

Practical steps for operators to avoid punitive defaults and reduce system risk include:

  • Document and publish ride‑through performance (ride‑through curves) and reconnection ramp rates.
  • Share protection settings and reclosing sequence information with planners and system operators.
  • Retire or redesign counting relays that latch a campus offline for hours without awareness of utility reclosing patterns.
  • Install fault recorders and perform commissioning tests that include ±10% voltage swings with live compute to validate behavior.
  • Coordinate reconnection procedures so large blocks of load return under managed ramps rather than instantaneously.

A Reciprocal Ask For Grid Operators

Publish reclosing sequences and protection practices on circuits serving clustered computational loads. Request verified ride‑through capability and reconnection ramps during interconnection and planning studies, and treat verified ride‑through curves like generator capability curves — documented, modeled, and contractual.

Costs, Risks, And The Strategic Choice

Achieving verified ride‑through and controlled reconnection has real costs: firmware and design changes, commissioning tests, instrumentation, and operational coordination. But the alternative is political and commercial risk. Each gigawatt‑scale uncontrolled event strengthens arguments against clustered data‑center growth, invites stricter defaults and standards, and undermines the industry's ability to monetize flexibility.

Conclusion: The industry has shown twice in two years that it can shed gigawatts in under a minute. The next time should be because a grid operator requested it and because the capability was compensated — not because uncoordinated customer protection disconnected clustered compute. Coordination, documentation, and verified testing are the practical path forward.

Author Note: Shalin Savalia is a senior electrical engineer at Amazon Web Services, working on data‑center power systems. He is a Senior IEEE member and an active contributor to the IEEE Power and Energy Society and the IEEE Industry Applications Society. The views expressed are his own and do not necessarily reflect those of his employer.

Help us improve.

Related Articles

Trending