# The Bug That Killed the Lights for 50 Million People

> How a race condition in power grid software helped cause the 2003 Northeast blackout, which cut power to an estimated 50 million people.

- Author: Shawn Ng (https://shawnng.com/about)
- Published: 2026-03-07
- Updated: 2026-09-30
- Canonical URL: https://shawnng.com/posts/the-bug-that-killed-the-lights
- Topics: software-bugs, systems, race-conditions, critical-infrastructure

## Key takeaways

- A race condition stalled FirstEnergy's XA/21 alarm system at 2:14 PM on August 14, 2003, with no crash and no error message.
- Operators watched normal-looking screens for 90 minutes while three 345 kV transmission lines sagged into trees and tripped offline.
- The cascade that began at 4:05 PM took more than 508 generating units at 265 power plants offline by 4:13 PM, darkening eight states and Ontario.
- Silent failures are the worst failures because a monitor that stops monitoring destroys your ability to detect every other problem.
- Redundant systems must be truly independent: FirstEnergy's backup server failed 13 minutes after the primary, because the stalled alarm process moved onto it intact.

On August 14, 2003, the lights went out for an estimated 50 million people across the northeastern United States and Ontario. It was the largest blackout in North American history, and in the middle of it sat a programming error that most software engineers learn about in their first concurrency course: a **race condition**.

Check out the full video on my YouTube channel [Divide and Quantum](https://www.youtube.com/watch?v=4ZipI4ivbsQ).

## What Actually Happened That Day

The story starts at FirstEnergy, an Ohio-based power utility. Their energy management system (EMS) was responsible for monitoring the state of the electrical grid in real time. Operators relied on this software to see alarms, track power line loads, and respond to problems before they cascaded.

Around 2:14 PM, a software bug in the XA/21 alarm and logging system caused the primary alarm system to stall. GE Energy, which made the XA/21, later traced it to a race condition: a timing flaw where two processes competed for the same shared resource, and the result depended on which one got there first. The window for it was measured in milliseconds. In this case, the wrong process won.

The alarm system went silent. No crash. No error message. No restart. It just stopped updating. Operators were staring at screens that looked perfectly normal while the grid was falling apart underneath.

Over the next 90 minutes, three high-voltage transmission lines sagged into overgrown trees and tripped offline, one by one. Each failure pushed more load onto the remaining lines. Without alarms, no one at FirstEnergy noticed. By the time they realized something was wrong, the cascade was unstoppable.

At 4:05 PM, one more line tripped and the cascade began. By 4:13 PM, more than 508 generating units at 265 power plants had gone offline. The lights went out from New York City to Toronto.

## What Is a Race Condition?

A race condition occurs when the behavior of a program depends on the relative timing of two or more concurrent operations. When the operations happen in the "expected" order, everything works fine. When they don't, the program enters an invalid state.

Here is a simple illustration. Imagine two threads both trying to increment a shared counter:

```
Thread A: read counter (value = 5)
Thread B: read counter (value = 5)
Thread A: write counter (value = 6)
Thread B: write counter (value = 6)  // Should be 7!
```

Both threads read the value before either writes, so one increment is lost. This is the classic "check-then-act" race condition. The result depends on which thread executes which instruction at which moment, something that can vary with CPU load, operating system scheduling, or even temperature. Timing that fine is a property of the hardware as much as the code, down to [whether a value is sitting in cache or in main memory](https://shawnng.com/posts/cache-misses-performance-killer).

In the FirstEnergy system, the race condition was more subtle. It occurred within the alarm processing software, where two processes contended for a common [data structure](https://shawnng.com/posts/data-structures-every-developer-should-master). "Through a software coding error in one of the application processes, they were both able to get write access to a data structure at the same time," GE Energy's Mike Unum told reporter Kevin Poulsen in 2004. "And that corruption lead to the alarm event application getting into an infinite loop and spinning." Under normal load, the timing worked out. Under the specific conditions that afternoon, it didn't.

## How the Bug Cascaded Into a Blackout

The race condition alone did not cause the blackout. What it did was **remove the safety net**. The power grid is designed to handle equipment failures; lines trip all the time. The system recovers because operators see the alarms, assess the situation, and take corrective action like rerouting power or shedding load.

Here is the cascade in slow motion:

1. **2:14 PM** - The race condition triggers. The alarm system freezes but does not crash. At 2:41 PM the primary control server fails, and at 2:54 PM the backup fails too, because the stalled alarm process moved onto it intact. Separately, MISO, the regional coordinator watching FirstEnergy's grid from outside, has its own state estimator failing on bad line-status data, so the one outside check is partly blind as well.

2. **3:05 PM** - The Chamberlin-Harding 345 kV line sags into a tree and trips. Operators do not see the alarm.

3. **3:32 PM** - The Hanna-Juniper 345 kV line trips. Still no alarm. Load redistributes across remaining lines.

4. **3:41 PM** - The Star-South Canton 345 kV line trips. By now, the remaining transmission corridors are dangerously overloaded.

5. **4:05 PM** - The Sammis-Star 345 kV line trips, and the cascade begins. Power surges and voltage collapses propagate at close to the speed of light across interconnected grids. Protective relays trip generation plants offline to prevent physical damage to turbines. Between 4:10:36 and 4:10:46 PM the grid breaks apart into islands, and by 4:13 PM more than 508 generating units at 265 power plants are offline across eight states and Ontario.

The task force put the cost in the United States alone at $4 billion to $10 billion. Water treatment plants lost power. Cellular networks went down. Hundreds of people were trapped in elevators and subway cars.

## Why Silent Failures Are the Worst Failures

The most dangerous aspect of this bug was that it produced a **silent failure**. The alarm system did not crash and restart. It did not display an error. It simply stopped doing its job while appearing to function normally.

This is a well-known pattern in reliability engineering. A monitoring system that crashes is annoying but recoverable. A monitoring system that silently stops monitoring is catastrophic because it destroys your ability to detect and respond to any other failure.

This is sometimes called a "Byzantine failure," where a component continues to run but produces incorrect or incomplete results. It is far harder to detect than a clean crash.

## Lessons for Engineers Building Critical Systems

The 2003 blackout is a case study taught in computer science and systems engineering programs worldwide. Here are the key takeaways.

### Design for Failure, Not Just Function

Every monitoring system should have a "watchdog" mechanism, a secondary process that verifies the primary system is actually working. If the alarm system had a heartbeat check that detected the stall, operators could have been notified through a backup channel.

### Race Conditions Must Be Eliminated, Not Just Mitigated

Race conditions are not theoretical curiosities. In critical infrastructure, they must be found and fixed through proper synchronization primitives like mutexes, semaphores, or lock-free data structures. Static analysis tools and formal verification methods can catch many race conditions before deployment.

### Defense in Depth Means Independent Layers

FirstEnergy had a backup server, and it failed anyway: the stalled alarm process moved onto it intact and took it down 13 minutes after the primary. A warm reboot at 3:08 PM brought the server back up with the alarm system still frozen, and nobody confirmed with the control room that alarms were working. Redundant systems must be truly independent. If your backup monitoring tool shares a database connection pool with the primary tool, they will fail together.

### Test Under Realistic Load

The race condition in the XA/21 system was not triggered under normal conditions. GE described it as a perfect storm of events and alarm conditions, and engineers had to inject deliberate delays into the code to reproduce it. Testing must include stress tests, fault injection, and chaotic scheduling to surface timing-dependent bugs.

### Software in Critical Infrastructure Is Critical Infrastructure

The U.S.-Canada Power System Outage Task Force named four groups of principal causes, and one of them was inadequate situational awareness, which is where the frozen alarms sit. The operators could not respond to what their software had stopped showing them. Software that controls physical systems carries physical consequences. It deserves the same rigor as the hardware it monitors.

## The Bigger Picture

Twenty-plus years later, our power grids, transportation networks, and financial systems are more software-dependent than ever. So are the devices we carry, right down to [the hidden computer inside a phone's SIM card](https://shawnng.com/posts/sim-card-hidden-computer). The 2003 blackout is a reminder that the most dangerous bugs are not the ones that crash your program. They are the ones that let your program keep running while it lies to you.

If you write software that other people depend on, especially software that monitors other systems, build in the assumption that your code will fail. Then build in the mechanisms to detect that failure when it happens.

The lights went out for 50 million people, and one reason was that two processes raced and the wrong one won. That is the stakes of concurrent programming in critical systems.

## Frequently asked questions

### What caused the 2003 Northeast blackout?

Several failures together. The one in software was a race condition in FirstEnergy's XA/21 energy management system: two processes got write access to the same data structure at once, and the alarm application stalled silently at 2:14 PM, so operators never saw three transmission lines trip over the next 90 minutes. The U.S.-Canada task force grouped the lost alarms under inadequate situational awareness, one of four groups of principal causes alongside poor tree trimming and weak system planning and oversight.

### What is a race condition?

A race condition is when a program's behavior depends on the relative timing of two or more concurrent operations. The classic case is two threads that both read a counter at 5 and both write 6, losing one increment. Whether it happens can vary with CPU load, operating system scheduling, or even temperature.

### How many people lost power in the 2003 blackout?

An estimated 50 million people across eight U.S. states and Ontario, according to the U.S.-Canada task force, which put the cost in the United States at $4 billion to $10 billion. Water treatment plants lost power, cellular networks went down, and hundreds of people were trapped in elevators and subway cars.

### How do you prevent race conditions in critical systems?

Use proper synchronization primitives such as mutexes, semaphores or lock-free data structures, and catch what remains with static analysis and formal verification. Give every monitoring system a watchdog that detects a stall, keep redundant systems genuinely independent, and test with stress, fault injection and chaotic scheduling.

## Sources

- U.S.-Canada Power System Outage Task Force (2004). [Final Report on the August 14, 2003 Blackout in the United States and Canada: Causes and Recommendations](https://www.energy.gov/sites/default/files/oeprod/DocumentsandMedia/BlackoutFinal-Web.pdf).
- Kevin Poulsen (2004). [Tracking the Blackout bug](https://www.theregister.com/2004/04/08/blackout_bug_report/). The Register.
