The 2003 Blackout Was Not Caused by a Software Bug
Two sentences carry this case: a tree in Ohio, and a software bug. Both are true, and neither is the cause. The final report of the U.S.-Canada Power System Outage Task Force names four groups of causes, and the software is not one of them. At 16:05:57 EDT a transmission line in northern Ohio took itself out of service because its protective relay had seen a short circuit. There was no short circuit. Seven minutes later fifty million people were dark — and essentially nothing in that sequence had broken.
Chapters
- 0:00 A relay that saw a fault that was not there
- 0:38 A tree and a bug — both true, neither the cause
- 1:22 An operator never looks at the grid
- 1:59 14:14 — the routine that stalled and said nothing
- 2:49 Why a quiet console is the dangerous kind
- 3:31 The failover copied the failure
- 4:34 15:08 — the repair confirmed the wrong thing
- 5:13 Three lines, three trees, and 44 percent
- 6:08 The phone kept ringing
- 6:52 The second blind room
- 7:44 Sammis-Star had no fault on it
- 8:44 Seven minutes, and twelve seconds
- 9:33 Nothing broke, and there was no shortage
- 10:22 Nobody was required to ask if it worked
- 11:23 Recommendation number one
- 12:20 Five ways of not knowing
Transcript
A relay that saw a fault that was not there 0:00
At five minutes and fifty-seven seconds past four, on the afternoon of the fourteenth of August, two thousand and three, a single transmission line in northern Ohio took itself out of service. Its protective relay had seen a short circuit. There was no short circuit. Seven minutes later, an area holding fifty million people was dark: eight American states and the province of Ontario, sixty-one thousand eight hundred megawatts of load. More than five hundred generating units at two hundred and sixty-five power plants had shut down. Most of that sequence happened in the final twelve seconds.
A tree and a bug — both true, neither the cause 0:38
You already know the two sentences that go here. A tree in Ohio. A software bug. Both of them are true. Neither of them is the cause. The final report of the United States and Canada Power System Outage Task Force runs to two hundred and thirty-eight pages and names four groups of causes, and the software is not one of them. You will usually hear that fifty-five million people lost power. That number adds two national figures together; the report's own count for the affected area is fifty million. And what the report describes is not a machine that broke. It is a room that could not tell whether it could see.
An operator never looks at the grid 1:22
To follow what happened next you need one fact about control rooms. An operator does not look at the grid. They look at a model of the grid, drawn from telemetry that arrives every few seconds, on a system called an energy management system. Nobody can watch every line, every bus and every generator at once, so the model does the watching, and it speaks through alarms. The report compares them to the warning lights in a car: the door that is open, the tank that is nearly empty. You do not read the dashboard continuously. You wait for it to tell you.
14:14 — the routine that stalled and said nothing 1:59
At two fourteen in the afternoon, the alarm and event processing routine on FirstEnergy's system began to malfunction. The company's own analysis says the process stalled while handling a single alarm event, and never finished handling it. After that, nothing. No sound at the consoles. Nothing new on any screen. Nothing printed, and nothing written to the alarm log. General Electric found the reason eight weeks later, after auditing about a million lines of code: two processes were competing for the same data structure, and a coding error let both of them write to it at once. The data was corrupted and the alarm application went into an infinite loop. What it did not do, at any point, was report an error.
Why a quiet console is the dangerous kind 2:49
This is the part that matters, and it is not really about software. An alarm that fails loudly is an annoyance. An alarm that fails silently is worse than no alarm at all, because its failure looks exactly like the thing you are hoping for. A quiet console means a healthy grid. That is what quiet is for. So for the next hour and fifty-one minutes, the operators in Akron read the silence the way every rule they had told them to read it, and the silence was a lie. The instrument was not lying about a value. It had stopped producing values, and the absence of a value was indistinguishable from good news.
The failover copied the failure 3:31
The system did keep telling them something, just not in a language anyone was reading. At twenty past two several remote consoles failed, and an engineer in the computer group was paged automatically. At twenty-seven minutes past two the Star to South Canton line tripped and reclosed itself. Five minutes later American Electric Power telephoned the control room to ask about it, and FirstEnergy had no alarm and no log of the trip at all. At two forty-one the primary server hosting the alarm function failed, and everything on it moved automatically to the hot standby. Here is the sentence from the report: the alarm application moved onto the backup intact, still stalled and still ineffective. Thirteen minutes later the backup died of the same thing. By two fifty-four both servers were gone, and screen refreshes that normally took one to three seconds were taking fifty-nine. The redundancy did exactly what it was built to do. It copied the failure faithfully.
15:08 — the repair confirmed the wrong thing 4:34
At eight minutes past three the primary server was restored. The control calculations came back onto it, and the room got its normal performance back. The alarm routine went on malfunctioning, and nobody knew, because nothing in the repair had been designed to ask. This is one of the causes the task force names, in its own words: FirstEnergy lacked procedures to test effectively the functional state of its monitoring tools after repairs were made. The fix corrected what could be seen. It confirmed nothing about what could not. By then the first line was three minutes from falling.
Three lines, three trees, and 44 percent 5:13
At five minutes past three the Harding to Chamberlin line shorted to ground against a tree and locked out. At half past three the Hanna to Juniper line did the same. At twenty minutes to four the Star to South Canton line, which had already brushed the same tree twice that afternoon, went out for good. Three transmission lines at three hundred and forty-five thousand volts, all three lost to vegetation nobody had cut back. And here is the number that kills the overload story. The report says none of them failed from sag caused by heavy current. The first line was carrying forty-four percent of its rating when it fell. The second, eighty-eight. The third, ninety-three percent of its emergency rating, after the loss of the first two had pushed it from eighty-two percent of normal to a hundred and twenty, where it sat for ten minutes.
The phone kept ringing 6:08
The room was not cut off from the world. The reliability coordinator in Indiana called. American Electric Power called. The operator of the neighboring grid to the east called. FirstEnergy's own field crews called. And the report describes what the control room did with all of it in a single clause: they used the outdated system information they had to discount the information other people were giving them. That is not stupidity, and it is worth being precise about why. An instrument exists so that you can stop arguing with anecdotes. Everyone in that room had been trained to trust the screens over the telephone, and on any other afternoon that training was correct.
The second blind room 6:52
There was a second blind room that afternoon, and its blindness had nothing to do with the first. At two minutes past two, a line belonging to Dayton Power and Light had tripped in southern Ohio. The reliability coordinator's state estimator, the program that reconciles all the telemetry into one consistent picture, could not solve. An engineer had gone to lunch at half past one without re-arming the program's automatic trigger. By eight minutes past three he had worked out that the missing line was the Dayton one, and he walked to the control room to say so. The operators checked their own outage schedule, told him the line was in service, and asked him to model it that way. He did, and the model failed again. Two rooms, two pictures of the grid, and neither picture was the grid.
Sammis-Star had no fault on it 7:44
At five minutes and fifty-five seconds past four, the Dale to West Canton line, loaded to somewhere between a hundred and sixty and a hundred and eighty percent of normal, gave out. Two seconds later the Sammis to Star line, now above a hundred and twenty percent, opened. And this one is different from the three before it. Those three had shorted to ground on trees. Sammis to Star had no fault on it at all. Its protective relays saw a low apparent impedance, which is depressed voltage divided by abnormally high current, and the relay reacted as if that flow were a short circuit. A distance relay measures volts divided by amps, because that ratio goes small when there is a fault nearby. A heavily loaded line with sagging voltage produces the same small number. The relay was not broken and it was not miscalibrated. It was answering its question correctly. It was the wrong question.
Seven minutes, and twelve seconds 8:44
That trip was the turning point. From five minutes and fifty-seven seconds past four to thirteen minutes past four, the blackout travelled from Akron across the north-east of the United States and into Canada. Seven minutes. More than four hundred transmission lines and more than five hundred and thirty generating units. The report is careful about what kind of event this was: between ten minutes past four and thirteen minutes past, thousands of things happened on the grid, driven by physics and by automatic equipment operating. Nobody decided any of it. Each line that opened pushed its power onto the next one, and the next relay saw the same picture and drew the same conclusion. Most of the whole sequence took place in the final twelve seconds.
Nothing broke, and there was no shortage 9:33
So it is worth saying clearly what did not fail. There was no shortage of generation. The report addresses the reactive power question head on, and says that insufficient reactive power was an issue in the blackout, but not a cause in itself. Voltage on the Star bus had been drifting down all day, from ninety-eight and a half percent at eleven in the morning to ninety-five point nine at five past three. FirstEnergy's operators described that as typical for a warm summer day on their system. They were right. It was typical. And in the cascade itself, essentially nothing broke. Four hundred lines and five hundred generators removed themselves from service correctly, obeying settings that engineers had chosen on purpose.
Nobody was required to ask if it worked 10:22
The root cause is written down, and it is not the race condition. It is the second of the four groups, item B: FirstEnergy lacked procedures to ensure that its operators were continually aware of the functional state of their critical monitoring tools. There was no check. There was no heartbeat, no watchdog, no question anybody was required to ask, and no way for a person at that console to find out whether the thing in front of them was still alive. And the bug tells the same story from the other end. The engineer who found it said the software had run more than three million operational hours in the field without anything ever exercising that defect. Which is the sentence from our last episode, written in software instead of metal. A protective function that has never once been exercised is indistinguishable, from the outside, from one that works. Except that this time, the protection that was never exercised was the instrument itself.
Recommendation number one 11:23
Recommendation number one in that report is one sentence long: make reliability standards mandatory and enforceable, with penalties for non-compliance. Until then they were voluntary. That is why, in November of two thousand and three, the United States Secretary of Energy said his department would not be seeking to punish FirstEnergy: there was no rule to have broken. The Energy Policy Act of two thousand and five created the electric reliability organization the report had asked for. The regulator certified the North American Electric Reliability Corporation in that role in July of two thousand and six, and the standards became enforceable. Tree clearance stopped being good practice and became a numbered standard. And the investigators asked for something for the relays too: better application of zone three impedance relays, which in plain language means teaching the protection to tell a heavy load from a fault.
Five ways of not knowing 12:20
Five episodes now, and five ways of not knowing. A theory that was right inside the range it had been tested in. A checklist that was complete for every tower built before it. A criterion that asked about frequency when the problem was stability. A safety device that had never once been asked to act. And now an instrument whose failure looked exactly like good news. Next time, a change on a drawing. In the summer of nineteen eighty-one, in a hotel atrium in Kansas City, a fabricator asked for one small change to a walkway hanger. Instead of one continuous steel rod running through two levels, there would be two shorter rods. Geometrically it is the same building. Structurally, it doubled the load on a single connection, and nobody did the arithmetic again.
Description and sources
At 14:14 EDT on 14 August 2003, FirstEnergy's alarm routine stalled on one event and never finished — no alarms, nothing on screen, nothing logged. GE later traced it to a race condition: two processes writing to the same data at once, looping the application forever.
What matters is the shape of the failure: an alarm that fails loudly is an annoyance; one that fails silently is worse than none, because its failure mode is identical to the success state of the system it watches. A quiet console means a healthy grid. For nearly two hours the operators in Akron read that silence the way every rule told them to, while redundancy made it worse: the stalled application moved onto the standby server intact and killed that too.
Three 345kV lines then shorted against untrimmed trees, none from overload. At 16:05:57 the Sammis–Star line opened with no fault at all: its relay saw depressed voltage over high current, read that as a short circuit — correct math, wrong question. Minutes later, hundreds of lines and generating units were down, most of it in the final twelve seconds.
The root cause isn't the race condition. It's that FirstEnergy had no procedure to keep operators aware of the functional state of their own monitoring tools — the same failure as the previous episode, in software instead of metal, except this time the instrument itself was the protection never exercised.
The fix was regulatory: standards, voluntary until then, became mandatory under the Energy Policy Act of 2005, enforced by NERC from July 2006.
PRINT-READY, FROM THIS CHANNEL
The Failure Atlas, Vol. 01 — Tacoma Narrows · Citicorp Center · Millennium Bridge · Apollo 13 · the 2003 blackout · Hyatt Regency
https://therepository.gumroad.com/l/failure-atlas
PRIMARY SOURCES
- U.S.-Canada Power System Outage Task Force, Final Report on the August 14, 2003 Blackout in the United States and Canada: Causes and Recommendations, April 2004 — chapter 1 (scale of the outage), chapter 2 (the four groups of causes), chapter 5 (how and why the blackout began in Ohio), chapter 6 (the cascade), chapter 10 (recommendations). https://www.nerc.com/globalassets/our-work/reports/event-reports/august_2003_blackout_final_report.pdf
- Same report, Appendix B — NERC Actions of 10 February 2004, recommendation 8 (zone 3 impedance relays, under-voltage load shedding, and the count of units and lines lost).
- Same report, Appendix D — sequence of events, including the FirstEnergy EMS timeline (14:14, 14:41, 14:54, 15:08) and the MISO state estimator timeline (13:30, 15:08, 15:09, 15:17).
- Kevin Poulsen, Tracking the Blackout bug, The Register, 8 April 2004 — GE Energy's account of the XA/21 race condition: the code audit, the contention over a shared data structure, and the three million operational hours that had never exercised the defect. https://www.theregister.com/2004/04/08/blackout_bug_report/
- Energy Policy Act of 2005, section 1211 (electric reliability organization); NERC certified by FERC as the ERO, July 2006.
The 21 technical plates in this video are illustrations generated for the channel by a diffusion image model, styled to match its cyanotype identity. They are diagrams of the system, not photographs of the hardware, and no person is depicted in any of them.
Root Cause investigates why engineered systems fail, using the official investigation reports and the primary technical literature. Sources for this episode are linked above.
The technical drawings in this video are cyanotype-style illustrations produced for the channel. They are diagrams, not photographs of the real hardware. The charts, timelines and dimensioned comparisons are drawn from the figures and text in the sources listed above.