U.S.-Canada Power System Outage Task Force Final Report on the August 14, 2003 Notes
Introduction
- On August 14, 2003, a blackout affected large portions of the Midwest and Northeast United States and Ontario, Canada.
- The outage impacted an area with approximately 50 million people and 61,800 MW of electric load across several states and Ontario.
- The blackout started shortly after 4:00 pm EDT, with power restoration taking up to 4 days in some U.S. areas and over a week in some parts of Ontario.
- Estimated total costs ranged from $4 billion to $10 billion (USD) in the U.S.
- In Canada, the GDP decreased by 0.7% in August, resulting in a loss of 18.9 million work hours and a \$2.3 billion (CAD) decrease in manufacturing shipments in Ontario.
- President George W. Bush and Prime Minister Jean Chrétien established a joint U.S.-Canada Power System Outage Task Force on August 15 to investigate the causes and recommend ways to prevent future outages.
- The Task Force was co-chaired by U.S. Secretary of Energy Spencer Abraham and Canadian Minister of Natural Resources Herb Dhaliwal (later succeeded by John Efford).
- The Task Force had two phases:
- Phase I: Investigate the causes and containment failures of the outage
- Phase II: Develop recommendations to reduce future outages and limit their scope.
- Three Working Groups assisted the Task Force:
- Electric System Working Group (ESWG)
- Nuclear Working Group (NWG)
- Security Working Group (SWG)
- The Task Force published an Interim Report on November 19, 2003, summarizing the investigation's findings.
- This Final Report updates and supersedes the Interim Report, including new information and detail.
Overview of the North American Electric Power System and Its Reliability Organizations
- The North American electricity system represents over \$1 trillion (U.S.) in asset value.
- It includes more than 200,000 miles (320,000 km) of transmission lines operating at 230,000 volts and greater.
- It possesses 950,000 MW of generating capacity.
- It serves well over 100 million customers and 283 million people.
- Modern society depends on reliable electricity for national security, communications, finance, transportation, food and water supply, and more.
- Providing reliable electricity is a complex technical challenge involving real-time assessment, control, and coordination.
- Electricity is produced at lower voltages (10,000 to 25,000 volts) and stepped up for transportation over transmission lines.
- High voltage (230,000 to 765,000 volts) reduces losses and allows economical long-distance power transmission.
- Transmission lines are interconnected at switching stations and substations to form a power "grid."
- Electricity flows from generators to loads along paths of least resistance.
- The bulk power system is predominantly an alternating current (AC) system.
- There are three distinct power grids or "interconnections" in North America:
- Eastern Interconnection: Eastern two-thirds of the continental United States and Canada
- Western Interconnection: Western third of the continental United States, the Canadian provinces of Alberta and British Columbia
- Texas Interconnection: Most of the state of Texas
- The three interconnections are electrically independent except for a few small direct current (DC) ties.
- Within each interconnection, electricity is produced the instant it is used and flows over transmission lines from generators to loads.
- Reliable operation of the power grid is demanding:
- Electricity flows at close to the speed of light (186,000 miles per second or 297,600 km/sec) and is not economically storable in large quantities.
- The flow of alternating current (AC) electricity cannot be controlled like a liquid or gas by opening or closing a valve.
- NERC and its ten Regional Reliability Councils have system operating and planning standards based on seven key concepts:
- Balance power generation and demand continuously. Failure to match generation to demand causes the frequency of an AC power system to increase (when generation exceeds demand) or decrease (when generation is less than demand).
- Balance reactive power supply and demand to maintain scheduled voltages. Reactive power sources, such as capacitor banks and generators, must be adjusted during the day to maintain voltages within a secure range pertaining to all system electrical equipment.
- Monitor flows over transmission lines to ensure that thermal (heating) limits are not exceeded. The flow must be limited to avoid overheating and damaging the equipment. Most transmission lines, transformers and other current-carrying devices are monitored continuously to ensure that they do not become overloaded or violate other operating constraints.
- Keep the system in a stable condition. Because the electric system is interconnected and dynamic electrical stability limits must be observed.
- Operate the system so that it remains in a reliable condition even if a contingency occurs, such as the loss of a key generator or transmission facility (the "N-1 criterion"). This principle is expressed by the requirement that the system must be operated at all times to ensure that it will remain in a secure condition (generally within emergency ratings for current and voltage and within established stability limits) following the loss of the most important generator or transmission facility (a “worst sin- gle contingency”).
- Plan, design, and maintain the system to operate reliably.
- Prepare for emergencies.
- Responsible system owners and operators practice "defense in depth," protecting the bulk power system through layers of safety-related practices and equipment.
Reliability Organizations
- NERC is a non-governmental entity ensuring the bulk electric system in North America is reliable, adequate, and secure.
- Established in 1968 after the Northeast blackout in 1965, NERC has operated as a voluntary organization.
- NERC's members are ten regional reliability councils that adapt NERC standards to their regions.
- The August 14 blackout affected three NERC regional reliability councils: ECAR, MAAC, and NPCC.
- "Control areas" are the primary entities subject to NERC and regional council reliability standards.
- A control area is a geographic area where an entity balances generation and loads in real-time.
- There are approximately 140 control areas in North America.
- Utility industry restructuring has led to Independent System Operators (ISOs) and Regional Transmission Organizations (RTOs).
- ISOs and RTOs manage real-time and day-ahead reliability and operate wholesale electricity markets.
- Reliability coordinators provide reliability oversight over a wide region, preparing assessments and coordinating emergency operations.
- There are currently 18 reliability coordinators in North America.
- The initiating events of the blackout involved two control areas, FirstEnergy (FE) and American Electric Power (AEP), and their respective reliability coordinators, MISO and PJM.
- Control area operators have primary responsibility for reliability
- They shall operate so that instability, uncontrolled separation, or cascading outages will not occur as a result of the most severe single contingency
- Emergency Preparedness and Emergency Response
- Action to correct an OPERATING SECURITY LIMIT violation shall not impose unacceptable stress on internal generation or transmission equipment, reduce system reliability beyond acceptable limits, or unduly impose voltage or reactive burdens on neighboring systems. If all other means fail, corrective action may require load reduction.
- Operating personnel and training
- Reliability Coordinators such as MISO and PJM are expected to comply with all aspects of NERC Operating Policies.
Causes of the Blackout and Violations of NERC Standards
- This chapter explains the causes of the blackout in Ohio and lists NERC’s findings concerning violations of its reliability policies.
- The blackout was caused by deficiencies in practices, equipment, and human decisions.
- Four major cause groups:
- Group 1: FirstEnergy and ECAR failed to assess and understand the inadequacies of FE’s system, The northeastern Power grid was highly vulnerable.
- Group 2: Inadequate situational awareness at FirstEnergy. FE did not recognize or understand the deteriorating condition of its system.
- Group 3: FE failed to manage adequately tree growth in its transmission rights-of-way.
- Group 4: Failure of the interconnected grid’s reliability organizations to provide effective real-time diagnostic support.
- The seven violations of NERC standards, as identified by NERC:
- Violation 1: Following the outage of the Chamberlin-Harding 345-kV line, FE operating personnel did not take the necessary action to return the system to a safe operating state as required by NERC Policy 2, Section A, Standard 1.
- Violation 2: FE operations personnel did not adequately communicate its emergency operating conditions to neighboring systems as required by NERC Policy 5, Section A.
- Violation 3: FE’s state estimation and contingency analysis tools were not used to assess system conditions, violating NERC Operating Policy 5, Section C, Requirement 3, and Policy 4, Section A, Requirement 5.
- Violation 4: MISO did not notify other reliability coordinators of potential system problems as required by NERC Policy 9, Section C, Requirement 2.
- Violation 5: MISO was using non-real-time data to support real-time operations, in violation of NERC Policy 9, Appendix D, Section A, Criteria 5.2.
- Violation 6: PJM and MISO as reliability coordinators lacked procedures or guidelines between their respective organizations regarding the coordination of actions to address an operating security limit violation observed by one of them in the other’s area due to a contingency near their common boundary, as required by Policy 9, Appendix C.
- Violation 7: The monitoring equipment provided to FE operators was not sufficient to bring the operators’ attention to the deviation on the system. (Policy 4, Section A, System Monitoring Requirements regarding resource availability and the use of monitoring equipment to alert operators to the need for corrective action.)
Context and Preconditions for the Blackout
- This chapter reviews the state of the northeast portion of the Eastern Interconnection during the days and hours before 16:00 EDT on August 14, 2003.
- At 15:05 EDT, the system was electrically secure and able to withstand any single one of more than 800 contingencies, before the trip- ping of FirstEnergy’s (FE) Harding-Chamberlin 345-kV transmission line.
- Electric Demands on August 14.
- Temperatures were hot in the northeast region typical for August, system operators has sucsessfully managed such temperatures/outputs in the pass.
- Generation Facilities Unavailable on August 14
- Several key generators in the region were out of service going into the day of August 14
- These were Davis-Besse Nuclear Unit, Sammis Unit 3, Eastlake Unit 4, Monroe Unit 1, Cook Nuclear Unit 2.
- Unanticipated Outages of Transmission and Generation on August 14.
- In addition to the generators out of service there were a series of unanticpated outtages,
- these included Cinergy transmission lines in south-central Indiana, FE’s Eastlake 5 generating unit and a line within the Dayton Power and Light (DPL) control area, the Stuart-Atlanta 345-kV line in southern Ohio,
- Key Parameters for the Cleveland-Akron Area at 15:05 EDT
- Cleveland-Akron area load = 6,715 MW and 2,402 MVAr
- Transmission losses = 189 MW and 2,514 MVAr
- Reactive power from fixed shunt capacitors (all voltage levels) = 2,585 MVAr
- Reactive power from line charging (all voltage levels) = 739 MVAr
- Network configuration = after the loss of Eastlake 5, before the loss of Harding- Chamberlin 345-kV line
- Area generation combined output: 3,000 MW and 1,200 MVAr.
- Power Flow Patterns
- Flows were heavy but still within ranges and well within estalished limits, flows did however increase steadily before the black out.
- Voltages and Voltage Criteria
- Voltages were depressed throughout Northern ohio, this was due to the hot temperatures and general draw of power from the grid, Voltages however were within normal tolerance.
- FirstEnergy use minimum acceptable normal voltages, which are lower than and incompatible with those used by its interconnected neighbors.
- Past System Events and Adequacy of System Studies
- Past events are not reflected in any FirstEnergy or ECAR seasonal or longer-term planning studies or operating protocols.
- Modle based analysis confirmed that as of 15:05 on August 14, the North Eastern power grid was in stable operating state.
How and Why the Blackout Began in Ohio
- This chapter explains the major events that occurred in Ohio in the hours leading up to the blackout.
- Several phases in this chapter describe how the events were related to one another:
- Phase 1: A normal afternoon degrades
- Eastlake Unit 5 tripped along the southwestern shore of Lake Erie tripped - 13:31 EDT
- Dayton Power & Light’s (DPL) Stuart-Atlanta 345-kV line tripped--
- - 14:02 EDT
- inaccurate input data rendered MISO’s state estimator (a system monitoring tool) ineffective.
- Phase 2: FE’s computer failures - After 14:14 EDT
- the alarm and logging system
- the EMS system lost a number of its remote control consoles
- Phase 3: Three FE 345-kV transmission line failures and Many Phone Calls - after 15:05 EDT
- FE’s 345-kV transmission lines began tripping out because the lines were contacting overgrown trees within the lines’ right-of-way areas.
- Harding-Chamberlin 345-kV line tripped - 15:05:41 EDT
- Hanna-Juniper 345-kV line tripped - 15:32:03 EDT
- Star-South Canton 345-kV tripped - 15:41:33-41 EDT
- FE’s 345-kV transmission lines began tripping out because the lines were contacting overgrown trees within the lines’ right-of-way areas.
- Phase 4: 138-kV Transmission System Collapse in Northern Ohio - - After 15:46 EDT
- the loss of some of FE’s key 345-kV lines in northern Ohio caused its underlying network of 138-kV lines to begin to fail
- FirstEnergy’s 345-kV line failure (sammis-Star at the end) at 16:06 EDT triggered the uncontrollable 345 kV cascade portion of the blackout sequence
- Phase 1: A normal afternoon degrades
The Cascade Stage of the Blackout
- This chapter details the reasons behind the reach and acceleration of the cascade beyond the Cleveland-Akron area
- In addition the chapter notes the reasons the blackout stopped where it did.
- The August 14 blackout's reach was due to three major factors:
- The Line failure of the Sammis-Star initiating numerous other line trips
- Many Key lines were opperated via zone 3 impendance relays that reacted to overloads
- relay protection settings were inappropriate, uncoordinated or intergrated.
- The blackout can be split into three phases:
- Phase 5: 345-kV Transmission System Cascade in Northern Ohio and South-Central Michigan - After Faliure of Sammis-star line
- Phase 6: The Full Cascade - After 16:10:36 EDT, the wave continued causing power surges, resulting in separations across the state lines and in canda in Ontario\Michigan
- Phase 7: Several Electrical Islands Formed in Northeast U.S. and Canada: 16:10:46 EDT to 16:12 EDT -- High power surges led to generation trips, The system broke into smaller and smaller islands ultimately causing the blackout.
- The blackout stopped where it did for various reasons
- The further any source is from the origin smaller the ripples become.
- More heavily layed electrical system such as those in PJN were able to adbsorb the surges
- Line trips made the grid break into islands effectively isolating the damage.
- Under-Frequency and Under-Voltage Load-Shedding also helped the electrical system not further the damage.
- Generators Tripping Off protected the generators from the damage caused from said oscillations to avoid damage..
August 14 Compared With Previous Blackouts
- Power System Outages. Short, localized outages occur on power systems fairly frequently.
- This chapter reviews seven previous outages and compares them with the August 14, 2003, blackout.
- Common Factors Among Major Outages
*Conductor contact with trees
*Over-estimation of dynamic reactive output of system generators
*Inability of system operators or coordinators to visualize events on the entire system
*Failure to ensure that system operation was within safe limits
*Lack of coordination on system protection
*Ineffective communication
*Lack of “safety nets
*Inadequate training of operating personnel - Comparisons With the August 14 Blackout
*Inadequate vegetation management
*Failure to ensure operation within secure limits
*Failure to identify emergency conditions and communicate that status to neighboring systems
*Inadequate operator training
*Inadequate regional-scale visibility over the power system
*Inadequate coordination of relays and other protective devices or systems.
Nuclear Power Plants Affected by Blackout:
Findings include:
- The affected nuclear power plants did not trigger the power outage or inappropriately contribute to its spread
- For U.S. nuclear plants: tripped due to main generator trips, and various reactor trips occurred across the various power plants.
- For Canadian nuclear plants: frequency and/or voltage fluctuations on the grid resulted in the automatic disconnection of generators from the grid.
- Recommendations for more details.
Physical and Cyber Security Aspects of the August 14 Blackout
The SWG found no evidence that malicious actors caused or contributed to the power outage. SWG also found no reason to amend, alter, or negate any of the informa-tion submitted to the Task Force for the Interim Report
Main findings from analysis include:there are potential opportunities for cyber system compromise of Energy Management Systems (EMS) and their supporting information technology (IT) infrastructure. Indications of procedural and technical IT management vulnerabilities
A failure in a software program not linked to malicious activity may have significantly contributed to the power outage. Internal and external links from Supervisory Control and Data Acquisi-tion (SCADA) networks to other systems introduced vulnerabilities.
There was also a lack of a system or process for some grid operators to adequately view the status of electric systems outside of their immediate control.
Recommendations made on these findings:Implement NERC IT standards.
Develop and deploy IT management procedures.
Develop corporate-level IT security governance and strategies.
Implement controls to manage system health, network monitoring, and incident management.
Initiate U.S.-Canada risk management study.
Improve IT forensic and diagnostic capabilities.
Assess IT risk and vulnerability at scheduled intervals.
Develop capability to detect wireless and remote wireline intrusion and surveillance.
Control access to operationally sensitive equipment.
NERC should provide guidance on employee background checks.
Confirm NERC ES-ISAC as the central point for sharing security information and analysis.
Establish clear authority for physical and cyber security.
Develop procedures to prevent or mitigate inappropriate disclosure of information.
Recommendations to Prevent Blackouts
- A comprehensive, detailed list of 46 recommendations designed to improve future incidents by improving the industry's protocols and equipment.
- The recommendaions can be sumed up into four parts:
- Making aherence to high reliability standards top proerity to the industry.
- Recognize that Relibility in not free, and there will need to be investments to follow such protoocols.
- Make efforts/Mechanisim to enforce perfromance based standards.
- The establishment of cyber protocols to ensure physical saftery.