ENGINEERING STORY 003

The fault was fixed.
The system wasn't.

Restoring the thing that failed is not necessarily the same as recovering the system around it.

WaypointHQ Engineering Story 003: an aviation system recovering after a technical fault.

The announcement everyone had been waiting for finally arrived.

The technical problem had been resolved.

So why were flights still being cancelled?

Why were passengers still stranded?

And why could disruption continue long after the equipment that caused the original problem was working again?

We confuse restoration with recovery.

It is an easy mistake to make.

Something fails. Engineers find the problem. The fault is corrected. The affected equipment returns to service.

From the viewpoint of that equipment, the incident may be over.

From the viewpoint of the wider system, recovery may only just have begun.

An aircraft cannot simply disappear from the queue.

Air traffic control is part of a much larger aviation system.

Airlines, airports, aircraft, crews, passengers, stands, baggage, engineering support, ground handling and airspace all have to work together for the timetable to function.

Much of that system is designed to use its available capacity efficiently.

Now imagine that part of the system would normally handle 100 movements during a period of operation.

A failure occurs and, to maintain safety, capacity has to be reduced. Only 60 can be handled.

The other 40 have not disappeared.
They have become part of the recovery problem.

Normal capacity cannot automatically clear abnormal demand.

Suppose the technical fault is fixed and the original capacity of 100 becomes available again.

That sounds like normal operation has returned.

But the system now has today's demand plus some of the demand it could not handle yesterday.

If the system was already designed to operate efficiently near its practical capacity, there may be very little spare capacity available to absorb that backlog.

When a system runs close to capacity, yesterday's failure becomes tomorrow's workload.

The backlog is not sitting neatly in a queue.

Aviation does not reset itself overnight.

An aircraft that should have ended the day in one city may now be somewhere else.

The crew scheduled to operate tomorrow's first flight may not be where tomorrow's schedule expects them to be.

Passengers may need rebooking. Airport stands may be occupied differently. Connections are missed. Working-time restrictions can begin to matter.

One technical failure has created consequences in parts of the system that never failed at all.

The original fault has gone. Its consequences are still travelling through the system.

Failure and recovery are different problems.

Engineers are usually very good at asking what happens when something fails.

What is the fallback? What capacity remains? What is the safe degraded mode? How do we restore the failed equipment?

Those are essential questions.

But there is another question that deserves the same attention.

What happens to the whole system after the failed equipment comes back?

A resilient system is not simply one that survives a failure. It is one that has been designed to recover from it.

Spare capacity can look wasteful until the day you need it.

There is an uncomfortable economic reality behind resilient systems.

Capacity that is not normally used can look inefficient.

Spare equipment costs money. Additional staff cost money. Alternative routes, facilities and recovery arrangements cost money.

If serious failures are rare, it can be tempting to remove that apparent inefficiency.

But when almost every available resource is required for normal operation, there may be nothing left with which to recover from abnormal operation.

Efficiency and resilience are not enemies, but pretending there is no trade-off between them is dangerous.

The same problem appears everywhere.

A hospital can restore an unavailable system while still facing a backlog of patients.

A telecommunications network can restore a failed service while customers and dependent systems are still recovering.

A manufacturer can repair a production line while orders continue to arrive faster than the backlog can be cleared.

An IT service can come back online while queued transactions, failed jobs and dependent services remain in an abnormal state.

Restoring the component is an engineering task.
Recovering the system is a systems engineering task.

Recovery needs an architecture too.

Good resilience engineering therefore looks beyond the instant of failure.

It considers how demand will be controlled, how backlogs will be managed, which services recover first, what resources are required and how the system moves safely from degraded operation back to normal.

And, importantly, it tests those assumptions before the real incident provides the test instead.

Recovery should not be something we improvise after restoration. It should be part of the design.

Listen carefully to what has actually been fixed.

The next time a major service disruption is followed by the reassuring announcement that the technical problem has been resolved, that statement may be completely true.

But it answers only one question.

Is the failed component working again?

The more important systems question may still remain.

How long until everything that depends upon it has recovered too?

The fault was fixed.
The system wasn't.

That's why resilience engineering must include the journey back.

What do you think?

Engineering gets better when ideas are challenged. Share your view, experience or disagreement.

Leave a comment

Comments are reviewed before publication. Your email address is optional and will never be published. How comments are handled.

This is the name that will appear if your comment is published.

Only used if we need to reply. It will never be published.

Share your thoughts about this story.