ENGINEERING STORY 005
How do you engineer
a system that doesn't exist?
Aviation behaves like a system, but nobody owns the whole thing. So how do we find the risks that live between its parts?
Aviation looks like a system.
To the passenger it certainly is. You book a ticket, arrive at an airport, board an aircraft and expect to arrive somewhere else.
But try drawing the system boundary.
Does it go around the airline? The airport? Air traffic control? Ground handling? Security? Fuel? Baggage? Border control? Meteorological services? Aircraft engineering?
The problem quickly becomes obvious.
There isn't really an aviation system in the conventional engineering sense. There is an aviation ecosystem made from independently owned systems, organisations and processes that have to cooperate sufficiently well to produce an apparently seamless service.
Nobody designed the whole thing. Nobody operates the whole thing. And nobody owns all of its risks.
THE QUESTION
How do you system-engineer something that isn't actually a system?
Before going any further, there is an obvious answer to get out of the way.
You don't try to model the whole thing.
Attempting to build a complete model of UK aviation would probably produce an enormous programme, require huge quantities of data, cross dozens of organisational boundaries and take so long that parts of the model would be obsolete before it was finished.
It would be an excellent way of proving that the problem was too difficult to solve.
So that isn't what I am suggesting.
START SMALL
Take one ordinary piece of aviation on one ordinary day.
Perhaps one aircraft rotation, a small group of flights, or one operational period at an airport.
Then ask a much more manageable question:
What was supposed to happen, what actually happened, and where was the difference absorbed?
That changes the problem completely.
We are no longer trying to model aviation. We are trying to trace operational margin across organisational boundaries.
If a small model tells us nothing useful, we stop.
But if it exposes an important dependency, disappearing margin or cross-boundary risk that wasn't visible when each organisation looked only within its own boundary, then we have demonstrated something worth pursuing.
Only then do we make the model bigger.
THE MODEL
Don't start with the technology.
The temptation would be to build an enormous technical architecture.
Don't.
Start with the outcome.
Move 180 passengers from London to Edinburgh at 10:00.
Now ask what has to be true.
An aircraft must be available. A suitably qualified crew must be present and within permitted hours. Passengers and baggage must be processed. Fuel and a stand must be available. Ground handling must be completed. The aircraft must be able to push back. Runway and airspace capacity must exist. The destination must be capable of receiving the aircraft.
Already we have stopped thinking about individual organisations and started thinking about the capability they collectively provide.
FOLLOW THE DEPENDENCY
Ask the same questions at every step.
What has to happen?
Who owns it?
What does it depend upon?
What constrains it?
How much capacity does it have?
What margin exists?
What happens elsewhere if that margin disappears?
That produces something more useful than an architecture diagram. It produces a dependency model.
And every time a dependency crosses an organisational boundary, mark it.
Because those boundaries may be where some of the most interesting risks live.
THE STRESS TEST
Let an ordinary day stress the model.
The first test should not be an emergency.
It should be a normal day's work.
Take what was planned and compare it with what actually happened.
Aircraft arrive a few minutes late. Turnarounds take slightly longer. Passenger loads vary. Weather changes. Crews accumulate duty time. Stands remain occupied longer than expected.
Nothing has necessarily failed.
That's simply aviation.
Yet thousands of small adjustments are being made to keep the operation working.
FOLLOW THE MARGIN
Where did reality depart from the plan?
Where was the variation absorbed?
Who absorbed it?
How much margin did doing so consume?
Was that margin subsequently restored?
And did solving one problem quietly reduce somebody else's ability to cope later?
This may tell us something far more interesting than injecting an artificial failure into the model.
It may tell us how aviation succeeds every day despite rarely operating exactly as planned.
And once we understand that, another question becomes possible.
How much disturbance is left for it to absorb?
TIME MATTERS
Margin isn't static.
An operation may have plenty at 06:00 and almost none at 08:45.
A delay absorbed easily in the morning might have very different consequences during the evening peak.
So the model needs to represent not just capacity, but capacity over time.
Demand → available capacity → remaining margin → margin consumed → margin restored
Perhaps a late aircraft is recovered through a quicker turnaround. Perhaps a stand remains occupied longer and another aircraft is moved elsewhere. Perhaps a crew's available duty time absorbs a delay. Perhaps another organisation quietly accommodates the variation.
Those are not necessarily failures. They may be the mechanisms keeping the whole ecosystem functioning.
But if the same margin is repeatedly being consumed during normal operations, we have found something worth understanding.
THE ORGANISATIONAL PROBLEM
Model the ecosystem, not the companies.
This is where the exercise could become uncomfortable.
A useful model may expose stretched processes, optimistic assumptions, dependencies that management hadn't appreciated, or margins that are routinely much smaller than believed.
Nobody particularly wants an engineering team arriving with a new method for discovering previously undocumented risks.
So that isn't how the work should be framed.
We are not auditing individual organisations. We are trying to understand the behaviour produced between them.
If a turnaround repeatedly has six minutes of recoverable margin, the initial conclusion isn't that an airline has inadequate resilience.
It is that this dependency repeatedly operates with approximately six minutes of recoverable margin.
Then investigate why.
The airline may contribute. So might the airport, ground handler, arrival pattern, stand allocation or another dependency entirely.
The six-minute margin may be an emergent property of all of them.
OWNERSHIP
Who owns a risk that exists only between organisations?
Once dependencies cross organisational boundaries, ownership becomes interesting.
Who controls the dependency? Who can observe it? Who can change it? Who consumes its margin? Who suffers the consequence when it is exhausted?
Those answers may not identify the same organisation.
One organisation could make a perfectly sensible efficiency improvement within its own boundary and unknowingly remove margin being relied upon somewhere else.
Organisation A changes something. Organisation B absorbs a little more variation. Organisation C loses some recovery capability. Organisation D eventually experiences the consequence.
A may not even know that D depends upon it.
That is precisely the sort of behaviour that conventional organisational risk registers struggle to expose.
GETTING PEOPLE TO PARTICIPATE
Don't ask organisations to reveal their weaknesses.
Some organisations may simply decline to participate.
The exercise might sound difficult, intrusive or potentially embarrassing.
So don't begin by asking:
Where are your weaknesses?
Ask:
How does a normal Tuesday work?
Much of the initial evidence can be operational rather than judgemental.
What was scheduled? What actually happened? Where were adjustments made? Where was additional capacity used? What recovered the operation?
The model can discover the interesting questions rather than requiring organisations to volunteer uncomfortable answers.
And there may be benefits beyond resilience.
One organisation may maintain expensive spare capacity because it cannot rely upon another organisation's process. Another may maintain additional margin because it cannot predict the first.
Better understanding of the dependency could potentially improve resilience and reduce unnecessary cost.
THE TEST
Know when to stop.
Perhaps the most important discipline is knowing when the idea has failed.
The first model should be deliberately small.
Its test is simple:
Can an ordinary day's operation reveal an important dependency, margin or risk that isn't obvious when each organisation looks only within its own boundary?
If not, stop.
We have learned something and haven't spent millions proving it.
But if it does, expand the model carefully.
Another operational thread. Another boundary. Another dependency.
THE POINT
We don't need to model the whole aviation ecosystem.
We need to discover whether looking across its boundaries tells us something that looking within them cannot.
Because perhaps the hardest risks in complex operations don't belong to any individual system.
They emerge from the relationships between them.
And perhaps the first step towards understanding them isn't building an enormous model of aviation.
It's taking one ordinary day's operation and asking:
What kept this working?
If the answer reveals something important that nobody could see before, then perhaps we haven't agreed to eat a blue whale after all.
We've simply found a way to eat it one bite at a time.
Sometimes the most important part of systems engineering is understanding the system that emerges between the systems we already know how to engineer.
DISCUSSION
What do you think?
Engineering gets better when ideas are challenged. Share your view, experience or disagreement.
ENGINEERING STORIES
Leave a comment
Comments are reviewed before publication. Your email address is optional and will never be published. How comments are handled.