ENGINEERING STORY 002

One was green.
One was red.
Both were right.

Apparently conflicting evidence does not always mean one source is wrong. Sometimes the evidence is supporting different engineering claims.

Two monitoring systems were observing the same operational service.

One was green.

The other was red.

Both were working correctly.

Two observations. Two different claims.

Imagine a service built from several components and dependencies.

Monitoring associated with one component is green. Its measurements are current, its checks have completed successfully and the evidence is valid.

Elsewhere, end-to-end monitoring of the service is red. The required outcome is not being achieved at the service boundary.

There has been no uncontrolled change. There is no defective monitor hiding conveniently in the story. Both observations mean exactly what their designers intended them to mean.

Both pieces of evidence are valid. They support different engineering claims.

The colour is not the claim.

The green evidence may support a claim about the operation of a component within a defined scope.

The red evidence may support a different claim about delivery of the end-to-end service.

Those two claims can coexist perfectly legitimately.

The contradiction appears only when we silently turn
“Component A is operating correctly”
into
“The service containing Component A must therefore be operating correctly”.

The evidence never supported that second claim.

We invented it.

Do not make one indication defeat the other.

When credible evidence appears to disagree, the instinctive response is often to resolve the discrepancy.

Which monitor is wrong? Which result should we trust? Which indication wins?

But if both observations are valid, choosing one prematurely can destroy useful information.

A better first question is: what engineering claim does each piece of evidence support?

The green observation tells us that one claim remains supported. The red observation tells us something different about another claim.

Together they constrain the investigation.

Apparently conflicting evidence may not be a problem with the evidence. It may be evidence about the problem.

Monitoring should collect evidence for a reason.

It is tempting to start with everything that can be measured and collect as much of it as possible.

For engineering assurance, the more useful direction is often the reverse.

Start with the engineering claim we need to support. Establish the argument behind that claim. Determine the evidence required by that argument. Then decide how that evidence should be obtained, including what should be automated.

Good monitoring does not collect the most data. It collects the right evidence for the decisions that matter.

Complex systems need more than engineering intuition.

On a small system, an experienced engineer may be able to reason through these relationships manually.

As complexity grows, that becomes increasingly difficult.

A claim may depend upon several other claims. Those may depend upon different evidence sources, assumptions, configurations or operating conditions. A single change in evidence can therefore have consequences some distance from where it was observed.

This is where an explicit engineering model becomes useful.

Evidence is associated with its subject and scope.

Claims identify their required support, assumptions and validity conditions.

Dependencies between claims are represented explicitly.

Now, when evidence changes, the impact can be assessed through those relationships.

Show the engineer where the reasoning changed.

The result should not be an automated diagnosis.

The tooling should not announce “Component B is faulty” unless sufficient evidence genuinely supports that conclusion.

Instead, it can expose something much more defensible:

This claim remains supported.

This claim is no longer sufficiently supported.

These dependent justifications are affected.

These areas require engineering reassessment.

The tool has not replaced the engineer. It has shown the engineer where the reasoning changed.

Unknown is an engineering result.

Sometimes the available evidence will not support a stronger conclusion.

That is not necessarily a failure of the model.

The correct result may simply be:
UNKNOWN. ASSESSMENT REQUIRED.

Unknown is uncomfortable, but inventing certainty is worse.

It tells the engineer that the available evidence is insufficient for the claim being considered. Additional evidence can then be sought deliberately rather than indiscriminately.

The objective is not to eliminate uncertainty by assertion. It is to reduce uncertainty with evidence.

Complex and safety-critical systems need an auditable argument.

If an assessment says that one claim remains supported while another requires reassessment, we need to be able to explain why.

Which evidence was considered? Which scope applied? Which assumptions were relied upon? Which dependencies were traversed? Which justification changed?

In other words, we need to show our working.

Formal methods can help.

An engineering model can be implemented using languages and verification tools such as Ada and SPARK. Important properties of the reasoning mechanism can then be formally specified and checked.

The specification can become something we prove.

Consider this deliberately small SPARK contract:

function Permissible_Claim
  (Claim    : Claim_Strength;
   Scope    : Scope_Status;
   Validity : Validity_Status) return Boolean
with
  Post =>
    (if Permissible_Claim'Result then
       Scope = In_Scope
       and Validity = Valid);

You do not need to know Ada to understand the important part.

The contract says that if this function reports a claim as permissible, the required scope must be in scope and its validity established.

SPARK can then be used to prove that the implementation satisfies that formally specified property.

Formal proof does not prove reality.

We have not proved that the physical system is healthy.

We have not proved that an observation about the real world is factually correct.

And we have not proved that our engineering model perfectly represents reality.

What formal verification can provide is evidence that the implemented reasoning obeys the rules we have specified.

The engineering model still has to be justified against the real system. The implementation has to be verified against the model.

Formal verification does not remove engineering judgement. It gives us stronger evidence about the machinery we are asking engineers to rely upon.

The assessment informs the decision. It does not own it.

Engineering analysis can establish what the available evidence currently supports and what operating options remain technically permissible.

That does not necessarily make the operational decision.

Operations may need to consider wider factors such as service demand, alternative resources, operational risk, staffing, degraded modes and the consequences of taking equipment out of service.

The engineering assessment informs that decision. It does not silently take ownership of it.

Neither indication needs to defeat the other.

One is green.

One is red.

Nothing needs to be averaged. Nothing needs to be discarded.

The green evidence remains valid support for its defined engineering claim. The red evidence remains valid support for another.

Once the claims, scopes and dependencies are made explicit, the apparent contradiction begins to disappear.

The objective of a good evidence model is not to make every indicator agree. It is to make clear what the available evidence legitimately allows us to say.

Green and red can both be right.

The engineering begins with understanding why.