Design for Testability: Finding Diagnostic Gaps Before Product Launch
Distinguish fault detection from isolation, find ambiguous diagnostic paths, and plan practical tests and access before a product reaches the field.
- Author
- Remy InfoSource Pte Ltd
- Published
- Reading time
- 5 min
- Testability
- Product Development
Detecting a fault is not the same as isolating it
A product can detect that a required function has failed while providing very little evidence about the responsible component. An alarm that reports low output is a detection capability. A set of tests that distinguishes the supply path from the drive path and the output assembly is an isolation capability.
Design for testability means considering that distinction while the product can still be changed. The goal is to make the required diagnostic evidence available, interpretable, and practical for the people who will use it.
Leaving this work until field service begins can expose a difficult tradeoff: replace several candidates, dismantle the equipment to reach a measurement point, or tolerate a lengthy ambiguous diagnosis.
Start with the diagnostic outcome you need
Before adding test points or more self-tests, define what the service organization needs to distinguish.
- Which functions and failure modes are in scope?
- Must isolation reach a module, a replaceable assembly, or an individual component?
- Which faults require an immediate safe-state response?
- What operating conditions must the tests support?
- What tools, access, and training will be available in the field?
- How much diagnostic time and disruption is acceptable?
An assembly-level answer may be sufficient for one product and inadequate for another. Coverage results are only meaningful when the required resolution and the modeled fault population are clear.
FMECA and failure-analysis work can help identify the faults and effects to consider. It does not remove the need to define field-executable tests.
Worked example: two faults with the same visible effect
Illustrative example, not a measured product result. A simplified pump system reports a missing response whenever either its upstream power path or its switched drive path fails.
If the only available test is the final flow indication, both failures produce the same observed result. The product can detect the problem, but cannot distinguish those two candidates through that evidence alone.
Now consider an approved measurement point between the power path and the drive path. Under defined operating conditions, it can establish whether the required upstream supply is available.
| Available evidence | What it can establish in this simplified example |
|---|---|
| Final flow indication only | A problem exists, but the upstream and switched-path faults remain ambiguous. |
| Final indication plus a valid upstream-supply test | The evidence can distinguish the modeled supply-path fault from a downstream candidate. |
| Additional valid input and response tests | The remaining drive-path and assembly candidates may be separated, subject to the model and test assumptions. |
The additional test is useful because it separates candidate explanations, not merely because it adds another reading. Its value disappears if it cannot be accessed, if the measurement is unreliable, or if the required operating state cannot safely be established.
Combine built-in and external tests deliberately
Built-in tests can make evidence available without special service equipment. However, a self-test may share a sensor, power supply, or communication path with the function it is assessing. That shared dependency can limit what its result proves.
External tests can provide independent evidence or reach a finer diagnostic resolution, but may require fixtures, access panels, trained personnel, and additional downtime.
For each test, evaluate:
- Discrimination: which candidate failure modes produce different outcomes?
- Prerequisites: what must already be working for the result to be interpretable?
- Independence: could a shared dependency make the observation misleading?
- Practicality: can the intended user perform it on the installed product?
- Safety: is it permitted under the approved servicing procedure?
- Disturbance: could reaching the test point alter or temporarily remove the fault?
Test-point count alone is not a testability measure. A small set of well-defined tests can be more useful than many measurements that all reflect the same downstream symptom.
Validate on representative configurations and faults
A model review can reveal missing relationships and ambiguous candidate groups before hardware testing. Prototype testing can then check whether the predicted evidence is obtainable and correct.
Use a controlled validation matrix: configuration, operating state, modeled fault, expected observations, executed tests, remaining candidates, and actual result. Include healthy-system cases as well as safely introduced or simulated failure cases. Representative intermittent and multiple-fault scenarios should be considered where required by the intended use.
Fault injection must follow approved engineering and safety methods; not every failure can or should be introduced physically.
The NASA Systems Engineering Handbook distinguishes verification from validation. Both matter here: a test can satisfy its technical specification yet still be impractical for the intended service environment.
Define coverage without hiding the denominator
If a team reports a detection or isolation percentage, state exactly what was counted. A percentage based on modeled failure modes is not automatically a percentage of all real-world failures. Weighted and unweighted methods also need to be distinguished.
For isolation, specify the permitted candidate-group size or the required replaceable-item resolution. Report excluded conditions, unsupported faults, and unavailable tests alongside the number. Do not imply a universal coverage calculation or an improvement percentage without the relevant method and evidence.
A practical pre-launch review
Before accepting the diagnostic design, check that:
- The supported model matches the release hardware, firmware, and configuration.
- Credible in-scope faults have defined expected observations.
- Important ambiguous groups have been reviewed, not hidden in an average.
- Test access, tools, and prerequisites match the actual service environment.
- Inconclusive or conflicting evidence leads to a documented next action.
- Repair verification is defined separately from initial fault detection.
- Ownership and regression checks are in place for later product changes.
Resolving an ambiguity may mean adding a measurement, revising an algorithm, improving a connector or access point, or accepting an assembly-level repair strategy. Make that decision explicitly while its costs and operational consequences can still be assessed.
For the reasoning behind candidate narrowing, read what model-based diagnostics is. See iTech's technology and validation approach or request a discussion of your testability requirements.
