Automated testing strategy for consequential systems
Practice · Automated Test Strategy
A green CI run is a signal, not verification evidence: a layered automated test strategy (unit, integration, system, acceptance) produces evidence only when 100% of requirements trace through the requirements traceability matrix to a passing test under IEC 62304, and a flaky test is treated as a defect in the verification chain.
When does automated testing produce verification evidence?
A green CI run is a signal, not verification evidence. A layered test strategy produces evidence only when 100% of requirements trace through the traceability matrix to a passing test under IEC 62304.
A flaky test is a defect in the verification chain, not an annoyance to quarantine and forget.
A test pyramid that produces evidence, not just a green build
"All tests pass" answers one question: did the code behave as the tests expected, right now, in this environment? It does not answer whether the tests covered the right things, whether a passing test is traceable to a specific requirement, or whether an intermittently failing test is a real defect being ignored. A layered strategy exists to make those three questions answerable.
The four layers, and what each one actually verifies
| Layer | What it verifies | Evidence it produces |
|---|---|---|
| Unit | One function or module's logic, in isolation from its real dependencies | Pass/fail plus a coverage delta against the specific lines changed |
| Integration | Two or more real components talking to each other — a database, an HTTP boundary — not a mock standing in for one | Pass/fail against an actual dependency, catching contract drift a mock can't |
| System / end-to-end | The deployed shape, exercised from outside as a black box | Behavior matches the specification from the caller's perspective, not the implementer's |
| Acceptance | Does the system satisfy the documented user need | Validation evidence, not verification — see Verification vs. Validation |
A flaky test is a verification-integrity problem
A test that intermittently fails is one of two things: it's catching a real, intermittent defect — a race condition, an unhandled timeout — or it's poorly written, asserting on something nondeterministic that was never actually part of the contract. Quarantining it without finding out which one is true means a requirement that was supposedly verified now has no reliably passing test behind it, which breaks the traceability chain silently — the traceability matrix still shows a link, but the link no longer means what it claims to mean.
From a test run to traceable evidence
Under IEC 62304, a test only becomes MedTech verification evidence once it has a stable identifier a traceability matrix can reference, and once its pass/fail result from a specific run is captured immutably in a test evidence log rather than overwritten by the next run. The first is a test-authoring discipline; the second is a pipeline-design discipline — see CI/CD Pipelines as Audit Evidence for what a pipeline needs to capture to make that true.
Engineering reference only. Specific coverage thresholds and layer proportions are a project-level risk decision, not a universal rule this page prescribes.
Provenance & review state
- Last reviewed
- Sources
-
- ISO/IEC/IEEE 29119-1:2022, Software and systems engineering — Software testing — International Organization for Standardization
- IEC 62304:2006+AMD1:2015 — International Electrotechnical Commission
- Ingested from
-