VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Common-Cause Failure Works | Why Redundant Parts Can Still Fail Together

Two backups are not two independent protections if one event can remove them both.

Redundancy is one of the oldest ways to improve reliability: add another path, machine, sensor, supplier or team. But the value of redundancy depends on independence.

A common-cause failure occurs when multiple supposedly separate components fail because they share one cause: the same power source, software defect, environmental hazard, maintenance error, credential, supplier, network path or design assumption.

This is a specialist branch beneath How Redundancy Works. The broad owner already explains alternate paths and common-cause risk; this article drills into the mechanism by which independence quietly disappears.


Redundancy Counts Components. Reliability Counts Causes.

Imagine two pumps. If both are fed by one electrical breaker, the breaker is a common cause. If both use the same firmware, one software defect may disable both. If both sit in one flooded room, geography becomes the common cause.

What looked like “two systems” is therefore one system with two visible components and one hidden shared dependency.

Common Causes Hide at Different Layers

  • Physical: fire, flood, heat, impact, vibration.
  • Power: shared breaker, feeder, fuel supply or battery system.
  • Network: one carrier, switch, DNS dependency or route.
  • Software: identical defect, configuration or update.
  • Human: one operator command, one maintenance procedure, one mistaken instruction.
  • Supply chain: identical component batch or sole supplier.
  • Governance: one permission, account or policy can disable all paths.

Diversity Is Different From Duplication

Duplicating the same design improves resilience against random individual component failure. Using diverse implementations can improve resilience against shared design faults.

But diversity brings cost: different spare parts, training, interfaces, testing and failure modes. It can also create new integration errors.

The wider mechanism belongs to How Diversity Works. Common-cause analysis asks when diversity is worth its complexity because duplicated sameness leaves one catastrophic dependency.

Maintenance Can Defeat Independence

Redundant systems are often most vulnerable during maintenance.

If both paths are updated at the same time, the maintenance window itself creates a common cause. If one technician applies the same wrong configuration to both, duplication amplifies the mistake.

Safe practice may stagger upgrades, preserve a known-good version, separate approval authority or validate one path before touching the other.

Worked Example: Data Service

Two server clusters run in different buildings. Both depend on the same cloud identity provider. The identity provider fails.

Compute redundancy survives; service still becomes unusable because authentication is common infrastructure.

The lesson is not “duplicate everything.” It is to identify the dependencies whose failure dominates end-to-end service.

Worked Example: Railway

Two trains have independent onboard systems but both rely on one section of traction power or one control authority. A local common cause can therefore remove several nominally independent vehicles at once.

Railway resilience needs fault domains that cross physical, electrical, signalling and operational layers, not only spare vehicles.

Worked Example: Finance

Two payment channels may appear redundant but use the same core ledger, same network route or same fraud decision service. If that shared dependency fails, customer-facing alternatives disappear together.

The finance owner remains How Finance Works; common-cause thinking reveals whether operational continuity is genuinely diversified.

A Careful Analogy: Education

A learner may use several study resources that all explain a concept through the same mistaken model. The apparent redundancy of examples does not protect against the shared misconception.

Independent explanation, varied representation and transfer testing can expose the common cause. The analogy is bounded but useful: several copies of the same error do not become independent evidence.

A Careful Analogy: Institutions

Several departments may appear independent while relying on one approval chain, one procurement vendor or one database. Organisational charts can hide operational common causes.

Resilience audits should therefore trace shared dependencies across formal boundaries.

A Common-Cause Checklist

  1. List the redundant paths.
  2. Trace every shared physical and digital dependency.
  3. Trace shared maintenance and operator actions.
  4. Trace shared suppliers and component batches.
  5. Trace shared credentials, policies and authority.
  6. Ask which one event could disable several paths together.
  7. Introduce separation or diversity where consequence justifies it.
  8. Test the common-cause scenario explicitly.

The CivDJ Rotation

  • Forward: shared dependency fails → several redundant paths fail together → expected resilience disappears.
  • Backward: start from a multi-path outage and find the earliest dependency common to all affected paths.
  • Rotate: compare architecture, procurement, operations, maintenance and receiver views of “independence.”

Common-cause failure is the moment redundancy reveals that it was only duplicated appearance resting on one shared weakness.

Continue through How Redundancy Works, How Fault Domains Work and the master How X Works hub. Next: latent failures — weaknesses that already exist but remain invisible until another condition exposes them.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading