Engineering decision analysis is the disciplined comparison of feasible alternatives, their consequences, the evidence behind them and the preferences that make one consequence more desirable than another. Its purpose is not to make a spreadsheet choose for us. It is to make the reasons for a choice visible enough to examine, challenge and revise.
At the centre of this guide is a fictional library deciding how to improve its book-return service. The equipment, observations, costs, scores, probabilities and conversations are invented teaching examples, not product specifications, quotations, measured library results or procurement advice. Adrian and Jo explore the problem with Ben, Aisha, Ryan, Mira, Clara and Ethan, the recurring fictional learners in the eduKate library. Their questions let us follow a decision from its first attractive idea to its eventual consequences.
The central question is simple: what should we choose when the fastest option is not the easiest to maintain, the cheapest purchase is not the cheapest service, and the most promising design depends on something we do not yet know?
This is an educational guide to reasoning, not a substitute for qualified engineering, safety assessment or the approvals required for a real installation. The numerical limits used in examples are deliberately hypothetical. A design that is inadmissible for safety, legal or technical reasons does not become acceptable because it receives a high preference score.
Choose a reading route
Understand the method: Sections 1–7 establish the decision. Follow the calculations: Sections 8–19 cover values, costs, uncertainty and information. Prepare a recommendation: Sections 20–25 connect the model to people, suppliers, implementation and learning. Practise and challenge: Sections 26–28 provide exercises, difficult questions and synthesis.
1–7: Frame the choice
- 1. The book on the counter
- 2. Give the decision a verb, a boundary and a date
- 3. Who may choose, and who has to live with the choice?
- 4. Some conditions are gates, not preferences
- 5. Generate alternatives that differ in mechanism, not just appearance
- 6. Keep observations, predictions and preferences in different drawers
- 7. Measure the service, not just the machine
8–14: Compare consequences
- 8. Convert performance into value without hiding the conversion
- 9. Weights describe trade-offs, not popularity
- 10. Find the point at which the answer changes
- 11. The quiet ways a matrix can choose its own winner
- 12. The purchase price is only the first chapter of cost
- 13. Probability weights are not preference weights
- 14. Two backups can share one problem
15–21: Choose under uncertainty
- 15. Compare futures without pretending to know their probabilities
- 16. Ask what information is worth before asking for more of it
- 17. Design a pilot that can actually answer the question
- 18. A staged decision is a different design
- 19. When the choice is a portfolio, ranking is not enough
- 20. Follow the work through people, queues and maintenance
- 21. The same discipline works across physical and digital choices
22–28: Act, learn and practise
- 22. Supplier claims belong in an evidence process
- 23. Make the recommendation reconstructable
- 24. Run a meeting that can hear an inconvenient result
- 25. Judge the decision after the result, but not only by the result
- 26. Workshop: rebuild the decision instead of repeating the answer
- 27. Difficult questions that a responsible analysis must answer
- 28. Return to the book on the counter
1. The book on the counter
At five minutes to closing, a book lies on the returns counter. A reader has brought it back. From the reader’s point of view, the job is almost over. From the library’s point of view, the book has only begun another journey. Its identity must be recognised, the loan record updated, any reservation noticed, the item routed and the physical book placed where someone can find it. A successful return is not merely a book passing through a slot. It is the restoration of an item to a usable place in a shared collection.
In our fictional library, that journey has become slow. Trolleys stand between the return point and the sorting shelves. Staff interrupt other work to clear them. Readers occasionally ask whether an item has really been returned because the record has not changed. The first proposal is immediate: buy a faster sorter. A brochure appears. Its headline number is impressive. Compared with the existing arrangement, the machine seems to offer an obvious improvement.
Ben asks the question that makes the discussion interesting. “If the machine can move the books faster, why is there still a decision?” Jo places the brochure beside a sketch of the room. The advertised rate describes one part of the journey, she explains. It does not yet describe the complete service. Someone must feed awkward items, clear exceptions, empty bins, manage accounts and maintain the equipment. A machine can be fast while the service around it remains slow. Moving a queue from the counter to the next room does not eliminate the queue.
That distinction changes the object of the decision. The group is no longer choosing a machine in isolation. It is choosing a way of delivering a return service. Equipment is one part of that way. Software, staffing, layout, maintenance and the treatment of unusual items are other parts. The question becomes harder, but also more faithful to the problem people actually experience. A narrow answer might be easy to defend with a single number while failing to solve the original difficulty.
Three possible arrangements begin to emerge. One uses a single large automated sorter. Another uses two smaller modules that can be serviced separately. A third automates routine returns but retains a staffed exception route and a simpler physical layout. None is a cartoon villain. The large system has strong processing potential. The modular arrangement offers flexibility. The assisted arrangement may cope well with unusual objects and changing work, but it depends on having people available when needed. Each has a reason to exist.
Adrian asks everyone to imagine a different afternoon. The large sorter is unavailable. A module has been removed for servicing. A staff member is absent. A software connection is interrupted. Now ask which arrangement still delivers a tolerable service. The preferred option may change with the question. That does not make analysis useless. It shows why the question must be stated before the answer is celebrated. Engineering decisions are often difficult because the alternatives are good at different things, not because nobody has found the one obviously perfect design.
NASA’s decision-analysis guidance treats the task as support for a decision authority facing alternatives and incomplete knowledge. It calls attention to objectives, criteria, uncertainty and documented reasoning rather than treating a numerical ranking as self-authorising. That is a useful starting point, but the worked cases in this guide develop their own figures and arguments. They are not NASA case studies. Source 1
The book on the counter gives us a concrete test for every abstraction that follows. Does the chosen arrangement help the reader complete the return, help the staff restore the book to circulation and help the organisation sustain that service? A score is useful only insofar as it represents those consequences. The point of calculation is not to disappear into calculation. It is to come back to the ordinary job with a clearer understanding of what a design will ask of the world.
A long analysis can lose this connection surprisingly easily. It acquires criteria because they are easy to measure, alternatives because vendors happened to offer them, and weights because a template requires percentages. Before long, the original problem becomes a preface nobody revisits. The remedy is to keep a plain-language statement beside the technical material. In this case: reduce avoidable return delays without creating an unmanageable burden elsewhere. That statement cannot replace requirements, but it can expose a beautifully calculated answer to the wrong question.
Engineering decision analysis begins here, before the matrix and before the meeting. It begins when someone notices that the object being selected and the outcome being sought are not necessarily the same thing. Once that difference is visible, alternatives can become more imaginative, evidence can become more relevant, and disagreements can become more precise. The next task is to write the decision in a form that actually permits a choice.
2. Give the decision a verb, a boundary and a date
“Improve the library” is a purpose, not a decision. “Modernise the return system” is somewhat narrower, but it still conceals what authority is being exercised. Are we selecting a concept for further investigation, authorising a pilot, choosing a supplier, committing to a building alteration or releasing the service for public use? Those are different decisions. They require different evidence because they expose the organisation to different consequences. The same concept can be promising enough to prototype and nowhere near ready to install.
A useful decision statement has a verb. Select one concept for detailed evaluation. Authorise a bounded trial. Choose between two qualified replacement arrangements. Retain the present system for another defined period under specified restrictions. The verb prevents the discussion from drifting between exploration and commitment. It also makes the output inspectable. At the end of the analysis, someone should be able to identify what changed in the organisation’s permission to act.
For the fictional library, the immediate decision is to select an arrangement for a limited, reversible pilot, not to certify a complete automated installation. That boundary matters. The group can compare anticipated service performance while acknowledging that detailed safety and installation design remain separate work. It must not turn a preference for an arrangement into permission to put an unassessed machine in front of the public. A recommendation can be strong within its scope and still have very clear limits.
The statement also needs a system boundary. Start with the reader placing a return and end with the item correctly recorded and routed. Include the exception path because exceptions are part of the service. Include the staff who empty bins because otherwise the machine’s output is treated as someone else’s unexplained problem. Include the connection to the circulation system because an unrecorded physical transfer does not complete the reader’s transaction. Exclude unrelated collection-acquisition decisions so the study does not expand into a redesign of the entire library.
This boundary is not a claim that excluded things are unimportant. It is a statement about which relationships the present decision must represent directly. Important external dependencies should still be recorded. Building power, network support and future collection growth might remain outside the pilot’s control, but assumptions about them must not remain outside the reasoning. The distinction between an external dependency and an irrelevant detail is one of the first serious judgements in a trade study.
Time enters twice. There is a deadline for making the decision and a horizon over which its consequences matter. The deadline may be driven by a lease, a procurement window or the need to reduce disruption before a busy period. The horizon may be several years of operation. Confusing them makes short-term urgency govern long-term value. A design that can arrive this month may still be a poor eight-year choice; a design that would be excellent in eight years may not solve an immediate operational problem.
Mira proposes an apparently sensible compromise: choose whatever arrives first and upgrade later. That can be a real option, but only when “later” is engineered into the present choice. Does the temporary arrangement leave space for the upgrade? Can data and interfaces migrate? Is the interim purchase reusable, recoverable or disposable at an understood cost? Without those conditions, the temporary choice can become the most permanent decision in the project. The analysis must compare an actual staged plan, not an optimistic sentence about future flexibility.
A decision statement should also identify what remains unchanged. The pilot might not alter the building’s public entrance, the existing catalogue rules or the staff’s authority over disputed returns. Those exclusions prevent a local change from quietly acquiring a wider mandate. They also help measure success. When too many parts of the service change at once, it becomes difficult to tell which alteration produced an improvement or a new problem.
There is another useful question: what would happen if no new purchase were made? The answer should not be an artificially disastrous baseline constructed to make investment look inevitable. Describe the likely path of the existing arrangement with realistic maintenance, staffing and small improvements. If a revised trolley route resolves most of the delay, the capital decision changes. If the backlog continues to grow even after inexpensive improvements, that evidence strengthens the case for a more substantial intervention. Honest baselines can support investment as well as challenge it.
The resulting statement might read: select one return-service concept for a supervised pilot, comparing end-to-end performance, continuity, maintainability, adaptability and whole-life cost over a common planning horizon, while keeping existing public-service obligations intact. The pilot must be reversible and must not be interpreted as permission for an unassessed production installation. That statement is longer than “buy a sorter”, but it creates a decision that can be analysed without pretending the unresolved engineering has vanished.
3. Who may choose, and who has to live with the choice?
A technical decision has at least two kinds of ownership. Someone is authorised to make it, and someone experiences its consequences. Those groups may overlap, but they are rarely identical. The person who approves a purchase may never clear a jam. The person who maintains the equipment may not attend the selection meeting. The reader who uses an accessible return point may not appear in the cost model. A decision process that sees only the people in the room can produce a very orderly form of exclusion.
The first practical step is to distinguish decision authority from technical expertise. An engineer can estimate whether a layout is feasible without having authority to set the library’s service priorities. A manager can approve expenditure without being able to certify mechanical safeguards. A user can identify a serious access difficulty without being responsible for designing the solution. Each contribution matters, but it matters in a different way. Blending them into a single undifferentiated vote conceals the kind of judgement being made.
The UK Government Analysis Function’s introductory MCDA guide distinguishes people whose preferences are represented in the decision model from subject-matter experts who provide evidence about performance. It also stresses proportionate competence and review. That distinction is useful beyond government, provided it is adapted to the actual organisation rather than treated as a universal organisational chart. Source 2
In the fictional discussion, Aisha asks whether all stakeholders should simply receive equal weight. Equal treatment sounds fair, but a numerical average of preferences does not settle questions of authority, rights or consequence. Suppose one group values shorter queues and another objects to losing a necessary accessible route. The latter concern may be a mandatory design obligation rather than an optional preference to be averaged. Fairness can require different kinds of treatment for different kinds of claims.
A stakeholder map should therefore identify the task each person performs, the benefit they seek, the burden they may receive and the decisions they legitimately influence. A return-desk worker might contribute evidence about exceptions and handling time. A maintenance specialist might identify the need to remove a module without closing the entire room. A reader with a mobility limitation might expose a route the layout team overlooked. The project sponsor might decide how much inconvenience during installation is acceptable, subject to obligations the sponsor cannot waive alone.
The map should include people downstream of the visible equipment. If faster processing creates a surge of unsorted trolleys elsewhere, another team may inherit the work. If a new arrangement depends on constant technical support, the support team should not discover that obligation after release. Such omissions often look like technical surprises even though the missing information was available in another person’s ordinary working day. Asking the right people early can be more valuable than refining a weak model later.
Disagreement should be classified before it is resolved. One person may believe the modular arrangement will fail more often; another may believe it will recover faster. That is a disagreement about predicted performance and should prompt evidence gathering. One may value maximum processing rate; another may value uninterrupted service. That is a preference difference and requires explicit discussion of trade-offs. A third may point to a non-negotiable access obligation. That is a constraint question. Treating all three as “different opinions” deprives the meeting of its most useful structure.
The analyst’s role is not to smuggle preferences into apparently objective numbers. A maintenance score can contain a judgement about how much independence from specialist support is worth. A reliability forecast can contain uncertain assumptions about operating conditions. Both can be legitimate, but they should be labelled. The model becomes more trustworthy when readers can separate what was observed, what was estimated and what was valued. It becomes less trustworthy when every entry is presented with the same confident decimal precision.
No stakeholder map is perfect. The practical objective is not to invite every conceivable person to every meeting. It is to identify perspectives whose omission could materially change the problem, the admissible options or the preferred choice. That judgement should be revisited when a pilot exposes a previously unseen burden. Discovering another stakeholder is not proof that the original team was careless; refusing to acknowledge the discovery would be a more serious problem.
The final authority should receive more than a winning label. It should see where stakeholders agree, where disagreement remains, which obligations limit discretion and which consequences accompany the recommendation. A decision can then be made honestly even when consensus is incomplete. What must not happen is for a clean final score to erase a significant disagreement that the organisation still needs to own. Good analysis does not eliminate responsibility. It gives responsibility a more accurate description of the choice being made.
4. Some conditions are gates, not preferences
A selection matrix is often drawn as though every desirable property can compensate for every undesirable one. The fastest option earns points. The quietest earns points. The cheapest earns points. Add enough points and the option wins. This arrangement is mathematically convenient, but it can be conceptually wrong. A machine that cannot be installed within the available space does not become installable because it is inexpensive. A design missing required safety evidence does not become acceptable because its processing rate is exceptional.
Start by separating feasibility conditions from preferences among feasible choices. A feasibility condition defines something that must hold for an option to remain under consideration at the present commitment level. Examples might include physical fit, compatibility with a required interface, an achievable installation window or an independently established safety obligation. The precise conditions depend on the real project. The numerical thresholds in this guide are not design guidance for a library installation.
An option can pass a gate, fail it or remain unresolved. The third state matters. “No evidence of failure” is not the same as “evidence of compliance”. If the supplier has not provided the information needed to assess a critical interface, the status is unknown. It should not receive half the points available for compatibility and proceed into the ranking. A weighted average is a poor place to hide a missing prerequisite.
For the fictional pilot, the group uses a simple eligibility register. Each arrangement must fit the pilot area, provide an agreed route for handling exceptions, permit an authorised assessment of public-use hazards, and allow recovery to the existing service if the trial is stopped. An option that cannot currently satisfy one of these conditions is not necessarily worthless. It may be sent back for redesign or considered for another site. But it is not ranked as though it were ready for this pilot.
NASA’s decision-analysis discussion explicitly distinguishes mandatory criteria from enhancing criteria and says that an option failing a mandatory condition should be removed from that comparison. The important lesson is the separation, not the adoption of any particular agency’s project rules. Source 1
Hard limits can be misused in the opposite direction. A team may describe a familiar preference as a non-negotiable constraint to eliminate competitors. “It must use our existing brand” may be a true compatibility requirement, or it may be a habit. “It must be fully automated” may express a legitimate operational need, or it may exclude a better assisted arrangement before anyone examines the service. Every gate should have a source and a rationale. Otherwise the shortlist can be engineered to produce the desired winner before the scoring begins.
The source does not need to be elaborate for every condition. A measured room width and an approved clearance requirement may justify a spatial constraint. A signed service agreement may define an installation window. A safety specialist may identify a requirement that the selection team is not competent to reinterpret. What matters is that another person can see why the condition exists and who may change it. A constraint without that traceability can become either dangerously weak or needlessly restrictive.
Ryan suggests assigning a very large negative score to failure of a critical gate. That can work inside some carefully designed optimisation models, but in an ordinary spreadsheet it creates avoidable risk. A mistaken weight, an additional criterion or an unusually high score elsewhere might offset the penalty. A separate admissibility check is clearer. The system first asks whether an option is eligible. Only eligible options enter the preference comparison. The architecture of the analysis then reflects the architecture of the decision.
Threshold uncertainty should also be visible. A design might appear to fit within a limit only because the estimate ignores tolerance or uncertainty. If the margin is smaller than the uncertainty in the estimate, “pass” may be too strong. The correct response might be another measurement, a revised design, a larger reserve or a conditional decision at a lower commitment level. The analysis should not confuse a nominally acceptable number with adequately established feasibility.
Finally, a gate does not authorise itself. The selection team may recommend that a concept remains eligible for investigation, but a specialist must still provide the approvals required for actual use. A preference study cannot manufacture regulatory authority. This distinction lets analysis remain useful without overclaiming. It can narrow choices, direct learning and reveal design weaknesses while respecting the separate work needed to make an engineered system fit for public operation.
5. Generate alternatives that differ in mechanism, not just appearance
When a team compares three nearly identical vendor offers, it may feel that it has explored alternatives. Sometimes it has. In other cases, it has explored three commercial variations of one unexamined idea. All may assume the same layout, the same staffing model, the same dependency on a central server and the same treatment of exceptions. The alternatives differ on the surface while sharing the assumptions most likely to determine success.
A useful alternative changes how the required outcome is produced. The fictional library’s large sorter concentrates processing in one installation. The modular arrangement distributes some capacity and maintenance exposure. The assisted arrangement allocates routine and unusual work differently between equipment and people. These are not merely different prices attached to the same picture. They distribute responsibility, failure and adaptation differently. That is why comparing them can teach the team something about the problem itself.
Alternative generation should include a credible do-minimum option. This might retain the existing equipment while changing the trolley route, improving exception recording and scheduling additional clearing during predictable peaks. It may not be sufficient, but it deserves a fair description. A weak baseline that receives no maintenance or process improvement while the new alternatives receive every conceivable enhancement is not a fair comparator. The baseline should represent what the organisation would realistically do without the major change.
It is also useful to consider a staged option. Rather than choosing between immediate full automation and no automation, the library might test one module, preserve a manual route and prepare the interface for later expansion. That is a distinct policy over time, not merely a smaller machine. Its value depends on what can be learned before the next commitment and whether the later expansion remains feasible. Staging must include transition costs and temporary limitations, not just the attractive promise of learning first.
A common trap is to define one alternative at much greater maturity than the others. The favoured option has a detailed layout, staffing plan and negotiated maintenance arrangement. Its rivals are rough sketches with obvious unresolved questions. The detailed option then wins partly because the team has invested more effort in making it coherent. Comparisons should either bring alternatives to a reasonably comparable level of definition or record the maturity difference explicitly as a reason for further work.
Comparable does not mean identical. Forcing each alternative to have the same number of machines, the same layout or the same automation level can erase the very differences the study needs to evaluate. The shared basis should be the required service and the planning assumptions, not an arbitrary component count. A manual exception route should be credited where it genuinely helps and charged where it genuinely consumes staff capacity. Neither romance about human flexibility nor enthusiasm about automation should exempt an arrangement from a complete description.
MIT’s teaching on concept selection and tradespace exploration emphasises looking at the relationships among competing designs rather than declaring one isolated concept optimal. The mathematical language is useful because it reminds us that the “best” option depends on what is being compared and on which criteria remain in view. The library example here is an original application, not a reproduction of the MIT lecture’s case material. Source 3
The group writes a one-page description for each concept. It includes the service path, principal dependencies, expected staffing, installation assumptions, maintenance approach, treatment of exceptions and possible future changes. It also includes a sentence about what the option does not solve. The large sorter does not eliminate the need to empty its outputs. The modular arrangement does not automatically remove a shared software dependency. The assisted arrangement does not make staff available by wishful thinking. These sentences are not attacks; they are the beginnings of an honest comparison.
An alternative should be removed for a stated reason. It may fail a mandatory condition, be dominated by another sufficiently evidenced option, or fall outside the decision scope. Keep a short record of the rejection. Otherwise a later participant may waste time reviving an already examined idea, or the team may forget that an option was excluded because of an assumption that has since changed. Rejection history is part of decision memory, not clutter to be discarded once a winner appears.
Clara asks whether this process will produce too many alternatives. It can. The answer is not to stop thinking early, but to separate broad mechanisms from minor variants. First compare fundamentally different ways of delivering the service. Then refine the promising families. There is little value in evaluating every possible bin colour while the team has not decided whether continuity during maintenance matters more than maximum processing rate. Good alternative generation expands the useful space of choice, then narrows it for defensible reasons.
6. Keep observations, predictions and preferences in different drawers
The first evidence table is often where an analysis begins to look more certain than it is. A measured dimension, a supplier’s advertised capacity, an engineer’s forecast of downtime and a manager’s preference for local maintenance can all appear as numbers in neighbouring cells. Their proximity makes them look comparable in authority. They are not. Before combining values, identify what kind of claim each cell contains.
An observation describes something recorded under particular conditions. An estimate infers a quantity from incomplete evidence. A prediction describes what the team expects to happen in a future situation. A preference describes how much the decision-maker values a consequence. An assumption temporarily supplies a condition needed to reason further. None of these categories is inherently illegitimate. The problem begins when a forecast is presented as a measurement, or a preference is presented as a physical fact.
Suppose a supplier reports that a module processes a certain number of items per hour in a demonstration. That is evidence about the demonstrated configuration and workload. It may support a prediction about the library, but the prediction requires additional reasoning. Were the books similar? Were exceptions included? Was the software connection real? Did staff empty the bins? The advertised number does not become a local service forecast merely by being copied into a spreadsheet.
Conversely, a local observation can be too narrow. Watching one quiet afternoon tells the team something about that afternoon. It does not necessarily describe seasonal peaks, unusual item mixtures or operation during maintenance. The evidence should be accompanied by its conditions. This does not diminish the observation; it protects its meaning. A modest, well-bounded observation can be more useful than a broad claim whose supporting conditions are unknown.
For each material estimate, record the source, configuration, conditions, method and limitations. A short evidence note might say that the processing estimate comes from a supervised demonstration with a specified item mix, that the estimate excludes some exception types, and that a pilot must test those exceptions before release. Another note might say that downtime is a planning estimate rather than a measured annual rate. These notes prevent numbers from acquiring unearned authority as they move from a working sheet into a presentation.
Ethan asks whether uncertain numbers should simply be left out. Leaving them out can be worse if the corresponding consequence is important. The better response is to represent the uncertainty and investigate whether it can change the decision. A rough but honest range for maintenance burden may be more useful than ignoring maintenance entirely. The analyst should not demand numerical perfection from inconvenient criteria while accepting optimistic single-point estimates for the favoured design.
Evidence quality is not always another score to add to the performance score. Multiplying every uncertain estimate by an arbitrary confidence percentage can mix different concepts and bias the comparison. A poorly evidenced option might be genuinely excellent, genuinely poor or simply unknown. Reducing its performance number does not describe that uncertainty accurately. Show the range, the source and the decision consequence. Then decide whether further evidence is needed or whether the uncertainty makes the option inappropriate for the present commitment.
A useful distinction is between uncertainty that affects the ranking and uncertainty that affects admissibility. If two feasible options exchange places when an operating-cost estimate changes slightly, the question is comparative. If uncertainty concerns whether a protective function exists at all, the question may be a gate. The same numerical treatment is not appropriate for both. The analysis must respect the role of the uncertain claim in the decision, not only its numerical size.
Preferences also need evidence, but of a different kind. The question is not whether the maintenance team can prove that independence from a supplier is objectively worth exactly fifteen points. The question is whether the decision authority understands the consequences and accepts the trade-off represented by the preference model. A value judgement becomes defensible through explicit reasoning and legitimate authority, not through pretending it was measured by an instrument.
The group marks every important input as observed, modelled, assumed or elicited. It does not colour every uncertain number red and stop. It asks which inputs deserve more work. That simple discipline prevents two common failures: an analysis that treats all inputs as equally trustworthy, and an analysis that refuses to move because nothing is absolutely certain. Engineering needs a middle position in which imperfect information is visible, useful and bounded.
7. Measure the service, not just the machine
The headline rate of the large sorter is attractive because it is easy to understand. More books per hour sounds better than fewer. But the library does not benefit from processing capacity that cannot be used. If the return point receives fewer items than the machine can handle, additional speed may have little value. If a later stage cannot accept the output, a faster sorter may create a larger downstream accumulation. Capacity is meaningful only in relation to demand and the complete path of work.
Define what counts as a completed job. In our example, a routine item must be correctly identified, its record updated and its physical route assigned. Items requiring an exception process must be recorded separately rather than quietly removed from the denominator. Otherwise a system that refuses difficult items can appear more capable than one that handles them. The definition of success is part of the measurement system, not a footnote to be added after comparing results.
The time basis matters too. An hourly rate measured during ten unusually smooth minutes may not describe sustained operation. A daily average can hide a short peak that produces unacceptable queues. A service that is fast while running but frequently unavailable can have a very different effective capacity from a slower, more continuous arrangement. Each measure answers a different question. The analysis should choose measures that correspond to the library’s actual difficulties rather than selecting the most flattering number available.
For the worked comparison, imagine that the group has constructed three planning estimates for complete routine returns under the same defined busy-period workload: 960, 900 and 820 items per hour for concepts A, B and C. These are invented figures used to demonstrate the mathematics. They are not observed equipment ratings. The common workload definition matters more than the particular values. If the figures came from incompatible conditions, even flawless arithmetic would produce an unfair comparison.
Now consider continuity. We use a separate planning measure: hours of service interruption per 1,000 scheduled service hours. The illustrative estimates are twelve hours for A, five for B and two for C. The measure must specify what interruption means. Is all return service unavailable, or only automated sorting? Can readers still obtain an immediate receipt through a fallback route? A machine-outage measure and a service-outage measure can differ substantially. The choice of denominator and event definition can change the apparent advantage.
Serviceability is more difficult to express in one natural unit. For the example, the group uses a constructed assessment of whether a defined set of routine recovery and maintenance tasks can be completed by the appropriately qualified local team with available tools and documentation. It produces illustrative value scores of 50, 80 and 85. These scores are not reliability probabilities or claims about staff competence. They summarise a stated rubric that would need real evidence before a real selection.
Adaptability is treated through a separate rubric: the ability to accommodate a specified set of future changes without disproportionate reconstruction. The changes might include a different item mix, a revised software interface or a reconfigured room. Scores of 70, 75 and 55 are assigned for teaching purposes. Without specifying the changes, “adaptability” would be little more than a compliment. An option cannot be meaningfully called flexible in the abstract; it is flexible with respect to particular changes.
These four criteria are deliberately limited. The example assumes that mandatory safety, access and compatibility obligations have already been handled through eligibility checks. It does not give them small preference weights. Financial cost is shown separately later rather than being folded into a benefit score without explanation. A real study could require additional criteria or a different structure. The purpose here is to make one transparent model whose limitations can be examined, not to prescribe a universal library procurement template.
The group also asks whether one measure is only a proxy for another. Maximum throughput might be used as a proxy for shorter queues, but the relationship depends on arrival patterns, exceptions and the rest of the workflow. If queue length is the real problem, a queueing or workload model may be needed. The proxy should remain provisional until its connection to the desired outcome is established. A measurable attribute can be technically precise and still be a weak representation of value.
At this point, the comparison has become more useful without becoming more complicated than necessary. The options are described on a common basis. The difference between machine and service is visible. Natural quantities are separated from constructed judgements. Uncertainty has not been erased, but neither has it been used as an excuse to avoid thinking. The next step is to explain how performance becomes value without pretending that all units can simply be added together.
8. Convert performance into value without hiding the conversion
An hourly rate cannot be added directly to an interruption duration. Nor can either be added to a maintenance rubric as though the result were a physical quantity. A value model creates a common comparison scale by describing how much the decision-maker values performance within a stated range. That transformation is a judgement. It can be transparent and useful, but it is not an automatic consequence of measurement.
Consider the capacity estimates. For this teaching example, the relevant range is 600 to 1,000 complete routine returns per hour. We assign zero value at 600 and 100 at 1,000, with a straight line between them. The value function is therefore 100 multiplied by the difference between the option’s rate and 600, divided by 400. At 960, the value is 90. At 900, it is 75. At 820, it is 55. The arithmetic is simple because the value assumption is simple.
Why use these endpoints? In a real study, they should come from the decision context and an explicit judgement about the range of interest. They should not be chosen secretly to favour one option. Here they are hypothetical anchors selected to make the example readable. Below the minimum acceptable performance, an option would need a separate feasibility decision; above the upper anchor, the team would need to decide whether additional capacity has value, whether the function should flatten, or whether the problem has been framed too narrowly.
A linear value function assumes that equal improvements within the range are equally valuable. Moving from 650 to 700 receives the same increase in value as moving from 900 to 950. That assumption may be wrong. The first improvement might reduce a persistent backlog, while the second adds capacity that the library rarely uses. In that case, a curved or piecewise value function would better represent the actual preference. The important question is not whether a straight line looks tidy, but whether it represents the consequence being valued.
For interruption duration, lower is better. The example assigns value 100 at zero interruption hours and zero at twenty interruption hours per 1,000 scheduled service hours. Between those anchors, the function is 100 minus five times the interruption duration. Twelve hours becomes 40, five becomes 75, and two becomes 90. Again, the scale describes a preference over a defined range. It does not turn an estimated outage measure into a probability that the equipment is good.
The serviceability and adaptability rubrics already produce values on a zero-to-100 scale. Their apparent similarity to the converted natural quantities should not disguise their different origin. A constructed scale needs operational descriptions for its levels. If two assessors would assign radically different numbers to the same evidence, the rubric is not yet doing enough work. In a consequential decision, resolve that ambiguity or represent it as a range rather than averaging the disagreement into a falsely precise score.
There is a deeper assumption behind adding separate criterion values. The relative value of an improvement on one criterion must be sufficiently stable across the levels of other criteria for the chosen additive model to be a reasonable approximation. This is not the same as statistical independence. Two physical measures can be correlated while the preference model remains useful, or statistically unrelated while the value of one depends strongly on the other. The issue is how consequences are valued, not only how they vary together.
For example, extra capacity may be nearly worthless when the service is unavailable at the only times readers can use it. A separate capacity score plus a separate continuity score could overstate the combined benefit if the two interact strongly. The team might instead model effective service capacity, introduce a joint criterion or use a more appropriate decision model. The additive approach is a tool with conditions, not a law stating that every engineering trade-off can be reduced to a weighted sum.
The public-sector MCDA guidance from HM Treasury warns against confusing a properly structured value model with casual scoring and weighting. Its specific rules govern its own appraisal context; the general warning is useful here. A calculation can be exact while its scales, preferences and scope are poorly defined. Source 4
Before moving on, Jo asks the group to explain the scales without using the word “score”. Ben describes the capacity range and what an improvement would do for the service. Aisha explains the interruption measure and why shorter disruption matters. Mira describes which maintenance tasks the rubric represents. Clara names the future changes included in adaptability. When those explanations are clear, the numbers have a chance of being useful. Without them, the matrix is only a set of decorations around an unstated opinion.
9. Weights describe trade-offs, not popularity
The next temptation is to ask everyone to rank the criteria by importance. Capacity first, continuity second, serviceability third, adaptability fourth. Such a ranking can begin a discussion, but it does not yet provide defensible weights. The value of an improvement depends on its size. A tiny improvement in an important attribute may matter less than a large improvement in a less prominent one. Weights must be considered with the scales they multiply.
Swing weighting makes that connection explicit. Imagine all criteria at the least preferred ends of their specified ranges. Which complete improvement from worst to best would be most valuable? How valuable would the other complete improvements be in comparison? The answer refers to changes across actual ranges, not to the attractiveness of criterion names. A weight on capacity cannot be interpreted independently of whether the capacity scale spans ten items per hour or several hundred.
For the fictional example, the decision authority accepts illustrative weights of 0.30 for capacity, 0.30 for continuity, 0.25 for serviceability and 0.15 for adaptability. They add to one. The weights are not survey results and not a recommendation for real libraries. They are a stated set of preferences chosen so we can inspect the model. Another legitimate authority, with different obligations or conditions, could adopt different weights and reach a different defensible choice.
The resulting value table is small enough to examine by hand.
| Criterion value, 0–100 | Weight | A: large sorter | B: two modules | C: assisted arrangement |
|---|---|---|---|---|
| Capacity within the stated range | 0.30 | 90 | 75 | 55 |
| Continuity of the complete service | 0.30 | 40 | 75 | 90 |
| Serviceability under the stated rubric | 0.25 | 50 | 80 | 85 |
| Adaptability to the named changes | 0.15 | 70 | 75 | 55 |
| Weighted value | 1.00 | 62.00 | 76.25 | 73.00 |
A’s total is 0.30 times 90, plus 0.30 times 40, plus 0.25 times 50, plus 0.15 times 70. That is 27 plus 12 plus 12.5 plus 10.5, giving 62. B produces 76.25 and C produces 73. Under this particular additive model, B has the highest value. That sentence is deliberately narrower than “B is the best engineering solution”. It identifies the conditions under which the ranking is true.
The result explains something that a single headline rate cannot. A is strongest on capacity, but its advantage is not large enough under the stated preferences to offset its weaker continuity and serviceability. B is not the most impressive option on every criterion. It is the most attractive combination in this model. C remains close because continuity and serviceability compensate for its lower capacity and adaptability values. The model has made the compromise visible rather than pretending it does not exist.
What does the difference between B and C mean? It is 3.25 value points on an artificial scale whose endpoints and weights have been defined for this decision. It is not a probability, a percentage improvement in real service or an amount of money. It does not imply that readers will be 3.25 per cent happier. Converting it into such a claim would be an unjustified change of meaning. The value difference is useful for comparison and sensitivity analysis within the model that created it.
The table also shows why preference elicitation should not be performed casually. If the team chose weights merely because 30, 30, 25 and 15 looked balanced, the output would be a numerical summary of that casual choice. If it discussed the consequences represented by each complete swing, the same numbers could carry a much more defensible meaning. Arithmetic cannot tell which process produced them. The record must explain that work.
The Government Analysis Function guide describes swing weighting and warns that weights belong to the ranges of criteria being considered. It also separates assessment of performance from elicitation of preferences. Those are the specific methodological ideas used here; the library values and calculations are original teaching material. Source 2
A useful test is to describe an implied trade-off in ordinary language. How much capacity would the authority give up for a specified reduction in interruption? Would it still accept that exchange near the low end of the capacity range? Does the maintenance advantage matter under the staffing arrangement actually available? When the answers conflict with the matrix, revise the model rather than asking people to obey a calculation that misrepresents their own stated priorities.
The ranking is now a finding, not yet a final recommendation. It needs an examination of evidence quality, uncertainty, cost, implementation and the possibility that another reasonable set of preferences changes the result. That is not an inconvenient extra step. It is where the analysis begins to reveal whether the ranking is stable enough to support the next commitment.
10. Find the point at which the answer changes
A sensitivity analysis asks a more revealing question than “What is the score?” It asks which assumptions would have to change for another option to become preferred. This helps distinguish a stable recommendation from a fragile numerical lead. It also tells the team where further discussion or evidence gathering may have value. Uncertainty that cannot affect the choice may deserve less attention than uncertainty near a switching point.
Start with the capacity weight. In the example it is 0.30. Suppose we vary that weight while keeping the other three weights in their original proportions. This condition is essential. Reducing capacity’s weight without explaining where the remaining weight goes leaves the analysis undefined. Here the continuity, serviceability and adaptability weights always share the remainder in the ratio 30:25:15.
Under that rule, B’s average value across the other three criteria is approximately 76.786. C’s is approximately 80.714. Let w represent the capacity weight. B’s total is 75w plus 76.786 times one minus w. C’s total is 55w plus 80.714 times one minus w. Subtract the second expression from the first and the difference becomes approximately 23.929w minus 3.929.
The two totals are equal when w is approximately 0.1642, or 16.42 per cent. Above that point, B is preferred to C under this particular rescaling rule. Below it, C is preferred. At the original 30 per cent, B leads by 3.25 points. This is a useful statement because it shows the preference change required to reverse the result. It is more informative than saying that the weights were “tested for sensitivity” without reporting what happened.
The threshold is conditional on the criterion values, the linear transformations and the way the other weights were rescaled. Change those assumptions and the threshold can move. A sensitivity result should therefore record its path through the model. There is no single universal “sensitivity to capacity” when several other quantities may change at the same time. The question must specify which inputs are varied, which remain fixed and why that variation is plausible.
The group can now ask whether a capacity weight below 16.42 per cent would reasonably represent the library’s priorities. If everyone agrees that the full capacity swing is more valuable than that, the ranking is relatively stable to this preference uncertainty. If legitimate decision stakeholders disagree across the threshold, the analysis has located a real value conflict. No additional measurement of machine speed will settle a disagreement about how much speed matters. The next step is a preference discussion, not another laboratory test.
Now vary evidence rather than preferences. B’s continuity value is 75. If its expected interruption is worse than assumed, that score will fall. Holding everything else fixed, B can lose 3.25 total points before tying C. Since continuity has weight 0.30, a fall of approximately 10.833 continuity points would erase the lead. Under the illustrative conversion of five value points per interruption hour, that corresponds to about 2.167 additional interruption hours per 1,000 scheduled hours. B’s estimate would rise from five to approximately 7.167 hours.
This second threshold identifies a potentially valuable learning question. Is B’s complete-service interruption likely to be near five hours, near seven or beyond that? A well-designed pilot might improve the estimate. But a short trial cannot reliably measure every rare failure mechanism. The team should ask what evidence would actually reduce the uncertainty: service records from comparable configurations, a maintenance demonstration, analysis of shared dependencies or a representative trial. Sensitivity tells us where to look; it does not tell us that any convenient test will answer the question.
Single-variable sensitivity is only a beginning. Capacity, continuity and serviceability may change together. A heavier item mix could lower throughput and increase interventions at the same time. A staffing shortage could affect both exception handling and recovery. Varying each input separately might understate the combined effect. Scenario analysis should examine plausible joint changes rather than assembling physically inconsistent combinations simply because a spreadsheet allows them.
The general reasoning connects with the library’s existing guide to sensitivity analysis, but the engineering task here is narrower: identify the thresholds that could reverse this particular design choice. A useful recommendation states those thresholds explicitly. “Select B unless evidence shows interruption above the stated switching region, or the authority assigns materially less value to capacity” is more honest than “B wins”. The first sentence tells future engineers when the reasoning should be reopened.
11. The quiet ways a matrix can choose its own winner
A decision matrix can be manipulated without changing a single arithmetic formula. The most influential choices may have been made earlier: which alternatives were included, which criteria were named, how the scales were anchored and which consequences were counted more than once. An honest analysis therefore examines the construction of the model as carefully as its final calculation. The danger is not always deliberate manipulation. Familiar habits can produce the same distortions accidentally.
Double counting is one of the easiest mistakes to make. Suppose the library gives separate weights to rapid processing, short queues and low reader waiting time. These may represent genuinely different outcomes, but they may also be three views of the same service benefit. If they are strongly overlapping consequences in this model, an option with high capacity receives repeated credit. Merely renaming one criterion does not make it independent. Trace each criterion to the outcome it is meant to represent.
The same issue appears in cost. If annual operating cost already includes staff time, and the final decision also monetises the same staff time as a separate benefit, the model may count one saving twice. That does not mean maintenance should disappear from the analysis. It means the team must distinguish the financial cost of the work from other consequences, such as whether the task can be performed locally, how much service interruption it causes and whether qualified support will remain available. Similar words can refer to different effects, but the distinction must be explicit.
Normalisation can create another problem. A common method gives the best observed option a score of 100 and the worst a score of zero for each criterion. This is convenient, but the scale then depends on the alternatives currently in the table. Add an extreme new option and the apparent distances between the original options can change, even though their physical performance has not changed. A ranking reversal may reflect the moving scale rather than a newly discovered engineering fact.
That is why the library example uses stated capacity and interruption ranges rather than silently deriving every endpoint from the current shortlist. Fixed anchors do not solve every problem, but they make the model’s reference frame easier to inspect. If the planning range changes for a good reason, update it openly and reconsider the weights. A weight assigned to an old range should not automatically be attached to a substantially different range as though its meaning were unchanged.
Dominance provides a useful preliminary check. If one eligible option is at least as good as another on every relevant criterion and strictly better on at least one, the weaker option is dominated within that comparison. It may be removed from further preference scoring if the evidence and criterion set are adequate. The phrase “within that comparison” matters. An apparently dominated option might have an omitted advantage, stronger evidence or a different implementation risk. Dominance is a conclusion about the represented attributes, not an omniscient judgement about everything in the world.
Likewise, a non-dominated option is not automatically globally optimal. It means that none of the alternatives examined has been shown to beat it on all the included criteria. There may be an unexamined design elsewhere in the feasible space. MIT’s concept-selection teaching makes the distinction between the explored set and the wider tradespace important. The practical lesson is to state the boundary of the search rather than describing a surviving shortlist as the final frontier of engineering possibility. Source 3
Rubrics can also encode the preferred architecture. A criterion called “degree of automation” rewards automation even if the actual objective is reliable service with manageable staff effort. A criterion called “number of modules” rewards modularity even when additional interfaces create new problems. The criterion should usually describe the consequence sought, not the mechanism favoured. Mechanisms belong in the alternative descriptions; outcomes belong in the value discussion. There are exceptions where an implementation is genuinely mandated, but the mandate should be treated as such.
False precision is a final warning sign. Reporting 76.2500 instead of 76.25 does not make the preference judgement more accurate. If criterion values are uncertain by several points, a long decimal expansion can make a fragile ranking look authoritative. Use enough precision to reproduce the arithmetic, then report the result at a level consistent with the evidence. It may be more useful to say that two options are close and identify the switching assumptions than to present a tiny numerical gap as decisive.
A matrix earns trust when someone can rebuild it from the stated evidence and preferences and understand why it is structured that way. It loses trust when the team can only say that the template has always used these columns. Before choosing a winner, ask whether a reasonable outsider could challenge the alternatives, criteria, scale definitions or double counting without being accused of resisting progress. The ability to inspect the model’s construction is part of the analysis, not a courtesy added after the result.
12. The purchase price is only the first chapter of cost
The large sorter has the lower purchase price than the modular arrangement. That fact is visible immediately because it appears on a quotation. Other costs are distributed through future years: energy, routine service, staff time, replacement components, software support, planned refurbishment and eventual removal. A decision based only on purchase price gives the present a privileged view of cost and leaves future operators to pay for what the comparison omitted.
Whole-life cost analysis places comparable costs on a common time basis. NIST’s life-cycle costing work provides an established reference for this way of thinking in its federal energy-management context. The arithmetic below is an independent teaching example using invented cost units and an assumed discount rate. It is not a quotation, a forecast of market prices or a recommendation to use that rate for an actual public project. Source 6
Use thousands of hypothetical cost units. Concept A costs 120 initially and 28 at the end of each of eight operating years. It also requires a planned refresh costing 30 at the end of year four and is assumed to have a recoverable residual value of five at the end of year eight. B costs 160 initially, twenty annually, fifteen for its year-four refresh and has a residual value of ten. C costs 95 initially, 34 annually, twenty for its refresh and has a residual value of three.
The example uses a real discount rate of five per cent and constant purchasing-power cost estimates. That pairing is important. If costs include general inflation while the discount rate is real, the model is inconsistent. A real project must follow its applicable financial and appraisal rules. Here the rate simply lets us demonstrate the time calculation. We also assume annual costs occur at year-end; changing the timing convention would change the values slightly.
A cost incurred at the end of year t is divided by 1.05 raised to the power t. Adding those discount factors for years one through eight gives approximately 6.4632. Thus A’s annual costs have a present value of 28 times 6.4632, or approximately 180.970. Its year-four refresh contributes approximately 24.681, and its year-eight residual value reduces cost by approximately 3.384. Adding the initial 120 gives a whole-life present cost of approximately 322.267.
| Hypothetical cost input | A | B | C |
|---|---|---|---|
| Initial cost | 120 | 160 | 95 |
| Annual cost, years 1–8 | 28 | 20 | 34 |
| Refresh at end of year 4 | 30 | 15 | 20 |
| Residual value at end of year 8 | 5 | 10 | 3 |
| Present cost under the stated assumptions | 322.267 | 294.836 | 329.173 |
B has the highest initial cost but the lowest present cost in this example. C has the lowest initial cost but the highest present cost. The lesson is not that expensive equipment is usually cheaper in the long run. It is that acquisition cost and service cost are different questions. Another set of operating assumptions could reverse the result. The model should reveal those assumptions rather than use “whole-life value” as a slogan for the preferred purchase.
Residual value deserves particular care. A theoretical resale value is not necessarily a recoverable cash flow. Removal costs, market access, contractual restrictions and the condition of the equipment can matter. Similarly, a planned refresh should not be omitted merely because it falls outside the first manager’s budget period. The comparison must use a common horizon and consistent treatment of remaining life, otherwise one option may appear cheap because its renewal obligation has been pushed just beyond the end of the table.
A simple break-even calculation can make the cost structure more intuitive. Consider a separate illustrative pair of systems. One costs 110,000 units initially and 0.20 units per completed job. The other costs 160,000 initially and 0.12 per job. Ignoring all other differences for this limited calculation, the second costs 50,000 more initially but saves 0.08 per job. The break-even volume is 50,000 divided by 0.08: 625,000 completed jobs. Below that volume the first has lower total cost; above it the second does.
That threshold is not a full recommendation. It assumes the variable costs remain constant, the systems can handle the required work, and other costs and consequences do not alter the comparison. Capacity limits, maintenance events and time discounting could all matter in a real case. Nevertheless, the threshold is useful because it turns an argument about whether the expensive option is “worth it” into a question about the volume and conditions at which its additional investment is recovered.
Costs and non-monetary benefits should now be viewed together without pretending that a value point has a price. Under the illustrative inputs, B has both the highest benefit score and the lowest present cost among A, B and C. That strengthens its position. But the cost advantage depends heavily on annual operating assumptions, while the benefit advantage depends on continuity and preference scales. A responsible recommendation reports both models and their sensitivity. It does not use agreement between two uncertain models as proof that uncertainty has disappeared.
13. Probability weights are not preference weights
The matrix weights described how much the decision authority values different improvements. Probabilities do a different job. They describe uncertainty about which state or outcome will occur. Both kinds of numbers can add to one, but that shared arithmetic does not make them interchangeable. A 30 per cent preference weight on continuity does not mean there is a 30 per cent chance that continuity matters. A 30 per cent chance of an interruption does not describe how much people are willing to tolerate it.
To make the distinction concrete, consider a separate, simplified component choice. Both candidates are assumed eligible under the relevant technical and safety requirements. The uncertain issue concerns non-safety cost exposure. One option costs eighty units if no difficult rework is needed and 240 units if rework is needed. For the example, the probability of rework is assumed to be ten per cent. The expected cost is 0.90 times eighty plus 0.10 times 240, which equals 96. A more predictable alternative costs 102.
Under an expected-cost objective alone, the uncertain option is preferred because 96 is below 102. But the expected cost is not necessarily an outcome anyone will actually observe. In this simplified model, the project pays either eighty or 240, not 96. The expectation describes a probability-weighted average across the modelled possibilities. It is a useful comparison quantity, not a promise that the project budget will land near that number.
The choice can change when the organisation’s exposure matters. If a cost of 240 would exceed an available reserve or prevent completion of other necessary work, expected cost alone may be an inadequate objective. The team may impose a budget-risk constraint, use a risk-sensitive utility model or compare the probability and severity of adverse outcomes directly. It should state that decision rule openly. It should not call the predictable alternative irrational merely because it has a slightly higher expected cost.
The switching probability is easy to calculate. Let p be the probability of rework. The uncertain option’s expected cost is eighty times one minus p, plus 240p, or eighty plus 160p. It equals 102 when p is 22 divided by 160, which is 0.1375. Above a 13.75 per cent rework probability, the predictable alternative has lower expected cost even before considering risk aversion or reserves. The original ten per cent estimate therefore sits only 3.75 percentage points below the expected-cost switching threshold.
Now ask where the ten per cent came from. It might come from comparable projects, a reliability model, expert elicitation or an assumption made for preliminary analysis. Those sources support different degrees of confidence. A numerical probability should carry a provenance note. It is particularly important not to treat a guess with two decimal places as though it were a measured frequency from a large, relevant population. The need for a probability in a model does not create evidence for its value.
If the probability is poorly known, show a range. If credible estimates span five to twenty per cent, the expected-cost ranking changes within that range. The recommendation should say so. This can justify further investigation, a reversible trial, negotiation of a cost cap or selection of the predictable option. It does not automatically justify delaying indefinitely. The question is whether additional information can change the decision enough to justify its cost and delay.
This distinction connects with the existing eduKate discussion of expected value, but the engineering application adds a specific discipline: identify the option, state, consequence, probability source and decision rule. A row of risk scores is not the same as a probability model. “Likelihood four, consequence five” does not automatically mean a twenty per cent failure probability or an expected loss of twenty units. Ordinal risk categories and quantitative decision models should not be mixed without a defensible translation.
Expected utility can represent preferences over uncertain outcomes when its assumptions and elicitation are appropriate. It need not value every extra unit of cost or benefit equally. But adding a utility curve does not make safety obligations negotiable or remove the need to understand rare severe outcomes. A model is useful only when it represents the real authority, constraints and consequences of the choice. More sophisticated mathematics can make a poor problem definition harder to notice rather than better.
The practical habit is to keep two questions beside one another. What may happen, and how much does each consequence matter? Evidence informs the first. Preferences, obligations and decision authority shape the second. Engineering judgement connects them. When those questions are kept distinct, disagreement becomes more manageable. The team can see whether it needs better forecasts, a clearer statement of values, or a different decision structure altogether.
14. Two backups can share one problem
The modular arrangement seems attractive because one module can continue when the other is unavailable. That is a plausible architectural advantage, but it is not yet a quantified reliability conclusion. The modules may share a controller, a power source, a network service or a maintenance procedure. If the shared element fails, both modules may stop. Counting boxes is not the same as demonstrating independence.
The NIST engineering statistics handbook describes the simple parallel reliability model under explicit assumptions, including independence and the ability of the system to function while at least one component remains operating. Those conditions matter as much as the multiplication formula. The calculation below is an original simplified illustration of what changes when a shared cause is added. It is not an estimate for the fictional sorter or any actual product. Source 9
Take one defined operating period with no repair during that period. Suppose each of two channels has a four per cent probability of failing through its own residual mechanism, independently of the other. If either channel can provide the required function alone, total loss occurs when both fail. Under those assumptions, the probability is 0.04 times 0.04, or 0.0016: 0.16 per cent. This looks like a strong improvement over a single channel.
Now add a shared event with probability one per cent that disables both channels. Conditional on that shared event not occurring, retain the independent four per cent residual failure probabilities. Total loss can occur either through the shared event or through both residual failures when the shared event is absent. The probability becomes 0.01 plus 0.99 times 0.04 times 0.04. The result is 0.011584, or 1.1584 per cent.
The difference is not an arithmetic correction at the edge of the model. It changes the interpretation of the architecture. The naive 0.16 per cent result described a system without the shared failure pathway. The revised result describes a different model with that pathway included. A design team cannot claim the smaller probability merely because the modules look separate. It must justify the assumptions connecting physical structure to the probabilistic structure.
Even using each channel’s marginal failure probability would not fix the problem by itself. In this model, one channel fails with probability 0.01 plus 0.99 times 0.04, which is 0.0496. Multiplying 0.0496 by itself gives approximately 0.246 per cent, still far below the actual joint loss probability of 1.1584 per cent. The missing information is dependence. Knowing the individual probabilities does not tell us how failures occur together.
The decision consequence is practical. The team might consider separating support paths, revising the controller architecture, preserving a manual service route or changing the maintenance arrangement. Each countermeasure has cost and new interfaces. Diversity is not automatically beneficial if it creates complexity that the organisation cannot support. The analysis should compare the resulting systems, including their dependencies, rather than treating “redundant” as a universal bonus word.
Uncertainty also comes in different forms. There may be variability in future workload, incomplete knowledge about a failure mechanism, uncertainty in a measurement and disagreement about a preference. These are not all reduced by the same activity. More measurements of routine throughput may do little to clarify a rare shared software failure. More stakeholder discussion may clarify preferences but cannot establish the strength of a mechanical joint. The learning action should match the kind of uncertainty that matters.
NIST’s guidance on evaluating measurement uncertainty illustrates another useful distinction: statistical evaluation of repeated observations is not the only legitimate basis for uncertainty estimates. Relevant calibration information, prior data and other scientific judgement may also contribute. That does not turn every judgement into strong evidence; it means the basis should be explicit and appropriate to the quantity being estimated. Source 5
A sophisticated simulation does not rescue an incomplete dependency model. If every simulated trial assumes independent failures, a million trials will reproduce that assumption very precisely. Numerical convergence means the calculation has settled under its model. It does not mean the model describes the real system. Before increasing computational effort, ask whether the shared pathways, conditional relationships and operational states have been represented. The most valuable engineering insight may be a missing arrow in the causal model, not another decimal place in its output.
15. Compare futures without pretending to know their probabilities
Sometimes the team cannot credibly assign one probability distribution to the future. Demand may change with policy, building use or behaviour. Suppliers may alter their products. A new service might attract a different item mix. The difficulty is not merely that the estimate is imprecise; reasonable people may disagree about the model itself. In that situation, assigning equal probabilities to a few scenarios because no better numbers are available can create a false sense of probabilistic knowledge.
Scenarios can still support a useful comparison. They describe internally coherent conditions under which each alternative is evaluated. The purpose is to expose vulnerabilities and trade-offs. The analyst should not call them forecasts unless they genuinely are forecasts, and should not describe the fraction of scenarios in which an option wins as its probability of success unless the scenario-generating process supports that interpretation.
RAND’s Robust Decision Making work is designed to support choices across multiple plausible futures, particularly where agreement on predictions is weak. The small table below is only an introductory regret calculation; it is not a complete implementation of RAND’s methodology. The distinction matters because a named framework involves more than one convenient mathematical step. Source 7
Consider three entirely separate design options, labelled A, B and C for this scenario exercise. Their hypothetical total costs under low, medium and high demand are shown below. These figures are not the life-cycle costs from the earlier library table. They are a new, deliberately small teaching model so the decision rules can be compared clearly.
| Scenario cost | Low demand | Medium demand | High demand |
|---|---|---|---|
| A | 80 | 120 | 210 |
| B | 110 | 125 | 155 |
| C | 95 | 130 | 175 |
| Lowest cost in that scenario | 80 | 120 | 155 |
A is cheapest under low and medium demand but becomes expensive under high demand. B is costly under low demand but handles high demand at the lowest cost. C is never the cheapest in any one scenario. It would be easy to dismiss C on that basis. Yet “never the cheapest scenario winner” and “never a useful choice” are different conclusions.
Regret measures how much worse an option performs than the best available option in the same scenario. For costs, subtract the lowest scenario cost from each option’s cost. A’s regrets are zero, zero and 55. B’s are thirty, five and zero. C’s are fifteen, ten and twenty. The maximum regret is therefore 55 for A, thirty for B and twenty for C. Under a minimax-regret rule, C is selected because it limits the worst modelled distance from the scenario-specific best choice.
That result does not prove C is universally best. It answers a particular question: which option minimises the largest regret across this stated scenario set? A different rule gives a different answer. If the objective is to minimise the worst absolute cost, B wins because its largest cost is 155, compared with 175 for C and 210 for A. Both rules are internally coherent, but they express different attitudes toward uncertain consequences.
Now suppose, for a separate exercise, that the authority accepts scenario probabilities of 0.20, 0.50 and 0.30. The expected costs become 139 for A, 131 for B and 136.5 for C. Expected-cost minimisation selects B. The probabilities were not discovered by the arithmetic; they were supplied as an additional assumption. Without that assumption, the expected-cost result has no basis. With it, the result is a legitimate conditional comparison.
The scenario set itself needs challenge. It may omit an important combination, such as high demand occurring alongside a staffing shortage. It may include impossible combinations, such as an expansion that assumes both reduced space and unchanged layout performance. A robust-looking choice can be an artefact of a weak scenario design. Ask which conditions cause each option to fail, whether those conditions are plausible, and whether the organisation can detect them early enough to respond.
For the library, this reasoning suggests a richer recommendation than choosing one fixed arrangement for every imaginable future. The team may prefer an initial concept that performs adequately across the near-term scenarios while preserving a credible expansion route. It may define demand thresholds that trigger another module, or support conditions that trigger a revised staffing arrangement. The existing guide to robust decision making explores that wider family of methods. Here the engineering lesson is to distinguish robustness across modelled futures from confidence in a single forecast.
A decision under deep uncertainty does not have to be vague. It can be precise about what is being protected, which futures have been considered, what would make the strategy vulnerable and which adaptation actions remain possible. Precision belongs in the structure of the response, even when precision about the future is unavailable. That is a more useful ambition than attaching unsupported probabilities to every uncertainty simply because the software has a place to enter them.
16. Ask what information is worth before asking for more of it
When the ranking is uncertain, the familiar response is to request more information. That can be wise, but information has a cost. Tests consume money, time and attention. A delay can postpone benefits or close an opportunity. The useful question is not whether more knowledge would be comforting. It is whether the information could change an available decision enough to justify obtaining it.
Consider a separate, fully hypothetical engineering choice between a standard module, S, and a novel module, N. Both are assumed to satisfy the mandatory requirements; the uncertainty concerns integration cost, not permission to accept an unsafe design. S costs 110 units regardless of the uncertain compatibility condition. N costs ninety if the surrounding interface is favourable and 170 if it is unfavourable. The current probability assigned to the favourable condition is 0.70.
Without further information, N’s expected cost is 0.70 times ninety plus 0.30 times 170, or 114. S costs 110, so an expected-cost decision selects S. The novel module is attractive in the favourable state but not attractive enough under the current uncertainty. The choice would reverse if the favourable-state probability exceeded 0.75, because N’s expected cost is 170 minus eighty times that probability.
Now imagine perfect information arriving before the choice is made. When the state is favourable, choose N at ninety. When it is unfavourable, choose S at 110. The expected cost with perfect information is 0.70 times ninety plus 0.30 times 110, which equals 96. The improvement over the current best choice is fourteen units. Under this model and objective, fourteen is the expected value of perfect information before paying for the information itself.
Perfect information is usually not what a test provides. It is a useful upper benchmark. A test might be noisy, incomplete or only partly representative of the eventual integration. The decision analysis must model what the test tells us rather than assuming that testing reveals the state without error. Otherwise it will systematically overstate the value of investigation.
For this example, suppose a proposed test gives a favourable signal with probability 0.90 when the interface is favourable, and with probability 0.20 when it is unfavourable. These signal properties are invented assumptions. They are not measured accuracy claims for a real test. The overall probability of a favourable signal is 0.70 times 0.90 plus 0.30 times 0.20, or 0.69.
Among the favourable-signal cases, the probability that the interface is actually favourable is 0.63 divided by 0.69, approximately 0.9130. After an unfavourable signal, the corresponding probability is 0.07 divided by 0.31, approximately 0.2258. This is a straightforward conditional-probability calculation: we separate the paths that can produce each signal, then compare the favourable-state path with the total probability of that signal. The arithmetic is shown so the conclusion does not depend on a mysterious software output.
After a favourable signal, N’s expected cost is approximately 96.9565, below S’s 110. Choose N. After an unfavourable signal, N’s expected cost is approximately 151.9355, so choose S. The expected cost of this signal-based policy, before paying for the test, is 0.69 times 96.9565 plus 0.31 times 110, which equals 101. Equivalently, add 0.63 times ninety, 0.06 times 170 and 0.31 times 110.
The test therefore has an expected information value of nine units: 110 minus 101. If it costs four, the total expected cost of testing and acting on the result is 105, an improvement of five over choosing S immediately. If it costs twelve, the total becomes 113, worse than acting without it under the stated assumptions. The test can be informative and still not be worth buying.
| Information position | Expected cost before any information charge |
|---|---|
| Choose immediately using the current belief | 110 |
| Receive perfect information, then choose | 96 |
| Receive the specified imperfect signal, then choose | 101 |
| Imperfect signal plus a test cost of 4 | 105 |
This example is related to the decision-tree and conditional-probability methods taught in quantitative decision-analysis courses. The particular modules, signal probabilities and cost figures here are original. MIT’s Data, Models, and Decisions materials provide a broader educational route into structured decision modelling. Source 8
Several limitations deserve attention. The model assumes the test arrives before the choice, that both actions remain available, and that the test does not itself change the underlying interface condition. It assumes the consequences and probabilities are adequately represented and that expected cost is the relevant objective after mandatory constraints have been satisfied. Change those conditions and the value calculation changes. A destructive test that consumes the only usable article is not merely an information purchase; it changes the available actions too.
Information is valuable when it supports a better policy, not merely when it narrows a number. A very accurate test may have little decision value if every possible result leads to the same action. Conversely, a modestly informative test near a switching threshold can be valuable. This does not mean mandatory verification may be skipped when its expected financial benefit is small. Required evidence serves obligations beyond this simplified discretionary-information model. Decision analysis must not turn a useful concept into an excuse for bypassing necessary engineering.
For the library, the lesson is to target the uncertainty that could alter the next commitment. If the modular arrangement’s service interruption is near the switching region, investigate the dependency and recovery behaviour that controls it. Do not spend the entire pilot budget measuring an already adequate throughput figure more precisely. The best next question is the one whose answer can change what the organisation should responsibly do.
17. Design a pilot that can actually answer the question
A pilot is often described as a small version of the final system. That description can be misleading. A useful pilot is a deliberately constructed opportunity to learn about a consequential uncertainty under bounded conditions. It may need realistic software but temporary furniture, realistic staff tasks but limited public exposure, or a representative item mix rather than full production capacity. The required fidelity follows the question.
Suppose the library wants to learn whether one module can be serviced while the other preserves an acceptable return service. A demonstration that both modules process clean, uniform items simultaneously does not answer that question. The pilot needs to exercise the relevant maintenance state, the actual fallback route and the shared support systems. It also needs an appropriate safety plan. The team should not create hazardous conditions merely to make the test look realistic.
The pilot plan should state what result would change the decision. If the concern is recovery time, define when the clock starts and stops, which tasks are included, who is qualified to perform them and what support is available. If the concern is queueing during a module outage, define the arrival workload and the remaining service path. If the concern is exception handling, specify the exception mix. Without these definitions, the pilot can generate impressive data while leaving the decisive uncertainty untouched.
Representativeness is selective but cannot be assumed. Expert developers standing beside the equipment may correct problems before ordinary staff even see them. A test network may be faster and more stable than the production connection. Temporary wiring may avoid a layout constraint that will exist in the final room. Record these differences. They may be harmless for one learning question and decisive for another. The word “pilot” does not automatically establish which conclusions can travel to the final system.
A short pilot is especially weak evidence about rare events. Suppose a simplified test consists of n independent, identically distributed demands, each with the same unknown failure probability p, and no failures are observed. The probability of that result is one minus p, raised to the power n. Setting this probability to 0.05 gives p equal to one minus 0.05 raised to the power one divided by n. This provides the familiar one-sided 95 per cent zero-failure upper confidence bound under the stated binomial model.
With twenty failure-free demands, the bound is approximately 13.9 per cent. With one hundred, it is approximately 2.95 per cent. With one thousand, it is approximately 0.299 per cent. Zero observed failures therefore does not establish zero failure probability. The result becomes more informative as the number of relevant, independent demands increases, but the assumptions remain crucial. Repeating the same easy action on the same article may not provide the independence or representativeness required for the intended claim.
The confidence statement must also be interpreted correctly. It does not mean there is a 95 per cent probability that the fixed true parameter lies below the computed bound in a frequentist interpretation. It describes the coverage property of the procedure under repeated sampling from the assumed model. For practical engineering communication, the central point is that a small set of successful demonstrations supports a bounded conclusion, not a broad guarantee of lifetime reliability.
The pilot should preserve unsuccessful and interrupted runs, not only the tidy successes. An invalid test can reveal a problem in the test setup, while a genuine product anomaly can reveal a design weakness. Both need context. Record the configuration, conditions and sequence before making changes. Otherwise the team may fix the immediate symptom but lose the evidence needed to understand whether the same weakness exists elsewhere.
The existing guides to engineering prototypes and engineering testing develop those evidence-generation questions in depth. Decision analysis uses their results for a narrower purpose: deciding whether the next commitment is justified. It should neither duplicate every test procedure nor pretend that a preference calculation can replace the work those procedures perform.
At the end of the pilot, report three things separately: what was learned, what remains unknown and what action the evidence supports. “The pilot was successful” is too compressed. It might mean the concept deserves another experiment, that a particular interface is credible, or that a defined operational scenario has been demonstrated. Those are valuable but different conclusions. The recommendation should preserve the distinction so a promising experiment does not become an accidental production authorisation.
18. A staged decision is a different design
A choice does not always have to commit the organisation to one complete configuration immediately. Sometimes it can commit to a first step that preserves later choices. This is especially useful when uncertainty can be reduced through operation or when demand will reveal itself over time. But staging is not automatically prudent. It introduces transition costs, temporary arrangements and the possibility that the hoped-for later option will no longer be available.
For the library, a staged plan might install one module in a reversible arrangement, preserve the staffed return route and reserve space and interface capacity for a second module. The first step provides useful service while gathering information about the actual item mix and support burden. The second step occurs only if the evidence and demand justify it. This is not equivalent to buying the cheapest first stage and vaguely promising to improve it later. The later path must be technically and organisationally credible.
The analysis should identify what must be prepared now to keep that path open. Physical space may need to be protected. The software interface may need to support another module without replacing the entire controller. Procurement terms may need to address future compatibility. Staff training may need a route from the initial operating arrangement to the expanded one. These preparations can cost money even if the expansion never occurs. That cost buys an option, not an immediate performance benefit.
A staged model must obey the information available at each decision point. The team cannot choose the perfect second-stage action separately in every scenario unless it will actually know which scenario has occurred before acting. This is sometimes called a non-anticipativity requirement: decisions made before uncertainty is resolved must be the same across futures that are still indistinguishable at that moment. Ignoring this condition creates a strategy with the powers of hindsight rather than a strategy an operator could execute.
Suppose the plan says “expand if demand becomes high”. What observation establishes high demand? A single busy afternoon, a sustained threshold over several weeks, or a forecast based on a confirmed service change? The trigger needs a definition. So does the lead time. If ordering and installing another module takes longer than the system can tolerate overload, a trigger at the point of failure is too late. Adaptation requires anticipation of the action’s implementation time, not merely detection of a problem.
The plan should also include a stopping or reversal condition. If the pilot reveals that the shared software dependency undermines the expected continuity advantage, the next action may be redesign rather than expansion. If the assisted route performs better than anticipated, additional automation may no longer be justified. A staged decision earns its value when future evidence can genuinely alter the path. If every possible result leads to the originally preferred purchase, the pilot may be theatre rather than learning.
Reversibility has degrees. Removing a temporary sensor is different from undoing a structural alteration or migrating a heavily integrated database. A choice can be physically reversible but economically difficult to reverse. It can be technically reversible but disruptive to readers or staff. The analysis should describe the relevant kind of reversibility rather than using the word as a general reassurance. What must be undone, at what cost, with what retained capability and under whose authority?
Staging can also transfer risk. A temporary manual process may protect against equipment uncertainty while increasing staff workload for months. A small first-stage purchase may avoid capital commitment while exposing the project to future supplier price or availability changes. Those consequences belong in the comparison. An option-preserving design is not costless; it is a deliberate exchange between present performance, present expense and future freedom.
The group therefore writes a staged recommendation as a sequence of decisions rather than a single purchase name. First, authorise the bounded pilot. Second, collect the evidence needed to assess the main continuity and workload assumptions. Third, compare the observed state with predefined triggers. Fourth, choose expansion, modification, continued operation or withdrawal. Each stage has its own evidence requirement and authority. This makes the strategy executable without pretending the future has already been observed.
The deeper lesson is that engineering alternatives can be policies over time. A flexible arrangement may be preferable not because it has the highest score on day one, but because it allows useful action while uncertainty resolves. To demonstrate that advantage, the analysis must include the costs of preserving and exercising options. Otherwise “flexibility” becomes another attractive label whose value is assumed rather than examined.
19. When the choice is a portfolio, ranking is not enough
Many engineering decisions are not about selecting one mutually exclusive option. The organisation may have a budget for several improvements: a sensor upgrade, a software change, staff training, a layout adjustment and a spare module. The task is to choose a combination. A ranking of individual projects can be useful, but it does not necessarily identify the best feasible portfolio.
Consider an intentionally small example. Project P costs four units and produces five units of agreed benefit. Q costs six and produces eight. R costs seven and produces ten. The budget is ten. Assume for this exercise that benefits are commensurate and additive, projects are indivisible, and there are no dependencies or interactions. Those are strong simplifying assumptions, stated so we can isolate one mathematical point.
The benefit-to-cost ratios are 1.25 for P, approximately 1.333 for Q and approximately 1.429 for R. A simple ratio ranking chooses R first. That spends seven, leaves three and prevents either remaining project from being added. Total benefit is ten. But choosing P and Q spends exactly ten and produces thirteen benefit units. The highest-ratio project does not lead to the best portfolio when projects are indivisible and the remaining budget cannot be used effectively.
| Project or combination | Cost | Benefit | Feasible within budget 10? |
|---|---|---|---|
| P | 4 | 5 | Yes |
| Q | 6 | 8 | Yes |
| R | 7 | 10 | Yes |
| P + Q | 10 | 13 | Yes |
| P + R | 11 | 15 | No |
| Q + R | 13 | 18 | No |
| P + Q + R | 17 | 23 | No |
The example is not an argument against efficiency ratios. It is an argument against using a heuristic outside the conditions that make it reliable. With divisible investments and particular mathematical assumptions, a ratio approach can be appropriate. With indivisible projects, dependencies or nonlinear benefits, the problem changes. The analyst should identify the decision structure before selecting the method.
Real engineering portfolios often contain dependencies. A monitoring upgrade may be useful only after an interface is changed. A second module may require a power or layout improvement. Training may need to precede deployment. These dependencies create feasible and infeasible combinations that are not visible in individual project scores. A high-value project that cannot operate without an omitted prerequisite is not a complete option.
Benefits can overlap. Two projects may both reduce the same bottleneck, so adding their separately estimated benefits overstates the total improvement. They can also reinforce each other. Better instrumentation may increase the value of a revised maintenance process because the process can now act on information it previously lacked. The portfolio model should represent these interactions where material. Adding individual benefits is an assumption, not a default truth.
The same applies to resources other than money. A combination may fit the capital budget but exceed the available installation window, specialist staff capacity or acceptable disruption. The library could afford both a layout change and a software migration yet be unable to perform them simultaneously without closing the service. A feasible portfolio must respect all binding resources, not just the most visible budget column.
Time creates another dimension. Choosing which projects to do is different from choosing their order. An early measurement project might improve the design of a later capital project. A maintenance improvement might preserve the existing system long enough to avoid a rushed replacement. A project with modest direct benefit may have high enabling value. That value should be described through the actual dependency or learning path rather than inflated through an unexplained bonus score.
When the portfolio becomes large, formal optimisation can help search the combinations. But the algorithm still depends on the model’s costs, benefits, constraints and interactions. It can produce an optimal solution to an incomplete formulation with perfect accuracy. Before celebrating the output, check whether the selected combination can actually be implemented and whether important shared consequences were omitted. The existing guide to mathematical optimisation explains the broader mathematical family; the engineering task is to formulate the right feasible problem.
For the fictional library, this prevents a false choice between “buy a machine” and “do nothing”. The best initial programme might combine a modest process change, a bounded pilot and better exception recording. Or the evidence might justify a more substantial installation with a smaller supporting portfolio. The point is not to prefer small changes automatically. It is to recognise when the unit of choice is a coordinated set of actions rather than a single item from a ranked list.
20. Follow the work through people, queues and maintenance
The assisted arrangement looks resilient because people can handle unusual items. The modular arrangement looks resilient because one module can remain available. Both claims depend on work that someone must actually perform. Human flexibility is not an unlimited resource, and a spare module is not useful if nobody can restore it or route work around it. Decision analysis should follow the practical work required to make each architectural advantage real.
Start with the shape of demand. Suppose, in a deterministic teaching example, returns arrive at an equivalent rate of 1,200 items per hour for fifteen minutes. That is three hundred items during the peak. A complete service processing nine hundred per hour can handle 225 in those fifteen minutes, leaving a backlog of 75 if there was no initial queue. The machine is not failing. Demand temporarily exceeds capacity. The backlog is a consequence of the operating pattern.
After the peak, suppose arrivals fall to six hundred per hour while service capacity remains nine hundred. The spare processing rate is three hundred per hour. Clearing the 75-item backlog therefore takes fifteen minutes under the simplified fluid model. This is not a stochastic waiting-time calculation, and it does not describe the distribution of individual reader waits. It demonstrates why an average daily arrival rate can conceal a meaningful short-term queue.
A faster service would clear the same peak differently, but the value of that speed depends on what the backlog does to people. Does it delay the update of loan records, obstruct circulation, force staff to interrupt other duties or simply wait harmlessly in a buffer? Two systems with the same backlog size can have different service consequences. The decision criteria should represent those consequences rather than assuming every item waiting anywhere is equally harmful.
Exception handling can become the controlling resource. If the automated route processes routine books quickly but diverts many unusual items to one staff member, the exception desk may determine the complete service capacity. Increasing the main conveyor speed could make the exception backlog worse. The relevant intervention might be improved classification, a clearer return process or additional qualified support during peaks. A component-centred comparison can miss this because the bottleneck lies outside the component’s advertised boundary.
Maintenance creates planned changes in capacity. One module may be removed, a software update may require a pause, or a staff member may need to perform an inspection. The question is not only how often maintenance occurs, but when it occurs relative to demand and what service remains. A longer intervention during a closed period can affect readers less than a short interruption at the busiest time. A single annual downtime number may therefore be insufficient for a consequential comparison.
The burden on staff should be described in tasks, not vague claims of convenience. Who observes an alarm? Who identifies the affected item? Who is authorised to intervene? What tools, access and information are needed? What happens if that person is unavailable? A model that assumes immediate expert response may overstate the benefit of a complex arrangement. The ordinary service organisation, not the demonstration team, must be able to sustain the chosen design.
Aisha notices that the assisted option’s attractive continuity score could disappear if it assumes staff who are already fully occupied. That is exactly the kind of cross-check the decision process should encourage. The answer is not to punish the option with an arbitrary lower score. Update the staffing assumption, estimate the resulting service consequences and recompute the comparison. A useful model can accept inconvenient information without collapsing into an argument about whose favourite option is being attacked.
This is also where distribution matters. A change may reduce total staff hours while making the remaining tasks more difficult, physically demanding or time-critical. It may shorten average queues while creating a worse experience for readers who need assistance. Those effects should not be erased by an aggregate benefit measure. Some may require design constraints, some separate criteria, and some a more careful description of who gains and who bears the burden.
The strongest engineering alternatives often emerge from this close attention to work. A layout change can reduce unnecessary handling. A clearer interface can reduce exception creation. Better access can shorten restoration. These improvements may be less visually dramatic than a faster machine, but they change the mechanism that produces service. Decision analysis becomes more valuable when it returns from abstract ratings to the sequence of real tasks that must happen for the book on the counter to become available again.
21. The same discipline works across physical and digital choices
The library case is deliberately ordinary. Its purpose is to make the reasoning visible, not to suggest that every engineering decision resembles a book sorter. The method transfers because the underlying structure recurs: define a required outcome, identify feasible alternatives, compare consequences, represent uncertainty and preserve the reasons for a choice. The technical evidence changes with the domain. The responsibility to keep the comparison honest does not.
Consider a pump replacement as a conceptual example. The question is not simply which pump has the highest efficiency on a catalogue curve. The receiving system requires a particular service across changing demand and operating conditions. Installation constraints, control behaviour, available power, maintenance support and compatibility with the existing network can matter. A component that is excellent at one operating point may be less attractive across the actual duty pattern. Qualified engineering is needed to establish the technical envelope before a preference study can compare eligible choices.
A decision analysis might identify three concepts: a direct replacement, a differently sized unit with revised controls, and a staged arrangement that changes how demand is met. The comparison should use the same service assumptions. It should not give one concept a carefully optimised control strategy while evaluating another under a poor inherited setting. It should include the cost and disruption of the supporting changes. The object of comparison is the complete workable alternative, not one catalogue line.
Now consider a digital records service. One option keeps a tightly integrated application, another separates functions into several services, and a third adopts a managed external platform. The most fashionable architecture is not automatically the best. More services can create useful boundaries, but they also create operational interfaces. A managed platform can reduce internal work while increasing dependence on an external provider. A simpler application can be easier to operate but harder to change in certain directions. The decision depends on the actual requirements and organisational capability.
Here the evidence may include workload tests, failure-recovery exercises, interface documentation, data-export trials and support arrangements. A benchmark measured on a developer’s machine is not the same as a production service forecast. A successful data export is not enough unless the exported data can be interpreted and used by the receiving system. The predecessor article on engineering technical data management explains why information must remain usable across tools and time. Decision analysis asks how those properties affect the choice among alternatives.
The two examples differ physically, but both contain a boundary problem. The pump depends on the network around it; the software depends on data, users and operations around it. Both contain a lifecycle problem: the system must remain supportable after the initial installation. Both contain an evidence problem: local tests do not automatically establish whole-system fitness. Both contain a preference problem: lower initial cost, greater flexibility and easier recovery may not be maximised together.
Transfer does not mean copying weights. A 30 per cent capacity weight from the fictional library has no authority in a pump replacement or a software migration. Even a criterion name such as reliability can refer to different required functions and time periods. What transfers is the habit of defining terms, tracing evidence and testing whether assumptions change the choice. The numerical model must be rebuilt for the new decision context.
It is also important not to confuse conceptual feasibility with professional approval. A classroom exercise can show why a modular arrangement might be attractive without establishing the safety of any actual installation. A trade study can recommend a software architecture without proving that its security controls are sufficient. Specialist design, verification and authorisation remain separate tasks. Decision analysis helps decide what to pursue and why; it does not replace the domain work required to make that pursuit responsible.
A useful transfer exercise is to identify the equivalent of the book on the counter in another system. For a pump network, it may be the required service arriving where needed. For a records platform, it may be a user retrieving the correct authorised information in time to act. For a manufacturing line, it may be an accepted part ready for the next process. Starting with that concrete outcome prevents the study from confusing internal technical activity with delivered capability.
The discipline becomes most powerful when it connects specialists without flattening their expertise. A mechanical engineer can explain why one operating assumption is implausible. A software engineer can identify a hidden consistency problem. A maintainer can expose an access burden. A decision authority can state which admissible trade-offs it is prepared to accept. The analysis creates a common structure for those contributions while preserving what each person is actually qualified and authorised to claim.
22. Supplier claims belong in an evidence process
A supplier is often the organisation that knows the most about an offered product. It may also have a commercial interest in how the product is compared. Neither fact should be ignored. Dismissing all supplier information would discard valuable expertise; accepting every claim at face value would confuse advocacy with independent evidence. The task is to make claims specific enough to examine and comparable enough to use fairly.
Ask what exactly a performance statement describes. Is it a maximum under selected conditions, a typical value across a defined population, a guaranteed contractual limit or an illustrative estimate? Does it refer to one component, a complete configuration or an installation with particular supporting equipment? A sentence such as “up to one thousand items per hour” answers a different question from a demonstrated minimum under a stated workload. The wording should remain attached to the number when it enters the comparison.
Evidence requests should follow the decision rather than become a contest in document volume. The library does not need every file a vendor has ever produced. It needs enough information to assess the requirements, interfaces, maintenance assumptions and uncertainties that could change the choice. A short, relevant test report can be more useful than a large generic brochure pack. Missing information should be recorded as missing, not silently replaced with the most favourable plausible assumption.
Where possible, give competing suppliers the same question and the same reference conditions. Ask them to describe how their complete arrangement would deliver the specified service. If one response includes a manual exception desk and another excludes exception handling entirely, the offers are not yet directly comparable. The remedy may be clarification or adjustment of the option definitions. It is not to compare the prices and assume the omitted work costs nothing.
A supplier’s proposed alternative can improve the study. Engineers should remain open to mechanisms they had not considered. But the revised concept should pass through the same eligibility and evidence process as the original options. Innovation should not receive a permanent exemption from scrutiny, and familiarity should not receive a permanent presumption of superiority. The comparison needs enough consistency that both can be evaluated on their actual consequences.
Long-term support deserves particular attention. Ask what information, tools, parts and rights the receiving organisation will need to operate, maintain or replace the system. A low initial price may be accompanied by a narrow maintenance channel or a difficult data-exit path. Those arrangements may still be acceptable, but they should be understood before the choice. Technical dependence is a lifecycle consequence, not merely a commercial detail to discover after installation.
Contractual and legal obligations vary, so real procurement must use the organisation’s authorised legal and commercial expertise. This guide does not prescribe tender rules. The general engineering point is narrower: make sure the technical alternative being evaluated is the one the supplier is actually proposing to deliver, with the supporting obligations needed to make it work. A preference score based on capabilities that are not included in the offer does not describe the offered product.
Clarifications and revisions need version identity. If a vendor changes the maintenance arrangement, software interface or included equipment during evaluation, update the affected costs and benefit assessments. Do not retain the best features of an earlier proposal while substituting the lower price of a later one. Such accidental hybrid offers can be more attractive than anything the supplier has actually committed to provide. Configuration discipline applies to the alternatives in a decision study as well as to the final product.
Conflict-of-interest handling should also be explicit within the applicable process. An assessor’s past experience with one vendor can provide valuable information, but it can also shape expectations. The answer is not to pretend that experts have no history. It is to disclose relevant interests, separate evidence from preference and ensure that consequential judgements can be reviewed. A transparent process protects both the organisation and suppliers whose offers deserve fair consideration.
By the end of evaluation, the team should be able to trace every material supplier-dependent claim to a defined source and proposal state. Some claims may remain uncertain; others may be resolved through tests or contractual commitments. The final recommendation should state which uncertainties remain with the purchaser, which responsibilities sit with the supplier and which assumptions must be checked before the next stage. That is a stronger basis for trust than either suspicion of every claim or enthusiasm for every demonstration.
23. Make the recommendation reconstructable
A decision report should allow a future person to understand the choice without having attended the meetings. That is a demanding standard. The report needs more than the winning score and a sentence saying that stakeholders were consulted. It should preserve the decision context, the alternatives, the relevant evidence, the preference model, the uncertainty and the reasons for accepting the consequences of the selected course.
For the fictional library, the recommendation begins with its scope: select concept B for a bounded pilot, not for unrestricted production deployment. The immediate purpose is to test the continuity and workload assumptions that materially affect its advantage over concept C. The recommendation does not claim that the illustrative figures are real equipment performance. In a real project, the equivalent report would replace those figures with traceable evidence and would identify the authorised professional assessments still required.
Next, record the alternatives actually compared. Each needs an identifier and a revision date or state. “B” is not enough if its staffing plan changed three times. The description should include the support systems, installation assumptions and exception route used in the analysis. Excluded alternatives should appear with concise reasons. This protects against a later argument that an option was ignored when it was actually considered under an assumption that may now need revisiting.
The evidence record should connect inputs to their sources and conditions. The value model should preserve its scales and weights. The cost model should preserve its horizon, timing, rate and cost categories. The uncertainty section should state the important ranges and dependencies. The report should distinguish information unavailable at the time from information available but not considered relevant. That distinction helps future reviewers assess the quality of the original decision without demanding impossible foresight.
The central findings can be expressed plainly. Under the illustrative value model, B scores 76.25 and C scores 73. Under the illustrative eight-year cost model, B has the lowest present cost. B’s preference advantage over C reverses if the capacity weight falls below approximately 16.42 per cent under the specified rescaling rule. Holding other inputs fixed, it also disappears if B’s interruption estimate rises from five to approximately 7.167 hours per 1,000 scheduled service hours. Those are not decorative details. They explain where the recommendation is vulnerable.
A useful record also states what is not being claimed. The comparison does not establish certified safety, prove long-term failure probabilities, settle every future demand scenario or authorise an implementation outside the stated scope. Such boundaries are not legalistic padding. They prevent a future reader from using the report as evidence for a decision it was never designed to support. A bounded conclusion is more durable than a broad claim that has to be qualified after something goes wrong.
The decision authority’s action should be recorded separately from the analyst’s recommendation. The authority may accept the recommendation, request another option, impose conditions or select a different admissible alternative for stated reasons. That distinction is essential. A model can support a choice without being the entity that makes it. The final record should show both the analytical result and the authorised decision, including any material departure from the recommendation.
Implementation conditions need owners and observable closure. If the pilot depends on confirming an interface, identify who must confirm it and what evidence is required. If a manual route must remain available, state how that availability will be demonstrated. If a cost estimate must be refreshed before expansion, name the trigger. A condition that has no owner or testable completion state is closer to a wish than a control.
The record should link to the relevant configuration and technical data. The previous guide to engineering configuration management explains why an approval belongs to a defined state. If the selected concept changes materially, the decision analysis may need to be updated. Old evidence should not be attached silently to a new arrangement merely because the project name remains the same.
Finally, include the conditions under which the decision should be reopened. These may concern demand, observed interruptions, supplier changes, support capability or new evidence about a critical assumption. A reopening rule is not a declaration of uncertainty so broad that the choice never settles. It is a way to preserve commitment while remaining corrigible. The organisation can act now and still know what future evidence would justify changing course.
24. Run a meeting that can hear an inconvenient result
A well-constructed model can still be defeated by the way a decision meeting is run. The favoured option may be presented as inevitable before anyone sees the evidence. Reviewers may receive the pack too late to examine it. Questions may be treated as disloyalty. The analysis then becomes an elaborate explanation of a choice already made. Good technical work requires a social process capable of using it.
Begin the meeting with the decision statement and its authority. Remind participants what is being chosen now and what is not. Present mandatory conditions before preference scores so an attractive total cannot distract from a missing prerequisite. Then show the alternatives, common assumptions and evidence quality. Only after that should the ranking appear. The sequence matters because it teaches the room how to interpret the result.
Ask for challenges to the model in distinct categories. Is the service boundary wrong? Is an alternative incomplete? Is a performance estimate weak? Is a preference scale inconsistent? Is a dependency omitted? Does an implementation condition make the recommendation infeasible? These questions produce more useful discussion than asking whether people “like” the recommended option. They also give specialists a clear route for contributing evidence without having to reject the entire analysis.
A dissenting view should be described accurately. Suppose a maintainer believes the modular concept’s recovery estimate is optimistic because two modules still share a controller. That is not resistance to automation. It is a claim about a dependency that could affect continuity. The appropriate response is to examine the architecture and evidence. If the concern is valid, update the model. If it is not supported, record why. Either outcome is better than translating the concern into a personality problem.
Preference disagreements should remain equally visible. A sponsor may value rapid peak processing more than the staff group does. That disagreement may be legitimate within the sponsor’s authority, but the consequences should be explicit. Show how the ranking changes across plausible preference sets. A single averaged weight can conceal a genuine conflict that the decision authority needs to resolve. The arithmetic should illuminate the disagreement, not hide it behind a consensus-looking decimal.
The room should also distinguish decision-critical uncertainty from minor uncertainty. An unresolved colour preference should not receive the same attention as an unverified service-recovery assumption. Conversely, a high number of closed minor items should not compensate for one blocking problem. The purpose is not to maximise the percentage of agenda items completed. It is to establish whether the next commitment is justified by the known state of the system.
The existing article on engineering reviews examines technical gates more broadly. Decision analysis contributes a specific input to that process: a defensible comparison of alternatives and the assumptions that could change the recommendation. A review should be able to request another comparison or a different learning action when the evidence does not support the proposed commitment.
Time pressure should be acknowledged rather than used as an unanswerable argument. Delaying can have real costs. Proceeding can also make future correction more expensive. Put both consequences in the decision. Sometimes the appropriate response is a limited action that preserves service while a critical uncertainty is resolved. Sometimes the uncertainty is immaterial to the choice and further delay is wasteful. The analysis should help distinguish those cases instead of treating either caution or speed as inherently virtuous.
At the end, state the decision in ordinary language. Which option is authorised for which purpose? What conditions apply? Which actions are prohibited until those conditions are met? Who owns the next evidence? What would require another review? If participants leave with different answers to those questions, the meeting has not produced a coherent decision even if everyone nodded at the final slide.
The quality of the meeting can be tested by asking whether an inconvenient but valid fact could still change the outcome. If not, the process is no longer using analysis as a decision aid. It is using analysis as a defence of momentum. A serious engineering culture does not require every recommendation to survive unchanged. It requires the final choice to remain answerable to the best available evidence and to the authority responsible for its consequences.
25. Judge the decision after the result, but not only by the result
Imagine that the fictional library pilots the modular arrangement. The first week goes smoothly. The group is pleased. It would be easy to declare the analysis correct. But a good early outcome does not prove that every assumption was sound. The workload may have been favourable, expert support may have been unusually available, or the weaknesses may require more time to appear. The result is evidence, but its meaning depends on what conditions produced it.
Now imagine a different first week. A shared software connection interrupts both modules. The service recovers, but the event undermines the expected continuity advantage. Does that mean the original decision was irrational? Not necessarily. The right question is what evidence and assumptions were available when the choice was made. A defensible decision under uncertainty can produce an unfavourable outcome. An indefensible decision can be lucky. Outcome alone cannot distinguish them.
The event should nevertheless change the engineering knowledge. Preserve the configuration, sequence and operational context. Determine whether the shared dependency was represented in the original model. If it was omitted despite available evidence, the decision process needs correction. If it was recognised but estimated poorly, update the estimate and examine why. If it was genuinely outside reasonable prior knowledge, ask how the new information should change the next action. In every case, the purpose is to improve the system and the decision method, not merely to assign a flattering or embarrassing label.
Separate three comparisons. First, compare observed performance with the prediction for the conditions actually experienced. Second, compare the selected option’s observed consequences with what can reasonably be inferred about the alternatives under those same conditions. Third, assess whether the original reasoning was appropriate given the information available at the time. These comparisons answer different questions. A simple before-and-after improvement does not establish that the chosen option was better than every alternative.
Counterfactual comparisons are difficult because the unchosen system was not operating beside the chosen one. Models, pilots, matched observations and controlled experiments can help, but each has limitations. Do not invent an exact performance history for the option that was not selected. State what is observed and what remains inferred. Decision learning becomes weaker when a successful project rewrites the unchosen alternatives as obviously bad, or a troubled project imagines that every alternative would have succeeded.
The record of switching conditions now becomes useful. If observed interruption under comparable conditions moves into the range where C would be preferred, reopen the comparison. But do not react to one event without considering exposure, uncertainty and the nature of the event. A rare but severe shared failure may deserve immediate architectural attention even before a stable frequency estimate exists. A harmless isolated fluctuation may not justify redesign. The response should follow consequence and mechanism, not merely the emotional force of the latest incident.
Observed success can also justify changing the model. Perhaps the assisted exception route handles a wider range of items than expected, or local maintenance proves easier after improved documentation. Those findings may alter the value of later expansion. Learning should not be reserved for failures. A decision process that only records anomalies misses opportunities to identify which conservative assumptions can responsibly be relaxed and which design features create more value than anticipated.
Beware of changing the success criteria after seeing the outcome. If the original purpose was to reduce end-to-end delay, a result showing higher machine throughput but unchanged reader delay is not complete success. The project may still have gained useful capacity, but the original outcome needs explanation. Preserve the difference between an intermediate technical improvement and the service benefit that motivated it. The book on the counter remains a useful reference precisely because it resists being replaced by a more convenient metric.
The guide to engineering failure develops causal investigation in greater depth. Decision analysis uses that learning to revise forecasts, preference models, alternative definitions and future evidence plans. The return path should reach the next project, not remain trapped in a maintenance report. Otherwise each team begins with the same optimism and repeats the same avoidable uncertainties.
A mature decision therefore has a life after approval. It is implemented, observed, compared with its assumptions and revised when justified. That does not mean endlessly reopening settled questions. It means maintaining a record strong enough to distinguish commitment from stubbornness. The organisation can say why it chose an option, what it expected, what actually happened and what it will do differently because of the difference.
26. Workshop: rebuild the decision instead of repeating the answer
Understanding a worked example is not the same as being able to construct one. The following workshop changes parts of the fictional library problem and asks you to decide what must be rebuilt. The aim is not to memorise that B won one table. It is to recognise when a conclusion still follows, when it needs new evidence and when the question itself has changed. All quantities in these exercises remain invented teaching inputs.
Exercise one: the option that wins before it is eligible
A revised scoring sheet gives the large sorter a total of 82, ahead of the other concepts. Its installation plan, however, has not demonstrated that the required accessible route can remain available. A student proposes awarding the sorter zero for accessibility and letting the other scores compensate. What is wrong with that approach?
The problem is not necessarily the number zero. It is the conversion of a mandatory condition into a compensatory preference. If the route is a genuine eligibility requirement, the candidate has not yet entered the set from which the preferred design can be selected. Mark it as unresolved or ineligible, depending on the evidence and the agreed process. The team can ask whether a revised installation would satisfy the condition. It should not pretend that excellent throughput buys permission to omit an essential route.
Notice a second distinction. Missing proof is not identical to proof of impossibility. The team might give the supplier a bounded opportunity to demonstrate an admissible arrangement. That opportunity has a cost and a schedule consequence. Those can be considered explicitly. Until the relevant condition is resolved, the candidate’s high preference score is not an authorisation to install it. This is how a decision process remains open to improvement without weakening its own boundary.
Exercise two: the new alternative that changes the measuring stick
Suppose the team normalises capacity scores by assigning zero to the slowest candidate and one hundred to the fastest candidate in the current list. A new, very slow concept is added, although nobody expects to select it. Scores and rankings among the original options now change. Has their actual service performance improved?
No. The machines, people and workflows are unchanged. The scale changed because its endpoints depended on the candidate set. That does not make every candidate-relative normalisation invalid, but it creates a consequence the analyst must understand and disclose. For a stable decision problem, a scale anchored to meaningful service consequences is often easier to interpret. The original worked example used fixed capacity anchors of 600 and 1,000 items per hour precisely so an irrelevant new candidate would not move the measuring stick automatically.
The repair is not simply to preserve the old winner. Re-examine the value function and explain why its endpoints represent the decision. If the newly considered design reveals that the original performance range was poorly framed, changing the scale may be justified. But the revision should be a conscious change to the preference model, followed by a fresh sensitivity check. It should not enter unnoticed as a spreadsheet side effect.
Exercise three: a saving that appears twice
A comparison includes eight-year lifecycle cost. It also includes annual electricity expense as a separate negatively weighted criterion. Both use the same energy forecast. Aisha notices that a design with lower electricity use receives an advantage in both columns. Is that necessarily wrong?
It is at least a warning that the same monetary consequence may be counted twice. If lifecycle cost already includes electricity expense, adding the same expense as another criterion gives it additional influence. The team should either remove the duplicate or justify a different underlying consequence. Energy use might matter independently because of a physical supply constraint or an environmental objective not adequately represented by the monetary bill. In that case, define that consequence in its own terms rather than relabelling the same financial saving.
The distinction matters when tariffs change. A design’s electricity bill can fall without its physical energy demand falling, or demand can fall while a higher tariff keeps the bill similar. Money, resource use and environmental consequence are related but not identical attributes. A good model states which one it values, how each is measured and why the combined structure does not reward one benefit twice by accident.
Exercise four: a pilot that cannot change the action
The team proposes spending four thousand cost units on a demonstration. Whatever the demonstration shows, management says it will purchase the same design immediately afterwards. Can the demonstration have decision value?
Under that stated policy, it has no value from changing which design is selected. It may still have other legitimate purposes: training staff, checking installation readiness, meeting an applicable requirement or identifying operating adjustments. Those purposes should be named and evaluated directly. Calling the event a decision experiment when no result can affect the decision hides its actual job. The earlier value-of-information calculation required a contingent policy: choose one option after one signal and a different option after the other.
A useful exercise is to write the policy before buying the test. Complete the sentence, “After result X we will do Y; after result Z we will do W.” Then ask whether Y and W genuinely differ in a way that matters. If they do not, the test may be measuring an uncertainty that is interesting but irrelevant to the present choice. That conclusion can save effort without denying the value of testing elsewhere in the programme.
Exercise five: the percentage without an exposure definition
A vendor reports that its unit has achieved 99 per cent reliability. Ben asks for the period, required function, workload, environment and meaning of failure. Another student objects that these questions are pedantic because 99 per cent is obviously high. How should Ben respond?
He should explain that the percentage is not yet a complete engineering claim. Ninety-nine successful starts out of one hundred trials describes something different from a probability of uninterrupted service over a year. A count of individual modules also differs from a claim about an integrated service. Without the exposure and function, the number cannot be compared fairly with another offer or inserted into the decision model. Precision in the printed percentage does not repair ambiguity in what was measured.
The next step is to request the definition and supporting evidence, not invent a conversion. If the information is unavailable, represent the input as uncertain or unusable for the intended comparison. A requirement for reliable service should not be satisfied by a statistic about an easier, narrower task merely because both use the word reliability. The practical lesson is to ask what has been counted before deciding how much confidence the count deserves.
Exercise six: choose the next question, not the next machine
Return to B’s 3.25-point advantage over C in the original value model. The team can investigate either the colour of a panel or the shared-controller dependency that may affect service interruption. Both questions are unresolved. Which deserves priority?
The dependency deserves attention because it can alter a consequence that materially controls the ranking and may affect other system obligations. The colour question might matter to the user experience, but the current comparison has not identified it as decision-critical. This is not a general rule that appearance never matters. It is a rule that investigation should follow the actual decision structure. In another project, visibility, contrast or visual comprehension could be central to safe and usable operation.
A complete answer would describe the dependency, the evidence needed, the possible findings and the action each finding would support. It would also check whether the test can be conducted safely and whether the observation will represent the intended configuration. Merely saying “test the controller” is still too vague. A learning action becomes useful when its result has an explicit route into the model and into the next authorised decision.
What a successful workshop answer looks like
A strong answer does not simply supply the expected label. It identifies the decision boundary, distinguishes an observation from an assumption, explains the mechanism by which the issue affects the choice and states what should happen next. Sometimes that next action is calculation. Sometimes it is clarification, a new alternative, a targeted test or a hold. The common discipline is refusing to let a convenient score stand in for a missing piece of reasoning.
For independent practice, choose a harmless everyday service such as organising classroom materials, designing a book-storage layout or selecting a process for returning borrowed equipment. Define three complete alternatives. Establish at least one genuine constraint, three distinct consequences, a small transparent value model and one uncertainty that could reverse the preference. Ask someone else to reconstruct your recommendation from the record. Their difficulty finding a source, interpreting a scale or identifying the authorised action is evidence about what your decision explanation still needs.
27. Difficult questions that a responsible analysis must answer
What happens when no alternative meets the constraints?
Do not choose the least noncompliant alternative and describe it as acceptable. The feasible set is empty under the current problem definition. That is a meaningful finding. The team may need a different architecture, a staged solution, additional resources, a narrower service commitment or a formally reconsidered requirement where reconsideration is legitimately available. Requirements imposed by applicable law, safety authority or other binding obligations cannot be relaxed merely by changing a spreadsheet.
An empty feasible set can also reveal a mistaken assumption. Perhaps the site boundary was drawn too narrowly, or a service can be scheduled differently. Explore those possibilities explicitly. The analyst’s job is to show which combinations of obligations cannot currently be satisfied and what changes would create a feasible choice. Pretending that an infeasible design is the best answer delays the conflict until implementation, when fewer options may remain.
Should the cheapest admissible option always win?
Only when cost is the agreed decision objective and the admissible alternatives are sufficiently equivalent in the consequences that matter. Meeting a minimum threshold does not make every design equally useful above that threshold. One may offer greater continuity, easier maintenance or more flexibility. Whether those differences justify additional cost is a preference and authority question supported by evidence, not a universal mathematical rule.
There are also contexts where a cost-minimisation rule is deliberately chosen. That can be perfectly coherent when the required outcomes are adequately specified and the remaining differences are immaterial or not valued. The important point is to state the rule before evaluating offers and to understand what it excludes. A decision method should serve the problem rather than force every problem into an elaborate multi-criteria model.
Does a higher value score mean a design is a certain percentage better?
Not automatically. A score of eighty versus forty does not usually mean that the first design delivers twice the service, twice the wellbeing or twice the engineering quality. The meaning depends on how the value scale was constructed. In the teaching example, the scores support comparisons under explicit anchors and weights. They are not universal physical quantities. Treating them as if they were measured output can exaggerate what the model says.
Explain the difference in the attributes that created it. A reader should be able to see whether the advantage came from capacity, interruption, serviceability or adaptability. Reporting only a total hides the trade-off. The total can be useful for decision support, but the consequence profile remains essential for understanding what will actually be experienced after the choice.
Can experts supply probabilities when little data exist?
Expert judgement can contribute information, but it should not be presented as observed frequency. Ask what experience supports the judgement, how the relevant event is defined, what uncertainty remains and whether different experts disagree. Where the decision is sensitive to a poorly supported probability, consider ranges, alternative assumptions or targeted evidence rather than printing a precise number that implies more knowledge than exists.
Sometimes probabilities cannot be credibly specified at the level needed. Scenario-based analysis can then explore consequences without assigning each future a probability. That is not an escape from reasoning. It changes the question from maximising a probability-weighted average to examining robustness, regret, vulnerability or adaptability. The decision authority should understand that these are different criteria and may recommend different actions.
What should happen when stakeholders disagree about weights?
First check whether the disagreement is truly about preferences. One person may believe an option is more reliable, while another may value reliability more strongly. The former is a disagreement about predicted performance; the latter concerns trade-offs. Resolving them requires different work. More evidence may help the first. A clearer discussion of consequences and authority may help the second.
Do not assume that averaging everyone’s weights produces a legitimate collective preference. It may hide distributional concerns, differences in responsibility or a conflict that cannot properly be reduced to one arithmetic average. Show the recommendations under relevant preference sets and identify the conditions under which they converge or diverge. The authorised decision can then acknowledge the conflict instead of claiming a false unanimity.
Can an apparently dominated option still matter?
Yes, if the dominance conclusion omitted a relevant attribute, relied on uncertain estimates or ignored a role in a staged strategy. An option that looks worse on cost and capacity might preserve a valuable exit path or require less irreversible construction. Add that consequence explicitly if it matters. Do not protect a favoured option by inventing vague benefits only after it loses, but do remain willing to improve an incomplete model.
Within a properly specified deterministic comparison, an alternative that is no better on any valued attribute and worse on at least one can be removed from further preference ranking. The qualification is important. Dominance is a statement about a particular consequence model, not an eternal property of a product name. A new use case or better evidence can change the relationship.
When should the team stop analysing and decide?
Stop when the remaining uncertainty is unlikely to change the action enough to justify the cost, delay and exposure of further investigation, while all required evidence and authority conditions are satisfied. That is not a demand for certainty. It is a judgement about the next commitment. A reversible pilot may require less evidence than an irreversible estate-wide installation, but it still needs appropriate safety and operational controls.
The stopping decision should be explainable. Identify which uncertainties remain, why they are acceptable at the proposed stage and what monitoring or later review will address them. Continuing analysis indefinitely can waste resources and delay useful service. Ending it prematurely can hide a decisive unknown. The discipline lies in connecting the amount of analysis to the consequences of acting, waiting and learning.
Can an automated tool make the recommendation?
A tool can calculate scores, explore scenarios, identify inconsistent inputs or generate a candidate recommendation. That does not give it the authority to define stakeholder rights, invent missing evidence or accept residual risk. Its output should remain linked to the model, data and assumptions that produced it. A fast calculation on an invalid scale remains an invalid calculation, however fluent the explanation around it sounds.
Automation can be particularly useful for checking arithmetic and testing many parameter combinations. Keep a record of the inputs and the exact logic used. For consequential work, arrange review proportionate to the decision. The question is not whether a human or a machine produced the first answer. It is whether the answer can be reconstructed, challenged and connected to an authorised action without losing sight of the people who will live with it.
28. Return to the book on the counter
At the beginning, a returned book waited while a group admired three different ways of moving it. By now, the book has become more than a prop. It is a reminder that the engineered system exists to complete a service. The sorter, modules, software, people, records and maintenance arrangements are means. The end is that a reader can return an item and another reader can use it again without avoidable confusion, delay or burden.
The group has not discovered a machine that wins everything. It has learned to define what winning means, to reject attractive but inadmissible comparisons and to expose the assumptions that control a preference. B leads under one explicit value model and one illustrative lifecycle-cost model. C becomes preferable under other priorities or scenarios. A might become attractive if the service need or its evidence changes. None of those statements requires pretending that the original arithmetic was universal truth.
That is the difference between a result and a decision argument. A result is a number or ranking. An argument explains the decision boundary, evidence, values, uncertainty, alternatives and authority that give the result meaning. It allows another person to disagree productively. They can challenge a throughput estimate, a value scale, a cost assumption or a staged policy without having to reject the entire idea of disciplined choice.
For a reader encountering this field for the first time, the most useful habit is to pause whenever someone says that one design is best. Best for which job? Compared with which complete alternatives? Under which constraints and future conditions? According to whose legitimate preferences? Supported by what evidence? Those questions do not make engineering indecisive. They make confidence accountable to something more durable than the enthusiasm of the person presenting the answer.
For the engineer preparing a real recommendation, the essential record can be stated in a few sentences even when the supporting work is extensive. We are deciding this action now. These alternatives are admissible. These consequences matter. These observations and assumptions support our forecasts. These trade-offs explain the preference. These uncertainties could change it. This further evidence is or is not worth obtaining. This authority permits the next step, with these conditions and these triggers for reconsideration.
The record should then survive the meeting. It should remain connected to what was actually built, tested and operated. When the world behaves differently, future people need enough information to learn whether the problem was a poor model, a missing alternative, an optimistic estimate, a legitimate uncertainty or a failure to follow the authorised plan. Without that memory, each new team inherits the conclusion while losing the reasons that made it defensible.
Engineering decision analysis is therefore neither a replacement for judgement nor an ornamental calculation placed around judgement after the event. It is a way to make judgement explicit enough to improve. It connects technical knowledge to human purposes while admitting that purposes can conflict and knowledge can remain incomplete. Its success is not that every choice looks inevitable. Its success is that the next commitment is understandable, challengeable and proportionate to what is actually known.
The book finally leaves the counter. The important question is not whether the chosen machine looked impressive when it started. It is whether the whole service keeps doing the job, whether people can support it and whether the organisation notices when the reasons for its choice stop holding. A good decision does not end the conversation with reality. It gives that conversation a structure strong enough to continue.