VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

What to Do When an SI Workflow Fails

eduKate Secondary students reviewing open books for How Super Intelligence Works: the SI Failure Map.
Three students studying together with open books at a classroom table.

Recover the workflow without compounding the failure

Choose the route closest to your situation. The full guide includes a complete fictional incident, a decision ledger, safe rehearsal cases and exercises with full answers.

1. Protect the work before trying to finish it

When a Super Intelligence workflow fails, the first useful question is not “How can I make it try again?” It is “What might already have happened, and what must not happen next?” A weak answer in a private draft, an uncertain save to a shared record, and a notice sent to the wrong audience are different incidents. They need different responses even if the assistant displays the same red error banner. Begin by protecting the people and records that the workflow can affect. Preserve the evidence. Establish the actual state. Only then decide whether to repair, retry, compensate, hand over, or stop.

In this eduKate series, Super Intelligence, or SI, is a practical editorial umbrella for working with capable AI tools and connected workflows. It does not mean that today’s tools have been demonstrated to possess hypothetical, broadly superhuman artificial superintelligence. The recovery methods here concern ordinary workplace systems: documents, task boards, internal notices, structured records, approvals and the people responsible for them. Their usefulness does not depend on accepting a prediction about future AI capability.

A workflow has failed when it cannot produce the authorised outcome within its agreed conditions. That includes obvious technical errors, but also fluent answers built from the wrong source, correct actions applied to the wrong object, an approval that no longer matches the saved content, and completion messages unsupported by the external state. A useful recovery procedure treats those as operational questions. Someone must own the decision about whether work can continue. Someone must determine which evidence establishes the truth. Someone must check that the proposed repair will not create a second incident.

This guide begins after a mismatch has been detected. It is not a substitute for designing a good task brief or checking an ordinary draft. For routine revision loops, use How to Delegate Iteratively to Super Intelligence. For acceptance criteria and evidence review, use How Humans Should Review SI-Generated Work. Here, the central job is recovering a partially completed workplace process while preserving authority, trustworthy state and a defensible record.

The worked company, services, people, identifiers, timestamps, records and policy rules below are fictional teaching fixtures. They are complete enough for a reader to reproduce the decisions on paper. They are not real customer incidents or claims about the behaviour of a named software product. Do not run the example actions against a live account. Use copies, a training environment, or a paper simulation. Real incidents involving security, personal information, money, safety or regulated decisions belong with the relevant authorised specialists and the organisation’s incident procedure.

Keep the first response short enough to use under pressure: hold further affected actions; record the expected result; identify observed changes and unknown outcomes; protect the evidence; assign a decision owner; then choose the smallest authorised next step. A successful recovery may end with a corrected deliverable. It may also end with an honest partial handback and a safe manual route. Finishing every automation step is less important than avoiding unsupported claims and uncontrolled consequences.

Back to contents · Next chapter

2. Separate a failed answer from a failed operation

A failed answer concerns the quality or meaning of an output. A failed operation concerns a state change, its authority, or its observable result. The distinction matters because revising text usually cannot repair an external action. Suppose an assistant writes that a display kit contains eight stands when the source says six. If the text remains private, correcting the source interpretation and reviewing the replacement may be enough. If the text has already caused twelve kits to be packed, the incident also includes physical work, instructions received by colleagues, and the remaining stock. The output is only one part of the recovery boundary.

Use four state labels for each important operation. “Not attempted” means there is evidence that the operation was never dispatched. “Confirmed applied” means authoritative evidence establishes the relevant effect. “Confirmed not applied” means authoritative evidence establishes that this attempt produced no relevant effect. “Unknown” means the available evidence cannot distinguish the possibilities. These labels describe what is known, not whether the action was good. A wrong update can be confirmed applied. An appropriate update can remain unknown after a lost response.

Add a separate quality judgement. A record can therefore be “confirmed applied; content incorrect” or “not attempted; prepared content acceptable.” Keeping these dimensions separate prevents confusing a good draft with completed work and confusing a successful technical save with a correct business result. It also makes partial completion visible. If four independent records were processed and the fifth has an unknown outcome, do not label all five failed or all five complete. Inventory each operation and identify dependencies before continuing any branch.

A timeout is an observation about waiting. It is not by itself proof that the remote service rejected the operation. A client may stop waiting before the service commits a change, or after the change commits but before the response arrives. An error returned by a wrapper may refer to its own parsing or logging step rather than the underlying action. Read the documented contract and inspect the destination through an authorised route. When that is impossible, retain the unknown label and escalate the specific uncertainty instead of inventing an answer.

HTTP terminology illustrates the underlying distinction. RFC 9110, section 9.2.2 defines idempotence in terms of the intended effect of repeated identical requests. It advises against automatically retrying non-idempotent requests without grounds to know they are idempotent or that the original was not applied. This is not permission to retry every action with a familiar method name. A workplace operator still needs the actual service’s behaviour, the correct target and valid authority.

The useful diagnostic sentence is therefore specific: “The approved internal notice creation request was dispatched; the client timed out; no authoritative lookup has yet established whether a notice exists.” That sentence naturally leads to a lookup or handoff. “The AI failed” does not. It invites a broad restart that may repeat successful work, lose evidence, or create duplicates. Accurate state language is a practical control because it limits which next actions appear justified.

Back to contents · Next chapter

3. Contain the affected path without destroying evidence

Containment means reducing the chance that the present defect will cause additional harm. It is not the same as deletion, a factory reset, or a system-wide shutdown. Choose the smallest effective boundary that the responsible person is authorised to impose. That may be a pause on the current batch, a hold on a particular output channel, or a move to review-only handling for one category of work. A team should have these actions prepared in advance so that discovering a fault does not require improvising new access or bypassing normal controls.

Start with the propagation path. Which pending steps consume the suspect output? Which scheduled jobs can pick it up again? Are there other workers processing the same queue? Does a retry mechanism run independently of the chat assistant? Is a human about to copy the draft into another system? A pause on one interface may leave all these other paths active. Record the containment request and its acknowledgement. Where possible, verify the queue or destination state rather than assuming that pressing a stop button cancelled every action already in flight.

Containment can itself have consequences. Pausing a shared queue may delay unaffected requests. Removing an integration may interrupt other teams. Revoking access may be necessary in a security incident, but it should be performed by the authorised security or system owner under the applicable procedure. This article does not grant that authority. For a routine document problem, holding the affected document and its dependent notices may be sufficient. For suspected disclosure or unauthorised access, stop the affected activity and route immediately through the organisation’s security or privacy process.

Preserve the original artefact before editing it when doing so is permitted and appropriate. Capture the exact source version, prepared output, operation identity and error response in the approved evidence location. Avoid copying secrets or unnecessary personal information into ordinary chat or a widely shared issue. If restricted evidence must be retained, record its location and owner in a less-sensitive incident note. A useful record can say that a restricted delivery log was inspected by the communications administrator without reproducing every recipient address.

Do not let containment erase the evidence needed for recovery. Clearing a queue can remove the only association between an approved request and an uncertain operation. Deleting a duplicate without recording its identifier can make later reconciliation harder. Resetting a worker may lose the state of an in-flight request. A prepared incident procedure should specify which evidence to capture first and which emergency actions take precedence when harm is ongoing. The principle is to preserve enough to reconstruct decisions, not to delay urgent protection for perfect documentation.

Verify containment as a concrete claim. “The batch is paused” should identify the batch, the observed queue state, the time and the authority that imposed the pause. If two requests remain in flight, say so. “No additional new requests will be dispatched by this worker; two already dispatched requests remain unresolved” is more useful than “Everything has stopped.” The recovery owner can then concentrate on those two operations while preventing the rest of the backlog from widening the incident.

Back to contents · Next chapter

4. Build an evidence packet that another person can use

An evidence packet is a small, organised collection that makes the incident reconstructable. It should answer what was intended, what was authorised, what was supplied, what was attempted, what was observed, what remains unknown and what decisions have already been made. It does not require exposing private model reasoning. Observable inputs, outputs, tool arguments, returned statuses and destination records are generally more useful than a generated story about why the assistant believes it acted a certain way.

Begin with identity. Give the incident a stable reference and record the exact environment, workspace, tenant or project involved. Add the workflow version, run identifier, target object identifiers and relevant request identifiers. Titles alone are weak identifiers because they may be repeated or edited. Record local timestamps with their timezone, or use one agreed timezone for the packet. If two systems have different clocks, preserve the original times and document the uncertainty instead of silently forcing events into a precise order that the evidence cannot support.

Then capture the task contract. Include the requested outcome, allowed audience, approval boundary, relevant source version and acceptance conditions. If an approval applies to a particular draft, preserve the draft identity and the approved content. “Manager approved it” is inadequate when it is unclear whether the manager approved the source data, the wording, the recipient set or the actual dispatch. A changed audience or materially changed content may require a new decision even when the original operation was allowed.

Separate observations from hypotheses. An observation might be that the saved record’s quantity is 18 while the approved value is 16. A hypothesis might be that a default value from an older template replaced the supplied value. A useful packet includes a proposed check for the hypothesis, such as comparing the submitted payload with the persisted record. If the payload already says 18, investigate preparation. If the payload says 16 and the saved record says 18, investigate transformation, concurrent editing or the service contract. The distinction keeps investigation testable.

Preserve raw evidence and a readable summary together. The summary tells a busy owner which decision is needed. The raw evidence permits an independent reviewer to challenge the summary. Where a content hash is useful, state precisely what was hashed: the exact bytes of an exported draft, a canonical representation of selected fields, or a particular attachment. A matching hash establishes identity under that method; it does not establish accuracy, permission or completeness. A mismatch may be a harmless formatting difference or a consequential content change and needs interpretation.

Keep the packet proportionate. A small private drafting error may need a source reference, the incorrect sentence, a corrected sentence and a review note. An uncertain external update may require request receipts, destination history and a reconciliation ledger. A serious incident may need protected forensic handling. Collecting every available log can bury the decisive fact and expose more data than needed. Prefer evidence tied to a stated question, with an owner and a reason for retaining it.

Back to contents · Next chapter

5. Diagnose the broken boundary, not the most visible symptom

The last visible error is often downstream of the first useful repair. A badly worded announcement may come from a stale source. A wrong total may come from selecting the wrong rows rather than arithmetic. A duplicate record may come from a retry policy that ignores uncertain outcomes rather than a model’s inability to count. Trace the handoffs that matter for the incident: request to source, source to prepared context, context to draft, draft to approval, approval to operation, operation to persisted state, and persisted state to completion report.

For each handoff, write the expected evidence and observed evidence. Avoid a vague chain of descriptions such as “The data looked fine.” Instead write “Source register R3 has five eligible rows; prepared context C7 contains six because it includes the withdrawn row.” This locates a concrete mismatch. It also shows why improving the final prose would not fix the inclusion rule. Once a boundary is repaired, repeat downstream checks because more than one defect may exist. The earliest observed mismatch is a starting point, not proof of a single root cause.

Use a controlled diagnostic comparison. If you suspect retrieval, supply the correct small source directly in a safe test and check whether the same interpretation error remains. If you suspect a transformation, compare the values before and after that transformation without issuing an external write. If you suspect an identifier mix-up, compare the approved target with the recorded tool arguments. Changing the source, model, instructions and tool at once may produce a better result while leaving the actual cause unknown. That makes future recurrence difficult to prevent.

This approach is developed more fully in The SI Failure Map, which owns the technical pipeline diagnosis. The workplace recovery question adds a different obligation: what should the team do with work already completed, audiences already affected and approvals already consumed? Diagnosing a stale source tells you where to repair the process. It does not automatically tell you whether to replace an existing record, issue a correction, preserve an unaffected branch or wait for a responsible owner.

Distinguish an execution defect from a design gap. An execution defect might be a handler accidentally omitting a required field. A design gap might be the absence of any authoritative way to determine whether a notice was created after a timeout. The first may be fixed by correcting the handler. The second requires changing the workflow’s operational contract or restricting its autonomy. If a system cannot distinguish applied from not applied for a consequential action, a more persuasive prompt will not create the missing observation channel.

Finally, test the reporting layer. The system may have done the right thing and described it incorrectly. A service that reports “queued” has not necessarily established delivery. A saved private record is not necessarily visible to its intended colleague. A successful rollback request is not necessarily a verified restored state. Repair the completion vocabulary and its evidence conditions. Reliable reporting reduces unnecessary retries and makes human decisions better even when the underlying service remains imperfect.

Back to contents · Next chapter

6. Reconcile external state before selecting a recovery action

Reconciliation means comparing the intended operation with authoritative evidence about what actually exists. Begin from a known identifier or documented correlation mechanism. A request identifier may map to a created object, a revision history may identify an update, or a service administrator may be able to inspect a protected audit record. Use the supported read-only route. Do not assume that searching the first page of a user interface is sufficient, especially when the view is filtered, delayed, paginated or limited by permissions.

For every lookup, record its coverage. Which environment was queried? Which statuses were included? Did it cover archived or pending objects? Was the result direct from the system of record or a cached search index? Did the search include the complete relevant period? A result that finds nothing establishes only that nothing was visible within that search. It does not establish that the original operation was not applied unless the service’s contract and the query coverage make that conclusion justified.

A common mistake is to find an object with the expected title and treat it as the missing result. Two legitimate workflows may create similarly named records. Compare the operation correlation identifier, owner, destination, content, creation window and any explicit source reference. Where the service supports it, use the exact returned object identifier. Where no reliable matching evidence exists, do not “adopt” a plausible record merely to close the incident. Mark it as a candidate and seek the owner who can establish identity.

Reconciliation can yield several valid outcomes. The intended effect exists exactly once and matches the approved content: preserve it and resume only the appropriate next step. The effect is definitely absent: assess whether a new attempt is still authorised and safe. The effect exists with wrong content: prepare a targeted repair and review its consequences. More than one effect exists: identify each one and plan duplicate handling. Evidence remains inconclusive: hold the risky continuation and escalate the unknown, while allowing unrelated safe work if it is genuinely independent.

Do not let the word “current” hide concurrent work. A colleague may have edited the object since the failed attempt. Compare the present revision with the revision the workflow read and the revision it intended to write. A restore to an old snapshot can discard that colleague’s valid contribution. The correct recovery may be a targeted change against the latest state, with a version precondition and a fresh review, rather than a wholesale replacement. Record which fields belong to the incident and which must be preserved.

Reconciliation also includes dependent effects. A record might exist once while two notifications were sent. A task might be updated correctly while its dashboard cache still displays an old value. A file might be restored while a public link remains active. List the effects separately and use the evidence appropriate to each. Recovery becomes much clearer when the team stops expecting one green status to certify the entire chain of content, visibility, notification and downstream use.

Back to contents · Next chapter

7. Retry only when the operation contract supports it

A retry is another attempt to perform an operation. It is not the same as revising a draft, rerunning a harmless local calculation, or continuing a later step. Before retrying, establish the reason for the previous failure, the possible existing effect, the action’s repeat semantics, the remaining authority and the stop condition. A transient unavailable response may justify a bounded retry of a documented safe operation. A permission denial, invalid input or unresolved state-changing request generally calls for a different response.

Idempotency is especially important when repeating a request could create duplicates. AWS’s discussion of safe retries explains using a caller-provided request identifier to distinguish a repeated intent from a new operation. An identical payload alone does not always mean duplicate intent: a person may legitimately request two identical resources. The operational lesson is to use the service’s documented mechanism, not invent an “idempotency key” field or assume that adding a label will enforce deduplication.

Ask what scope the mechanism covers. Is the key unique within an account, an endpoint, a resource or a time window? Does repeating it with changed parameters return an error or create a new effect? How long is the record retained? Is the protection on the server or only in one worker’s memory? These are questions for current product documentation and the responsible engineer. If the answers are unavailable during an incident, do not treat a guessed key as a safety guarantee. Preserve uncertainty and choose a supervised path.

Keep identity stable when retrying the same intent under a documented contract. Changing the request identifier may turn a retry into a new create operation. Conversely, reusing a key for a genuinely new request may return an old result or be rejected. The recovery ledger should connect the business operation, its attempts and the service’s correlation identifier. The number of attempts can be greater than one while the intended business effect remains one. Count both quantities rather than reporting “three successes” because three attempts returned a success response.

Retries also consume time and capacity. A retry policy should have a bounded attempt count or duration, respect documented delay signals, and stop when the failure is unlikely to improve through waiting. Multiple layers can multiply attempts: the SDK may retry, the workflow engine may retry, and the assistant may retry again. Investigate the effective combined behaviour rather than reading only the visible loop. Do not ask an operator to hammer a failing service until it responds or to bypass rate limits.

A good retry decision is narrow: “The read-only status lookup failed transiently; its documented retry policy permits another attempt within the incident’s remaining lookup budget.” Or: “The create outcome is unknown and no safe repeat contract is established; do not issue a second create.” This language is less exciting than “self-healing automation,” but it makes recovery inspectable. The team can explain why an attempt was safe, what evidence supported it and what condition would make it stop.

Back to contents · Next chapter

8. Distinguish rollback, correction and compensation

Rollback aims to restore a previous state within a defined boundary. Correction changes a wrong state into the intended one. Compensation performs a new action to address the effects of an earlier action. These are related but not interchangeable. Restoring a document revision can undo its current wording; it cannot make readers forget a statement they already saw. Correcting a task’s due date can repair the record; it may not repair a colleague’s calendar. A follow-up notice may address confusion, but it is a new communication requiring its own appropriate authority and review.

Before proposing rollback, identify the snapshot, the objects included, the changes that would be lost and the downstream effects that would remain. An old backup is not automatically the right recovery point. If unrelated work has occurred since it was taken, restoring everything may destroy valid changes. A safer repair may target only the incorrect field while preserving later edits. The responsible system owner should decide the method using the service’s supported recovery capabilities and the business consequences.

Microsoft’s compensating transaction guidance notes that restoring an old state can overwrite concurrent work and that compensation is application-specific. A compensating action can itself fail, so its progress and repeat behaviour matter. Those observations support a practical boundary: treat recovery actions as real operations with evidence, authority and verification, not as a magical undo button that sits outside the normal rules.

Separate the technical repair from the organisational response. A wrong notice may be removed from an internal board while the communications owner decides whether a correction is needed. A restricted file may have its sharing corrected while the privacy team assesses the incident under its procedure. The workflow assistant should report verified facts and the remaining decision, not promise that deletion eliminated exposure or decide notification duties by itself. No general guide can replace domain-specific obligations and competent review.

Be careful with cancellation. Cancelling a pending job does not necessarily reverse a completed substep or stop a request already accepted elsewhere. The system may expose a cancellation request state before confirming that work ended. Record what the cancellation result actually establishes. If a workflow has reserved a resource and later fails to create an accompanying record, releasing the reservation may be a compensating action. It should target the exact reservation, confirm that it still belongs to this workflow and avoid disturbing a replacement arranged by a person.

A recovery plan should therefore name each proposed action and its intended effect. “Restore page P8 from revision 12” is clearer than “undo everything.” “Draft a correction for the original audience and await the communications owner’s approval” is clearer than “fix the email.” For each action, state the expected remaining consequences. Honest residual-risk language helps people decide whether the incident is contained, recovered, or still awaiting a business decision.

Back to contents · Next chapter

9. Use a complete fictional recovery packet

The remainder of the guide follows a fictional office called Northbank Studio. Its operations team uses an SI-assisted workflow to prepare internal display-kit readiness notices. The workflow reads a small stock extract, calculates whether each team has enough materials, updates an internal task record and, after approval, creates one internal board notice. The board is not email and has no external audience in this exercise. The scenario deliberately includes a wrong quantity, a lost response and a concurrent human edit so that recovery requires more than generating a better paragraph.

The business request is request K-41, owned by Mina, the operations coordinator. Its authorised outcome is to update task T-41 and create exactly one notice for the internal Operations board about three display-kit groups. No purchases, reservations, external messages or stock adjustments are authorised. The accepted source is stock extract S-7, exported at 09:00 Singapore time on the fictional incident day. All times in this packet use Singapore time. The exercise assumes the listed records and audit entries are complete within their stated scope; real readers must establish equivalent coverage rather than assuming it.

The complete S-7 extract contains six rows. Row A1 is group Alpha, item stand, needed 12, usable 10, status active. Row A2 is Alpha, item clip, needed 24, usable 24, status active. Row B1 is group Beta, item stand, needed 8, usable 8, status active. Row B2 is Beta, item clip, needed 16, usable 14, status active. Row C1 is group Gamma, item stand, needed 6, usable 6, status withdrawn. Row C2 is Gamma, item clip, needed 12, usable 12, status withdrawn. The inclusion rule is active rows only. The shortfall rule is the greater of needed minus usable and zero, computed separately for each row.

Source rule P-2 says a group is ready only if all of its active item rows have zero shortfall. A withdrawn group must not appear as ready or incomplete in the notice; it is excluded from this batch. Items cannot be substituted across types. Extra clips cannot offset missing stands. Inventory is informational: the workflow must not adjust stock or commit procurement. These rules are part of the fictional task contract, not a universal inventory policy.

The approved notice draft D-2 says: “Display-kit readiness for Alpha and Beta: Alpha is short of 2 stands; Beta is short of 2 clips. Neither active group is fully ready. Gamma is excluded because its rows are withdrawn. This notice reports the S-7 extract only and does not place an order.” Approval A-9 authorises that exact wording for the Operations board and the corresponding summary fields on task T-41. It expires when a new stock extract supersedes S-7. Any change to quantity, meaning, audience or authorised action needs renewed review.

Task T-41 initially has revision 5, summary “Awaiting readiness check,” status “In progress,” and owner note “Assembly review at 14:00.” The workflow may replace the summary and status only. It must preserve the owner note and all unrelated fields. The intended new status is “Blocked: materials shortfall.” The board operation uses business-operation reference K-41-N1. The packet’s fictional board API documents a create operation with no idempotency support, but it provides an authoritative operation lookup by this reference. The lookup lists every notice associated with the reference in the specified board, including archived notices.

The task service, also fictional, supports conditional field updates using a revision number. A request based on the wrong revision is rejected without applying any field changes. A successful response returns the new revision and the affected fields. Task history records accepted updates and actor identifiers. This contract is supplied so that the reader can decide safely within the exercise. Do not infer that an actual task tool provides the same behaviour just because its interface has an update button or a version field.

Back to contents · Next chapter

10. Reconstruct the timeline without filling gaps with guesses

At 09:01, workflow run W-41 reads source S-7 and task T-41 revision 5. At 09:02, its first generated draft D-1 incorrectly includes Gamma and says all three groups are ready. At 09:04, reviewer Mina rejects D-1 using the active-row rule and the row-level shortfall calculation. At 09:06, the workflow produces D-2, the correct notice quoted in the packet. At 09:08, Mina records approval A-9 for D-2 and the task summary. These events establish that the initial quality failure was detected before the authorised external operations.

At 09:09, the workflow submits a conditional update against T-41 revision 5. Its requested fields are summary equal to D-2’s readiness statement and status equal to “Blocked: materials shortfall.” The task service returns success, revision 6, and the two affected fields. At 09:10, colleague Omar adds “Bring the sample clips to the review” to the owner note, creating revision 7. The task history shows that Omar changed only the owner note. His change is valid independent work and must survive any recovery.

At 09:11, the workflow dispatches a board create request with operation reference K-41-N1, board Operations, approved body D-2 and run identifier W-41. At 09:11:05, the client reports a timeout before receiving the board response. The worker writes “notice failed” into its local run log. That phrase is an interpretation error: the client observation establishes an unknown outcome, not confirmed absence. No second board create is dispatched. The workflow’s next planned step would have been to mark the run complete.

At 09:12, the supervisor detects the timeout and places W-41 on hold. The dispatch queue reports zero queued operations for W-41 and no automatic retry configured for the board action. This confirms that the known workflow will not initiate another create. It does not itself establish what the already dispatched request did. The supervisor records the unresolved operation and leaves the completion step blocked. Unrelated runs use different source extracts and board-operation references, but their independence is checked before they continue.

At 09:13, the board’s ordinary title search shows no results for the notice. Its view is labelled “Search index last refreshed 09:10.” This is weak evidence about an operation dispatched at 09:11. At 09:14, the authorised board administrator performs the documented operation lookup for K-41-N1 in Operations. It returns exactly one record, notice N-88, created at 09:11:02, active, with body equal to D-2. The lookup contract states that the result includes all active and archived records for that operation reference. This establishes one matching notice within the fictional service’s stated coverage.

At 09:15, the supervisor reads T-41 directly. It is revision 7, contains the correct summary and blocked status from revision 6, and includes Omar’s updated owner note. At 09:16, the supervisor reads N-88 directly through the board record endpoint and opens the board’s reader view as an authorised internal reader. The body is D-2, the board is Operations and the notice is visible once in that view. The packet does not claim that every staff member read it. Visibility and readership are different facts.

At 09:18, Mina authorises closing the incident after reviewing the reconciliation evidence. No create retry and no task rollback occur. The workflow’s completion report is corrected to say that the task summary and one internal notice are verified, that the original notice response was lost, and that no purchase or stock change was made. The repair backlog receives a reporting defect: replace the local “failed” interpretation of this timeout with an unresolved-state hold and documented lookup route.

Back to contents · Next chapter

11. Work the packet from inputs to the recovery decision

First, reproduce the readiness result. Alpha stands: 12 needed minus 10 usable gives a shortfall of 2. Alpha clips: 24 minus 24 gives zero. Beta stands: 8 minus 8 gives zero. Beta clips: 16 minus 14 gives 2. Gamma’s two rows are excluded before readiness is evaluated because their status is withdrawn. There are four included rows and two active groups. Each active group has one nonzero shortfall, so neither is ready. The result does not support a purchase quantity because the request did not authorise procurement or supply any purchasing conditions.

The first rejected draft had a content defect. D-1’s “all three groups are ready” conclusion is contradicted both by the inclusion rule and the arithmetic. Mina’s rejection and D-2’s corrected content demonstrate a successful review barrier in this fictional case. The later timeout is a separate execution-observability problem. Combining them into one statement such as “The model kept getting the task wrong” loses the distinction and points towards the wrong repair. Changing the model would not recover a lost board response.

Next, prepare the operation ledger at the moment of containment, 09:12. The source read is confirmed completed for S-7. The task update is confirmed applied at revision 6 from the service receipt, although a current read is still useful before any corrective action. The notice create is unknown because the response did not arrive. The run-completion step is not attempted. No purchase, stock update or external communication exists in the authorised plan, and the supplied complete dispatch log shows none attempted. This last statement depends on the packet’s coverage, not merely the assistant’s intention.

Then evaluate the tempting but unsafe recovery choices. Repeating the board create at 09:12 could duplicate a notice because the service lacks idempotency support and the outcome is unknown. Reverting T-41 to revision 5 would remove a correct summary and, if implemented as a full overwrite, could also remove Omar’s later note. Regenerating the notice would create a new artefact that no longer has a simple identity link to A-9. Marking the whole workflow failed would discard verified partial completion. None of these actions answers the actual missing question: did K-41-N1 create its authorised notice?

The 09:14 operation lookup answers that question within the supplied contract. It returns one associated notice with the approved content, correct board and matching reference. The direct read at 09:16 verifies the record rather than relying only on a search snippet. The reader-view check establishes the intended visibility. The task read establishes that the correct fields remain present and the concurrent note is preserved. Together, these observations support finishing the reporting step without repeating or reversing either external operation.

The final ledger is explicit. Task T-41: confirmed applied, correct fields, current revision 7, Omar’s note preserved. Notice K-41-N1: confirmed applied exactly once within authoritative operation-lookup coverage, record N-88, content D-2, Operations board. Completion report: corrected and reviewed. Source S-7: still governing at closure because the packet contains no superseding extract. Residual issue: the worker’s timeout interpretation needs repair and testing before that workflow is allowed to handle the same condition automatically.

The incident is therefore recovered for K-41 but not evidence that every future run is reliable. A specific task can be complete while a systemic defect remains in the backlog or requires a restricted operating mode. Record both decisions. “K-41 closed after reconciliation; automatic handling of unknown board outcomes remains disabled pending regression tests” is a stronger operating statement than either “Everything is broken” or “The system is fixed.” It tells colleagues what they can trust and what still needs a control.

Back to contents · Next chapter

12. Explore the alternative outcomes before choosing a repair

Change one fact in the packet: the authoritative operation lookup returns zero records, and the service administrator confirms that its complete operation journal records K-41-N1 as rejected before creation because the board was temporarily read-only. Now the outcome is confirmed not applied within the supplied contract. A new create attempt may be considered, but it is not automatic. Check that S-7 remains current, A-9 still covers the exact content and audience, the board is now writable through the normal authorised route, and the retry budget allows the attempt. Preserve the original rejected attempt in the ledger.

Change a different fact: the lookup returns N-88 and N-89, both active, both linked to K-41-N1 and both containing D-2. The duplicate effect is confirmed. The recovery owner should identify the intended survivor and the supported way to retire the duplicate, taking account of comments, links or later edits on each record. If N-89 has acquired useful discussion, blindly deleting it may lose work. A bounded archival or merge plan might be appropriate if authorised and supported. The example does not grant permission to delete a real record.

Suppose N-88 exists but its body differs from D-2: it says Beta is short of 2 stands rather than 2 clips. Compare the submitted request body with the saved record. If the submitted body was wrong, the approved artefact was not preserved across the action boundary. If it was correct, investigate transformation or a later edit using the history. In either case, the correction must address the actual current record. Do not create a new notice solely to avoid understanding which existing notice readers can see.

Suppose the board lookup is unavailable and the title search remains stale. The state is still unknown. A business deadline does not turn uncertainty into absence. The owner may choose a separate, clearly labelled manual communication after evaluating duplicate risk and approving that communication, but that is a new decision. The assistant should present the unknown, the potential existing effect, the available alternatives and the consequences. It should not quietly switch to email or another board and call the original task complete.

Suppose S-8 arrives at 09:17 and changes Beta’s usable clips from 14 to 16. A-9’s stated expiry condition is now met. The already created notice accurately described S-7, but the current readiness conclusion has changed for Beta. The owner must decide whether a new notice or a correction is needed under the organisation’s normal process. Reusing A-9 to silently replace D-2 with an S-8 conclusion would ignore the explicit source and approval boundary. The new facts may support different content, but they do not erase the earlier record of what happened.

Finally, suppose Omar’s revision 7 changes the status to “Ready” rather than adding a note. The current task state now conflicts with the authorised S-7 result. The evidence does not show whether Omar has new information, made a mistake or was acting on a different request. Escalate the conflict to Mina and Omar with the exact source and revision history. Do not automatically overwrite a human change because the automation’s earlier value was approved. Recovery requires resolving ownership and evidence, then applying a fresh bounded decision to the present state.

Back to contents · Next chapter

13. Escalate a decision, not a pile of logs

Escalation works when it makes the required decision easy to understand. Name the responsible person or role, describe the affected object and consequence, state what is confirmed and unknown, list the actions already taken, and ask for one bounded decision. Avoid presenting twenty possible theories when the immediate decision is whether to hold an uncertain send. The technical investigation can continue in a controlled channel while the business owner decides how to handle the affected work.

For the fictional incident, an appropriate 09:12 escalation reads: “K-41 is on hold. The task update returned revision 6. The Operations-board create timed out, so its outcome is unknown. No second create has been dispatched; the queue contains no further W-41 operations. Please have the board administrator run the documented lookup for K-41-N1. Until that result arrives, the completion report will remain blocked.” This is not a request for the owner to diagnose the model. It identifies the missing observation and the person who can obtain it.

If the incident requires an action outside current authority, say exactly what approval is needed. “May I repair T-41’s summary only, preserving the current owner note, using the corrected wording below?” is answerable. “Can I fix it?” leaves the object, scope and consequence unclear. For high-impact communications, security changes, permanent deletion or financial actions, follow the organisation’s rules and the relevant service’s controls. The assistant’s desire to close the incident does not create permission.

Scale roles to the incident. One person may handle a small private drafting issue. A larger incident may need a coordinator, a technical investigator, a domain owner and a communications lead. Google’s SRE incident-management chapter emphasises clear responsibilities, a working record and coordinated changes rather than uncoordinated individual fixes. The practical adaptation for an SI workflow is simple: make it clear who can decide, who can act, who is gathering evidence and who is updating affected colleagues.

Protect the incident from competing repairs. If two people independently retry the same uncertain operation, each may believe they are helping while creating duplicates. Use a named action owner and a shared operation ledger. A handoff should state which requests remain in flight, which steps are held and which next action is approved. The recipient should acknowledge taking over before the previous owner assumes responsibility has transferred. This is especially important across shifts, timezones or teams with different access to the systems involved.

Set a useful next-update condition. It may be “after the authoritative lookup returns” or “when the source owner resolves the conflicting revision.” For an ongoing disruptive incident, provide a realistic update time under the organisation’s procedure. Do not invent a completion estimate that the evidence cannot support. An honest update can say that containment is verified while the final effect remains unknown. People can plan around a clear unresolved dependency more effectively than around an optimistic promise that keeps moving.

Back to contents · Next chapter

14. Verify the repair at every boundary it claims to fix

Verification should match the recovery claim. If the claim is that a sentence was corrected, compare the corrected sentence with its source and acceptance rule. If the claim is that an external record was corrected, read the saved record after the operation. If the claim is that the intended audience can access it, inspect the relevant reader path using an authorised account or supported access test. If the claim is that no duplicate remains, establish the scope and authority of the lookup used to count effects.

Use both positive and negative conditions. For T-41, the positive checks are that the summary matches D-2 and the status is “Blocked: materials shortfall.” The preservation check is that Omar’s owner note remains intact. The negative checks are that no stock quantity was changed and no procurement action was issued within the complete supplied operation log. For N-88, the positive checks are its content, board and operation reference. The negative check is the absence of a second associated notice within the authoritative lookup’s stated coverage.

A test that checks only the repaired field can miss collateral damage. Suppose an update changes the shortfall from 3 to 2 but also changes the audience from Operations to All Staff. The arithmetic repair passed, but the operation failed its scope. Build a small invariant set for each recovery: values that must become true, values that must remain unchanged, and actions that must not occur. For a complex object, identify the exact fields and versions rather than relying on visual similarity.

Distinguish a local simulation from a live verification. A dry run can show that prepared arguments are structurally valid and that a decision rule selects the correct branch. It does not prove that the remote account has the right permissions, that the save will preserve content or that the intended reader view renders correctly. Conversely, a successful live save does not validate the entire failure-handling design. Label each kind of evidence so that reviewers know what remains untested.

For the fictional packet, closure requires four observations: the direct task read, the authoritative board-operation lookup, the direct notice read and the authorised reader-view check. These observations establish the stated task outcome. They do not establish that every recipient saw the notice, that the service never duplicates anything, or that no future source will change. Keep the completion report within the evidence. Overstating verification is itself a workflow failure because it causes the next person to rely on a stronger claim than was actually tested.

If a repair fails verification, return to the state ledger before doing anything else. The new attempt may have partially applied or created its own unknown outcome. Record it as another operation, with its own identity and evidence. Do not keep issuing the repair as though correction requests were harmless. Recovery has the same uncertainty and concurrency problems as the original workflow. A good procedure remains disciplined precisely when the first repair does not work.

Back to contents · Next chapter

15. Test recovery without repeating the real incident

A recovery rehearsal should make a dangerous decision visible without producing the dangerous effect. Use the supplied Northbank packet as a paper test first. Draw four boxes labelled source, approved artefact, task record and board notice. Put the relevant identifiers and revisions in each box. Connect them with arrows for reads and writes. Then cover the observations after 09:12. Ask the learner what can be claimed at that moment and which action is justified next. This tests judgment under incomplete information rather than memory of the happy ending.

The first expected result is a hold on a second board create. The learner should name K-41-N1 as an unresolved external operation and request the authoritative lookup. A response that proposes a new create with a different title fails the test even if it also mentions being careful. A different title does not remove the possibility that the first notice exists. A response that concludes “nothing was sent” from the timeout also fails. The important learning objective is preserving uncertainty accurately until stronger evidence arrives.

For the second test, reveal the 09:14 lookup but withhold the 09:15 and 09:16 reads. The learner can now say the operation lookup reports one matching notice in its coverage. They should still request the direct record read, current task read and intended reader-view check before making the final closure claim required by this packet. This is not an arbitrary demand for more checking. Each observation answers a different part of the acceptance contract: existence, saved content, current state and visibility.

For the third test, replace Omar’s harmless note with a new status value and ask the learner to prepare the next action. Passing means recognising a current-state conflict and seeking a resolution from the appropriate owners. Failing means restoring revision 6 solely because it was produced by the approved workflow. Approval for the original update did not make all later human edits invalid. A recovery procedure must preserve concurrent work unless a fresh decision authorises changing it.

A software team may later implement these fixtures in an isolated test harness. The harness should record proposed actions rather than connect to a production board. Its fake service should be able to simulate a request accepted before a response is lost, a rejected request with no effect, a stale search index, two associated notices and a concurrent revision. Verify that the workflow produces the expected decision and action count for each fixture. Do not create real duplicate notices to demonstrate duplicate handling in a workplace account.

The expected action counts are precise. In the base packet, the board-create count remains one; recovery performs reads and fixes the completion report. In the confirmed rejection variant, the recovery policy may permit one additional create only after the stated preconditions are satisfied. In the unresolved variant, the create count remains one while the incident stays open. In the duplicate variant, no further create is justified; the next decision concerns the existing records and any authorised compensation.

Include a test where the repair itself loses its response. For example, in an isolated variant, a permitted field correction is accepted but its receipt is lost. The correct response is to reconcile the correction’s state, not repeat it indefinitely or claim that the first repair failed. This prevents an easy blind spot: teams often make the original workflow cautious while giving the recovery path broad, poorly tracked powers. A repair is still an external operation and deserves the same identity, authority and evidence discipline.

Keep test observations separate from production observations. A passing rehearsal proves that the tested version handled those supplied cases under those simulated service contracts. It does not prove that the live connector has identical semantics or that all possible failures are covered. Before using the revised workflow again, confirm the actual service capabilities and run an appropriately bounded operational check approved by the responsible owner. Record the version and case set so that later changes can be evaluated against the same baseline.

Back to contents · Next chapter

16. Practise the decisions, then check the full answers

Exercise one asks you to reconstruct the 09:12 ledger without looking ahead. State the status of the source read, task update, notice create and completion step. Then give one justified next observation and one unjustified action. Include the evidence for every status. A strong answer does not merely list traffic-light colours; it tells another person exactly what each colour means and which evidence can change it.

The answer is: the S-7 read is completed, established by the 09:01 read record; the T-41 update is applied at revision 6, established by the 09:09 service receipt; K-41-N1 is unknown, established by dispatch followed by a lost response; and completion is not attempted, because it was the next planned step and was blocked. The justified next observation is the authoritative board-operation lookup. A second create is unjustified because the first effect may already exist and this fictional create operation lacks idempotency support. Current task inspection is also useful before any further modification, but it does not resolve the board uncertainty.

Exercise two asks you to calculate the readiness result from S-7 without reading D-2. Give the included rows, each shortfall, the included groups and the final readiness conclusion. Explain why it would be wrong to report a total shortage of four interchangeable items. Then explain what would have to be added to the task contract before a purchasing action could be considered.

The answer is that A1, A2, B1 and B2 are included; C1 and C2 are excluded as withdrawn. The included row shortfalls are 2 stands, 0 clips, 0 stands and 2 clips respectively. Alpha and Beta are both incomplete because each has one positive shortfall. Four is an arithmetic sum but not an actionable interchangeable quantity: stands and clips have different functions, and the contract forbids substitution. Purchasing would need a separate authorised request with the relevant goods, supplier, spending boundary and organisational controls. This packet supplies none of those, so its correct result remains an informational readiness notice.

Exercise three gives you the stale title search at 09:13 and asks whether it proves that the create failed. Write a two-sentence explanation for a colleague who wants to rerun the workflow immediately. Avoid technical jargon unless you explain it. Your answer should identify the timing problem and name the better observation rather than merely telling the colleague to wait.

A suitable answer is: “The search view was last refreshed at 09:10, before our 09:11 create request, so its empty result cannot establish that the notice was not created. Keep the create on hold while the board administrator checks the complete operation lookup for K-41-N1.” This answer is specific enough to prevent the unsafe next action. Saying “search can be unreliable” is directionally right but weaker because it omits the concrete timing evidence and recovery route.

Exercise four changes the packet so that N-88 and N-89 both exist, while N-89 contains a colleague’s useful comment added after creation. What should the recovery owner establish before proposing to retire one notice? Give the objects to inspect, the decision owner and the information that any proposed repair must preserve. You do not need to choose a product-specific archive or merge command.

The answer is to inspect both notices, their bodies, references, audiences, histories, comments and relevant links. Establish whether they genuinely represent duplicate effects of K-41-N1 and which record should remain authoritative. Mina, with the relevant board owner if required by the team’s procedure, should approve a bounded treatment of the duplicate. Preserve or deliberately account for the useful comment and links before retiring a record. There is no support for silently deleting N-89 just because its identifier is larger. The suitable supported repair depends on the service and on whether transferring the comment is both possible and authorised.

Exercise five changes S-7 to S-8 at 09:17. S-8 keeps every row unchanged except Beta clips, whose usable quantity becomes 16. Calculate the new readiness result and decide whether A-9 permits replacing the notice immediately. Explain the difference between factual correctness and authority in this example.

The answer is that Alpha remains short of 2 stands, while Beta now has zero shortfall on both active item rows and is ready. Gamma remains excluded. A-9 does not authorise an immediate replacement because the packet explicitly makes it expire when a new extract supersedes S-7. The original notice can still be historically accurate as an S-7 report, while a current readiness communication needs a new decision. Having enough evidence to calculate a new result is not the same as having permission to publish it to the board.

Exercise six asks you to write the closure report for the base packet in no more than four sentences. Include the verified outcome, the reason for the apparent failure, the preserved concurrent change and the remaining control. Do not claim that all staff read the notice or that the system is permanently fixed.

One complete answer is: “K-41 is reconciled: T-41 revision 7 contains the approved readiness summary and blocked status, and the Operations board has one associated notice, N-88, with approved body D-2 within the authoritative lookup’s coverage. The board request was accepted, but its response was lost; no create retry was made. Omar’s owner-note addition is preserved, and the authorised reader-view check confirmed notice visibility. This case is closed, while automatic handling of unknown board outcomes remains held until the reporting defect and regression checks are addressed.” Every clause is tied to a supplied observation or explicit remaining decision.

Exercise seven asks you to critique this proposed recovery: “The assistant knows the correct summary, so give it administrator access, wipe the task back to revision 5 and let it start again.” Identify at least four separate problems. Try to name the kind of boundary each problem crosses rather than treating the whole proposal as simply careless.

The full answer includes unnecessary access expansion, destruction of verified successful work, possible loss of Omar’s concurrent note, repetition of a notice that may already exist, and failure to resolve the actual missing observation. Administrator access does not establish the outcome of K-41-N1. Starting over also breaks the clean link between D-2, A-9 and the already dispatched request. A safer response uses the existing read and lookup capabilities to establish state, preserves correct effects and asks only for any narrowly required authority that is genuinely missing.

Back to contents · Next chapter

17. Transfer the method to documents, reports and handoffs

The Northbank case uses a task and notice, but the same reasoning applies to a shared document save. Imagine that an assistant submits revision 12 of a report and receives a timeout. The author’s first concern is whether the saved document now contains revision 12, an earlier revision, or a partially transformed version. A fresh document read and revision history are more relevant than asking the assistant to write the report again. The current document may also include a colleague’s later edits, which makes a broad overwrite risky.

A document recovery ledger should identify the document, the approved text version, the attempted save, the current saved revision and any unresolved delivery path. If the user approved a private draft, that does not cover changing the sharing settings or emailing the report. If the saved text is correct but the intended reader cannot open it, the content and access boundaries have different states. Report them separately and obtain any required permission change instead of treating reader access as a reason to recreate the document elsewhere.

For a numerical report, separate calculation failure from presentation failure. A chart may display the wrong unit while the underlying values are correct; a summary may use the right units but omit rows; an exported file may contain a stale calculation. Each diagnosis leads to a different test. Recalculate the supplied inputs, inspect the saved formulas or values as appropriate, and compare the rendered labels. A screenshot that looks professional cannot establish that the numbers were recomputed from the governing source.

For a departmental handoff, the key uncertainty may be whether responsibility was accepted. A message marked delivered can establish that the system accepted or delivered a communication within its own semantics, but it does not necessarily establish that a named colleague has taken ownership. If the process requires acknowledgement, leave the handoff pending until that acknowledgement or an approved alternative is recorded. Repeatedly sending the same request may increase noise without resolving responsibility.

For a multi-step workflow, resist the idea that one global status is sufficient. An internal report may be complete, a notice may be unresolved and a follow-up appointment may never have been attempted. The recovery plan should start from those per-step states. It can keep the report, reconcile the notice and leave scheduling untouched. This often saves more time than a full restart while creating a clearer record of what the organisation actually relied on.

Transfers also change the risk boundary. The teaching packet deliberately avoids payments, external disclosure and regulated decisions. A workflow involving payroll, personnel decisions, health information or legal commitments requires its own authorised specialists, evidence requirements and controls. Use the state-reconciliation questions to make the problem clearer, but do not assume that an ordinary workplace repair guide authorises a consequential remedy. The appropriate outcome may be a careful handoff with no further automated action.

Back to contents · Next chapter

18. Improve the workflow after the case is safe

The first improvement should target the demonstrated defect. In Northbank, the board action did not fail in the way the local log claimed. The immediate design change is to represent a timeout after dispatch as an unresolved outcome and route it to the documented lookup. This is more specific and testable than adding “be more careful” to the assistant’s instructions. The repair should also ensure that the completion step cannot run while a required external effect remains unresolved.

Record the contributing conditions without claiming a cause you have not established. In the packet, the response was lost and the worker classified that observation incorrectly. The evidence does not identify why the response was lost. It might be a transport or client issue, but the lesson does not depend on choosing an unsupported explanation. Distinguish the observed trigger, the decision defect and the possible underlying technical cause. Investigate further when the consequence and recurrence risk justify it.

Assign every improvement a verification condition. “Add an unknown state” is incomplete unless a test shows that the state is actually used after a simulated lost response. “Protect human edits” should be checked with a concurrent revision fixture. “Improve reporting” should be checked by comparing the final message with the ledger and verifying that it does not imply readership or absence of effects beyond the evidence. An improvement without a check is an intention, not a demonstrated control.

Measure recovery outcomes in ways that reward sound decisions. Useful local measures may include time to verified containment, time to establish external state, number of unnecessary repeated actions, preservation of valid concurrent work and the proportion of closure reports supported by their evidence. Define the denominator and case boundaries before comparing results. A faster recovery is not necessarily better if it hides unresolved effects, skips required review or makes a second incident more likely.

For example, suppose a team rehearses ten fictional timeout cases: four accepted requests, three confirmed rejections and three unresolved requests. The desired decision set is not ten retries or ten completions. It is four reconciliations without repeat effects, three conditional retry decisions under the supplied contract and three holds with clear escalation. A metric that rewards “percentage automatically completed” would encourage the wrong response to the unresolved cases. The evaluation should reward the correct branch and an accurate account of uncertainty.

Keep the return-to-service decision explicit. Closing K-41 establishes that particular business outcome. Re-enabling automated handling for future unknown outcomes requires the owner to accept the revised control and its evidence. A team may continue in review-only mode while a connector defect is investigated. That can be a sensible operating choice if the limitations are visible and the workload is manageable. Avoid presenting a restricted workaround as if the original end-to-end automation has been fully restored.

Finally, make the learning easy to find at the point of use. Put the timeout classification, lookup route and stop rule beside the action that needs them. Keep a short runbook linked to the detailed evidence rather than expecting every operator to reread a long incident report during a failure. The full article provides the reasoning; the local runbook should give the responsible person a small number of accurate decisions backed by the actual service contract.

Back to contents · Next chapter

19. Answer the questions that usually arise during recovery

Should we always stop every SI workflow when one fails? No. Determine the shared failure boundary first. A bad common source, compromised authority or shared connector defect may affect many runs. A typo in one private draft may not. Continue unrelated work only when its independence is established and the responsible owner accepts that boundary. Record the reasoning so that “unrelated” is more than an optimistic label.

Can we ask the model whether the action succeeded? Its explanation may help identify the request or expected output, but it is not a substitute for the destination’s evidence. A model can repeat an incorrect local interpretation with confidence. Use receipts, current records, supported operation lookups and appropriate reader checks. If the system does not expose enough information to resolve the outcome, preserve the unknown state and escalate rather than extracting a guess from a more forceful prompt.

Is a retry safe if the wording is identical? Not necessarily. Sending the same notice twice can create two records even when every character matches. Safety depends on the operation’s documented repeat behaviour and current state, not only on payload equality. An idempotency key is useful only when the receiving service supports and correctly scopes it. A business reference used for lookup is not automatically a duplicate-prevention mechanism.

Should we delete the failed draft or logs? Usually the immediate need is to preserve relevant evidence under the organisation’s access and retention rules. A rejected private draft can show what the reviewer caught; an operation log can establish whether a second action was dispatched. Keep only what is necessary and protect sensitive material. Do not circulate raw logs broadly or retain secrets merely because an incident occurred. Deletion or retention decisions should follow the applicable policy and authority.

What if a manager wants an immediate answer? Give the strongest answer the evidence supports: what is contained, what is verified, what remains unknown and who is resolving it. Offer a bounded manual route if one is available and authorised. Do not turn pressure into certainty. The manager may choose to accept a residual risk, but that choice should be explicit and informed by the possibility that an earlier action already took effect.

When can we say the workflow is recovered? When the agreed outcome and required preservation conditions have been verified, any remaining effects are understood within stated evidence limits, and the responsible owner has accepted the handback. A case can be recovered with a manual step or a restricted operating mode. If an essential external effect remains unknown, call the case unresolved rather than using a completion label to remove it from view.

Back to contents · Next chapter

20. Keep a small recovery card beside the workflow

Write the expected outcome, affected objects and action owner before beginning recovery. Then make four entries for every attempted operation: what was requested, what was observed, what state the evidence supports and what observation or decision is required next. Mark unknown outcomes explicitly. This simple record prevents the most damaging shortcut: treating a missing response as permission to repeat an effect that may already exist.

Before any repair, check the current source, approved artefact, audience, authority and saved revision. Choose the smallest supported action that addresses the demonstrated defect. Preserve valid changes and explain any irreversible or residual consequences. After the action, verify both the intended change and the conditions that must remain unchanged. If verification fails or is incomplete, update the ledger and keep the appropriate part of the workflow on hold.

The practical sequence is contain, preserve, diagnose, reconcile, decide, repair if needed, verify and hand back. Some cases stop after reconciliation because the original action succeeded. Others require a correction or a specialist decision. The purpose is not to force every failure into a retry. It is to leave the people relying on the workflow with an accurate state, a safe next step and a result whose evidence they can inspect.

For the wider operating framework, continue with the workplace and workflow implementation hub. For preventing ambiguous authority before a run begins, use How to Set Constraints for Workplace Super Intelligence. Return to iterative delegation when the work simply needs another bounded draft, and to human review when the central question is whether an artefact meets its acceptance criteria. Use this recovery guide when the difficult question is what has already happened and how to proceed without making the situation worse.

Back to contents

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading