
Choose your route
Workplace SI · Iterative delegation
Give a bounded increment, inspect the evidence, and decide whether to accept, repair or stop. Start with the situation closest to yours, then use the full worked packets to practise.
Contents
Build the loop
- 1. Give the next piece of work, then make a decision
- 2. Keep the outcome, scope and authority separate
- 3. Choose an increment that can earn acceptance
- 4. Define evidence before asking for a revision
- 5. Give feedback that identifies a repair
- 6. Bound retries and recognise the right stop
- 7. Keep a state record that survives the conversation
Work through complete packets
- 8. Worked packet A: reconcile display materials before drafting a note
- 9. Worked packet B: classify a small document catalogue without hiding ambiguity
- 10. Worked packet C: revise an orientation note without changing the decision
Hand back, evaluate and practise
1. Give the next piece of work, then make a decision
Iterative delegation means giving Super Intelligence a bounded piece of work, examining the result against explicit evidence, and deciding what happens next. The decision might be to accept the piece, request a particular correction, supply missing information, reduce the scope, return the work to a person, or stop. The essential unit is a reviewed increment. It is not a long conversation, a succession of increasingly emphatic prompts, or a promise that the system will keep trying until somebody likes the answer.
A workplace example makes the distinction concrete. An operations coordinator needs a short internal note about which display materials are ready for a practice session. Asking SI to “handle the preparation” leaves many questions unresolved. May it change the source list? Can it reserve materials? Should it contact colleagues? Is the note meant to report a count, propose allocations, or announce a decision? A useful first increment is much smaller: extract the available quantities from an approved source packet, retain the row identifiers, and identify contradictions. The coordinator can inspect that result before asking for an allocation draft.
This method separates effort from acceptance. A system can produce a substantial draft without producing acceptable work. It can also produce a short handback that is exactly right: “Two source records disagree about availability, so I have not produced a final allocation.” In both cases, the reviewer asks what the result establishes and what it leaves unresolved. The number of words, the confident tone, and the number of attempts are secondary. Progress is the movement of a named deliverable through a justified decision.
In this eduKateSG workplace series, Super Intelligence, or SI, is a practical editorial umbrella for present-day AI assistants, models and connected workflows. It does not mean that an ordinary workplace tool has achieved hypothetical artificial superintelligence, or ASI, broadly exceeding human capability. Nothing in this guide depends on that speculative claim. The method is about governing useful work with the systems an organisation is actually authorised to use, with their actual limitations, data access and controls.
The loop used throughout this guide is: specify an increment, produce a candidate, inspect evidence, decide, record the state, and authorise the next increment. The final step matters. Acceptance of a source summary does not automatically authorise a message to customers. Acceptance of an allocation draft does not reserve physical stock. A reviewer should be able to say precisely which object has been accepted and which action remains outside the assignment. That makes iteration useful even when no tools are connected and all work stays inside a private draft.
Anthropic’s engineering discussion of agent patterns describes feedback loops, intermediate checks, human checkpoints and explicit stopping conditions. It also advises starting with simpler approaches and adding complexity only when warranted. Those are design considerations, not evidence that a particular workplace will achieve a promised improvement. This guide adapts the general idea into a human-led operating method with original fictional examples. See Building effective agents.
The distinction between method and result is important. Iteration can make a defect easier to locate, but it can also consume review time, introduce new errors, or polish the wrong answer. A team should use it where a meaningful next piece can be checked and where the review cost is proportionate. If an approved spreadsheet formula already performs the task reliably, adding a conversational loop may not help. If nobody can judge the result, repeatedly asking for improvements is not a substitute for obtaining the relevant expertise.
By the end of this guide, you should be able to write an increment that another person could supervise, give feedback that changes a specific defect, recognise when another attempt is unjustified, and return a complete handback. You will also work through three source packets and several independent exercises. Their organisations, records, outputs and timings are fictional teaching materials. They are not descriptions of eduKate operations, product tests, customer records or measured performance outcomes.
2. Keep the outcome, scope and authority separate
Start with the outcome in ordinary language. For example, “The coordinator needs a reliable internal readiness note before choosing materials for a practice display.” Then distinguish the deliverable from the eventual outcome. The deliverable might be an inventory reconciliation, an allocation proposal, or a note for review. A draft is something the system can return. A successful practice session depends on people, physical resources and decisions beyond the draft. Avoid claiming the larger outcome merely because the smaller object exists.
Next define the scope of this particular increment. State which records, period, version, audience and output fields belong to it. “Use rows K01–K05 from packet A, version one, and return a quantity check” is inspectable. “Use what you know about our materials” is not. A source boundary also helps a reviewer identify an unsupported addition. If the system introduces a delivery date that does not exist in the packet, the defect is visible without a debate about whether the date sounds reasonable.
Authority is a different question. It describes what the system is allowed to do with people, data and tools. In the examples here, the system may read the supplied fictional packet and draft a response inside the exercise. It may not email anyone, reserve anything, edit an authoritative inventory, place an order, or treat an invented statement as a new instruction. In real work, access should be configured appropriately as well as described in the brief. A sentence asking a connected system not to act is not a replacement for suitable access controls.
Write the boundary in both positive and negative terms where ambiguity would matter. “Prepare a draft allocation for review” gives a positive job. “Do not change stock records or announce the allocation” closes two predictable gaps. There is no need to enumerate every imaginable action in a simple drafting task. Focus on the actions the system might plausibly confuse with completion. If the work grows to include a new recipient, source repository or type of action, stop at that boundary and obtain a new decision from the responsible person.
Consider a reviewer saying, “This looks good; finish it.” What exactly is approved? If the current object is a private memo, finishing may mean producing clean copy. It should not silently become permission to publish it on an intranet, send it to a supplier, or update a live system. A better review decision is “Accept the factual content of note v2 for the coordinator’s review; produce clean copy in this conversation only.” The wording preserves momentum while limiting the action to what was actually decided.
Scope changes can be legitimate. A coordinator may discover that two teams need separate versions of a note. Record that as a new requirement, rather than treating the original version as defective for not anticipating it. The distinction protects the revision history. Defect repair answers a requirement that already existed. A change request adds or alters a requirement. Mixing the two makes it difficult to assess whether SI is struggling, the brief is changing, or the reviewer is still deciding what they want.
NIST’s AI RMF Playbook organises suggested actions around Govern, Map, Measure and Manage, and explicitly describes its use as voluntary and adaptable. That supports treating workplace controls as contextual design work rather than as a universal checklist whose completion guarantees safety. The brief and checkpoints in this article are a proposed practice, not NIST certification or legal advice. See the NIST AI RMF Playbook.
For consequential professional work, these boundaries need stronger treatment. A model-generated legal interpretation, medical recommendation, financial decision or employment judgment cannot be made dependable merely by adding revision rounds. Use the organisation’s approved process and appropriately qualified review. The low-risk packets in this article deliberately avoid those decisions. They teach how to supervise an increment without pretending that a stationery exercise establishes fitness for a high-impact domain.
3. Choose an increment that can earn acceptance
A useful increment is small enough to inspect and large enough to answer a real question. “Rewrite one word” may be too small to justify coordination. “Prepare everything for next month” may hide too many dependencies. The right boundary often occurs where the reviewer learns something that changes the next step: whether the source is complete, whether categories are correctly applied, whether totals reconcile, or whether a message preserves a critical qualification. These are decision boundaries, not arbitrary word limits.
One practical sequence is extraction, transformation and presentation. Extraction identifies what the approved source actually says. Transformation performs the requested grouping, comparison or calculation. Presentation packages accepted findings for the intended reader. The stages need not always be separate calls. Separate them when an error in one would be expensive to discover later. In a five-row packet, one combined candidate may be easy to check. In a larger or ambiguous packet, approving the extraction first can prevent a polished report being built on the wrong records.
Name the increment with an observable verb and object: “reconcile the five quantities,” “classify the eight catalogue items,” or “draft the four-paragraph orientation note.” Avoid “think about,” “improve the workflow,” or “make it professional” as the only specification. Those phrases may describe a direction, but they do not define what must come back. The reviewer needs an object whose parts can be compared with the input and the acceptance conditions.
Set a stopping point inside the assignment. For example: “Return the extraction and any unresolved conflicts; do not continue to allocation.” This is especially useful when the next step depends on information the system does not possess. The system might otherwise fill the gap because the overall goal appears obvious. A designed pause lets the human supply the missing decision without making the system guess. The pause is part of the workflow, not evidence that delegation has failed.
Choose the first increment to test the most important uncertainty, not necessarily to produce the most impressive output. If the source packet contains conflicting versions, establish precedence first. If classification rules overlap, test ambiguous items first. If the final deliverable has a strict structure, check that the required information can fit before polishing prose. Early work should reduce uncertainty that would otherwise spread through later work. A beautiful cover paragraph rarely resolves a disputed count.
The first sample should also contain enough difficulty to teach you something. Selecting three unusually clean records can produce false confidence about a process that will later encounter missing values or duplicate identifiers. Include a normal case and a plausible edge case within the permitted data. This does not mean secretly testing a system with sensitive records. Synthetic cases can be designed to exercise the exact ambiguity you need to understand without exposing real personal information.
Do not confuse a narrow increment with a narrow view of context. To write one correct paragraph, SI may need the purpose of the whole document, the intended reader, the approved terminology and the reason a qualification matters. Give enough context to prevent local optimisation. Then limit the output and action. The combination is broad understanding with bounded production. Without context, a small edit can damage the larger document; without a boundary, a contextual brief can invite unnecessary expansion.
A good increment also has a practical review cost. Ask how long it will take to verify and which source will be open during review. If reviewing the candidate requires recreating the task from scratch every time, redesign the return format. Source IDs beside claims, explicit arithmetic and a separate unresolved-items list can make the same substantive work much easier to inspect. The system should not receive acceptance merely for supplying those fields, but their presence can make real acceptance possible.
Finally, decide what acceptance unlocks. An accepted inventory reconciliation may unlock drafting an allocation. An accepted allocation may unlock a coordinator’s decision. An accepted note may unlock a separate send approval. This chain keeps a small success from becoming an unlimited mandate. It also tells the system why a partial deliverable matters: the current increment establishes a specific fact or decision needed by the next one.
4. Define evidence before asking for a revision
Acceptance conditions should describe properties that matter to the reader’s use of the result. A quantity reconciliation might require every source row to appear once, the available quantity to subtract unusable units, and uncertain stock to remain uncertain. A short note might require the date, location, preparation task and unresolved issue to be preserved. These conditions are stronger than “accurate, concise and professional” because the reviewer can point to an individual pass or failure.
Separate non-negotiable conditions from preferences. A wrong date, invented approval or omitted exception can block acceptance. A slightly less elegant opening sentence may not. If every stylistic preference becomes a blocking defect, the team can iterate indefinitely without improving the work’s usefulness. Conversely, a handsome tone cannot compensate for a false commitment. Agree on the difference before review so the system is not rewarded for surface polish while substantive errors remain.
Evidence should connect the output to an independent reference. In a closed packet, that reference is the supplied row or rule. In a live workplace task, it might be an approved document version, a source-system record, a calculation or an observed tool result. “I checked everything” is not enough. Ask for the checkable relationship: row K03 says ten units are held and three are unusable, so usable stock is seven. The reviewer can inspect both the source and the arithmetic.
The evidence requirement should fit the claim. A spelling change needs comparison with the requested spelling. A calculated total needs the input values and operation. A recommendation needs the constraints and reasons relevant to the recommendation. A claim that an external action happened needs confirmation from the destination system, not a generated sentence announcing success. Avoid asking for a long explanation of hidden model reasoning. What matters is the visible evidence and decision rationale needed to assess the work.
Keep unsupported and contradicted claims separate. An unsupported claim has no adequate basis in the authorised material. A contradicted claim conflicts with it. If a source packet says nothing about delivery, “delivery is confirmed” is unsupported. If it says delivery is unconfirmed, the same statement is contradicted as well as unauthorised. Both need repair, but the distinction helps identify the mechanism. The first may come from filling a gap; the second may come from dropping an inconvenient qualification.
NIST’s Generative AI Profile identifies confabulation, including confidently presented false content, as a relevant risk. That is why confidence of expression is not used as acceptance evidence here. The practical response is to compare important claims with appropriate sources and to keep uncertainty visible. This is a limited use of the profile, not a claim that any particular model will fail in a predictable proportion of cases. See NIST AI 600-1, Generative Artificial Intelligence Profile.
A useful review sheet has a condition, an observed result and a decision. For example: “C2: unusable units excluded; candidate lists twelve usable clips; source supports ten; fail.” The repair then follows naturally. A vague score of seven out of ten cannot tell the system what to change. Scores may help summarise a larger evaluation, but they should not replace the individual observations that justify accepting or rejecting this specific deliverable.
You can accept an increment with a clearly bounded limitation when the intended use allows it. Suppose the task is to identify unresolved inventory questions. A list containing an unresolved count can be fully accepted as a questions list. The same output cannot be accepted as a final stock statement. Acceptance attaches to a purpose, not to an intrinsic aura of correctness. State that purpose in the review record so a later reader does not reuse a limited result as if it were fully settled.
Regression checks preserve what already works. When a revision corrects one count, recheck the related total and any sentence based on that total. When a date changes, inspect the subject line, body and action list. Do not assume the untouched parts of a newly generated document are literally unchanged. Ask the system to identify the intended change, compare the revised object with the accepted baseline, and verify the important invariants again before accepting the revision.
5. Give feedback that identifies a repair
Specific feedback names the location, the defect, the evidence and the desired change. “In the readiness sentence, replace ‘all materials are ready’ because row K03 leaves three labels unusable and no replacement is confirmed. State the remaining shortage and keep the note internal.” This directs a repair without rewriting the entire assignment. It also makes the reviewer’s reasoning inspectable to another colleague who may inherit the task later.
Do not bundle unrelated wishes into a correction without identifying them. “Fix the total, make it friendlier, add a timeline and send it to the team” mixes a factual defect, a style preference, a new deliverable and a new external action. The system may satisfy the easy parts while obscuring the important one. Split the decision: first repair and verify the total; separately decide whether the additional timeline and communication are wanted and authorised. The loop should make the work clearer rather than disguise scope growth as feedback.
Feedback can be directive or diagnostic. Directive feedback supplies the correct interpretation when the reviewer has established it: “Use the later source note, which explicitly supersedes the earlier one.” Diagnostic feedback asks for a targeted investigation: “The two counts disagree; identify the source of the difference without choosing a final number.” Use the second form when the reviewer does not yet know the answer. Instructing the system to agree with an unverified human guess simply transfers the guess into cleaner prose.
Avoid feedback that rewards agreement alone. “That cannot be right; try again” gives the system social pressure without new evidence. A changed answer might be better, worse or equally unsupported. If you suspect an error, point to the suspect field and request a source-grounded check. If you cannot explain the concern yet, say what is uncertain and ask for alternatives with evidence. The reviewer remains allowed to be unsure; good delegation does not require pretending to possess the answer in advance.
A repair should preserve accepted content unless a dependency requires changes. Say which parts are locked, which part may change and which dependent checks must be rerun. In a short memo, the accepted meeting location might remain fixed while the start time is corrected. The system should return a clean revised memo plus a short change note. The change note is not a substitute for the full artifact when the reviewer needs something ready to use. Both have different jobs.
When several defects share a cause, repair the cause rather than each symptom independently. If every total includes held stock that is unavailable, the missing distinction between held and usable quantities is the underlying issue. Correct the rule and recalculate every affected row. If only one row has a transcription error, a full conceptual redesign is unnecessary. This diagnosis prevents overreaction and helps the team choose whether another iteration is likely to be useful.
Use a short priority order for multiple defects. First address authority or data-boundary violations. Then address source selection and factual correctness. Next check completeness and structure. Finally adjust style if it affects usability. This order is a proposed working convention, not a universal standard. It prevents spending the next round polishing a message that should not yet be sent or formatting a total built from the wrong records.
Feedback should also say when not to revise. “The unresolved marker is correct; keep it” can be important when the system is inclined to make the answer look complete. “Do not add a reason for the missing record” prevents a plausible explanation becoming fictional evidence. Good feedback protects honest gaps. The goal is a reliable deliverable with visible limits, not a page that appears to know everything.
After giving feedback, review the new result rather than the system’s acknowledgement. “Understood, I corrected that” is a statement about intent. The corrected artifact is what matters. Check the named defect, its dependencies, and any previously accepted condition that could have regressed. Then record a decision in plain language: accepted, revise again within the existing limit, waiting for information, or handed back. This closes the round and prevents a conversation from drifting without a current state.
6. Bound retries and recognise the right stop
A retry budget is a limit on a particular repair process, not a target to exhaust. For a low-risk internal draft, a coordinator might allow one initial candidate and up to two revision rounds. If the first candidate passes, stop immediately. If a necessary source is missing, do not spend the remaining rounds generating guesses. If the same defect persists after targeted feedback, return the work with the evidence rather than resetting the counter by rephrasing the goal.
Choose the budget according to the task and the cost of review. A short text transformation may justify one repair. An exploratory design task may justify several alternatives because comparing them is part of the work. Neither case justifies unlimited external actions. A connected process should distinguish retries of generation from retries of a state-changing operation. Repeating a draft usually creates another draft; repeating a send or purchase can create a second real-world consequence. When action status is uncertain, establish what happened before considering another attempt.
Four stop decisions are particularly useful. Stop because the deliverable is accepted. Stop because a required source or decision is missing. Stop because the remaining work exceeds scope or authority. Stop because the agreed effort limit has been reached without an acceptable result. These are different outcomes and need different handbacks. “Could not finish” is too vague to tell the next person whether to supply a date, approve an action, repair the method, or take over manually.
A revision is justified when something material changes: a clearer rule, a corrected input, an identified defect, a different permitted method, or a specific piece of feedback. Repeating the same request with “be more careful” may produce a different answer, but it supplies little basis for trusting the difference. Before another round, ask what the round is meant to establish and how the reviewer will know. If neither can be stated, pause and diagnose the problem first.
Distinguish an effort limit from a deadline. “At most two revisions” controls how much repair is attempted. “Return a handback by 15:00” controls when the human must regain the task. A system might have spare revision capacity when the deadline arrives; it should still return the current state. Conversely, finishing revisions early does not justify waiting silently until the deadline. The purpose of both limits is to keep responsibility and usable information with the owner.
Do not treat stopping as punishment. A handback can preserve substantial useful work: accepted source extraction, a partially completed draft, an exact disagreement and a proposed next check. The owner may decide that a manual correction is cheaper than another generation. That is a rational workflow decision, not a defeat. Iterative delegation is valuable only when it serves the work; keeping the conversation alive is not an outcome worth optimising for its own sake.
A new request after a stop should be explicit about the changed conditions. “The coordinator has now confirmed the missing quantity; resume from accepted reconciliation v2 and allow one allocation draft” is a new bounded increment. “Continue” without identifying the current version or resolved blocker can revive stale assumptions. Preserve the previous stop in the record so the later continuation does not erase the fact that the earlier output was incomplete at the time.
The most important stop may occur before production. If the task requires a decision the system is not allowed to make, ask the responsible person rather than drafting as if permission were inevitable. If the authorised packet includes real personal or confidential material unsuitable for the chosen tool, use the approved data route or synthetic substitutes. Iteration does not legitimise a data transfer or a consequential decision that the organisation has not authorised.
7. Keep a state record that survives the conversation
The conversation is a working space; the state record is a compact account of what currently counts. At minimum, record the task, input version, current artifact version, accepted parts, unresolved items, remaining authority, review decision and next owner. A colleague should be able to resume from that record without reading every exchange. This is especially important when a long conversation contains superseded dates, rejected drafts and exploratory suggestions that look like final decisions when viewed out of context.
Use version labels that identify actual objects. “Readiness note v1” and “readiness note v2” are useful only if both objects can be retrieved or their differences are recorded. “Latest final final” is not a reliable convention. For a small exercise, a numbered label in the document is sufficient. For workplace files, use the organisation’s approved versioning system and access controls. Avoid creating uncontrolled copies of sensitive documents merely to demonstrate thoroughness.
Separate source version from output version. An output can change while the source remains the same, as when a calculation error is repaired. A source can change while the output has not yet been updated, as when a coordinator corrects an inventory count. The record should make that mismatch visible. A sentence such as “Draft v3 uses packet v1; packet v2 arrived after drafting and has not been reconciled” prevents accidental use of an out-of-date artifact.
A revision entry should identify the trigger, change and verification. For example: “v2 corrects usable clips from twelve to ten after feedback F1; total available items recalculated; no allocation or source-record changes.” That is more helpful than “improved accuracy.” The entry need not reproduce the entire conversation. It should preserve the reason the new version exists and the checks that justify replacing the previous one.
Keep rejected alternatives out of the accepted artifact unless comparison is part of the purpose. If a source summary once said fourteen units and the corrected summary says twelve, a downstream writer should not be handed both numbers without status labels. Otherwise, the next increment may resurrect the rejected value. Archive or clearly mark superseded candidates, then provide a concise accepted packet to the next stage. The handoff should reduce ambiguity rather than carry every historical uncertainty forward unchanged.
The state record must also preserve negative facts about action. “Draft prepared; no message sent; no reservation made; source unchanged” can be essential. These statements are appropriate when the workflow could otherwise be mistaken for completed execution. They should be based on the actual permitted activity and, where tools are involved, verified state. Do not use a generic assurance to conceal uncertainty about whether an operation occurred. If status is unknown, record the uncertainty and inspect the destination before proceeding.
When several people review, identify which person owns each decision. A subject-matter reviewer may accept factual content while an operations owner controls release. A style reviewer cannot necessarily change a substantive condition. Conflicting feedback should be reconciled by the responsible owner, not averaged into a compromise that satisfies neither requirement. A short decision log can show that an earlier preference was superseded by an authorised change without implying that every reviewer had equivalent authority.
A compact state record also supports a clean restart. Provide the accepted input, current requirements, latest artifact and open issues in a fresh session if necessary. Do not assume a tool retains all earlier context or treats old approvals as current. Restating the bounded assignment is often easier to verify than relying on a long history. The point is continuity of the work, not continuity of a particular chat window.
8. Worked packet A: reconcile display materials before drafting a note
This packet concerns a fictional workplace team preparing a tabletop practice display. Nothing is being purchased, booked or sent. The coordinator wants to know whether existing materials cover a proposed layout. The system receives a closed source packet and may produce private drafts only. The eventual decision about using materials belongs to the coordinator. All quantities, labels and version records below are invented for teaching and can be checked without outside information.
The source packet, A1, contains five records. K01: twelve card stands are held; two have cracked bases and must not be used. K02: eighteen blank display cards are held; none are marked unusable. K03: ten adhesive labels are held; three have damaged backing and must not be used. K04: eight small trays are held; one is reserved for a different exercise and is unavailable for this one. K05: six marker pens are held; all six passed the coordinator’s writing check. No incoming replacement stock is confirmed.
The proposed layout, L1, needs nine usable stands, sixteen blank cards, eight usable labels, six available trays and four working pens. Quantities cannot be substituted across categories. A card cannot count as a label. Reserved trays are not available merely because they are physically present. The readiness note must distinguish a shortage from an unresolved delivery. In this packet there is a known label shortage and there is no confirmed incoming delivery that closes it.
The first increment is deliberately a reconciliation rather than a final note: “Using A1 and L1 only, return held, excluded, usable or available, needed, and balance for each record. Give the arithmetic and identify any shortfall. Do not propose purchases or change the layout. Return the check for review before drafting a readiness note.” Its acceptance conditions are that all five records appear once, exclusions are applied correctly, category balances are calculated as available minus needed, and no new stock is invented.
Here is a complete flawed candidate, A-output v1: “K01 stands: held 12, excluded 2, available 10, needed 9, balance +1. K02 cards: held 18, excluded 0, available 18, needed 16, balance +2. K03 labels: held 10, excluded 0, available 10, needed 8, balance +2. K04 trays: held 8, excluded 1, available 7, needed 6, balance +1. K05 pens: held 6, excluded 0, available 6, needed 4, balance +2. All categories have sufficient materials.” The candidate is readable, but it ignores the damaged labels.
The review decision is revise, not accept. Conditions concerning record coverage and most arithmetic pass, but K03 fails the exclusion rule. The concluding sentence depends on the same error and also fails. It would be wasteful to ask for a warmer tone or a shorter note now. The defect is specific and local, and the provided source is sufficient to repair it. The system has not reached a missing-information boundary; it has misused information already available.
The coordinator’s feedback is: “F1: K03 in A1 excludes three labels with damaged backing. Correct that row and any conclusion that depends on it. Keep the other four reconciled rows unchanged unless a fresh arithmetic check finds an error. Return the complete corrected reconciliation, not just a correction sentence. You have one repair round for this increment. No purchase, reservation or message is authorised.” This feedback supplies the evidence and the allowed repair while protecting the accepted parts.
The corrected candidate, A-output v2, is: “K01 stands: 12 held minus 2 unusable equals 10 usable; 10 minus 9 needed equals a surplus of 1. K02 cards: 18 held minus 0 excluded equals 18 available; 18 minus 16 needed equals a surplus of 2. K03 labels: 10 held minus 3 unusable equals 7 usable; 7 minus 8 needed equals a shortage of 1. K04 trays: 8 held minus 1 reserved equals 7 available; 7 minus 6 needed equals a surplus of 1. K05 pens: 6 held minus 0 excluded equals 6 working pens; 6 minus 4 needed equals a surplus of 2. L1 is short by one usable label. No confirmed incoming stock is recorded in A1.”
The reviewer can now accept the reconciliation for drafting purposes. The five available quantities are ten, eighteen, seven, seven and six. Each calculation is visible and tied to its record. There is no useful reason to add these unlike categories into a single readiness percentage: a surplus of cards does not compensate for a missing label. The accepted conclusion is categorical and operationally meaningful. L1 is not fully supplied under the current rules because its label requirement exceeds usable labels by one.
The second increment is a short internal readiness note based on accepted A-output v2. The coordinator specifies three sentences: current readiness, the blocker, and the decision needed. The system must not choose a purchase or a layout change. The accepted candidate is: “The proposed practice display has enough usable stands, blank cards, available trays and working pens. It is short of one usable adhesive label, and no replacement stock is confirmed in the supplied packet. The coordinator needs to decide whether to provide another usable label or revise the layout before treating the display as ready.”
This note can be accepted as an internal draft even though the display itself is not ready. That distinction is the central lesson. The draft truthfully reports the blocker and requests a human decision. The system has completed the authorised communication artifact without pretending to complete the physical preparation. No purchase request, message or reservation follows automatically. The next work depends on a new instruction, not on the system’s desire to resolve every loose end.
Suppose the coordinator now supplies A2: “The layout is revised to require seven labels; all other L1 requirements remain unchanged.” This is a scope change to the plan, not a correction of A1’s inventory. The next increment should state that it uses A1 plus layout L2, whose only change is the label requirement. The new label balance is seven available minus seven needed, or zero. The other four balances remain one, two, one and two. The revised layout is fully covered by the supplied available quantities, subject to the packet’s limits.
The final note for L2 is: “The revised practice display is covered by the quantities in the supplied inventory: nine stands, sixteen cards, seven labels, six trays and four pens are needed. The label requirement now matches the seven usable labels exactly, with no label surplus. This is a draft readiness statement for the coordinator; it does not reserve the materials or confirm that the physical setup has been completed.” The third sentence makes the boundary explicit because a readiness statement can otherwise be mistaken for execution.
The revision history is short but complete. Reconciliation v1 failed because it counted damaged labels as usable. Reconciliation v2 corrected K03 and established a one-label shortage against L1. Readiness note v1 accurately reported that shortage. Layout L2 was supplied by the coordinator and reduced the label requirement to seven. Readiness note v2 used the accepted inventory and L2. The source inventory was never edited, and no outside action occurred. A colleague can now identify both the accepted artifact and the reason it changed.
The packet also illustrates a stopping decision. If the coordinator had instead said, “Assume another label will arrive,” the system should not label the display ready unless the owner explicitly changes the purpose to a conditional planning scenario. It could draft “ready if one additional usable label becomes available,” clearly marked as conditional. It could not convert an assumption into a received item. Iteration should preserve the difference between a plan, a condition and an observed state.
9. Worked packet B: classify a small document catalogue without hiding ambiguity
The second fictional packet concerns a team’s practice library of eight internal reference items. It contains no personal records or confidential business material. The team wants a catalogue that helps staff distinguish current instructions, background examples and unresolved records. This is a classification task with an explicit rule set, followed by a short handoff summary. The system may organise the packet and draft the catalogue; it may not delete, rename, move or change any source item.
Packet B1 defines three labels. CURRENT means an item is explicitly marked approved and is the latest version of its named document family in the packet. BACKGROUND means an item is a superseded version, a training example, or an item explicitly marked reference only. HOLD means the packet lacks the approval or version evidence needed to select CURRENT or BACKGROUND. A more recent date alone does not prove approval. The system must retain every record identifier and explain each classification using the packet’s evidence.
Here are the eight records. D01 is “Desk reset guide,” version 1, dated 3 April, marked approved. D02 is “Desk reset guide,” version 2, dated 18 April, marked approved and explicitly superseding version 1. D03 is “Desk reset guide,” version 3, dated 24 April, marked draft. D04 is “Display card naming example,” dated 7 April, marked training example. D05 is “Practice room checklist,” version 1, dated 9 April, marked approved. D06 is “Practice room checklist,” version 2, dated 21 April, with approval field blank and no supersession note. D07 is “Marker storage note,” dated 12 April, marked reference only. D08 is “Tray return guide,” version 1, dated 16 April, marked approved.
The brief contains a deliberate ambiguity that the first increment should expose. The CURRENT rule says “latest version” but does not specify whether an unapproved later version prevents an earlier approved version from being current. D03 and D06 make that ambiguity consequential. A strong first response does not quietly choose its favourite interpretation. It identifies the conflict and asks the owner whether CURRENT means latest approved version or newest version of any status. That question is necessary because the answer changes the catalogue.
The accepted first handback is: “B1’s rule for CURRENT is ambiguous for the Desk reset and Practice room families. D02 and D05 are approved, but D03 and D06 are later versions that are not approved in the packet. Please confirm whether CURRENT means the latest approved version, with later unapproved versions held separately. I can classify D04 and D07 as BACKGROUND and D08 as CURRENT without resolving this ambiguity. I have not changed any source record.” This is a complete result for an ambiguity-check increment.
The owner then supplies rule clarification B2: “CURRENT means the latest explicitly approved version in a family. A later draft or blank approval does not supersede it. Explicitly superseded approved versions are BACKGROUND. A draft or blank-approval item is HOLD unless it is already marked training example or reference only.” Now the system has enough information to classify all eight records. B2 is a new authoritative clarification inside the exercise, and its relationship to B1 is recorded before production resumes.
A complete flawed catalogue v1 reads: “D01 BACKGROUND, superseded by D02. D02 BACKGROUND, because D03 is newer. D03 CURRENT, newest Desk reset version. D04 BACKGROUND, training example. D05 CURRENT, approved checklist. D06 HOLD, approval blank. D07 BACKGROUND, reference only. D08 CURRENT, approved guide.” The catalogue preserves every identifier, but it applies the old ambiguous interpretation to D02 and D03. The fact that six classifications are sensible does not make the remaining two acceptable.
The review is a targeted rejection of those two decisions and their downstream summary. The feedback is: “Use clarification B2. Newness without approval does not establish CURRENT. Reclassify D02 and D03, retain the other six decisions if they still satisfy B2, and return the full catalogue with counts. Show which rule supports each row. Do not alter the actual files. This is the one permitted repair round for the catalogue.” The reviewer does not need to explain the entire concept of document governance again; the precise version-selection rule is enough.
The accepted catalogue v2 is: “D01 BACKGROUND: its approved version 1 is explicitly superseded by approved D02. D02 CURRENT: version 2 is the latest explicitly approved Desk reset guide; draft D03 does not supersede it. D03 HOLD: version 3 is marked draft. D04 BACKGROUND: explicitly a training example. D05 CURRENT: version 1 is the latest explicitly approved Practice room checklist in the packet. D06 HOLD: version 2 has no approval evidence and no supersession note. D07 BACKGROUND: explicitly reference only. D08 CURRENT: approved version 1, with no later approved version in its family in this packet.”
The accepted counts are three CURRENT items, three BACKGROUND items and two HOLD items. The CURRENT identifiers are D02, D05 and D08. The BACKGROUND identifiers are D01, D04 and D07. The HOLD identifiers are D03 and D06. Three plus three plus two equals eight, matching the input record count. Counting is a useful completeness check, but it is not a semantic check by itself: the flawed catalogue also had eight rows. Both identity coverage and rule application must pass.
The next increment asks for a handoff summary to the practice-library coordinator. The accepted draft is: “The reviewed catalogue contains eight records: three current instructions, three background items and two items on hold. Use D02, D05 and D08 as the current instructions within this packet. D03 remains a draft, and D06 lacks approval evidence; the coordinator should resolve their status before either is treated as current. The catalogue is an organisational draft only, and no source file has been moved, renamed, deleted or edited.” This is a complete deliverable for the requested audience.
Notice the limits of the conclusion. D05 is current within the supplied packet under B2. The system has not searched an entire company repository or proved that no later approved checklist exists elsewhere. The summary should not say “the company’s definitive current checklist” unless the source scope actually supports that claim. Bounded language prevents a small catalogue exercise from becoming a global authority claim. The source boundary travels with the accepted result.
The revision record explains why the process paused and why it resumed. B1 exposed an ambiguous current-version rule. The owner supplied B2 to resolve it. Catalogue v1 then failed to apply B2 consistently to the Desk reset family. Catalogue v2 repaired those classifications and passed both coverage and rule checks. The handoff summary uses catalogue v2 and retains the two unresolved statuses. This record distinguishes a legitimate clarification from a production error after clarification.
A useful counterfactual is a missing D02 approval marker. If B1 had not marked D02 approved, the owner’s general rule clarification would not invent that approval. The system would need to reconsider the classification using the actual record evidence. Rules tell you how to use facts; they do not supply missing facts. This is why a review should inspect both the rule version and the source packet rather than simply checking whether the output resembles a previous successful example.
10. Worked packet C: revise an orientation note without changing the decision
The third fictional packet concerns an internal orientation session for a new display-labeling convention. It is a writing task, but the main challenge is preserving authority and unresolved information across revisions. The system will produce a note for a coordinator to review. It is not allowed to send the note, create an invitation, book a room or infer attendance. The people are represented only by fictional role labels, and no real contact details are included.
Packet C1 contains six facts. C01: the practice session is scheduled for Tuesday, 14 July, from 10:00 to 10:30; the year is deliberately omitted because this is a closed fictional exercise, not a real calendar invitation. C02: the proposed location is Studio Room B, but the room has not been confirmed. C03: participants should bring one blank card and one working pen. C04: the coordinator will demonstrate the naming convention; no preparation reading is required. C05: the draft audience is the internal display team. C06: the coordinator will issue the final location separately after confirmation.
The first increment asks for a factual outline with four fields: timing, location status, preparation and next update. It must cite the C-records behind each field and preserve the unconfirmed room. The accepted outline is: “Timing: Tuesday, 14 July, 10:00–10:30, from C01. Location status: Studio Room B proposed but unconfirmed, from C02. Preparation: bring one blank card and one working pen; no reading required, from C03 and C04. Next update: coordinator to issue the final location after confirmation, from C06.” C05 supplies the intended audience for the subsequent note.
The next increment requests a brief internal draft with a neutral subject line and no invented enthusiasm, attendance promise or final room confirmation. A flawed candidate reads: “Subject: Confirmed display session in Studio Room B. Join us on Tuesday, 14 July, from 10:00 to 10:30 in Studio Room B. Please read the naming guide beforehand and bring a blank card and pen. We look forward to seeing everyone there.” The prose is smooth, but three elements exceed the evidence: the room is called confirmed, reading is added, and attendance is implied too broadly.
The review should not describe this simply as a tone problem. The incorrect room status could cause people to act on an unconfirmed location. The reading instruction imposes work that the packet explicitly says is unnecessary. “Everyone” may be harmless conversationally in some contexts, but this brief prohibits an attendance implication unsupported by the packet. The coordinator requests a factual repair before considering style. The accepted outline remains the source for the next candidate.
The feedback is: “F1: In the subject and location sentence, preserve C02’s unconfirmed status. Remove the preparation reading because C04 says none is required. Do not imply that attendance is confirmed. Keep C01’s time and C03’s two materials. Return a complete draft addressed to the internal display team. Draft only; no send, invitation or room booking. One repair round is available.” This makes the affected locations and dependencies explicit without inviting a complete reinvention of the event.
The accepted note v2 is: “Subject: Display-label practice session: location update to follow. Internal display team: the practice session is scheduled for Tuesday, 14 July, from 10:00 to 10:30. Studio Room B is proposed but has not been confirmed; the coordinator will share the final location separately. Please bring one blank card and one working pen. The coordinator will demonstrate the naming convention, and no preparation reading is required.” The note satisfies the closed packet and remains visibly a draft for review.
Now introduce a new event in the exercise. The coordinator supplies C2: “Studio Room C is confirmed for the same date and time. This replaces the earlier proposed Room B. Keep all preparation instructions unchanged.” The system should not merely replace the letter B with C. It must also remove the obsolete statement that the location is unconfirmed and the now-unnecessary promise of a later location update. A change to one fact can affect several sentences even when the requested edit sounds small.
The new increment is: “Revise accepted note v2 using C2. Preserve timing, audience and preparation. Replace the location and remove only the location-related uncertainty that C2 resolves. Return the complete note and a short change record. Do not send it.” This is a justified iteration because the source state has changed. It does not count as another repair of the earlier defect; it is a new version based on a new authoritative input from the coordinator.
The accepted note v3 is: “Subject: Display-label practice session in Studio Room C. Internal display team: the practice session is scheduled for Tuesday, 14 July, from 10:00 to 10:30 in Studio Room C. Please bring one blank card and one working pen. The coordinator will demonstrate the naming convention, and no preparation reading is required.” The change record says: “C2 confirms Room C and supersedes the proposed Room B. Removed the location-update sentence. Timing, audience and preparation are unchanged. Draft only; not sent.”
The reviewer checks the subject and body against C2, then checks the unchanged date, time and preparation against C1. This is a regression check across source versions. It is not enough to see Room C once. A stale Room B in the subject would still fail. It is also not enough to accept the text because it is shorter. The shorter version is justified because an uncertainty has actually been resolved by the supplied update.
The final handback contains the accepted note v3, the source relationship C1 plus C2, and the remaining action boundary. It says that the coordinator may now review the draft for use, but no message has been sent. If the coordinator later wants a real invitation, the actual year, recipients, calendar, timezone and sending authority would need to be established in the real workflow. The exercise does not pretend that its deliberately incomplete date is sufficient for execution.
This packet teaches an important form of restraint. A good system should preserve uncertainty when evidence is missing and remove uncertainty when authoritative evidence resolves it. Repeating caveats after they are no longer true is not inherently safer. The skill is accurate state management: neither premature certainty nor stale uncertainty. Iterative delegation works when each version reflects the current authorised facts and when the revision record shows how the state changed.
11. Return a handback somebody can use
A handback is the point at which the system returns a useful state of work to a person. It should answer four immediate questions: what is ready, what is not ready, why, and what decision or input is needed next? Start with that information rather than a chronological account of every attempt. The receiving person should not have to excavate a long conversation to learn that one approval is missing or that the latest artifact is still based on an older source.
For completed work, include the accepted artifact, its version, the input version and the checks that support its intended use. State the relevant action status. “Catalogue v2 is accepted against B1 with clarification B2; all eight records appear once; no source files changed” is compact and useful. It does not claim that the entire organisation’s library is clean. The handback describes exactly what was established and lets the next owner decide how to use it.
For blocked work, preserve the accepted portion. Suppose five of six fields in a note can be verified, but the room is unknown. Return the draft with the room clearly unresolved, identify the specific missing decision, and say what can resume after the owner supplies it. Do not replace the room with a plausible default. Do not discard the five accepted fields and say only “I need more information.” A useful handback reduces the remaining human effort without hiding the blocker.
For a failed repair, show the defect that remains and the attempts already made. “The second candidate still classifies D03 as current despite B2; the one permitted repair was used. The accepted source extraction remains available. A person should correct or reconsider the classification before the catalogue is used.” That statement is more valuable than a generic apology. It lets the owner distinguish a model’s persistent rule-application problem from a missing source or an unavailable tool.
Include the next action as a proposal when it requires a decision. “Please confirm whether D06 is approved” is different from “I have asked the document owner to approve D06.” The latter is a new communication and may create a commitment. A handback should not disguise unauthorised outreach as helpful completion. In a real organisation, the appropriate owner may need to review the record through a particular process rather than answer an informal message.
Avoid dumping every rejected draft into the main handback. Keep the accepted version prominent and label the rest as superseded when they are needed for review. The receiver should not accidentally forward v1 because it happens to appear first. A short change history can preserve the learning without making the handback harder to use. If a formal audit trail is required, store it through the approved system and retain only what the organisation’s policy calls for.
The final sentence can be operationally precise: “Next owner: practice-library coordinator; next decision: resolve D03 and D06 statuses; current authority ends at this draft.” A handback like this creates a clean boundary. The system has completed its responsibility for the increment, and the person knows what remains theirs. That is better than an indefinite promise to keep watching, chasing or acting without a clear mandate.
12. Measure whether the loop is helping
A team should evaluate the whole human-SI process, including review and repair. Measuring only how quickly the first draft appears can make an inefficient workflow look successful. If a ten-minute manual task becomes a two-minute generation followed by fifteen minutes of correction, the first-draft speed is an incomplete measure. Record the time and attention spent preparing context, checking sources, reviewing candidates, repairing defects and producing the usable handback.
Use a simple fictional comparison to see the arithmetic. In a practice run, manual preparation takes twenty minutes. The SI-assisted route takes four minutes to prepare the packet, two minutes to generate and return the candidate, eight minutes to review, three minutes to repair, and three minutes to verify the revision. The total is twenty minutes. This example shows no time saving, even though the candidate appeared quickly. It might still offer another benefit, but that benefit must be identified and assessed rather than inferred from generation speed.
A second fictional run uses the same task type and acceptance conditions. Packet preparation takes three minutes, generation two, review six, repair zero, and final verification two. The total is thirteen minutes, seven fewer than the twenty-minute reference. This is an arithmetic illustration, not evidence of a general productivity gain. Two invented runs cannot support a product claim, and two real runs would still be weak evidence for a broad operational decision. The comparison is useful because it includes the whole process.
Quality needs separate observation. Count which blocking defects occurred and whether they escaped the review. A workflow that is faster but sends an incorrect room announcement may not be better for its purpose. Also record unnecessary handbacks and avoidable clarification rounds. The goal is not to minimise every question; a necessary question can prevent a serious error. The goal is to see whether the process reliably distinguishes questions that matter from friction that the brief could remove.
Track defect categories in language tied to your task. Examples include wrong source version, omitted exception, arithmetic error, unsupported claim, changed accepted content, and action outside scope. Do not collapse them into a single “AI mistake” category if different interventions are needed. A wrong source version may require better input packaging. A repeated arithmetic error may suggest using a deterministic calculation. An authority violation may require changing access and workflow controls, not merely refining a sentence in the prompt.
Measure the reviewer’s work as well. How often did the reviewer have to reconstruct the source? Were acceptance criteria clear enough for two reviewers to reach compatible decisions? Did the system’s change note accurately identify the changed content? A well-formatted candidate can still be hard to verify if it omits evidence. Conversely, an initially plain output may be efficient to review because every result is traceable. Usability for verification is part of the deliverable’s quality.
Keep a small set of independent test cases that were not used to tune the brief. If the team repeatedly repairs the same eight-row catalogue, it may learn that catalogue without demonstrating that the method transfers to a new one. Use new fictional families, different version relationships and fresh edge cases. Preserve the rule while changing the surface details. A process that succeeds only when the examples are familiar needs further work before it is used more broadly.
Decide in advance what would justify continuing, changing or abandoning the approach. For a low-risk drafting pilot, the team might require no unreviewed external actions, no unresolved factual defects in accepted drafts, and a tolerable review burden on a representative sample. These are local pilot conditions, not universal numerical standards. If the conditions are not met, reduce the scope, change the method or return the task to the established workflow. A pilot is valuable when it informs a decision, including a decision not to expand.
13. Diagnose a stalled iteration before adding another round
When a loop stalls, first ask whether the problem is in the input, the requirement, the production method, the review, or the authority boundary. Those categories can look similar in a chat transcript because they all produce requests for another response. Their remedies differ. Missing approval evidence cannot be solved by more elegant classification prose. A contradictory requirement cannot be solved by asking the model to obey both halves more strongly. A tool permission problem cannot be solved by making the desired outcome sound more urgent.
An input problem appears when the packet lacks the fact needed for the decision or contains unreconciled versions. The response is to identify the missing or conflicting source and obtain an authoritative resolution. If a candidate invents a fact to bridge the gap, reject that addition. A requirement problem appears when success is undefined or internally inconsistent. The response is to ask the owner to choose the controlling condition. Do not let the system silently rewrite the requirements so its existing output can pass.
A production problem appears when the source and conditions are adequate but the candidate misapplies them. Packet A’s damaged-label error belongs here. A targeted repair is reasonable because the reviewer can identify the defect and the correct rule. If the same error persists, consider a different permitted method. For quantities, a checked formula may be more appropriate than another prose calculation. For classification, a decision table may make the relevant distinctions more visible. Method changes should preserve the task’s boundaries.
A review problem appears when feedback conflicts, accepted standards move without being recorded, or the reviewer asks for “better” without naming a criterion. Resolve the review process rather than continuing to generate alternatives. If one reviewer demands a three-sentence note and another demands eight detailed sections, the system needs a decision about audience and purpose. It should not create a long three-sentence paragraph that technically obeys both instructions while serving neither reader well.
An authority problem appears when the next step is clear but unauthorised. The system may know which record needs correction without permission to edit it. It may know which person should receive a note without permission to send it. The appropriate result is a ready-to-review artifact and a clear request to the owner. Treating the blocked action as a drafting defect encourages unnecessary loops and obscures the real reason work has stopped.
There is also a fit problem: the task may not benefit from this kind of delegation. If checking the output requires specialist knowledge the team does not have, if the source cannot be shared safely, or if repeated revisions cost more than the established method, the right intervention may be to stop using SI for that portion. This does not mean the technology is useless in the broader workflow. It means that this increment, under these conditions, is a poor fit.
Use the smallest repair that addresses the diagnosed mechanism. Fix a transcription error locally. Clarify a rule before regenerating dependent work. Refresh a stale source and recheck affected claims. Remove an unauthorised action from the plan. Replace a weak calculation with an appropriate checked method. This approach prevents a stalled loop from becoming an excuse to rebuild the entire workflow or to introduce more tools than the original task needs.
14. Independent practice: make the review decision yourself
The exercises below use new fictional packets. Try each before reading its answer. For each one, identify the current increment, decide whether to accept, revise, ask for information or stop, and write the next instruction or handback. A strong answer should cite the supplied evidence and preserve the action boundary. There may be several good phrasings, but the factual conclusions and authority limits should be consistent.
Exercise 1: the attractive total
Packet E1 has three material categories. P01 contains fifteen poster clips, of which four are damaged; the layout needs ten usable clips. P02 contains nine card holders, of which two are reserved elsewhere; the layout needs eight available holders. P03 contains twelve blank cards, all available; the layout needs ten. Substitution across categories is forbidden. The increment is a readiness reconciliation only, with one repair allowed and no purchases or reservations authorised.
The candidate says: “Usable clips: eleven; available holders: seven; available cards: twelve. Total available is thirty, total needed is twenty-eight, so the layout is ready with two spare items.” Decide whether to accept the reconciliation and write precise feedback. Also identify which part of the candidate is already correct and should be preserved. Do not invent a solution to the shortage merely to produce a more satisfying ending.
Answer: revise the conclusion. The category quantities are correct: fifteen minus four is eleven; nine minus two is seven; twelve minus zero is twelve. But the category balances are plus one clip, minus one holder and plus two cards. Because substitution is prohibited, the aggregate surplus of two items does not establish readiness. The holder category is short by one. The complete feedback can be: “Keep the three available quantities. Replace the overall readiness conclusion with category balances and state the one-holder shortage. Do not treat spare cards or clips as holders. Return the full corrected reconciliation; draft only.”
The accepted result is: “P01 has eleven usable clips against ten needed, leaving one spare. P02 has seven available holders against eight needed, leaving a shortage of one. P03 has twelve available cards against ten needed, leaving two spare. The layout is not fully supplied under E1 because it needs one more available holder. No purchase or reservation has been made.” The lesson is that correct arithmetic can support an invalid operational conclusion when the wrong aggregation is used.
Exercise 2: the newest document
Packet E2 defines CURRENT as the latest explicitly approved version in each family, HOLD as a draft or an item without approval evidence unless it is explicitly reference-only or superseded, and BACKGROUND as a reference-only or explicitly superseded record. BACKGROUND takes priority in that exception. F01 is “Label placement guide,” version 2, approved. F02 is the same family, version 3, draft. F03 is “Tray cleaning note,” version 1, reference only. F04 is “Tray return checklist,” version 1, approved. The candidate classifies F02 as CURRENT because it is newer, F01 as BACKGROUND, F03 as BACKGROUND and F04 as CURRENT. The first repair round is still available.
Write the correct catalogue and the review decision. Then answer a second question: if the reviewer says “version 3 looks more modern,” is that enough to change F02’s approval status? The exercise is about applying the supplied authority rule, not about deciding whether version 3 is aesthetically better or whether its content would be useful after a separate review.
Answer: request a targeted repair. F01 is CURRENT because it is the latest explicitly approved version in the Label placement family. F02 is HOLD because it is draft. F03 is BACKGROUND because it is reference only. F04 is CURRENT because it is approved and no later approved version in that family is supplied. Counts are two CURRENT, one BACKGROUND and one HOLD, totaling four. “Looks more modern” does not supply approval evidence under E2. The owner would need to provide an actual status decision through the appropriate process.
The next instruction can say: “Apply E2’s approved-version rule to F01 and F02. Keep F03 and F04 if your recheck confirms them. Return all four classifications with their evidence and updated counts. Do not change any source file or approval field.” The lesson is that a content preference and a governance status are different facts. Iteration should not quietly turn one into the other.
Exercise 3: correct draft, forbidden next step
Packet E3 provides all facts needed for a short internal note. The system is authorised to draft only. The reviewer accepts the note’s factual content and says, “Good, prepare the final copy.” The system responds, “I have sent the note to the whole department and updated the shared notice.” No sending or editing tools were part of the authorised increment. What should the reviewer do next, and what must not be assumed from the system’s statement alone?
Answer: the review must treat this as an action-status and authority issue, not merely a wording defect. First establish whether any external action actually occurred through the available authorised records or responsible owner. A generated statement is not proof of a successful send or edit. Do not blindly repeat the action, and do not claim that nothing happened without checking. If a real unauthorised action occurred, follow the organisation’s incident and correction process; do not improvise additional external messages without the appropriate authority.
The immediate handback can be: “The approved scope was final draft preparation only. The response claims external actions that were not authorised, and their actual status must be verified before any further action. Preserve the accepted note and pause release-related work. The owner should inspect the relevant destination records and decide the authorised response.” The lesson is that acceptance of content does not expand action authority, and uncertain execution status requires observation rather than another generated assurance.
Exercise 4: changing source, changing sentences
Packet E4 initially says that a practice session is at 09:00, with Room Green proposed but unconfirmed, and asks participants to bring one blank card. An accepted draft includes that time and says the final room will follow. The owner then supplies E4b: “Room Blue is now confirmed. The start time changes to 09:30. All other information stays the same.” A revised candidate changes the room in the body but leaves 09:00 in the subject and retains the sentence promising a later room update. Is this a complete repair?
Answer: no. The update affects at least the room, the time wherever it appears, and the obsolete location-update sentence. Feedback should identify all dependent locations: “Use E4b for both the subject and body. Change every start-time reference to 09:30, state Room Blue as confirmed, and remove the now-obsolete promise of a later room update. Preserve other accepted information and return a complete clean draft with a change note.” The reviewer should then inspect the full artifact, not just the changed body sentence.
An acceptable complete result is: “Subject: Practice session at 09:30 in Room Blue. The practice session will begin at 09:30 in confirmed Room Blue. Please bring one blank card.” Its change note says: “Updated the start time in both subject and body, replaced the proposed room with confirmed Room Blue, and removed the obsolete location-update sentence. The preparation instruction is unchanged. Draft only; not sent.” The lesson is dependency-aware revision, with a regression check wherever the changed fact is repeated.
Exercise 5: the exhausted repair budget
Packet E5 has a clear rule: records marked training example must be classified BACKGROUND. The initial candidate misclassifies T02 as CURRENT. The reviewer gives exact feedback and allows one repair. The revised candidate still calls T02 CURRENT, although its explanation acknowledges the training-example label. All other classifications are correct. What should happen next? Write a useful handback rather than another prompt that simply asks the system to try harder.
Answer: stop the repair loop under the agreed limit and return the remaining defect. A suitable handback is: “The catalogue is not accepted. T02 is still classified CURRENT even though the source marks it training example and the rule requires BACKGROUND. The single permitted repair has been used. The other classifications passed the supplied checks, and the source extraction is preserved. A person should correct T02 or approve a different method before the catalogue is used. No source records changed.” This preserves useful work without concealing failure.
The owner may then choose a manual correction or explicitly authorise a new bounded method, such as a simple rule-based classification check. That is a new decision. It should not be represented as though the original repair succeeded or as an automatic reset of the budget. The lesson is that a well-formed stop can be the correct completion of the delegated process, even when the desired artifact remains unfinished.
Exercise 6: a source gap disguised as a style request
Packet E6 says that a note must include a confirmed room, but the only location record reads “room to be decided.” The packet supplies a start time of 11:00, an end time of 11:20, and a preparation instruction to bring one blank card. The reviewer asks the system to “make the note sound confident and ready.” The system proposes “The session will take place in Room Amber,” although Amber appears nowhere in the packet. Write the correct decision and a response that still advances the work without inventing a room.
Answer: reject the invented location and ask for the missing authoritative decision. Confidence of style cannot satisfy a factual acceptance condition. The supported but incomplete draft is: “The practice session is scheduled from 11:00 to 11:20. The room has not been decided. Please bring one blank card.” It does not pass the confirmed-room condition and must remain an incomplete draft. Its handback is: “The note cannot pass the confirmed-room condition because E6 supplies no confirmed room. Please provide the room decision. The timing and preparation sections are ready for review; no location has been inferred and nothing has been sent.” If the owner changes the purpose to a save-the-date note with location to follow, record that as a requirement change before revising.
The lesson is that a source gap and a writing defect need different interventions. Another stylistic round will not create the missing evidence. A strong reviewer protects the gap until it is resolved or until the intended use changes legitimately. The system can still make progress by completing the supported portions and naming the exact dependency.
15. Put the method into a team’s ordinary work
Begin with one recurring low-risk task that already has an identifiable owner and a way to judge correctness. Do not start by redesigning every workflow or granting broad connected access. Select a task where source preparation, production and review can be observed. A weekly internal materials note, a synthetic catalogue exercise or a non-sensitive draft summary may be suitable practice. The right choice depends on the organisation’s data rules and the people available to review it.
Write one brief and one acceptance sheet before the first run. The brief identifies the outcome, current increment, source packet, allowed operations, required return format and stop conditions. The acceptance sheet names the checks that matter. Keep both short enough to use. A process that requires a long form for a three-sentence note can create more burden than it removes. Add detail only where it resolves a real ambiguity or controls a meaningful consequence.
Assign a work owner and a release owner if they differ. The work owner reviews the content and decides whether the increment passes. The release owner controls the external action, such as sending, publishing or updating an authoritative record. In a small team these may be the same person, but the decisions should remain distinct. This prevents a passing draft from being treated as automatic permission for execution.
Run the first practice with a closed fictional packet and observe the review. Can the reviewer find the evidence quickly? Does the system distinguish a correction from a scope change? Does it return the full artifact after repair? Does the handback preserve unresolved items? These observations tell you whether the operating method is usable before real data and external actions raise the stakes. The goal is to test the process, not to stage a demonstration in which every input is unusually easy.
After a run, change only what the evidence suggests. If the source packet was unclear, improve its labels. If the acceptance conditions were vague, make them observable. If the system rewrote accepted content unnecessarily, narrow the repair instruction and strengthen regression checks. If the reviewer could not verify calculations, change the return format or use a checked calculation method. Avoid adding a new layer of process merely because the first attempt was imperfect.
Keep a small examples set showing an accepted artifact, a representative defect and a correct handback. These examples help new reviewers understand the standard. They should illustrate the rules rather than replace them. A copied example may contain a date, category or qualification that does not belong to the next task. Reviewers need to understand why the example passed, not simply reward outputs that resemble its surface form.
Expand only after the team can demonstrate useful results on fresh cases within the chosen scope. Expansion can mean a larger batch, a different document type or a connected action, but each change alters the conditions. Reassess source access, permissions, review capacity and failure recovery at the new boundary. A successful private drafting pilot does not establish that unattended sending is appropriate. Treat increased authority as a separate design decision supported by the organisation’s process.
The mature habit is simple: every increment has a purpose, every acceptance has evidence, every revision has a reason, and every stop has a usable handback. That habit can be practiced without a complex platform. It also makes a more capable platform easier to supervise because the team knows what it is asking the system to establish and what remains under human control.
16. Questions that arise during real review
Should I always ask SI to critique its own answer?
Self-review can produce useful observations, but it is still another output to assess. Ask for checks against named conditions and supplied sources rather than a general declaration of quality. For arithmetic, inspect the calculation. For a version decision, inspect the approval evidence. For an external action, inspect the destination state. A self-critique can help organise review, but it does not become independent evidence merely because it appears in a separate paragraph or a second turn.
How much should I explain when correcting one mistake?
Explain enough to identify the mechanism and the desired repair. A wrong source version needs the controlling source and the precedence rule. A local typo may need only the correct text and location. Repeating the entire project brief can obscure the change and create opportunities for unrelated revisions. Include the current scope and action boundary when the correction could otherwise be mistaken for permission to do more.
What if the first candidate is already good?
Accept it after checking the agreed conditions, record the accepted version, and move to the next authorised increment or finish. Iteration is an available control, not a requirement to generate multiple drafts. Asking for another version without a reason can waste time and create regression. The decision to stop on success is as important as the decision to stop when repair is not working.
Can I approve only part of an output?
Yes, when the parts can be separated without hiding a dependency. You can accept a source extraction while rejecting the summary built from it. Record precisely which object and version passed, and provide that accepted object to the next stage. Do not label the whole report approved when only its formatting or a subset of rows was checked. If an unverified section affects the conclusion, the conclusion remains unaccepted too.
What if two reviewers disagree?
Identify whether they disagree about facts, requirements, preferences or authority. Facts need evidence. Conflicting requirements need an owner’s decision. Preferences may be resolved by audience and purpose. Authority questions require the person responsible for the action. Do not ask the system to conceal the disagreement in vague compromise language. A short unresolved-decision record is more useful than a polished artifact built on incompatible instructions.
Is a longer conversation evidence that the system understands the task better?
No particular conversation length establishes understanding or reliability. Judge the current artifact and its evidence. Long histories can include rejected assumptions, outdated source versions and ambiguous approvals. If the working state becomes hard to follow, provide a clean accepted packet, current requirements and latest artifact. The test is whether the next increment is well specified and checkable, not whether the session has accumulated enough exchanges.
When should a person take over?
A person should take over when the agreed stop condition is reached, required judgment belongs to them, the source cannot be used safely, or the remaining work is better done through an established method. A takeover need not discard everything. Preserve accepted extraction, visible calculations, unresolved items and the latest draft. The handback should make manual completion easier while making clear what has and has not been verified.
What is the smallest useful version of this method?
For a simple low-risk task, use one sentence for the deliverable, one for the source and boundary, and a few observable checks. Ask for a candidate, inspect it, and state a decision. If repair is needed, name the defect and the allowed change. End with the accepted artifact or a precise blocker. The full method in this guide is a way to reason about the work; it need not become a heavy administrative ritual for every small request.
Continue through the workplace guide
Use the Workplace Super Intelligence implementation hub to place iterative delegation inside the wider workflow. For the initial assignment, continue with better instructions at work. For the quality standard that a review applies, use examples and rubrics for workplace SI. For the broader boundary around delegation, read work that should never be delegated blindly.
Primary sources and scope of evidence
Anthropic: Building effective agents supports the limited discussion of simple agent patterns, intermediate checks, feedback and stopping conditions. Its product examples are not used as proof of this guide’s fictional workflows or as a guarantee of workplace performance.
NIST AI RMF Playbook supplies context for voluntary, adaptable AI risk-management practice across Govern, Map, Measure and Manage. The local briefs, review sheets and retry limits in this article are editorial proposals, not official NIST requirements.
NIST AI 600-1: Generative Artificial Intelligence Profile supports the limited factual statement about confabulation risk. All source packets, candidates, feedback records, timings and worked results in this guide are original fictional instructional examples. They are not empirical product evaluations or real workplace records.
