VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How to Improve a Weak Super Intelligence Answer | Diagnose, Repair and Verify AI Responses

eduKate Secondary small-group study for How Super Intelligence Works: Parameters and Weights.
Secondary students checking and improving written work together

How do you improve a weak Super Intelligence answer? Do not immediately ask for “a better answer.” First identify what is weak: wrong facts, missing context, unsupported claims, poor structure, incorrect calculation, unsuitable tone, broken constraints, weak reasoning or an external action that never actually happened.

A weak SI answer is useful evidence. It shows how the current task specification, source set, tool access or verification process failed. The fastest repair is usually targeted: preserve what is already correct, fix the first important failure and test the repair on a fresh case.

This eduKateSG guide explains a systematic answer-repair process for writing, research, learning, data, coding and tool-connected workflows. It follows How to Get Structured Answers From Super Intelligence in Stage 2 of the How to Learn Super Intelligence Quickly curriculum.

Terminology: SI is our editorial term for practical contemporary AI learning. Improving an answer does not mean making it longer or more confident. It means making the work more correct, useful, traceable and appropriate to the task.


The First Principle: Diagnose Before You Regenerate

When a response disappoints you, the temptation is to say “try again”. That can produce a different answer without teaching you why the first one failed.

Diagnosis asks: which requirement was violated first? Was the source missing? Was the instruction ambiguous? Did the calculation use the wrong denominator? Did the answer solve a different problem?

Once the earliest meaningful failure is identified, the repair can target that layer instead of changing everything at once.

Failure Type 1 — Wrong Task

The answer may be competent but solve the wrong problem. This usually begins with an unclear objective.

Example: you ask for “help with a report” and receive a rewritten report when you actually wanted a fact-check. The repair is not stronger writing. It is a clearer task: “Check every numerical claim against the attached source; do not rewrite the prose yet.”

Always distinguish task failure from output-quality failure.

Failure Type 2 — Missing Context

The answer may be generic because the system does not know the audience, source, deadline, current state or constraint that changes the job.

Repair by adding only the missing information. Do not dump every historical detail into the prompt. Use the context-engineering principles from Article 13.

Then rerun the same task so you can observe whether the context change caused the improvement.

Failure Type 3 — Bad or Outdated Source

A beautifully reasoned answer can still be wrong when the source is outdated, incomplete or not authoritative for the claim.

Repair the source layer. Locate the current or appropriate evidence, label versions and make source priority explicit.

Do not try to prompt your way out of a source-quality problem.

Failure Type 4 — Unsupported Addition

SI may fill missing details with plausible content. A summary invents a room, a research answer invents a statistic or an event draft assumes a time.

Repair with explicit unknown handling: “If the source does not state the room, write Unknown. Do not infer it from previous examples.”

Then test with another incomplete source to make sure the behaviour transfers.

Failure Type 5 — Meaning Drift

During rewriting, optional can become compulsory, tentative can become confirmed or a limitation can disappear.

Repair by extracting protected facts and semantic relationships before editing. State them as invariants: may remains optional; must remains required; unknown remains unknown.

Compare the revised output with the original source line by line for high-consequence details.

Failure Type 6 — Missing Required Information

The answer may omit a field or section even though everything it includes is correct.

Repair by using structure: define the required sections or fields explicitly. If the output feeds a recurring workflow, add a schema or checklist.

Do not confuse omission with factual error. The repair mechanism is different.

Failure Type 7 — Wrong Structure

A useful answer can still be difficult to use when the receiver cannot find the information. A manager needs risks and decisions; a student needs explanation and practice; software needs predictable fields.

Repair by changing output structure rather than regenerating the substantive analysis.

Keep the correct content and reorganise it around the receiver’s next action.

Failure Type 8 — Wrong Tone

Tone is a presentation layer. Fix it after meaning and evidence are stable.

Example: a factual parent notice is accurate but too formal. Ask for a warmer rewrite while preserving dates, obligations and uncertainty.

Avoid asking for a complete rewrite before protecting the content that is already correct.

Failure Type 9 — Too Much Detail

An answer can be correct and still unusable because important information is buried.

Repair by ranking information. Ask for a one-paragraph executive summary followed by optional detail, or define a word limit with protected facts.

Compression should not remove evidence or caveats that change the conclusion.

Failure Type 10 — Too Little Detail

A short answer can omit the mechanism needed to understand or verify the result.

Repair by asking for the specific missing layer: formula, evidence, assumptions, example, failure case or next step.

Do not simply ask for “more detail”. Name what kind of detail would make the answer useful.

Failure Type 11 — Weak Reasoning

The answer may jump from evidence to conclusion without showing assumptions or alternatives.

Repair by asking: “List the assumptions required for this conclusion. Give one alternative explanation. Identify what evidence would distinguish them.”

The goal is not to force a longer internal monologue. It is to expose the reasoning artifacts needed for human evaluation.

Failure Type 12 — Wrong Comparison

A comparison may use different criteria for each option or ignore missing information.

Repair by defining common criteria and a structured table. Keep unknown cells visible.

The human should still decide which criteria matter most.

Failure Type 13 — Wrong Calculation

A numerical answer can fail through arithmetic, wrong inputs, wrong units or wrong interpretation.

First check inputs. Then reproduce the operation with a deterministic tool where appropriate. Finally ask whether the calculation answers the actual question.

A correct percentage using the wrong denominator is still a wrong answer.

Failure Type 14 — Wrong Data Interpretation

The computation may be correct while the story is not. A correlation is presented as causation, a mean hides skew or a small sample is generalised broadly.

Repair the interpretation layer. State what the metric establishes and what it does not establish.

Keep calculation and interpretation in separate fields when the workflow repeats.

Failure Type 15 — Weak Source Traceability

An answer cites sources without making it clear which source supports which claim.

Repair with an evidence map: Claim, Source, Passage or Locator, Date and Limitation.

A citation list at the bottom is not enough when the relationship between claim and evidence remains ambiguous.

Failure Type 16 — Wrong Time Frame

Historical evidence can be presented as current, or a current answer may ignore the user’s requested period.

Repair by defining the time window and requiring date labels for current claims.

For fast-changing information, retrieve current evidence rather than relying on remembered knowledge.

Failure Type 17 — Wrong Audience

The answer may assume expertise the receiver does not have or oversimplify material for an expert audience.

Repair by specifying prerequisite knowledge and the receiver’s next action.

Audience is a task constraint, not merely a tone preference.

Failure Type 18 — Tool Not Used

A current or external task may require search, file access, calculation or another tool. If the system answers from memory instead, the result may be weak even though the prose is plausible.

Repair by identifying the required capability and confirming it is available. If the tool is unavailable, change the task or obtain the information another way.

Do not pretend a tool ran when it did not.

Failure Type 19 — Tool Failure Hidden by Prose

The model may continue after a tool error and generate a plausible completion. This is dangerous in operational workflows.

Repair by requiring explicit tool-result states and stop conditions. “If the calendar call fails, report the error and do not claim the event exists.”

Check the destination system for consequential actions.

Failure Type 20 — Overreach Beyond Authority

A response may make a recommendation or action that exceeds the intended human role, organisational permission or professional scope.

Repair by separating preparation, analysis, recommendation and final authority.

The workflow should make human approval points visible rather than treating them as friction to eliminate.

The Repair Loop

  • Preserve the original task and first answer.
  • Identify the earliest important failure.
  • Name the failure category.
  • Locate the layer that owns the failure: task, context, source, prompt, tool, structure, reasoning or verification.
  • Make one targeted change.
  • Regenerate or repair only the affected part where possible.
  • Verify the revised result.
  • Test the fix on a fresh example.
  • Record the lesson if the task recurs.

The loop converts a disappointing answer into a training signal.

Preserve What Is Already Correct

A common repair mistake is rewriting the whole answer. Large rewrites can destroy facts, wording or structure that were already correct.

Ask for a bounded change: “Keep paragraphs 1 and 2 unchanged. Replace only the unsupported claim in paragraph 3 using the supplied source.”

This makes the repair easier to inspect and reduces regression risk.

Use a Diff Mindset

Think in differences between the accepted state and the current state. Which sentence, field or calculation must change? Which parts must remain invariant?

This mindset is valuable in writing, code, data and workflow repair because it limits collateral change.

For important documents, compare versions explicitly.

Repair With Evidence, Not Preference

If you say an answer is wrong, identify the evidence. “The source says Monday, but the answer says Tuesday” is actionable. “This feels off” is weaker until you identify what creates the problem.

For subjective tasks, define evaluation criteria: clarity, audience fit, structure or brand voice.

Evidence makes revision reproducible.

Self-Critique: Useful but Not Independent Verification

You can ask SI to review its answer for unsupported claims, missed constraints or alternative interpretations. This is useful for generating checking questions.

However, another generated response is not independent proof. Important claims still need sources, calculations, tests or appropriate human expertise.

Use self-critique as a diagnostic assistant, not a truth oracle.

A Worked Example: Repair a School Notice

Source: “Students must arrive by 7.30 am. Parents may enter from 8 am. Venue: Hall A.” Weak answer: “Everyone should arrive at 7.30 am in Hall A.”

Diagnosis: parent optional attendance and later entry time were collapsed. Repair: preserve distinct audiences and obligations.

Revised output should state students’ required time and parents’ optional entry separately. Verify all three logistics against the source.

A Worked Example: Repair a Research Summary

Source reports an association between sleep duration and test scores in one school. Weak answer: “More sleep causes better grades for all students.”

Diagnosis: causal overclaim and population generalisation. Repair: “In this school sample, longer reported sleep was associated with higher test scores; the study does not by itself establish causation or universal generalisation.”

The repair makes the conclusion no stronger than the evidence.

A Worked Example: Repair a Percentage

Data: 18 attendees from 24 registrations. Weak answer: 18/30 = 60% because 30 was room capacity.

Diagnosis: wrong denominator. Repair: 18 ÷ 24 × 100 = 75% attendance relative to registrations.

Then explain that room capacity answers a different question.

A Worked Example: Repair Code

A generated function fails on empty input. Weak response rewrites the entire module.

Diagnosis: scope overreach. Repair: reproduce the empty-input failure, add a regression test and make the smallest change that satisfies the specification.

Run the full relevant test suite afterward.

A Worked Example: Repair a Tool Workflow

SI says “Your event has been booked” after the calendar tool returns an error.

Diagnosis: tool-result blindness and false completion. Repair: stop on tool error, report the error and keep the approved event draft available for retry.

Confirm the event in the destination calendar after a successful later tool call.

A Worked Example: Repair Student Feedback

Student makes an arithmetic error. Weak feedback gives a full algebra lesson. Diagnosis: wrong layer and excessive assistance.

Repair: identify the first arithmetic mistake, give one targeted hint and provide a fresh transfer question.

The repair supports learning rather than replacing it.

Repairing Structure Without Rewriting Meaning

If the content is correct but hard to use, extract it into a better structure rather than regenerating the analysis.

Example: convert a long meeting summary into Decision, Owner, Deadline, Risk and Unknowns. Keep every factual statement traceable to the accepted version.

This is often faster and safer than “write the whole thing again”.

Repairing Tone Without Reopening Facts

Lock the factual layer first. Then ask for tone change only.

Example: “Keep all facts and obligations unchanged. Make the language warmer and more concise for parents.”

Review commitments after the edit because tone changes can accidentally create promises.

Repairing Research Without Losing Traceability

When a source changes, update claims that depend on it and preserve unaffected sections.

Do not replace the whole evidence map unless necessary. Mark which claims were revalidated.

This makes maintenance easier in long-lived articles and reports.

Repairing Long Conversations

If the conversation itself has become confused, create a current-state brief: goal, accepted facts, active constraints, rejected ideas, unresolved questions and next step.

Start the repair from that brief rather than trying to correct dozens of old turns.

Article 20 develops this technique for long productive conversations.

The Error Log

For repeated work, keep a small error log: Task, Failure, Cause, Repair, Transfer Result.

Patterns emerge over time. You may discover that most failures come from source versions, not prompting. Or that missing-value handling causes repeated automation errors.

Let the error log choose the next learning target.

Historical Failures Become Regression Tests

When a recurring workflow produces an important failure, save an anonymised or synthetic version as a test case.

After prompt, model or workflow changes, rerun that case. A good repair should prevent the old failure without breaking normal cases.

This turns experience into system memory.

When to Stop Repairing and Restart

Sometimes the answer is so far from the task that targeted revision is inefficient. Restart when the objective changed, the source set was wrong, the conversation contains many contradictory assumptions or the structure is fundamentally unsuitable.

A restart should use a clean brief containing only the current task state.

Restarting is not failure. It is choosing a lower-cost recovery path.

When to Stop Using SI for the Task

If the task cannot be verified, requires expertise you do not have or involves information that should not be shared with the system, route the work elsewhere.

A strong SI user knows when the correct repair is another tool, an authoritative source or a qualified human expert.

Tool choice is part of answer quality.

A Weak-Answer Diagnostic Checklist

  • Did it answer the correct task?
  • Did it use the right source?
  • Is the source current enough?
  • Did it preserve names, dates, quantities and obligations?
  • Did it invent missing information?
  • Did it omit a required field or section?
  • Is the structure usable by the receiver?
  • Are calculations reproducible?
  • Are interpretations stronger than the evidence?
  • Did required tools actually run?
  • Did external actions actually succeed?
  • Did it exceed permission or authority?
  • What is the earliest failure worth fixing?

A Practice Lab: Repair One Answer Three Ways

Take one weak answer. First repair factual fidelity only. Second repair structure only while preserving accepted facts. Third repair tone only while preserving the first two repairs.

Compare the versions. This exercise teaches that quality has layers and that one repair does not require reopening every layer.

The final answer should pass all three checks without losing earlier improvements.

A Practice Lab: Diagnose Before Prompting

Collect five weak outputs from previous work. Do not rewrite them. Label each failure: task, context, source, structure, reasoning, calculation, tool or permission.

Then write the smallest repair instruction for each.

This builds diagnostic fluency and reduces random prompt iteration.

A Practice Lab: Transfer the Repair

After fixing one failure, create a fresh example with different surface details. Apply the same repair method.

If the failure returns, the method may be overfitted to the original case. Improve the invariant or add a clearer constraint.

Transfer converts one correction into a real skill.

Frequently Asked Questions

Should I ask SI to “try again”?

Only when variation itself is useful. For recurring failures, diagnose what was wrong and make a targeted change.

Should I rewrite the whole answer?

Not if large parts are already correct. Bounded repairs are easier to verify and reduce regression risk.

Can SI critique its own answer?

Yes, as a diagnostic aid. But important claims still need independent evidence or testing.

How many revisions are too many?

When revisions stop improving the task or the underlying source/specification is still unclear, stop and repair the task definition or restart from a clean brief.

What if the answer is factually correct but not useful?

Repair receiver fit: structure, prioritisation, tone or next action. Correctness is necessary but not always sufficient.

What if every new revision breaks something else?

Use preservation constraints, diff-based editing and regression tests. The workflow may also need stronger structure or decomposition.

What comes next?

Continue with How to Have Long, Productive Conversations With Super Intelligence, where repair expands from one answer to the state of an entire conversation.

The Answer-Quality Stack

A weak answer can fail at several layers at once. The answer-quality stack helps you inspect those layers in order so you do not polish the surface while leaving the foundation broken.

  • Task: Did the system solve the job the user actually intended?
  • Context: Did it have the information needed to solve that job?
  • Source: Was the evidence current, authoritative and relevant?
  • Reasoning: Did the conclusion follow from the evidence and assumptions?
  • Structure: Can the receiver find and use the required information?
  • Presentation: Is tone, length and wording appropriate?
  • Tool state: Did required external capabilities actually run?
  • Verification: Is there independent evidence that the important result is correct?

Repair from the bottom of this stack upward. If the source is wrong, polishing structure does not help. If the task itself is wrong, better reasoning on the wrong task still produces the wrong result.

Severity: Not All Weaknesses Deserve the Same Repair Effort

Classify weaknesses by consequence. A spelling issue may be low severity. A changed deadline, unsupported legal statement or false external-action claim can be high severity.

Prioritise the failures that would mislead the receiver or create irreversible consequences. Do not spend equal review effort on every sentence.

A practical severity ladder is: cosmetic, usability, factual, decision-critical and action-critical. The stronger the consequence, the stronger the independent check.

The First-Failure Rule

When several errors appear, find the earliest one that explains the rest. A wrong source version can create multiple wrong facts. A vague task can create irrelevant structure and tone problems.

Repairing the root failure may eliminate several downstream symptoms at once.

This is similar to debugging software: fix the upstream cause before patching every downstream error individually.

Answer Repair Versus Prompt Repair

Sometimes the answer needs editing; sometimes the prompt or workflow needs redesign. Distinguish one-off output failure from recurring system failure.

If one sentence is wrong due to unusual wording, repair the answer. If the same type of mistake appears repeatedly across inputs, repair the prompt, examples, schema or source process.

Recurring failure belongs to the system, not only the output.

Answer Repair Versus Source Repair

If the answer reflects an outdated document correctly, the model may not be the main problem. Replace or relabel the source and regenerate only the dependent claims.

Keep source-version information with long-lived workflows. A weak answer caused by stale evidence will keep returning until the evidence layer changes.

This is why source provenance is part of answer quality.

Answer Repair Versus Tool Repair

If a tool call fails or returns incomplete data, rewriting the natural-language answer may conceal the problem. Repair the tool step or the tool-result handling.

Example: a web search returns no current source. The answer should not fabricate one from memory. The repair may be a different search query, another authorised source or an explicit limitation.

Tool failures deserve operational recovery, not rhetorical recovery.

Answer Repair Versus Human Decision

Some apparent weaknesses are actually unresolved value judgments. A system may present two reasonable options and the user dislikes that it does not choose one.

If the evidence does not determine the choice, the repair is not “make the AI decide”. Define the human criteria and preserve the decision owner.

A good answer can remain noncommittal when the task genuinely requires a human trade-off.

The Minimal Repair Principle

Make the smallest change that fixes the diagnosed failure while preserving accepted content. This improves traceability and reduces regression risk.

Example: one paragraph overstates causality. Replace that paragraph rather than rewriting the whole report. One JSON field uses the wrong source. Repair that field and rerun validation.

Minimal repair is especially important in long documents, code and operational workflows where unrelated changes are costly.

The Preservation Set

Before revising, mark what must remain unchanged. The preservation set may include accepted facts, citations, formatting, approved wording, tested code or validated calculations.

State it explicitly: “Keep sections 1–3 unchanged. Preserve all existing citations. Revise only the conclusion to reflect the new source.”

The preservation set protects good work from collateral damage.

The Repair Spec

A repair instruction is stronger when it contains four parts: failure, evidence, required change and preserved content.

Example: “Failure: paragraph 4 says the policy begins 1 October. Evidence: official notice says 15 October. Change: replace the date and update dependent wording. Preserve: all other paragraphs and citations.”

This is more useful than “fix paragraph 4”.

Repairing Unsupported Claims

List the unsupported claim and ask what source would be capable of supporting it. If none is available, remove or qualify the claim.

Do not replace one unsupported claim with a different plausible claim. The repair should improve evidence, not wording alone.

For public content, verify the source date and scope before publication.

Repairing Missing Qualifications

An answer may be broadly correct but omit a condition that changes applicability. Examples include “for this age group”, “under these assumptions” or “in the tested environment”.

Add the qualification close to the claim it limits. Do not bury it in a generic disclaimer at the end.

Precise scope is part of correctness.

Repairing Overly Hedged Answers

The opposite problem also occurs: a system may qualify everything so heavily that the receiver cannot identify the supported conclusion.

Separate what is established from what remains uncertain. State the strongest bounded conclusion first, then list the limitations that genuinely matter.

Useful caution is specific, not vague.`

Repairing Repetition

Long SI answers can repeat the same point under different headings. Repetition increases reading time and can make the answer seem deeper than it is.

Identify the invariant idea and keep the clearest version. Merge duplicated examples or sections.

After compression, confirm that no unique evidence or exception was lost.

Repairing Poor Prioritisation

An answer may include the right information in the wrong order. High-consequence facts are buried beneath background.

Repair by ranking information according to receiver need: decision, critical evidence, risk, unknowns and supporting detail.

Prioritisation is a receiver problem rather than a factual problem.

Repairing Weak Openings

A weak answer may begin with generic background instead of addressing the user’s intent. Rewrite the opening to state the direct answer, task definition or key distinction.

For SEO content, the opening should align with the query while remaining accurate and useful. Avoid keyword repetition that adds no meaning.

The rest of the article can then expand mechanism, examples and edge cases.

Repairing Weak Conclusions

A conclusion should close the task. It can summarise the decision, state the next action or identify the unresolved blocker.

Weak conclusions merely repeat the introduction or add motivational language. Repair by making closure observable.

For educational content, the conclusion can specify the transfer task. For professional work, it can specify the approved next step.

Repairing FAQ Sections

FAQ sections become weak when they repeat material already answered or introduce new claims without evidence.

Keep questions that reflect real user uncertainty and answer them directly. Use FAQ to cover boundary cases, misconceptions and next actions.

Remove filler questions written only to increase page length.

Repairing Tables

A table can hide uneven evidence. One option may have current data while another has guesses, yet both cells look equally authoritative.

Add Evidence and Unknown fields or markers. Keep criteria consistent across rows.

A table is useful only when comparable cells actually represent comparable information.

Repairing Structured Outputs

When one field is wrong, inspect whether the problem is extraction, field definition or schema design. Do not regenerate the whole object automatically.

Check missing-value rules and cross-field consistency. A Confirmed status cannot coexist with Evidence that says “tentative”.

Structured repair should preserve valid fields.

Repairing Numerical Explanations

Separate arithmetic from interpretation. First reproduce the calculation. Then test whether the chosen metric answers the requested question.

Add units, numerator and denominator where ambiguity exists.

If the source values are uncertain, the final numerical result should reflect that uncertainty rather than display false precision.

Repairing Charts and Visualisations

A chart can be visually polished but misleading through scale, missing labels, wrong aggregation or inappropriate chart type.

Repair the data and encoding before aesthetics. Confirm axes, units, categories and denominators.

A visualisation is an argument about data; its design should preserve the data’s meaning.

Repairing Code Explanations

An explanation can sound correct while not matching the actual code. Compare claims with the implementation and tests.

If the explanation says a function validates input but the code does not, repair the code or the explanation depending on the specification.

Never allow comments or generated documentation to become the source of truth when executable behaviour disagrees.

Repairing Generated Tests

Generated tests can mirror the implementation’s assumptions and miss the bug. Add tests derived from the specification and historical failures.

Include boundary and invalid-input cases, not only happy paths.

A passing test suite is meaningful only when the tests represent the intended behaviour.

Repairing Tool-Generated Plans

A plan may assume tools or permissions that are unavailable. Mark each operational step with required capability and permission.

Remove or reroute steps that cannot be executed in the real environment.

A plan should describe the system that exists, not the system the model imagines.

Repairing Agent Traces

If an agent takes an unexpected path, inspect the state transitions and tool results rather than only the final answer.

Identify where the agent’s belief diverged from reality. Add validation or a stop condition at that handoff.

Agent repair is often workflow repair rather than language repair.

Repairing Long Documents With Section Ownership

Assign each section a purpose and source set. When one section fails, repair within its owner boundary where possible.

This reduces cross-document drift. A change to a research section should not silently alter unrelated policy or sales sections.

Section ownership also makes collaborative editing easier.

Repairing Internal Links and Navigation

For published knowledge systems, a content answer can be factually good but structurally weak if it links to wrong, duplicate or unpublished pages.

Repair navigation by preserving one canonical owner per intent, linking only live destinations and updating hubs when new pages publish.

Information architecture is part of answer quality for large content ecosystems.

Repairing SEO Without Weakening the Article

SEO repair should improve discoverability while preserving reader usefulness. Align title, opening, headings and internal links with the real search intent.

Do not add repetitive keyword paragraphs that dilute the mechanism. Search-oriented language should remain natural and specific.

A page that ranks but fails the reader’s task is not a strong long-term asset.

Repairing for Mobile Reading

A long article may be difficult on phones if paragraphs are dense, headings are rare or tables are too wide.

Repair by shortening paragraphs, adding meaningful headings and using lists where scanning matters. Preserve substantive depth while reducing visual load.

Mobile readability is presentation repair, not content thinning.

The Repair Decision Tree

  • Wrong task? → redefine objective.
  • Missing context? → add only task-relevant context.
  • Wrong source? → replace or reprioritise evidence.
  • Unsupported claim? → source, qualify or remove.
  • Wrong calculation? → check inputs, units and operation.
  • Wrong structure? → reorganise accepted content.
  • Wrong tone? → edit presentation after facts are locked.
  • Tool failure? → repair tool step and confirm result.
  • Permission problem? → stop or escalate.
  • Recurring failure? → repair the workflow, not only the answer.

A Repair Scorecard

  • Task fit.
  • Source fidelity.
  • Factual accuracy.
  • Uncertainty handling.
  • Reasoning scope.
  • Numerical validity.
  • Structure.
  • Receiver usefulness.
  • Tool-result integrity.
  • Permission compliance.
  • Verification.
  • Transfer to fresh examples.

Use the scorecard qualitatively. One high-severity failure can outweigh several cosmetic strengths.

A Full Repair Case Study: Research Brief

A brief says a programme is available to all students, cites an old article and recommends immediate application. The current official page actually restricts eligibility.

Diagnosis: wrong source authority, outdated evidence and overconfident action recommendation. Repair source first. Update eligibility. Then revise recommendation to match the user’s status and identify any missing requirements.

Preserve unaffected background sections. Rerun internal links and citations. The final brief becomes both shorter and more accurate because unsupported material is removed.

A Full Repair Case Study: Student Study Plan

A plan schedules two hours every night despite the student having two fixed commitments. It also assigns advanced practice before a prerequisite is stable.

Diagnosis: context omission and dependency error. Repair availability and skill sequence. Preserve useful retrieval-practice elements.

Test the new plan against the actual calendar and a fresh diagnostic question. The answer improves because the task model improves.

A Full Repair Case Study: Operational Automation

An agent creates tasks from meeting notes but sometimes assigns tentative ideas as confirmed work. A surface repair would change wording. The real repair is classification plus approval.

Add Status = Proposed/Confirmed, source evidence and a rule that only Confirmed actions with Owner may be written to the project system.

Test historical tentative phrases and missing-owner cases. The repair becomes part of the workflow, not a one-off answer edit.

A Full Repair Case Study: Public Article

An article has strong content but weak opening, duplicated sections and a header image from the wrong media family.

Repair query alignment in the opening, merge repeated explanations, preserve unique worked examples, verify internal links and replace the image without altering the article’s substantive claims.

Run a final audit for word count, schema, indexability, featured image and canonical slug. Publishing quality includes both content and page state.

A Final Answer-Repair Examination

Take one SI output you consider mediocre. Before editing, write the task, source, receiver and acceptance test. Mark every failure but rank them by consequence.

Fix only the highest-ranked root failure. Re-evaluate the answer. Continue one layer at a time until the acceptance test passes.

Then run the repaired method on a new input. If the same failure returns, update the system-level rule, example, schema or workflow.

The examination is complete when you can explain not only what changed, but why that change repaired the underlying failure without damaging accepted work.

Answer Repair as a Maintenance Discipline

Recurring SI workflows need maintenance because model behaviour, source data, tools and user requirements change. An answer that was acceptable last month can become weak when the context or environment changes.

Treat repair as an ongoing discipline rather than an emergency response. Keep a small set of representative tasks and known historical failures. After significant prompt, model, source or tool changes, rerun them.

Maintenance turns answer quality from a one-time judgement into a controlled process.

Regression Testing for Answers

A regression occurs when a change fixes one problem but breaks something that previously worked. This can happen in prose as easily as in software.

Example: adding stronger brevity instructions stops repetition but causes the system to omit deadlines. Adding more examples improves classification but causes the output format to drift.

Keep tests for protected behaviour: factual preservation, unknown handling, required fields, tool boundaries and receiver usefulness.

Repair Ownership

Every recurring failure should have an owner layer. Source failures belong to content or data maintenance. Prompt failures belong to instruction design. Tool failures belong to integration. Approval failures belong to workflow governance.

Without ownership, the same symptom is repeatedly patched in the final answer while the root cause remains.

Even in a one-person workflow, naming the owner layer improves diagnosis.

Repair Priorities for High-Consequence Work

  • Stop external harm or unauthorised action.
  • Correct false or unsupported facts.
  • Correct wrong calculations and units.
  • Restore source traceability.
  • Restore missing uncertainty and qualifications.
  • Restore required fields and receiver usability.
  • Repair style, tone and cosmetic presentation last.

This order is not universal, but it prevents cosmetic repair from consuming attention while decision-critical failures remain.

Repair Escalation

Some problems should be escalated rather than repeatedly regenerated. Examples include conflicting authoritative sources, professional judgement outside the user’s competence, permission uncertainty and tool errors that cannot be resolved safely.

Escalation can mean asking the responsible human, consulting the official source, obtaining a specialist review or pausing the workflow.

A strong repair system knows when not to continue autonomously.

Repair Documentation

For important recurring workflows, record major repairs: what failed, why, what changed and which test now protects against recurrence.

Documentation does not need to be long. One concise repair note can preserve the lesson for future maintainers.

The record is especially valuable when a future change accidentally reintroduces an old failure.

Answer Repair and Receiver Feedback

Sometimes the strongest evidence of weakness comes from the receiver. A manager asks where the source came from. A student cannot apply the explanation. A colleague misreads the deadline.

Treat repeated receiver confusion as a design signal. Repair the answer structure or handoff, not just the wording.

Receiver feedback reveals failures that internal review may miss.

A Final Repair Gate

Before accepting a repaired answer, ask four questions: did the repair fix the root failure, did it preserve accepted content, did it create any new failure and does the result still work on a fresh case?

Then confirm the external state if the task included tools or actions. A repaired conversational answer is not enough when the destination system remains wrong.

The strongest repair is one that survives the next example and reduces the chance that the same failure returns.

This is the point where answer improvement becomes system improvement: the lesson is captured, tested and carried forward.

The Repair Transfer and Retirement Gate

A repair is not fully proven until it works on a fresh case. Change names, dates, source wording or domain while preserving the failure type. If the old error returns, the repair may be tied too closely to the original example.

When the repair transfers, decide whether it belongs in the permanent workflow. A frequent or high-consequence failure may deserve a new constraint, example, validation rule or test. A one-off anomaly may need only a note.

Also retire obsolete repair rules. If a stronger technical control now prevents the failure, remove redundant prompt text. If a new schema makes an old checklist unnecessary, keep the regression test but simplify the instructions.

The mature answer-repair system therefore does three things: it fixes the current result, protects future results from the same failure and removes outdated controls when they no longer add value.

A Final Receiver Check

Before accepting the repaired answer, give it one receiver test: can the intended reader, learner, colleague or downstream system use the result correctly without reconstructing missing context?

If the repair fixed the internal logic but the receiver still cannot identify the decision, evidence, deadline, next action or uncertainty, the answer is not finished. Repair usefulness as well as correctness.

This final check protects the real purpose of answer improvement: not producing prettier text, but creating work that survives contact with the person or system that must use it next.

Weak Answers Are Diagnostic Data

A weak answer is not just something to discard. It reveals the first place where task, evidence, structure, tool use or checking became unstable.

Strong SI users learn to preserve the correct parts, repair the failing layer and test the change on new material.

Use the complete SI learning hub to continue. The final article in Stage 2 shows how to preserve that discipline across long conversations and projects.