LEARNING WITH SUPER INTELLIGENCE · PRACTICAL CODING
Debug one cause at a time
Reproduce the failure. Test the explanation. Make a small repair. Prove what changed.
How do you debug code with Super Intelligence and still learn to solve the next problem yourself? Start by making the failure observable. State what should happen, reproduce what actually happens, identify competing explanations, and run a small test that separates them. Ask SI to help you design those tests. Accept a patch only after its explanation matches the evidence and the relevant checks pass.
The important change is from asking for a replacement program to asking for a better investigation. A long answer can feel productive while leaving the cause unknown. A ten-line example, a precise prediction and one failing assertion can teach much more. This guide is designed for learners who know some variables, conditions, functions and lists but become unsure when their own code behaves unexpectedly.
In this eduKateSG series, SI is an editorial umbrella for practical contemporary AI assistance. It does not mean that a current coding assistant has been established as artificial superintelligence, or that its confidence gives it authority over your program. You remain responsible for deciding the intended behaviour, protecting private information and checking the result. The aim here is independent debugging ability, with assistance gradually reduced as your understanding grows.
The worked cases are original, deliberately small teaching programs. Their inputs are fictional and contain no real student records, credentials or customer information. The Python examples and checks were run locally with Python 3.12.14; the JavaScript examples and checks were run with Node.js v24.19.0. These are the environments used for this article, not a claim about the latest releases. No external AI service was needed to run them. The suggested assistant responses are illustrative examples written for teaching, not transcripts of measured model behaviour.
You can follow the guide in order or choose the route that matches your difficulty. If you cannot yet explain a function call or a loop, begin with How to Learn Coding With Super Intelligence. If you are planning a new feature rather than repairing a failure, How to Write Code With Super Intelligence covers the broader development workflow. This page concentrates on the investigation between “something is wrong” and “I can show why this repair is justified”.
Choose your reading route
- I do not know where to start: identify the missing evidence
- My program gives the wrong answer: boundary and missing-value cases
- Later actions change earlier results: sharing, mutation and timing
- The suggested fix seems too broad: reject incomplete repairs
- I want to practise independently: exercises, answers and transfer
Contents
Understand the failure
- 1. Choose the next diagnostic move
- 2. Write the contract before diagnosing the code
- 3. Build a minimum complete evidence packet
- 4. Turn a plausible explanation into a testable hypothesis
Work through five cases
- 5. Case one: the passing score that disappears
- 6. Case two: one call changes another call’s list
- 7. Case three: a real zero is mistaken for missing data
- 8. Case four: sorting fixes the display but changes the caller
- 9. Case five: an older response overwrites newer results
Inspect, protect and practise
- 10. Read traces and inspect the earliest wrong state
- 11. Reject repairs that remove the evidence
- 12. Protect private information while preserving the bug
- 13. Practise independently: three investigations before the answers
- 14. Full answers and changed-condition checks
Verify, transfer and continue
1. Choose the next diagnostic move
The first useful question is not “Which AI tool should I use?” It is “What kind of evidence is missing?” A syntax error, an incorrect total, a result that changes on the second call and a delayed response overwriting the screen are different situations. Treating them all as a request to rewrite the program makes the investigation less precise. Choose a route that tells you what to observe next.
If the program stops with an exception, capture the error type, message and relevant call frames. Find the first frame in code you control that helps explain the failure. Do not assume that the final displayed line is the original cause: an earlier function may have supplied a value of the wrong shape. Read the exception from the bottom to identify its type, then inspect the call chain and the value at the failing operation. Python’s errors and exceptions tutorial explains the distinction between syntax errors and runtime exceptions.
If the program runs but returns the wrong answer, write one concrete expected result before changing anything. A statement such as “the average seems wrong” is difficult to test. “For recorded scores zero and eighty, ignoring one missing entry, the average must be forty” is checkable. Work out the expected result by a method independent of the implementation. Otherwise, you risk copying the same mistake into the test and congratulating both versions for agreeing.
If the failure appears only after another action, investigate state. Note the full sequence, not only the last click or function call. Does a list retain an earlier entry? Does sorting change an object owned by the caller? Does a previous request finish after a newer request? Repeating the same final step in a fresh process may not reproduce a sequence-dependent bug. The earlier steps are part of the input to your investigation.
If the same source code behaves differently in two places, compare the environment before editing the algorithm. Record the interpreter or runtime version, command, working directory, relevant configuration and dependency versions. Do not export your whole environment; it may contain secrets. A small difference such as reading a different file can explain behaviour that looks like a complicated language problem. Ask the assistant which observation would distinguish an environment mismatch from a logic error.
Finally, decide whether this is an appropriate learning exercise. A local calculation over invented values is a good place to practise. A production payment failure, permission defect, unexplained data deletion or security incident requires the responsible maintainer and an established response process. You can still prepare a clear, redacted report. You should not experiment with irreversible operations merely because an assistant proposes a command that sounds plausible.
Contents · Next: 2. Write the contract before diagnosing the code
2. Write the contract before diagnosing the code
A contract is a short statement of the behaviour the program is meant to provide. It does not need legal language or a formal specification system. For a beginner’s function, it can identify the accepted input, the returned output, what happens at boundaries and whether the function may change objects passed to it. These decisions are essential because the same code can be correct for one contract and wrong for another.
Consider a function that counts passing scores. Does passing mean strictly more than fifty, or fifty and above? Are duplicate scores separate attempts? Is an empty list permitted? Can the threshold change? If those questions are unanswered, an assistant can produce a polished fix that silently chooses the wrong rule. For our first case, each entry is one score, every score at or above the supplied threshold counts, and an empty list produces zero. We assume numeric inputs; input validation is outside this tiny function’s stated lesson.
The contract also protects behaviour that already works. Suppose a helper appends a topic to a list supplied by its caller. If the caller intentionally passes a list, the helper must append to that very list. If the caller supplies nothing, each call must receive a new list. Both parts matter. A repair that returns an isolated list for every call may fix unwanted sharing while breaking the caller’s intentional accumulation. Debugging is not simply making one red assertion green.
Write examples on both sides of a decision boundary. For an inclusive threshold of fifty, the scores forty-nine, fifty and fifty-one are more informative than a list of comfortable passes. For missing data, compare a real zero with an absent entry. For sorting, choose numbers whose textual order differs from numeric order. Test design begins with the rule you care about; it should not be an afterthought added to whatever patch the assistant happens to propose.
A useful learner prompt is: “Before suggesting code, restate the contract. Mark any assumption that is not supported by my description. Give me one example whose expected result depends on that assumption.” This turns ambiguity into a visible decision. If the assistant cannot tell whether empty input should return a default or raise an error, that is not necessarily a model failure. The requirement may genuinely be missing. Resolve it before judging an implementation.
For assessed schoolwork, the assignment and teacher’s rules remain the relevant contract. Do not quietly change required interfaces, forbidden libraries or the amount of permitted AI assistance. You can ask for an explanation of an error or for a separate practice example when that is allowed. Being able to describe what the program should do is part of the learning, and outsourcing that decision can conceal the very misunderstanding you need to repair.
Contents · Next: 3. Build a minimum complete evidence packet
3. Build a minimum complete evidence packet
A useful debugging packet is small enough to inspect and complete enough to reproduce the problem. It contains the contract, the smallest relevant code, an exact input or action sequence, the expected result, the observed result, the command used and the relevant environment. It also states what you have already tried. The purpose is not to impress the assistant with detail. It is to prevent the assistant from filling important gaps with guesses.
“It does not work” provides no stable target. “Running python boundary.py prints one for the list forty-nine, fifty, fifty-one, but the contract says fifty and above must count, so I expect two” identifies the disagreement. Include the function as well as its call. A function can be correct when called one way and wrong when its caller supplies a string instead of a number. Hiding the call can make both you and the assistant investigate the wrong layer.
Reduce the example one change at a time. Remove unrelated display code and run it again. Replace private data with invented values that preserve the relevant shape and run it again. Remove an unrelated dependency and run it again. If the failure disappears, the removed part may matter; restore it or record the changed condition. A reduced example is useful because you demonstrated that it retains the failure, not because it happens to be short.
Do not clean up the code while reducing it. Renaming everything, changing loops and replacing functions can accidentally repair or transform the bug. Keep an unchanged failing copy beside your working copy. If you use version control, review the differences before applying a patch, and preserve unrelated work. The investigation should be reversible. You do not need a complex branch strategy for these exercises, but you do need a reliable way to compare before and after.
For an exception, preserve the meaningful traceback structure while redacting sensitive paths or values. Do not remove the error type because it looks alarming. A TypeError and a ValueError point toward different questions. Equally, do not paste hundreds of lines of unrelated logs. Identify the first relevant event and include enough surrounding context to show order. If the order is uncertain, say so rather than manufacturing a neat sequence from separate sources.
The following packet is a complete teaching prompt you can adapt. Its most important restriction is that the assistant must help discriminate between explanations before producing a replacement. If your tool cannot execute code, ask it to label all proposed outputs as predictions. You then run the program yourself and supply the observed result. A prediction and an execution result can be useful in different ways, but they must not be presented as the same kind of evidence.
Task: Help me investigate this local teaching example.
Contract: Count every score at or above threshold; empty input returns 0.
Environment: Python 3.12.14, no external packages.
Code:
def count_passes(scores, threshold=50):
return sum(score > threshold for score in scores)
Input: count_passes([49, 50, 51])
Expected: 2
Observed: 1
Already checked: the three input values reach this function unchanged.
First response: propose two possible explanations and a small test
that separates them. Label predictions. Do not rewrite the whole program.
Contents · Next: 4. Turn a plausible explanation into a testable hypothesis
4. Turn a plausible explanation into a testable hypothesis
A hypothesis connects a proposed cause to an observable consequence. “It is probably an off-by-one problem” is a category, not yet a useful prediction. “The comparison excludes a score equal to the threshold, so a one-item list containing fifty will return zero” is testable. Before changing the comparison, run that case. If it returns something else, your understanding is incomplete and the proposed repair deserves another look.
Keep at least one alternative explanation alive long enough to test it. Perhaps the comparison is wrong; perhaps the function receives different values from those displayed; perhaps the threshold is being overridden by the caller. You do not need twenty hypotheses. Two or three credible alternatives usually reveal the next useful observation. Ask SI to rank them using the evidence already available, while requiring it to state what would change that ranking.
Use discriminating tests. Suppose both the “wrong comparison” and “missing final list element” explanations predict one for a particular three-item input. Repeating that input will not separate them. A one-item list at the threshold, followed by a one-item list above it, provides more information. A good test is not merely another example. It is chosen because competing explanations predict different results under its conditions.
Make the prediction before the run. It is easy to see an output and produce a story that fits it afterwards. A written prediction creates a small commitment your explanation can fail. If the result contradicts the hypothesis, record that as progress. You have eliminated an explanation. Do not edit the prediction afterwards to make the investigation look smoother; that removes the evidence of how your understanding actually changed.
When an assistant proposes a diagnosis, ask it to identify the exact expression or state transition that creates the failure. “The asynchronous code has a race” is less useful than “both completed requests assign to the same visible value, and the older request can finish last”. The second explanation suggests a controlled experiment. It also clarifies the difference between fixing response ordering and making the network faster, which are not the same intervention.
A compact notebook entry can contain four sentences: “I suspect this cause. If it is true, this input should produce this result. I observed this result. Therefore I will keep, revise or reject the hypothesis.” Use it for one bug rather than turning your session into administrative work. The habit matters because it makes reasoning visible enough to review. Over time, you should need fewer hints to design the next discriminating test.
Avoid treating the assistant’s agreement as an additional independent test. If you show it your preferred diagnosis and ask whether you are right, it may simply organise your explanation more fluently. Instead, ask for a counterexample or for the strongest competing explanation consistent with the evidence. Then choose a test you can execute. The program, specification and independently checked output should carry more weight than the confidence of the conversation.
Contents · Next: 5. Case one: the passing score that disappears
5. Case one: the passing score that disappears
Our first case is intentionally small. The function is meant to count scores at or above a threshold. The failing version uses a strict comparison. Save the following complete program as boundary.py and run it with Python. It has no files to download, no packages to install and no connection to a school system. Its assertions are learning checks, not a substitute for validating external inputs in a real application.
def count_passes_broken(scores, threshold=50):
return sum(score > threshold for score in scores)
def count_passes(scores, threshold=50):
return sum(score >= threshold for score in scores)
print("broken:", count_passes_broken([49, 50, 51]))
assert count_passes_broken([50]) == 0
assert count_passes_broken([51]) == 1
assert count_passes([49, 50, 51]) == 2
assert count_passes([]) == 0
assert count_passes([50, 50]) == 2
assert count_passes([59, 60, 61], 60) == 2
assert count_passes([0], 0) == 1
print("boundary regression checks passed")
A direct contract test against the broken function is also important: the assertion count_passes_broken([49, 50, 51]) == 2 must fail. In the local check, that assertion was given the message “expected 2, observed 1”. It raised the following failure line. This is an expected failure demonstrating that the test detects the defect, not a failure of the repaired function.
AssertionError: expected 2, observed 1
The observed first line is “broken: 1”. The two assertions about the broken function document its behaviour; they are not claiming that behaviour is correct. The later assertions describe the intended contract. If you replace the fixed function’s inclusive comparison with the original strict comparison, the first contract assertion fails. That deliberate check shows the test can detect the defect rather than merely exercise the function without examining its result.
Trace the three comparisons by hand. Forty-nine is below fifty, so it contributes zero to the count. Fifty is equal to fifty and should contribute one under the contract. Fifty-one is above fifty and contributes one. The failing function excludes the middle entry. Its total is therefore one, while the intended total is two. This explanation matches both the original failure and the one-item experiments; it does not depend on the assistant’s confidence.
An illustrative bad suggestion is “lower the default threshold to forty-nine”. That makes the original three-item example produce two, but it changes the meaning of the threshold instead of implementing the stated rule. Reject it with a call that explicitly passes sixty and includes a score of sixty. The problem returns because the comparison remains strict. This is a valuable pattern: a patch can fit the original example while preserving the underlying mistake.
Another tempting suggestion is “add one to the final count”. Reject it using an empty list and a list containing only scores below the threshold. Neither should acquire an imaginary passing score. A fix should follow from the mechanism of the failure. Here the repair belongs at the comparison, where equality is incorrectly excluded, rather than at the total, where the mistake has already been accumulated.
The regression set protects several aspects of the contract: an ordinary mixed list, no entries, repeated boundary values, a changed threshold and a zero threshold. Each case has a reason. Ten random high scores would mostly confirm a path that already worked. Explain the purpose of each assertion to a partner before asking SI to generate more. If you cannot say which mistaken implementation an assertion would reject, its value may be less clear than it appears.
For transfer, imagine the rule changes to “strictly above the threshold”. The original comparison would then fit the new contract, and the tests would need to change. Do not describe the operator itself as universally wrong. The bug is the disagreement between behaviour and a particular requirement. Learning that distinction prevents you from memorising patches without understanding the decision each patch expresses.
Contents · Next: 6. Case two: one call changes another call’s list
6. Case two: one call changes another call’s list
The next failure cannot be understood from a single call alone. We want a helper that appends a topic. When no list is supplied, each call should start a separate list. When an explicit list is supplied, the helper should append to that list and return it. The broken version seems reasonable during its first call, so a test that runs it only once can miss the defect completely.
def add_topic_broken(topic, topics=[]):
topics.append(topic)
return topics
def add_topic(topic, topics=None):
if topics is None:
topics = []
topics.append(topic)
return topics
first = add_topic_broken("loops")
second = add_topic_broken("tests")
print("first after second call:", first)
print("second:", second)
print("same object:", first is second)
one = add_topic("loops")
two = add_topic("tests")
assert one == ["loops"]
assert two == ["tests"]
assert one is not two
shared = []
assert add_topic("loops", shared) is shared
assert add_topic("tests", shared) is shared
assert shared == ["loops", "tests"]
print("state regression checks passed")
The broken program prints both lists as [‘loops’, ‘tests’] and reports that they are the same object. The second call did not mysteriously reach backwards in time. Both names refer to the same list, so appending through one reference changes what the other reference sees. Python evaluates a default argument when the function is defined, rather than creating that default anew on every call. The official default-argument explanation documents this behaviour.
The discriminating observation is identity, not only equality. Two independent lists can contain equal values, so comparing their contents does not always reveal sharing. In this small example, “first is second” asks whether the names refer to the same object. You can also watch the first list before and after the second call. Both observations support the shared-default explanation and weaken a vague claim that printing is somehow combining results.
An illustrative bad suggestion is “clear the list at the beginning of the function”. That may remove old entries from the new return value, but it also changes the earlier caller’s list because the list is still shared. It can also erase a deliberately supplied list. The repair hides the accumulated symptom by deleting information. Reject it using the contract’s explicit-list case, where existing content must survive and the new topic must be appended.
The minimal repair uses a sentinel, None, to distinguish an omitted list from an explicit list. It creates a list inside the function only when the argument is None. Be careful with an assistant suggestion that uses “if not topics” instead. An explicitly supplied empty list is false in a Boolean test, so that version can replace the caller’s list with a different one. Our identity assertion for the supplied empty list catches this changed behaviour.
This case teaches you to ask a more specific question: “Who owns this object, who can change it, and how long does it live?” Those questions apply beyond Python defaults. A helper may mutate a list supplied by a caller, a module may retain a cache, or a test may leave shared state for the next test. The right repair depends on which sharing is intended. Removing all sharing without checking the interface is not automatically an improvement.
For a changed-condition test, call the fixed helper twice with the same explicit list and once without a list. The explicit calls should accumulate together; the omitted call should remain independent. Predict all three results before running them. If you can explain why those two behaviours coexist in the same implementation, you have learned more than the slogan “never use an empty list as a default”. You understand the contract the sentinel preserves.
Contents · Next: 7. Case three: a real zero is mistaken for missing data
7. Case three: a real zero is mistaken for missing data
Our third case produces a believable number without raising an exception. The contract accepts a list of numeric scores and None for missing entries. It must average every recorded score, including zero, and raise a ValueError when no recorded scores exist. This is a data-meaning problem disguised as a convenient filter. A language’s idea of a false value is not necessarily the same as your application’s idea of an absent value.
def mean_score_broken(scores):
present = [score for score in scores if score]
if not present:
raise ValueError("no recorded scores")
return sum(present) / len(present)
def mean_score(scores):
present = [score for score in scores if score is not None]
if not present:
raise ValueError("no recorded scores")
return sum(present) / len(present)
print("broken mean:", mean_score_broken([0, 80, None]))
assert mean_score([0, 80, None]) == 40
assert mean_score([0, 0]) == 0
assert mean_score([60, 80]) == 70
assert mean_score([None, 100]) == 100
for missing in ([], [None, None]):
try:
mean_score(missing)
except ValueError as error:
assert str(error) == "no recorded scores"
else:
raise AssertionError("missing scores must raise ValueError")
print("mean regression checks passed")
The observed broken result is 80.0. The expected mean is forty because the recorded values are zero and eighty, whose sum is eighty and whose count is two. None contributes neither a score nor an entry to that count. In the failing comprehension, zero and None are both excluded by the truth-value condition. The Python built-in types documentation lists zero and None among false values; the application must still decide which of them means missing.
A useful diagnostic probe is to examine the intermediate list, not just the final average. For the failing input, the broken function keeps only eighty. That narrows the problem to selection before the division occurs. A hypothesis about floating-point rounding is poorly supported here: the wrong denominator already explains the large difference. Ask SI to identify the earliest intermediate value inconsistent with the contract, and then verify that value yourself.
An illustrative bad suggestion is “replace None with zero before averaging”. That changes the meaning of missingness. A missing attempt would now lower the average as if a real zero had been recorded. Reject it with [None, 100], for which the contract requires one hundred. Another bad suggestion is to return zero when there are no recorded values. That may be a valid requirement in another system, but it contradicts this case’s explicit choice to signal the absence of data.
Notice why the fixed function still uses “if not present” after the filter. At that point, present is a list, and the question really is whether the list is empty. A list containing a numeric zero is not empty. The same language feature can therefore be appropriate in one place and inappropriate in another. Debugging requires examining the meaning of the condition, not banning a syntax pattern wherever it appears.
The all-zero regression is especially valuable. Without it, you might repair the mixed example while leaving a special case that falsely reports no scores. The missing-only cases protect the distinction between a measured result of zero and no measurable result. In real data work, that distinction can affect decisions. These teaching values are invented; do not paste actual pupil scores or personal records into an external assistant to reproduce the lesson.
For transfer, suppose the input also permits a string marker such as “absent”. The current contract no longer describes the input completely. Decide whether the parser should convert that marker to None or reject it before it reaches the averaging function. Do not keep expanding a truthiness filter until it happens to silence each error. Make the boundary between parsing, validation and calculation explicit, and test that boundary with invented records.
Contents · Next: 8. Case four: sorting fixes the display but changes the caller
8. Case four: sorting fixes the display but changes the caller
The JavaScript case has two separate defects. We want a function that returns a new array of numeric times in ascending order while preserving the original array. The input is a dense array of finite numbers; strings, missing values and special numeric values are outside this teaching contract. The broken function uses the default sort and sorts the caller’s array in place. Either problem alone can make a later part of a program behave strangely.
const assert = require('node:assert/strict');
function orderedTimesBroken(times) {
return times.sort();
}
function orderedTimes(times) {
return [...times].sort((a, b) => a - b);
}
const original = [3, 12, 9];
console.log('broken result:', JSON.stringify(orderedTimesBroken(original)));
console.log('broken caller:', JSON.stringify(original));
const source = [3, 12, 9];
const result = orderedTimes(source);
assert.deepEqual(result, [3, 9, 12]);
assert.deepEqual(source, [3, 12, 9]);
assert.notStrictEqual(result, source);
assert.deepEqual(orderedTimes([]), []);
assert.deepEqual(orderedTimes([4, 4, -2]), [-2, 4, 4]);
console.log('sorting regression checks passed');
Save this program as sorting.cjs and run it with Node.js. The observed broken result is [12,3,9], and the caller’s original array has the same reordered contents. JavaScript’s default array sort compares string representations when no comparison function is supplied. It also changes the array being sorted. These behaviours are documented in MDN’s Array.prototype.sort reference. They are language behaviour, not evidence that the runtime is malfunctioning.
An illustrative assistant suggestion might add the numeric comparison and stop there. That produces the desired order, but it still violates the no-mutation requirement. This is why the test checks both the returned result and the original source. If you only display the returned array, the remaining defect is easy to miss. A function’s correctness includes promised side effects and promised absence of side effects, not merely the value printed on the screen.
The fixed version makes a shallow array copy and sorts that copy numerically. For this contract, the elements are numbers, so that is enough to separate the array containers. Do not generalise the example into a claim that spreading an array deeply copies arbitrary objects. If the elements were mutable objects and a later step edited one, both arrays could still refer to that object. A changed data shape demands another investigation of ownership.
An illustrative bad suggestion is “convert every number to a string before sorting”. That moves further away from the numeric ordering requirement. Another is “use a comparison function that returns true or false”. A comparison function needs the appropriate negative, zero or positive relationship, rather than a Boolean answer to only one direction of comparison. For these finite numbers, subtracting the second from the first supplies the ordering information we need.
Node’s assertion documentation explains the assertion methods used in this example. Deep equality checks the array contents; the separate identity assertion checks that the returned array is a different container. An empty array tests the simplest boundary. Repeated and negative values test ordering beyond the initial three positive values. Each check protects a stated property of this small function, while making no claim that an arbitrary sorting application is fully verified.
For independent practice, ask SI to propose a test that would pass for a numeric in-place repair but fail for the complete contract. You should recognise that preserving source order is the relevant check. Then hide the answer and write that test yourself. The learning milestone is not remembering the spread syntax. It is noticing that the caller’s state must be part of the observation whenever a function might mutate its input.
Contents · Next: 9. Case five: an older response overwrites newer results
9. Case five: an older response overwrites newer results
Asynchronous bugs can look unpredictable when you depend on real network timing. This example removes the network and controls completion order directly. The contract is narrow: after two searches begin, only the most recently started search may update the visible result. We deliberately supply fulfilled promises only. Rejection, loading indicators and cancellation are separate behaviours to specify and test before using a similar pattern in an application.
const assert = require('node:assert/strict');
function deferred() {
let resolve;
const promise = new Promise(done => { resolve = done; });
return { promise, resolve };
}
function createSearchBroken(fetcher) {
const state = { value: null };
return {
state,
async search(query) {
state.value = await fetcher(query);
}
};
}
function createSearch(fetcher) {
const state = { value: null };
let latest = 0;
return {
state,
async search(query) {
const mine = ++latest;
const value = await fetcher(query);
if (mine === latest) state.value = value;
}
};
}
async function runOrder(factory, newerFirst) {
const oldRequest = deferred();
const newRequest = deferred();
const app = factory(q => q === 'old' ? oldRequest.promise : newRequest.promise);
const oldRun = app.search('old');
const newRun = app.search('new');
if (newerFirst) {
newRequest.resolve('new results');
await newRun;
oldRequest.resolve('old results');
await oldRun;
} else {
oldRequest.resolve('old results');
await oldRun;
newRequest.resolve('new results');
await newRun;
}
return app.state.value;
}
async function main() {
console.log('broken:', await runOrder(createSearchBroken, true));
assert.equal(await runOrder(createSearch, true), 'new results');
assert.equal(await runOrder(createSearch, false), 'new results');
console.log('async regression checks passed');
}
main().catch(error => { console.error(error); process.exitCode = 1; });
Save the program as ordering.cjs and run it with Node.js. The observed broken line is “broken: old results”. The newer request finishes first and writes its result. The older request then finishes and overwrites that value. No random delay is necessary to demonstrate the problem. Our small helper gives the test direct control over when each promise resolves. MDN’s Promise reference explains promises and their settlement behaviour.
The fixed version increments a local request counter when a search begins. Each call remembers its own number. After the await, it checks whether that number still matches the latest started request. If another request has begun, the older result is no longer permitted to replace the visible value. The two tests run the completions in opposite orders. Both must end with the newer search’s result under our contract.
An illustrative bad suggestion is “wait half a second before assigning the result”. A delay may change how often a failure is observed, but it does not establish which request owns the display. Another suggestion is “make the requests finish in order”. That may be appropriate for a different application, but it can unnecessarily delay the latest request behind an obsolete one. Here the requirement concerns which completed result may update state, not the speed of completion.
The counter does not cancel the old operation. It ignores an obsolete fulfilled result. That distinction matters if the operation consumes resources or has side effects. You should not describe this miniature solution as complete request management. It is a verified repair for a specific visible-state rule in a controlled example. If an assistant claims that it solves every network race, ask which test establishes cancellation, error handling, resource cleanup or multiple independent search panels.
For changed conditions, ask what should happen if the latest request fails while an older one later succeeds. Does the interface show the latest error, retain previously displayed results or fall back to the older response? There is no single answer without a product requirement. Write the rule and then extend the controlled test with rejection. Do not assume that the fulfilled-only example already protects a failure path it never exercises.
This case’s most transferable lesson is to control the variable that makes a bug hard to reproduce. Timing can be replaced with explicitly ordered completions. A random value can sometimes be replaced with a known sequence. A file reader can be given a small fixture instead of a large live folder. Each substitution needs to preserve the behaviour under investigation. After the local repair, separately check the real integration where that behaviour matters.
Contents · Next: 10. Read traces and inspect the earliest wrong state
10. Read traces and inspect the earliest wrong state
A traceback tells you where execution encountered a problem; it does not automatically tell you why the program reached that state. A learner who treats every final line as the root cause can keep adding guards around symptoms. Instead, follow the relevant value through the small example. Ask where it first stops matching the contract, and inspect immediately before and after that transformation.
In the mean-score case, the final division is mathematically consistent with the already-filtered list. Looking only at that line can invite an unnecessary rounding repair. Inspecting the intermediate list reveals that zero disappeared earlier. In the sorting case, the displayed order is one observation, while the changed source array is another. In the async case, the critical fact is which request performs the final assignment. Good inspection follows the cause that the hypothesis predicts.
A print statement can be sufficient for a tiny local example. Label the value so that multiple outputs remain understandable, and include its type when a type mismatch is plausible. Avoid dumping entire objects by default. A large object may contain private fields, and its volume can obscure the single attribute that matters. In a controlled exercise, print the length, a synthetic identifier or a carefully selected value rather than a complete session or request object.
An interactive debugger provides another way to pause and inspect. Python’s pdb documentation describes breakpoints, stepping and inspecting the current stack. For your own local learning script, pause before the suspect expression, inspect the input and then step through it. Choose a question first. Aimless stepping through every line can consume time without distinguishing any hypothesis, especially when library code takes you far away from the relevant transformation.
Treat observation itself carefully. Adding logging can affect the timing of a concurrent program, and inspecting a live object later may not reveal what it contained at an earlier moment. For the deterministic examples here, the simplest approach is to record values at the relevant step or use assertions directly after it. For larger systems, work with the project’s existing logging and debugging practices rather than inventing unrestricted diagnostic output.
When you ask SI to interpret a trace, ask it to separate three things: the facts literally present, the inferences those facts support and the information still missing. If it names a file or function not in your supplied example, require it to explain where that name came from. An invented path can send you on a long search through a repository that never contained the supposed culprit. A useful explanation stays anchored to evidence you can inspect.
Contents · Next: 11. Reject repairs that remove the evidence
11. Reject repairs that remove the evidence
A common failure pattern is an edit that makes the error disappear without restoring the intended behaviour. Catching every exception and returning an empty result can make a failing test look quieter, while converting a visible defect into incorrect data. Removing the assertion can turn a red test green without changing the program at all. Disabling a check may sometimes be justified by a changed requirement, but it must not be used to avoid understanding a failure.
Before accepting a suggested patch, ask which observation it explains. If the proposed edit affects an unrelated function, require a causal path from that function to the failure. If it changes several behaviours at once, ask whether a smaller repair can test the leading hypothesis first. A broad rewrite can accidentally fix one case while creating new ones, leaving you unable to say which change mattered. For learning, that uncertainty reduces the value of the whole exercise.
Look for silent contract changes. Does the patch treat missing data as zero? Does it sort the caller’s array even though the interface promises to preserve it? Does it catch invalid configuration and quietly enable a feature? Does it lower a threshold so that one example passes? These are not merely stylistic differences. They alter what the program means. Compare the patch with the written examples rather than with the assistant’s summary of its own intentions.
Use a “defeat the patch” question: what small input would expose this repair if it only fits the original example? You are not trying to be adversarial toward a tool. You are testing whether the reasoning generalises. The original threshold failure rejects a strict comparison; an explicit changed threshold rejects a hard-coded adjustment; an empty input rejects adding one to the count. Together these cases tell a much stronger story than the first happy result alone.
Also reject unnecessary dependencies when the current contract can be met with the language features already available. A package may be appropriate for a larger problem, but installing it is not a substitute for understanding the failure. Check the official documentation for any new API and confirm that it exists in the environment you actually use. Do not run an unexplained installation script, destructive cleanup command or security-setting change merely because it appears in a proposed debugging answer.
Finally, preserve a difference between “plausible repair”, “passed these local checks” and “ready for the real application”. The worked programs establish specific behaviours for controlled inputs. They do not verify an entire deployed service, its data handling or its security. A concise completion note should say what was tested and what was not. Honest scope is part of strong debugging, because it tells the next reader where confidence comes from and where further work remains.
Contents · Next: 12. Protect private information while preserving the bug
12. Protect private information while preserving the bug
Debugging often creates pressure to paste everything quickly. Resist that pressure. A traceback, configuration dump, browser request or screenshot can contain passwords, tokens, personal details, private source code and internal addresses. Use an assistant only with information you are permitted to share through that tool. If the code belongs to a school or employer, follow its rules; being able to copy a file does not itself grant permission to disclose it.
The safest teaching packet usually uses synthetic input. Replace a real name with a fictional label, a genuine email with an obviously invented example address, and a private record with a tiny object that preserves the relevant types. If the bug concerns an empty value, preserve the emptiness. If it concerns the difference between zero and missing, retain that distinction. Redaction should protect identity while keeping the property that triggers the failure available for investigation.
Do not replace every secret with a string that resembles a valid credential and then use it against a real service. These exercises need no service connection. For a configuration example, use a local mapping of non-sensitive values or a fixture. When debugging a genuine integration, keep credentials in the approved local mechanism and share only the error category and safe metadata. An assistant does not need a working password to explain the difference between a missing variable and a malformed Boolean string.
OWASP’s Logging Cheat Sheet identifies categories such as access tokens, passwords and sensitive personal data that generally should not be recorded directly in logs. It also discusses sanitising event data and limiting access. Apply that guidance before you create extra logs, not only before you paste them into a conversation. A temporary debug statement can still leave a sensitive record behind after the immediate session ends.
A redacted report should say what was changed. For example: “Identifiers and text are synthetic; the list length, missing-value marker and numeric zero are preserved.” That sentence prevents the reader from assuming that a fabricated name is a real person or that the displayed data came from a production incident. If redaction changes the behaviour so much that the bug no longer reproduces, report that limitation and continue investigating locally with an authorised maintainer.
Do not ask the model to infer or reconstruct redacted secrets. The useful question is whether the visible structure supports a hypothesis. If a pasted error accidentally exposes a credential, stop sharing it and follow the relevant provider’s or organisation’s incident process, including revocation or rotation where appropriate. Removing it from a later message does not establish that the earlier disclosure has been undone. For this guide, all evidence can remain within harmless invented examples.
Contents · Next: 13. Practise independently: three investigations before the answers
13. Practise independently: three investigations before the answers
Attempt the following exercises before reading the solutions. For each one, write the contract in your own words, predict a failing result, identify at least one competing explanation and choose a discriminating test. Then make the smallest repair you can explain. Ask SI for a hint only after you have recorded your first attempt. If you need help, ask for one observation to make rather than for the final function.
Exercise A concerns a text configuration flag. A function receives either None or a string. None means the feature is disabled. After removing surrounding spaces and ignoring letter case, “true” and “1” mean enabled; “false” and “0” mean disabled. Every other string is invalid and should raise a ValueError. The broken implementation below makes the wrong decision for “false”. Predict what it returns for “0”, an empty string and “ TRUE ” before running it.
def flag_broken(raw):
return bool(raw)
for raw in (None, "false", "0", "", " TRUE "):
print(repr(raw), flag_broken(raw))
Exercise B concerns topic labels. Input is a list of strings containing only ordinary ASCII letters and surrounding spaces. Return non-empty labels in lowercase, without duplicates, preserving the order of their first appearance after normalisation. Do not change the caller’s list. For [“ B ”, “a”, “b”, “”, “ A ”], the answer must be [“b”, “a”]. The proposed repair below looks compact, but it does not meet the full contract. Identify more than one issue without relying on the order a set happens to display.
def topics_broken(values):
return sorted(set(value.lower() for value in values))
print(topics_broken([" B ", "a", "b", "", " A "]))
Exercise C is an explanation challenge using the earlier list helper. Someone changes its fixed condition from “if topics is None” to “if not topics”. They argue that both versions create a list only when there is nothing in it, so the change is harmless. Construct a two-line caller that exposes the contract difference. Your check must examine which list object receives the appended value, not merely whether the returned list contains the right word.
Do not measure success solely by how fast you reach working code. A complete answer explains why the original fails, why the chosen test reveals it and why the repair preserves the rest of the contract. If your first patch fails a new case, use that failure to improve the explanation. An independent debugging attempt can be useful even when it begins with an incorrect hypothesis, provided you let the evidence change your view.
For a second round, ask SI to generate a fresh input for each contract without supplying a solution. Check its expected answer yourself before using it as a test. Model-generated tests can also be wrong, especially when the requirement has several clauses. Keep one or two independently calculated examples as anchors. The aim is to practise judging evidence, not to build a closed loop in which the same assistant invents both the program and every criterion for success.
Contents · Next: 14. Full answers and changed-condition checks
14. Full answers and changed-condition checks
For Exercise A, a non-empty string has a true truth value even when its text says “false” or “0”. The broken function is testing whether a value is truthy, not interpreting the allowed vocabulary of the flag. It returns false for None and the empty string, and true for the other three displayed strings. The empty string is still wrong under the contract because it should be rejected, not silently treated as a valid disabled setting.
The complete repair below parses the specified tokens explicitly. Its error message describes the permitted format without echoing arbitrary supplied content. None is handled before string operations. The contract permits strings or None only; a different input type would require a separate decision about validation. Python’s os.getenv documentation is relevant when a real application obtains such a value from its environment, but this practice function reads no actual environment variables.
def read_flag(raw):
if raw is None:
return False
value = raw.strip().lower()
if value in ("true", "1"):
return True
if value in ("false", "0"):
return False
raise ValueError("flag must be true, false, 1 or 0")
for value in ("true", "1", " TRUE "):
assert read_flag(value) is True
for value in (None, "false", "0", " FALSE "):
assert read_flag(value) is False
for value in ("", "yes", "sometimes"):
try:
read_flag(value)
except ValueError:
pass
else:
raise AssertionError("invalid flag accepted")
print("flag regression checks passed")
A changed-condition test might introduce “yes” as a newly permitted enabled token. That requires an intentional contract change and a corresponding parser update, not a general truthiness fallback. Another change might require an omitted variable to be an error instead of disabled. In that case, update the None branch and its test together. The lesson is to make configuration meaning explicit, particularly when an assistant proposes a convenient default that could conceal a setup mistake.
For Exercise B, lowercasing does not remove surrounding spaces, and the proposed function retains the empty label. Sorting also replaces first-appearance order with alphabetical order. A set can help track membership, but it does not by itself express the desired output sequence. The complete answer separates the normalised key, membership tracking and ordered result. It does not assign into or reorder the caller’s input list.
def unique_topics(values):
result = []
seen = set()
for value in values:
topic = value.strip().lower()
if topic and topic not in seen:
seen.add(topic)
result.append(topic)
return result
source = [" B ", "a", "b", "", " A "]
before = source.copy()
assert unique_topics(source) == ["b", "a"]
assert source == before
assert unique_topics([]) == []
assert unique_topics([" ", ""]) == []
assert unique_topics(["c", "c", "C"]) == ["c"]
print("topic regression checks passed")
Why choose a first label of B followed by a? Because the expected first-appearance order differs from alphabetical order. If you test only a then b, a sorting-based implementation may pass accidentally. Why include both a blank and a space-only label? They exercise normalisation before exclusion. Why copy the source before the call? The saved value gives the test an independent reference for the no-mutation promise. Each input is designed to expose a particular incomplete repair.
Now change the requirement: duplicate detection should still ignore case, but the output should retain the trimmed spelling of the first occurrence. The previous answer no longer satisfies that changed contract because it always returns lowercase. A complete revised answer follows. We keep the example’s ASCII scope; deciding equivalence for international names or all Unicode text requires a more careful requirement and relevant documentation, not an unexamined claim that lowercasing solves every language problem.
def display_topics(values):
result = []
seen = set()
for value in values:
display = value.strip()
key = display.lower()
if key and key not in seen:
seen.add(key)
result.append(display)
return result
assert display_topics([" B ", "a", "b", "A"]) == ["B", "a"]
assert display_topics([]) == []
print("changed-contract checks passed")
For Exercise C, supply an explicit empty list, call the function and then assert that the returned object is the supplied object. With the “if not topics” variant, a fresh list replaces the empty one inside the function, so the caller’s list remains empty. The return contents can still look correct, which is why checking only the word “loops” is insufficient. The fixed sentinel version from Case Two passes both the identity check and the content check.
def add_topic_wrong(topic, topics=None):
if not topics:
topics = []
topics.append(topic)
return topics
target = []
returned = add_topic_wrong("loops", target)
assert returned == ["loops"]
assert target == []
assert returned is not target
print("wrong repair exposed: caller's list was not updated")
These answers are complete for their stated contracts, but they are not an invitation to paste them into unrelated applications. After reading, close the solutions and reconstruct one function from its contract and tests. Then explain a changed requirement without looking back. If you can only reproduce the exact text, practise another example. If you can select a revealing input and defend the resulting repair, your debugging skill is becoming more transferable.
Contents · Next: 15. Build a regression set that earns its confidence
15. Build a regression set that earns its confidence
A regression test preserves evidence of a failure so that a later change does not quietly restore it. Begin with the smallest input that demonstrates the original problem. Run that test against the broken version and confirm that it fails for the intended reason. Then run it against the repaired version. If both versions pass, either the test is not detecting the bug or the setup differs from the original failure. Investigate before adding more tests.
Add neighbouring cases around the boundary the repair changes. A threshold bug deserves below, equal and above cases. A missing-data filter deserves zero, missing and ordinary values. A state-sharing bug deserves consecutive calls and an intentionally shared object. These cases give the repair a shape. They help you see whether it addresses the mechanism or merely patches one input. The goal is meaningful coverage of the rule, not a large count of assertions that all exercise the same easy path.
Protect invariants as well as outputs. An invariant is a property that should remain true across the relevant operation. In the sorting example, the source array stays unchanged. In the deduplication exercise, output labels are non-empty and no normalised label appears twice. In the async case, an obsolete response cannot replace the latest search’s result. Writing those properties in plain language often reveals a missing test before you need a sophisticated testing framework.
Check error behaviour explicitly. If empty recorded data should raise a ValueError, a test must distinguish that from returning zero or failing with an unrelated TypeError. A catch-all test that accepts any exception can pass for the wrong reason. The standard library’s unittest documentation describes assertions for values and exceptions, and a way to organise repeatable test cases. A small script is enough to begin, while a test framework becomes useful as the checks grow.
Do not overstate what a passing suite proves. It establishes that the tested behaviour matched expectations under the tested conditions. It does not establish that the expected answers were correct, that untested inputs are safe or that an integration behaves identically. Review the test oracle, meaning the source of the expected result. For a simple average, manual arithmetic is a useful independent oracle. For a complex domain rule, you may need a specification owner or an authoritative reference.
Make one deliberate negative check after the fix: temporarily restore the relevant broken expression in a disposable copy and confirm that the regression catches it. Do not damage your main working version. This simple exercise helps you notice tests that never reach their assertion, inspect the wrong function or merely check that execution completes. A test that demonstrably rejects the defect is more persuasive than a green output whose sensitivity you have never examined.
Contents · Next: 16. Handle environment and integration differences without guessing
16. Handle environment and integration differences without guessing
Sometimes the small example passes while the application still fails. That does not mean the example was pointless. It establishes a narrower result: the isolated logic behaves as expected with those inputs in that runtime. The next question is where the real application differs. Compare the value entering the function, the version actually loaded and the surrounding state. Do not repeatedly rewrite a locally verified function while leaving its caller and environment unexamined.
Begin with identity of execution. Are you running the file you edited? Is the test importing an installed copy rather than the local one? Does the command use the expected interpreter? A learner can spend a long time improving source that the running process never reads. Record the command and safe version information, restart the relevant local process when appropriate and observe whether a deliberately harmless visible change appears. Avoid broad cleanup or deletion as a first diagnostic move.
Then compare inputs at the boundary. A form may produce strings while your local test supplies numbers. A missing field may become an empty string rather than None. A parser may already have normalised values, or a caller may have reordered them. Ask SI to draw a short chain of transformations from input to failing output, but verify the actual values at the important boundaries. A diagram based on imagined data flow is only another hypothesis.
For the async case, the controlled test verifies response ownership under two completion orders. The real interface may add loading state, error state and component lifetime. Those additional behaviours need their own contract. If the view disappears before a request completes, what may still update? If two independent panels search at once, do they need separate counters? Use the small case to understand the mechanism, then test the application’s actual boundaries rather than stretching one example into a universal guarantee.
When official documentation and local behaviour appear to disagree, verify the version and the exact API being called. Search the documentation for that version where possible. A generated answer may describe a newer feature or a different library with a similar name. Ask for a direct primary-source link, inspect the documented signature and create a minimal call. If you still cannot reconcile the behaviour, preserve the uncertainty and ask an experienced maintainer with the evidence packet.
Know when to stop local experimentation. If the next step would involve real private data, account permissions, a database migration or a production deployment, the learning exercise has reached a different risk boundary. Prepare the diagnosis, proposed test and expected result for the responsible person. A good debugger does not have to press every button personally. Knowing which evidence can be gathered safely, and which action requires additional authority, is part of competent practice.
Contents · Next: 17. Use SI as a tutor, then reduce the help
17. Use SI as a tutor, then reduce the help
The most useful assistance level changes during learning. At first, you may need help interpreting an exception or distinguishing equality from identity. Later, you should be able to generate a hypothesis and ask only whether your test separates it from an alternative. Eventually, you should solve a comparable problem without SI and use it only for a review after your own explanation is complete. Keeping the assistant permanently at the highest level of intervention can hide a lack of transfer.
Use a hint ladder. First ask for a question about the failing behaviour. If you remain stuck, ask which intermediate value would be informative. Next ask for a competing hypothesis or a small test. Request a patch only after you can state why the test supports a particular cause. This is a proposed practice structure, not a claim that every learner must follow a fixed timetable. Move up or down the ladder according to the skill you are trying to build.
A twenty-five-minute session can be enough for one focused case. Spend the opening minutes reading the contract and predicting behaviour. Reproduce the failure and design a discriminating test. Make a small repair, run the relevant checks and explain the result without the conversation visible. If time runs out, stop with a clear record of what is known and the next observation needed. Do not turn the final minutes into a rush of unreviewed patches just to produce a green screen.
Keep a concise error notebook organised by misunderstanding rather than by embarrassing incident. A useful entry might say: “I confused a false value with a missing value; the zero-and-None test separated them.” Another might say: “I checked returned contents but forgot input mutation; comparing the caller’s array exposed the incomplete fix.” Record the discriminating test and the principle that transferred, not pages of chat. The notebook should help you recognise a future situation without encouraging rote patch copying.
Repeat a case under changed conditions after a break. Change the threshold, preserve display spelling while deduplicating, or reverse completion order. Start from the requirement rather than from yesterday’s solution. If you immediately paste the old code, you cannot tell whether you remembered the reasoning or merely the text. Ask SI to withhold the answer until after your attempt, and then compare your explanation with the actual output and the contract.
For pair work, one learner can predict and the other can run the test. Switch roles before the repair. Require both people to explain why the chosen input was informative. A teacher or tutor can assess the explanation, the test and the minimal diff, rather than only the final program. Use invented values and respect the class’s AI rules. The learning outcome is a person who can investigate, not simply a shared screen containing working code.
Measure progress with observable behaviours. Can you produce a minimal reproduction without losing the bug? Can you name a test that would reject your preferred explanation? Can you catch a patch that changes the contract? Can you explain an unfamiliar but related failure without asking for the full answer? Those are stronger signals than the number of prompts sent or the speed with which an assistant generates a replacement function.
Contents · Next: 18. Finish with a repair record another person can check
18. Finish with a repair record another person can check
A completed debugging exercise should leave a short account that survives beyond the conversation. State the symptom, the agreed requirement, the demonstrated cause, the changed expression or state rule and the checks performed. Include the important limitation. For the async example, a good record says that both controlled fulfillment orders passed and that rejection and cancellation were outside the example. It should not say that all asynchronous bugs have been fixed.
For the boundary case, the record might read: “Scores equal to the threshold were excluded because the comparison was strict. The contract requires equality to count. I changed that comparison and checked mixed, empty, repeated-boundary, changed-threshold and zero-threshold inputs. The function still assumes numeric inputs.” This is short enough to be useful and specific enough that another learner can reproduce the reasoning. It distinguishes the repair from unrelated possible improvements.
Include evidence at the appropriate level. A screenshot of a green terminal is less useful than the command, relevant test names and result when someone needs to rerun the checks. For a tiny program, save the code and assertions together. For an existing project, preserve the regression in its normal test suite. Avoid claiming “fully tested” when you ran one example, and avoid presenting an assistant’s forecast of test results as if it were an execution log.
Check the diff once more before calling the work finished. Remove temporary prints that expose unnecessary information, retain the useful regression and confirm that no unrelated behaviour changed. If you adjusted a requirement during the investigation, make that decision explicit. The next person should not have to infer from the implementation that missing data now has a different meaning. Good debugging repairs both the program and the explanation of what the program is intended to do.
The final independence check is simple: close the assistant conversation and describe the cause in your own words. Then identify one new input that would matter if the requirement changed. If you cannot do either, return to the relevant case and reconstruct the investigation. You have reached the learning goal when you can produce justified confidence from evidence, rather than borrowing confidence from the fluency of an answer.
Contents · Next: 19. Frequently asked questions
19. Frequently asked questions
Can SI debug my code if I do not understand it yet?
It can help explain the code and suggest observations, but you need enough understanding to judge the intended behaviour and the proposed change. Begin with a small function and ask what each value represents. Predict one output before running it. If a patch contains unfamiliar operations, request an explanation and a separate practice example. Do not treat code you cannot explain as independently verified merely because it executes once.
Should I paste the entire error log?
Usually start with the relevant error and enough surrounding context to show the call or event sequence. Inspect it for secrets and private information first. Preserve the exception type and the meaningful structure, and say what you redacted. If the first packet is insufficient, add the next relevant observation. Sending a larger log is not automatically more informative, and unnecessary detail can both obscure the cause and expose information you should not share.
What if the assistant keeps proposing different fixes?
Stop collecting patches and return to hypotheses. Ask what each proposed cause predicts and choose an input that makes those predictions differ. Keep the original failure reproducible. If successive suggestions do not explain the same observed evidence, the missing step is probably investigation rather than another rewrite. You can also ask a human reviewer to inspect the contract and evidence packet, especially when the code or domain is unfamiliar to you.
Is a passing test enough to accept a repair?
It is evidence for the case it checks. Confirm that it fails against the original defect, that its expected result is independently justified and that important neighbouring behaviours remain intact. The sorting example needs a no-mutation assertion as well as an ordering assertion. A real application may also need integration and security checks. State the tested scope clearly rather than turning one successful run into a claim of complete correctness.
Should I ask SI to write the tests or the fix first?
Write the intended behaviour and at least one expected result first. SI can then suggest additional tests or a diagnosis, but you should review both. For learning, predicting a failing case before seeing the patch makes your understanding more visible. If the assistant writes the function and all its tests from the same mistaken assumption, the whole package can agree with itself while contradicting the requirement.
Why do these examples use small programs instead of a full application?
Small programs make the causal relationship visible and let you run the examples without accounts, services or private data. They isolate boundary logic, shared state, missingness, mutation and completion order. After you understand one mechanism, test how it appears in a larger application. The small program is a learning instrument and a diagnostic aid; it does not replace checking real interfaces and environmental differences.
What if a bug disappears when I add print statements?
Record that observation rather than concluding the bug is solved. The added work may have changed timing or the conditions of execution. Try a controlled reproduction that makes the relevant order explicit, as in the deferred-promise case. For a larger concurrent system, work with its maintainers and established tools. Repeatedly adding arbitrary delays can conceal the symptom while leaving the underlying ordering rule undefined.
Can I use this method for school coding assignments?
Yes, within the assignment’s rules about assistance. You can practise on separate invented examples, learn to read errors and explain your tests. If the assessed work must be independent, do not submit an assistant’s repair as your own unaided work. Keep a record of permitted assistance when required. The most valuable outcome is being able to perform the reasoning again without the assistant present.
What should I do when the correct behaviour is unclear?
Treat that as a specification question. Write the alternatives and an example where they produce different results, then ask the teacher, owner or relevant decision-maker. For example, an empty data set might need an error, an omitted result or a documented default. An assistant can help reveal the ambiguity, but it cannot invent the responsible person’s requirement. Do not label one implementation defective until you know which rule it must satisfy.
How do I know that I am becoming less dependent on SI?
Look for transfer. You should be able to choose a revealing input, interpret its result, reject an incomplete repair and explain the cause in a changed example. Reduce assistance from full solutions to hints and then to after-the-fact review. If you can only repeat a familiar patch, practise a nearby case with a different contract. Independence is visible in the investigation you can perform, not in how quickly you can obtain code.
Contents · Next: 20. Sources and your next learning route
20. Sources and your next learning route
The language and tool references below support the specific technical points linked in the article. The teaching cases, contracts, exercises and practice sequence are original applications, not copied examples or claims about a measured AI model. Documentation can change; when applying a pattern to a real project, check the version and interface used by that project.
- Python: Errors and Exceptions
- Python: Default Argument Values
- Python: Built-in Types and Truth Value Testing
- Python: pdb debugger
- Python: unittest
- Python: os.getenv
- MDN: Array.prototype.sort
- MDN: Promise
- Node.js: Assert
- OWASP: Logging Cheat Sheet
Return to the Super Intelligence learning hub to choose the next skill. If your next task is a new feature, continue with writing code with SI. If the language itself is still unfamiliar, use the coding learning guide and practise one small function at a time.
Before you leave, choose one case and repeat it without looking at the repair. State the contract, make the failure happen, write a prediction, test it, make the smallest justified change and check a changed condition. That sequence is the skill you are building. The next bug will look different, but you will have a dependable way to begin.
