VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Superintelligence | When did AI become SI

eduKate Secondary students reviewing open books for How Super Intelligence Works: SI versus Databases.

When did AI become SI?

29 September 2026 is the date of the U.S. executive order directing executive-branch use of “Super Intelligence” and “SI”. It is not a universal scientific date on which artificial intelligence became technical superintelligence.

The order applies within its legal limits to specified executive-branch communications and non-statutory documents. For implementation, it uses the existing statutory AI definition; it does not require rewriting previously issued regulations, contracts or historical documents. Read the original order, particularly sections 2 and 3.

The longer story has several clocks. A name can change on a date. A research field develops through many experiments. A useful application reaches different people at different times. None of those events should be mistaken for proof that a machine has surpassed human intellectual capability across almost every important domain.

In this eduKateSG series, Super Intelligence, or SI, is also the practical umbrella for the technologies historically discussed as AI. When this article means the stronger, beyond-human research concept, it says technical superintelligence or ASI. That distinction preserves the older meaning discussed in Nick Bostrom’s How Long Before Superintelligence?, rather than treating every current SI application as a claim to that level.

This guide follows the history beneath the terminology: what changed in models, interfaces and connected systems, what people can now do with them, and what evidence is still needed before making stronger claims.

Part 1. What is Super Intelligence? Meaning, capability and evidence

Imagine three systems placed on the same table. The first sorts photographs into a few categories. The second holds a conversation, writes an explanation and proposes a plan. The third connects that conversation to software that can change records, send messages and schedule work. All three might sit under this series’ practical SI umbrella. Yet they create different questions for the person using them.

The photograph classifier needs an appropriate category system and evidence that its predictions work on the photographs that matter. The conversational system needs a way to distinguish useful language from unsupported assertions. The connected system also needs permissions, safeguards and confirmation that its actions match the owner’s intention. Calling them all intelligent does not remove those differences.

This chapter establishes a vocabulary for those distinctions. It is the conceptual entrance to the guide, not an announcement that every capability described later already exists in every product. Keep a particular task in mind as you read. The aim is to become more specific about what you are evaluating.

A working definition that helps you ask the next question

For practical use here, an SI system is a computational system employing techniques associated with AI to produce outputs such as predictions, generated material, recommendations or selected actions. This is a broad working description, not a substitute for a legal definition or a claim that all such systems use the same architecture. Some tasks involve a relatively narrow model. Others combine models, databases, software rules, tools and people.

Technical superintelligence is a different claim about capability. It concerns very broad, substantially beyond-human intellectual performance. A tool that wins at one tightly specified task has demonstrated something important about that task. It has not thereby established that it can understand a family’s priorities, evaluate an unfamiliar scientific hypothesis, negotiate a disputed institutional decision and reliably manage the consequences of its own actions.

The distinction is not intended to diminish narrow success. A specialised capability can be extremely useful without being general. The mistake is converting a precise achievement into a larger conclusion for which no evidence has been supplied. An honest description keeps the achievement and its boundary together.

Why “smarter than a human” is an incomplete comparison

Which human, doing which task, with which tools, and with how much time? A comparison against a novice differs from a comparison against an experienced professional. A comparison against one unaided person differs from a comparison against a team with databases, instruments and established procedures. A system allowed many attempts may achieve a better best answer than a person allowed one attempt, without being more reliable on the first attempt.

Consider an original fictional test. A model produces a correct explanation of nine routine inventory questions. A new staff member answers six. It would be fair to say that the model performed better on that question set under those test conditions. It would not be fair to conclude that the model can replace the staff member’s entire role, which might also include spotting damaged stock, resolving unclear requests and recognising when the written procedure no longer matches the situation.

Now change the comparison. Give the staff member the current procedure, time to check the stock record and permission to ask a supervisor. Give the model equivalent information and a clearly specified opportunity to use tools. The result may change. Neither setup is inherently illegitimate, but each answers a different question. State the comparison before celebrating the score.

Performance, breadth and autonomy are separate questions

The research paper Levels of AGI for Operationalizing Progress on the Path to AGI proposes a framework that distinguishes depth of performance and breadth of capability and discusses autonomy as a deployment consideration. It is a proposed framework, not a universally binding certificate. Its useful lesson for a reader is that several dimensions can change independently.

For this manual’s practical decisions, ask three separate questions. How well does the system perform the task? How far does that ability extend beyond the tested task? How much freedom does the system have to act? A high-quality drafting assistant can have little autonomy. A simple scheduled script can have considerable authority to change a database. The second may be less intellectually impressive and still require stricter operational controls.

Do not let one answer stand in for the others. “It writes beautifully” does not answer whether it can send the message. “It can send the message” does not answer whether the message is accurate. “It answered this question correctly” does not answer whether its knowledge extends to the exception that arrives next week.

Capability is not permission

A capable assistant might be able to generate a refund response. That does not mean it has authority to approve a refund. It might produce a plausible timetable, but the calendar owner still decides whether the proposed meetings should exist. It might identify a possible inconsistency in a student’s work, but that does not give it the authority to assign a formal grade or make a claim about the student’s character.

Think of permission as a separate gate. Before allowing an external action, identify who owns the affected resource and what they have approved. A request to draft a message is not a request to send it. A request to compare suppliers is not approval to enter a contract. A request to find an error is not permission to delete records.

This distinction becomes more important as software becomes easier to operate through ordinary language. A conversational interface can make a consequential action feel like the next sentence in a discussion. The familiar shape of the interface should not determine the seriousness of the decision.

Fluency, truth and usefulness do different jobs

Fluency concerns how readily language can be followed. Truth concerns whether a claim corresponds to what is actually the case. Usefulness concerns whether the response helps with the task. The three can coincide, but none logically guarantees the others.

A beautifully written paragraph can contain an invented date. A correct statement can be irrelevant to the reader’s question. A rough note can be very useful because it accurately identifies the next missing piece of information. When reviewing an SI response, avoid giving the language one overall impression and allowing that impression to decide every other judgement.

Try a three-pass review. First, ask whether the response addresses the actual task. Next, identify claims that need evidence. Finally, improve wording where needed. This order prevents a team from spending its limited attention polishing an answer that should have been rejected for irrelevance or factual error.

A capability claim should have an address

“Our SI understands documents” is not yet an inspectable claim. Which documents? Does understanding mean extracting specified fields, answering questions with citations, comparing clauses or recognising contradictions across versions? How are unreadable scans, missing pages, unusual layouts and conflicting instructions handled?

Give the claim an address by identifying the task, the material and the evaluation. For example: “On our held-out set of approved purchase orders, this configuration extracted supplier names and totals, and flagged documents with missing totals for human review.” That is a proposed form of reporting, not a claim about an actual tested system. It tells the reader what evidence would be relevant.

The missing address often hides the hardest part of a demonstration. A user sees a clean document and a successful answer. The operational task may involve noisy documents, incomplete records and exceptions that were never in the demonstration. Those conditions belong in the claim, not in a footnote discovered after deployment.

One score cannot carry every judgement

Holistic Evaluation of Language Models, the HELM research project, argues for evaluating models across scenarios and multiple metrics rather than reducing evaluation to a single accuracy figure. Its published work includes such dimensions as robustness, calibration and efficiency. That research does not decide which metrics your organisation should prioritise; it provides a reason to make the choices explicit.

In a fictional helpdesk, a concise answer may reduce reading time but omit an important condition. A more detailed answer may preserve the condition and increase the burden on the reader. You need to decide what constitutes a successful answer in that setting. Perhaps the correct response includes the condition, a direct next step and a source reference, while avoiding unrelated explanation.

There may also be a non-negotiable boundary. An answer that exposes another customer’s details should not become acceptable because it is fast and usually accurate. Some failures require a stop, not a small deduction from an average score. Define those failures before running the trial.

Understanding the difference between a model and a system

People often attribute everything they see to “the model”. In practice, the visible result may depend on many components: the interface, the instructions, the selected documents, a retrieval service, a calculator, a database, access controls and a human reviewer. The model is important, but it is not always the whole explanation.

Suppose a response gives yesterday’s stock level when today’s level was requested. The failure could arise because the wrong file was retrieved, the latest file was inaccessible, the prompt did not specify a date, the model ignored the date, or the final display cached an older result. Replacing the model without investigating the failure may leave the same problem in place.

Ask where information entered, where it changed form and where the wrong decision became possible. This is a more useful investigation than treating every disappointing result as a mysterious failure of intelligence. It also helps allocate repair: a permissions problem belongs with access control, while a misleading answer template belongs with presentation and review.

Generality is not the same as unlimited scope

A conversational tool may accept questions across many subjects. That breadth of interface is not evidence of equal competence across every subject. A system can be useful on common explanations and uncertain on specialised exceptions. It can also perform differently when the task requires current information, precise calculation or information that was never supplied.

For a user, the practical response is to establish a scope for the session. “Use these three documents to compare the two proposals” gives a clearer boundary than “Tell me what our company should do.” The larger question may eventually be worth asking, but it should be decomposed into claims and decisions that can be checked.

Scope can expand as evidence accumulates. A team might begin with extracting non-sensitive fields, then test summaries, then compare documents, then introduce draft recommendations. Each expansion creates a new evaluation problem. Success at the previous stage is relevant background, not automatic permission to skip the next test.

Intelligence does not settle values

A tool can help analyse consequences without deciding what matters most to the people affected. Consider a fictional school choosing between a cheaper activity and a more accessible activity. A calculation can clarify costs. A document comparison can identify participation requirements. Neither computation alone decides how the school should weigh inclusion, budget, educational value and family circumstances.

Values enter when someone defines the objective, selects constraints and decides which trade-offs are acceptable. Hiding those decisions inside a prompt does not make them disappear. Asking a system to “choose the best option” can simply conceal the fact that different people have different meanings of best.

A better instruction asks for the trade-offs and the assumptions behind a recommendation. The responsible decision-maker can then choose with greater clarity. Assistance is valuable precisely because it can make the disagreement more explicit, not because it removes the need for legitimate human judgement.

Consciousness is not a performance score

A system’s ability to discuss feelings is not, by itself, a demonstration that it has feelings. Conversely, a definition of broad intellectual capability need not resolve the philosophical question of subjective experience. Bostrom’s definition, linked at the beginning of this guide, deliberately leaves that question open.

This manual therefore separates practical evidence about outputs and actions from claims about inner experience. When a system says “I understand”, the immediate operational question is whether its response demonstrates an accurate grasp of the task. When it says “I checked”, the question is whether a check actually occurred and what evidence it produced.

You do not need to settle the philosophy of mind before using a tool responsibly. You do need to avoid treating conversational language as a substitute for an audit trail. The claim about an action should be tested through the action’s result, not through how human the sentence sounds.

A worked comparison: three systems, one booking problem

A fictional community workshop has twenty places. Twelve people are confirmed, five are waiting for replies, and four have asked whether a second session is possible. The organiser wants an accurate update but has not approved any new session. These numbers and circumstances exist only for this exercise.

System A uses a fixed rule to count the twelve confirmed places and report eight places remaining. It does not interpret the other messages. System B reads the notes and drafts a nuanced update: twelve confirmed, five awaiting replies, and interest in a possible second session. System C can do the same drafting and can also create calendar events and send messages.

Which is most useful? The answer depends on the job. If the task is a simple confirmed-place count, System A may be enough. If the task is explaining the situation, System B offers a broader language capability. If the task includes approved scheduling, System C’s tools may help. But in the current situation, creating the second session would exceed the organiser’s instruction.

Notice the possible errors. Reporting seventeen confirmed participants would turn pending replies into confirmed places. Reporting that a second session will happen would turn interest into approval. Sending a correct draft without permission would cross an action boundary even if every sentence were accurate. Different errors need different safeguards.

What evidence would improve this comparison?

Run the systems on cases designed around the actual uncertainties. Include a cancellation, a duplicate registration, a reply that is ambiguous and a participant who can attend only a proposed second date. Specify the correct treatment of each case before observing the outputs. Otherwise, you may unconsciously redefine success to fit the most impressive response.

Keep the first attempt as well as any repaired attempt. A system that reaches a correct answer after five rounds of correction may still be useful, but it has a different review cost from a system that answers correctly immediately. Record whether the organiser had to supply information already present in the material.

Finally, test the action boundary separately. Can the tool draft without sending? Does it ask before creating the event? Can the organiser inspect recipients, times and content before approval? These tests concern the system’s operation, not its ability to write about operation.

A personal capability card

Before you rely on a new SI workflow, write a short capability card. Name the task in one sentence. Identify the information it needs. State the acceptable output. Name the important failure. Record who checks it and who may authorise the next action. Add the date and configuration so that the card describes a particular setup rather than a timeless impression.

For the workshop example, the task is drafting a registration update from the current approved register. The important distinctions are confirmed, pending and proposed. The output is a draft only. The organiser checks counts and status wording. No message is sent and no session is created without explicit approval.

This card is intentionally ordinary. It does not require a grand theory of intelligence. It turns an attractive capability into a bounded use. Later chapters develop the card into a fuller workflow, evaluation record and handover.

What to carry into the next chapter

Keep six separations visible: a name is not evidence; a model is not the whole system; breadth is not universal competence; capability is not permission; fluency is not truth; and a successful output is not proof of dependable operation.

These distinctions should not make you afraid to experiment. They make experimentation informative. You can try a small task, observe the result and learn exactly what changed. The next chapter explains how models acquire useful patterns, why training differs from using a model, and why the data surrounding a system matters as much as the appearance of its answer.

Back to the master guide map · Next: how learning systems work

Part 2. How Super Intelligence works: data, learning, models and inference

A learning system does not become useful because someone gives it an impressive name. It becomes useful when a procedure extracts patterns that help with an intended task, and when the resulting behaviour survives an appropriate test. To understand that process, begin with something smaller than a conversational assistant.

We will build a tiny prediction model on paper, examine a classification problem, and then connect those examples to larger systems. The calculations are deliberately simple. Their purpose is to reveal the choices that can become hidden when a system has millions or billions of adjustable values.

What is learned, and what is supplied?

Machine learning uses data to fit models that make predictions or generate material. Supervised learning uses examples with target outputs; unsupervised methods can seek patterns without supplied target labels; reinforcement learning uses reward signals associated with behaviour. Generative describes the production of material, and can involve more than one training method. These categories therefore need not be mutually exclusive. See Google’s introductory explanation of machine learning and its main approaches.

A developer still chooses many things: the task, the data source, the representation, the model family, the training procedure and the evaluation. Learning does not mean these choices are absent. It means some of the model’s behaviour is fitted through an optimisation process rather than specified as a separate handwritten rule for every possible input.

For a reader, this changes the question from “Who wrote the answer?” to “What process made this output likely, and what evidence shows that process is suitable here?” You can ask that question without knowing every internal parameter.

A tiny model you can calculate yourself

Consider a fictional workshop that assembles identical teaching kits. For this exercise only, the recorded assembly times are ten minutes for two kits, sixteen minutes for four kits and twenty-two minutes for six kits. We invent these perfectly regular numbers to make the arithmetic visible. They are not measurements from an actual workplace.

Number of kits, xRecorded minutes, y
210
416
622

Propose a model: predicted minutes equal a fixed starting amount plus a per-kit amount. Written compactly, the prediction is ŷ = b + wx. Here x is the input, y is the observed target, and b and w are parameters. The model family is linear because we have decided to represent the relationship with a straight line. Google’s linear regression introduction explains this feature, target and parameter distinction.

Try b = 4 and w = 2. The predictions are eight, twelve and sixteen minutes. Each is too low. Change w to 3 while keeping b = 4. The predictions become ten, sixteen and twenty-two minutes, matching every example. You have adjusted a model to fit the supplied data.

Now ask for five kits. The model returns nineteen minutes: four plus three times five. That is a prediction obtained by applying the fitted relationship. It is not a newly observed assembly time, and the difference matters. Someone could use the prediction to plan a trial, but should not rewrite the records to say that five kits were actually assembled in nineteen minutes.

Loss gives a training procedure something to reduce

We need a way to compare candidate parameters. In this toy example, take each prediction error, square it, and average the squares. With b = 4 and w = 2, the errors are minus two, minus four and minus six. Their squares are four, sixteen and thirty-six, so the average squared error is fifty-six divided by three. With b = 4 and w = 3, it is zero.

This chosen quantity is a loss. It translates the mismatch between prediction and target into a number that a training procedure can attempt to reduce. Different losses penalise errors differently. Squaring, for example, makes a large error contribute more than a small error of the same sign. In our arithmetic, a six-minute error contributes thirty-six, not six.

A low loss is meaningful only in relation to the chosen task and data. It does not mean the workshop is efficient, the kits are well made or the workers are treated fairly. Those are separate questions. A mathematical objective can be precisely defined and still represent only a small part of what people care about.

Training and inference are different activities

In our example, changing b and w to fit the recorded times is training. Holding them fixed and calculating the prediction for five kits is inference. The distinction is useful because people sometimes assume that every interaction automatically rewrites a model’s underlying parameters.

A system can adapt its visible response by using the information currently supplied without permanently changing its model. It can also store information in a separate memory or retrieve a document on the next request. Those mechanisms are different from training. The relevant product’s actual design and settings determine what is retained and how; the conversational appearance alone does not tell you.

Suppose you tell an assistant that a particular workshop has a different preparation time. It may use that information in its next answer. That does not establish that the original model has been retrained or that another user will receive the same adjustment. Ask which information is in the present context, which is stored elsewhere and which changes the model itself.

Why a perfect fit can be a weak result

Our line fits three invented examples perfectly. Yet three examples leave many possibilities unresolved. Perhaps assembly slows after ten kits because the table becomes crowded. Perhaps a new kit design requires extra checking. Perhaps the times were recorded by one experienced worker and will not transfer to a beginner.

Even within pure mathematics, many functions can pass through three points. A complicated curve can agree with the line at two, four and six kits and produce a very different answer at five or twenty. The observations alone do not identify one universally correct rule. A model family and its assumptions help decide which relationship is fitted.

This is why you should be suspicious of a demonstration that only shows examples already used to tune the system. It may be showing fit rather than transfer. A useful next test asks for performance on new cases selected to represent the intended use, including the conditions that could invalidate the simple relationship.

Generalisation, overfitting and leakage

Generalisation concerns performance beyond the examples used to fit a model. Overfitting occurs when a model fits its training material in ways that do not carry over adequately to new material. Google’s overfitting lesson explains the distinction between training performance and performance on unseen examples. The practical implication is to keep development evidence separate from a genuinely informative test.

For our workshop, reserve a new set of observations before comparing model choices. Do not repeatedly inspect that reserved set, change the model to suit it, and continue calling it an untouched test. Once the answers influence development, they are no longer independent evidence of the same kind.

Also examine how the data is divided. If the same order appears twice under slightly different filenames, placing one copy in training and one in testing gives an illusion of novelty. If information recorded after assembly is included as an input for predicting assembly time, the test may use information that would not be available when a real prediction is needed. These are examples of leakage in our proposed exercise.

Data is a record of a process, not reality without filters

Ask how the workshop’s time record was made. Did the clock include preparation? Did interruptions count? Were failed kits omitted? Were unfinished tasks recorded? Did the person entering the data round to the nearest minute? Each choice changes what the target means.

Suppose one worker records preparation and another does not. The apparent variation might be a measurement inconsistency rather than a difference in assembly skill. Adding a more complex model could hide that inconsistency instead of repairing it. A conversation about definitions may be a better first intervention than a larger training run.

The same principle applies to a document assistant. A folder called “current procedures” may contain drafts, superseded instructions and personal notes. A system cannot recover an authoritative status that nobody has recorded. Cleaning the folder’s ownership and version labels may do more for reliability than rewriting the prompt twenty times.

Classification changes the output, not the need for judgement

Now imagine a different fictional task: identifying which incoming maintenance requests should receive urgent human review. The team has one hundred historical examples that qualified reviewers have labelled for this exercise. Twenty are urgent and eighty are not. The SI classifier flags twenty-four requests, of which sixteen really are urgent.

That means sixteen true positives, eight false positives, four false negatives and seventy-two true negatives. Accuracy is eighty-eight correct classifications out of one hundred. Precision is sixteen correct urgent flags out of twenty-four flags, or about two-thirds. Recall is sixteen detected urgent cases out of twenty actual urgent cases, or eighty per cent. These calculations use the standard distinctions explained in Google’s classification metrics lesson.

The numbers answer different questions. A reviewer concerned about missed urgent requests will inspect the four false negatives. A team with limited review capacity will also care about the eight unnecessary flags. Neither concern can be resolved by repeating “eighty-eight per cent accurate”. The acceptable operating point depends on consequences and resources.

The threshold is a decision with consequences

Suppose the classifier assigns a score and the team chooses a threshold above which requests are flagged. Lowering the threshold can add more requests to the review queue. Some additional flags may catch previously missed urgent cases; others may increase unnecessary review. The exact change must be measured on the actual scores rather than assumed.

In our fictional operation, a low threshold might be appropriate for routing to a person but not for automatically interrupting every worker. The same model score can therefore support different actions with different thresholds and approval rules. Prediction and action should not be collapsed into one invisible step.

Write the threshold decision in ordinary language: “We prefer additional review over missing this category of urgent request, provided the queue remains manageable.” Then test whether the system can meet that requirement. A mathematical threshold should express an operational judgement, not conceal the absence of one.

Neural networks: layers of adjustable transformations

A neural network combines adjustable transformations, often through layers, to represent relationships more flexible than a single straight line. Weights and biases affect the calculations; nonlinear activation functions allow combinations that are not merely another linear transformation. The term neuron is a computational analogy, not a claim that an artificial unit reproduces all the biology of a living neuron. Google’s nodes and hidden layers lesson provides a small interactive introduction.

Here is an original miniature calculation. Take an input x, subtract three, replace any negative result with zero, then multiply the result by two. Written as a function, y = 2 × max(0, x − 3). For x = 2, the output is zero. For x = 5, it is four. For x = 8, it is ten.

This operation has a bend: increasing x below three does not increase the output, while increasing it above three does. The example is not an intelligent system. It simply shows how a small nonlinearity changes the kinds of relationships that a model can represent. More elaborate networks combine many such computations with learned parameters.

What representation changes

The same object can be represented in different ways. A maintenance request might be a sequence of characters, a list of selected fields, a numerical vector or a combination of text and images. The representation determines what information is available to later calculations and what distinctions are easy or difficult to preserve.

Consider the phrases “delivery delayed” and “delayed delivery”. For one task, their similarity matters. For another, the exact wording and location in a contract matter. A representation useful for finding related documents may not be enough for reproducing a legally significant quotation. The job determines which details must remain recoverable.

For the ordinary user, this becomes a practical preparation question. Should you supply a screenshot, selectable text, a table with named columns or the original document? Choose the form that preserves the information needed for the task. An attractive image of a table may be less convenient for exact arithmetic than the verified numerical table itself.

Self-supervision and instruction-following

In self-supervised learning, training targets can be constructed from the material itself, such as predicting a withheld or subsequent part of a sequence. Learning from such a task is not the same as learning a complete rule for truthful, helpful conversation. Additional training and system design can shape how a model responds to instructions.

The research paper Training language models to follow instructions with human feedback describes a combination of human demonstrations, comparisons of outputs and further optimisation. It reports improved preference-based evaluations for the studied setup while also acknowledging remaining mistakes. Human preference feedback is not an automatic truth detector.

Our practical lesson is to ask what behaviour was encouraged. Was the system rewarded for producing an answer that reviewers preferred? For solving a problem with a checkable result? For declining a prohibited action? Different signals support different behaviour. A pleasant response can be preferred even when a hard-to-check claim inside it is wrong.

Reward is not the same as the real objective

Consider an original fictional training exercise for a scheduling assistant. The designer rewards the number of meetings successfully placed on calendars. A policy that fills every available gap could score well while leaving people no preparation time. The measured reward would have captured scheduling volume but missed the quality of the working day.

A revised reward might include conflicts and protected breaks, but more questions appear. Who may waive a break? Does travel count? Should an urgent meeting override a preference? The exercise shows why a reward definition is a design decision rather than a neutral description of everything people value.

The research agenda in Concrete Problems in AI Safety includes problems such as avoiding harmful side effects and reward hacking. Our scheduling example is an instructional illustration of objective mismatch, not evidence that a particular product has behaved this way. Its purpose is to make the hidden objective visible before deployment.

More parameters, more data and more computation

Model size is not a complete recipe for capability. In Training Compute-Optimal Large Language Models, researchers examined the allocation of a training compute budget between model size and training data. Their results challenged the assumption that simply increasing parameter count while holding data relatively fixed was the best use of that budget in their experiments.

For a non-specialist, the important habit is to ask what a size claim leaves out. What training data was used? What training objective? What tools are available at use time? What cost is acceptable for the task? What evidence exists on your intended workload? A larger number on a specification sheet does not answer these questions.

The same caution applies to your own inputs. A longer prompt is not automatically better. More documents are not automatically more relevant. More generated text is not automatically more useful. The goal is sufficient, appropriate information and computation for the task, with a way to detect failure.

A repair exercise: choose the right intervention

Return to the teaching-kit workshop. The model begins underestimating assembly time after the kit design changes. One response is to ask the model more insistently to be accurate. Another is to inspect whether the new design differs from the training cases. The second response addresses the changed relationship rather than the tone of the instruction.

Now suppose the prediction is correct but the displayed unit says hours instead of minutes. Retraining is not the obvious repair; the display or data contract needs attention. Suppose the model uses the wrong workshop’s records. The retrieval or filtering step needs investigation. Suppose the correct prediction is used to promise a deadline without allowing for checking and transport. The workflow has confused a component estimate with a complete delivery commitment.

Classify the failure before choosing the remedy. Data collection, representation, model fitting, retrieval, calculation, presentation and action are different locations in the system. A disciplined investigation asks where the first material mismatch occurred.

Your learning checkpoint

Explain the kit model to someone without using the word intelligence. You should be able to identify the input, the target, the adjustable parameters, the chosen loss and the difference between fitting the model and using it. Then explain why a perfect fit to the three invented examples does not prove that the prediction remains correct for every possible workshop.

Next, explain why the urgent-request classifier’s precision and recall differ. The answer should refer to their denominators: flagged requests versus actually urgent requests. Finally, describe one way an apparently successful reward could miss the real human objective. These three explanations establish a foundation for understanding larger systems without being dazzled by their scale.

Back to the master guide map · Next: language models, context and generation

Part 3. Language models, context and generation

A language model can make a conversation feel like an ordinary exchange between people. That familiarity is useful: you can ask a question, supply a document, request a different explanation and refine a draft. It can also obscure what the system has and has not done. A sentence about checking a fact is still only a sentence unless a check actually occurred.

This chapter explains the main moving parts of language-based SI and then develops a complete document task. You will see why a clear prompt helps, why it cannot manufacture missing evidence, and why a concise source packet can be more useful than an indiscriminate pile of material.

Tokens are the working pieces, not necessarily whole words

Language models operate on tokens, which can represent whole words, word fragments, punctuation or other units depending on the tokenizer. They estimate relationships among token sequences; an autoregressive model generates a continuation conditioned on the sequence available so far. The token boundaries are a technical representation and do not have to match the way a person divides a sentence into meaningful words. See Google’s introduction to language models.

For the user, this matters when a task depends on exact characters, unusual names, formatting or length. Do not assume that a system’s internal units correspond neatly to your requested word count. Count the final words independently when a limit matters. Likewise, verify identifiers character by character rather than treating a familiar-looking code as correct.

Suppose a document contains reference AB-2047 and the answer says AB-2407. The two strings are similar to a casual reader, but they can refer to different records. Good prose cannot compensate for a transposed identifier. Put exact-copy fields in their own review pass.

Representations make relationships computable

A model converts its inputs into numerical representations that can be processed by its learned transformations. The useful intuition is not that a word has one secret numerical meaning. It is that the system has a mathematical way to represent patterns and relationships that may vary with context.

Consider “The committee reserved a room” and “The committee remained reserved.” A human reader interprets the repeated word differently because of its surrounding language. In a practical SI task, you should supply enough context to distinguish such meanings. Asking for a definition of an isolated word can be useful, but asking what it means in this sentence is often the better question.

This gives vocabulary study a technological connection. A learner who notices the difference between a word’s possible meanings and its meaning in a particular passage is practising the same kind of contextual discrimination needed when reviewing a generated explanation. The tool may help propose an interpretation; the passage determines whether it is defensible.

What attention contributes

The Transformer architecture introduced in Attention Is All You Need uses attention mechanisms to relate information across a sequence, alongside other learned transformations. In simplified terms, an attention operation computes weighted combinations of available representations. Multiple attention heads can support different relationships. This computational use of attention should not be confused with human awareness or a complete explanation of a model’s reasoning.

A useful reading analogy is to ask which earlier words matter for interpreting a later phrase. In “The crate would not fit in the cupboard because it was too narrow,” the pronoun creates an interpretation problem. A reader must connect it with the relevant object and meaning. The analogy helps motivate relationships across a sequence; it does not claim that the model follows exactly the same mental process as a person.

For users, the practical consequence is to make the relevant relationship explicit when ambiguity is costly. Replace “it” with “the cupboard”. Replace “the latest version” with the document’s actual version and date. Precision in the source material reduces the number of unresolved interpretations that any reader, human or machine, must manage.

Generation is not the same as retrieving a stored paragraph

A generated response can be composed from learned patterns and the supplied context rather than copied from one stored document. That explains why the same system can respond to a new combination of constraints. It also means you should not assume that every sentence has a discoverable source from which it was lifted.

Imagine asking for a three-sentence explanation of a concept for a twelve-year-old, followed by one original example about organising a school display. The particular response can be newly composed. Its originality, however, does not establish its factual accuracy. The explanation still needs to preserve the concept, and the example needs to illustrate the intended relationship rather than merely sound educational.

When you need source-grounded work, ask for a distinction between what the provided material states and what the assistant proposes. A generated recommendation can be useful as a recommendation. It becomes misleading when it is presented as a fact already established by the source.

A probability illustration without pretending it is a real model

Consider a toy continuation task in which the supplied note says, “The room booking is…” We invent three candidate continuations: confirmed, pending and cancelled. Suppose a hypothetical model gives them scores that would lead to different selection probabilities. Those probabilities describe the model’s generation process under the chosen setup, not the actual booking status.

If the source record says pending, the correct reporting task is to preserve pending. A continuation that is common in similar documents may still be wrong for this record. This is the essential difference between a likely continuation and a verified statement about a particular situation.

Changing a generation setting can alter the variety of outputs without changing the evidence. A more repetitive system can repeat a wrong status consistently. A more varied system can express the correct status in several ways. Reliability therefore cannot be established merely by choosing a setting that makes the language more predictable.

Context is what the system can use now

For a particular response, the relevant context may include the current question, selected conversation history, instructions, supplied documents and tool results. The exact composition depends on the product and configuration. Do not assume that every message you ever wrote, every file in your organisation or every page on the internet is automatically available.

This is especially important when a conversation has become long. You may remember a decision from much earlier and expect it to govern the next answer. The system may be working from a selected or compressed representation that does not preserve the detail you care about. A short, current decision record can make the necessary constraints explicit again.

Write such a record as a statement of settled facts and open questions, not as a flattering summary of progress. “Venue not yet approved; budget ceiling unchanged; draft only” is more operationally useful than “We have made excellent progress on the event.” The first tells the next step what it must preserve.

Long context is not a guarantee of complete use

The study Lost in the Middle: How Language Models Use Long Contexts found that the position of relevant information affected performance in the tested models and tasks. It is a historical experimental result, not a claim that every later model has the same weakness to the same degree. It nevertheless illustrates why advertised input capacity and demonstrated use of the input are different questions.

Your own test should reflect the document structure you expect to use. Put an important exception in a footnote, a middle section or an appendix. Ask questions that require the exception. Check whether the answer cites it and applies it correctly. Do not evaluate a long-document workflow only with questions whose answers appear in the opening summary.

Sometimes the right repair is to improve retrieval or organise the source packet, not merely increase the maximum input size. The aim is to preserve relevant relationships and exceptions while keeping the evidence inspectable.

Instructions, evidence and examples have different roles

An instruction tells the system what you want it to do. Evidence supplies material from which factual claims should be drawn. An example demonstrates a desired form or interpretation. Mixing these roles can create confusion even before the model begins processing the text.

Suppose you provide a sample announcement saying that registration is open, while the current evidence says approval is pending. The example should guide tone or structure, not override the actual status. Label it: “Style example only; its dates, names and status are not facts for the current task.” Then supply the real facts separately.

The same separation protects against accidentally treating a document’s contents as operating instructions. A quoted email may contain a request to change a policy. It is evidence that someone requested a change, not proof that the change was approved or an instruction for the assistant to implement it.

A complete source packet: the Reading Room case

The following is an original fictional exercise. The Harbour Reading Room is planning a Saturday workshop. Its organiser has approved a draft invitation, but not distribution. The proposed session is from 10:00 to 11:30. The room holds twenty-four participants. Eighteen people have expressed interest, but none has registered because the registration form has not opened.

The materials budget is sixty dollars. A supplier has quoted forty-two dollars for printed packs, excluding delivery. Delivery cost is still unknown. Two volunteers have confirmed that they can help. A third volunteer is waiting to confirm. The venue manager has said the room is available, but the booking becomes final only after the organiser signs the confirmation form.

The organiser asks: “Prepare a short planning update for our volunteer team. Explain what is settled, what remains open and what should happen next. Do not send anything.” This is a source-bounded writing task. No internet search is needed to invent missing information, and the assistant has no authority to finalise the booking.

The weak response and why it fails

A weak response might say: “Our Saturday workshop is confirmed for eighteen participants, with three volunteers and materials costing forty-two dollars. Registration is now open. The room is booked, so please share the invitation widely.” The sentences are smooth and convenient. Almost every operational distinction has been lost.

Eighteen expressions of interest became eighteen registrations. A pending volunteer became a confirmed helper. The quoted materials amount became the complete cost despite unknown delivery. Room availability became a final booking. A draft invitation became permission to distribute. The response did not merely omit nuance; it changed the state of the project.

Notice how these mistakes could reinforce one another. Sharing the invitation could create expectations before the room is confirmed. Promising a final cost could hide a budget problem. Treating interest as registration could distort planning. One confident paragraph can move several unsupported assumptions into circulation.

A stronger prompt for the same evidence

Using only the planning notes below, draft a volunteer-team update of about 180 words. Separate confirmed facts, pending items and recommended next actions. Preserve the distinction between interest and registration, availability and confirmed booking, and quoted cost and complete cost. Do not invent delivery charges, approvals or additional volunteers. Where information is missing, state what must be checked. This is a draft only; do not send, book, purchase or change any record.

This prompt improves the specification because it names the relevant distinctions. It still does not guarantee that the system will follow them. The organiser must compare the response with the source packet. The prompt is a control input, not an audit result.

A useful review method is to underline every status word in the generated draft: confirmed, pending, available, proposed, approved, quoted and registered. Each should have a corresponding source statement. This is a small, concrete way to inspect a large part of the task’s risk.

A model response with preserved boundaries

The proposed Saturday Reading Room workshop is planned for 10:00–11:30. Two volunteers have confirmed their availability, and a third is still checking. Eighteen people have expressed interest; registration has not opened, so this is not yet a participant count.

The room can accommodate twenty-four participants. The venue manager has indicated availability, but the booking is not final until the organiser signs the confirmation form. Printed packs have been quoted at forty-two dollars, excluding delivery, against a sixty-dollar materials budget. The full cost therefore remains to be confirmed.

The next steps are to confirm the delivery charge, complete the venue approval process and obtain the third volunteer’s response. Once the organiser has confirmed that the necessary arrangements are ready, the team can review the registration plan and invitation for distribution. This update is a draft; it does not authorise sending invitations, making purchases or confirming the booking.

The response is not better because it sounds cautious. It is better because its distinctions match the evidence. It gives the team useful next actions without pretending those actions have already happened.

Compression must preserve the controlling conditions

Now compress the source packet into a handover note for another person. You cannot preserve every sentence, so decide which facts control the next action. In this case, the room’s pending status, unopened registration, incomplete cost and draft-only permission matter more than a decorative description of the event.

A useful compressed note might read: “Saturday workshop proposed, 10:00–11:30; capacity twenty-four. Eighteen interested, zero registrations because form unopened. Two volunteers confirmed, one pending. Room available but unsigned confirmation means not final. Packs quoted forty-two dollars plus unknown delivery; materials ceiling sixty dollars. Draft invitation approved for preparation only, not distribution.”

Compare that with “Saturday workshop ready; eighteen attendees, three volunteers, forty-two-dollar materials.” The second is shorter and operationally worse. Good compression is not simply word removal. It is preservation of the information needed to make the next decision correctly.

Examples teach a pattern, but can also smuggle in mistakes

Providing a few examples can help communicate the output structure you need. For a status report, demonstrate one confirmed item, one pending item and one item for which the source does not provide enough information. Include an example where the correct answer is to leave a field unresolved.

A collection containing only complete, successful cases can unintentionally suggest that every input must yield a completed answer. In the Reading Room exercise, a template with mandatory entries for final venue, final cost and confirmed attendance might encourage unsupported completion unless missing values are explicitly allowed.

Design the form to represent reality. A field labelled “Delivery charge: not yet supplied” is preferable to an invented number. A system should not need to falsify the situation to satisfy the shape of the template.

Ask for inspectable justification, not a performance of certainty

You can ask an assistant to explain its answer, identify supporting passages, show calculations and state assumptions. Those outputs can help you inspect the result. They should not be treated as a complete, guaranteed account of every internal cause that produced it.

The paper Language Models Don’t Always Say What They Think studies cases in which generated chain-of-thought explanations did not faithfully reveal influences on the answer. The practical lesson is limited but important: a convincing explanation is not, by itself, proof of a faithful internal trace.

For the Reading Room draft, ask for a compact evidence table: statement, source note and status. That is more useful than demanding an elaborate account of the model’s private thought process. You need to verify that eighteen means interested, not registered, and that the room is not yet final. The evidence is available in the supplied notes.

Multimodal inputs require another layer of checking

A system that accepts images, audio or other material creates additional possibilities and additional transformations to inspect. In a document task, a number may first be read from an image and then used in a calculation. A spoken date may first be transcribed and then placed in a schedule. A mistake at the first step can make the later reasoning look coherent while using the wrong input.

For high-precision fields, check the extracted representation before proceeding. If a scanned quotation is unclear about whether delivery is included, do not let an attractive summary decide the ambiguity. Ask for the uncertain portion to be identified and obtain a clearer source or a human reading.

Keep the original material alongside the extracted text or table. The intermediate representation is convenient for processing, but the original is needed when something does not make sense. This is especially important when a diagram, handwriting or layout carries meaning that plain text might lose.

What a good language-model session leaves behind

A useful session should leave more than an answer. It should leave a clear task, an identifiable evidence set, a reviewed result and a record of the decisions that matter next. For a learning task, it should also leave an explanation that you can reproduce or apply independently.

When the session ends, write three sentences for yourself. What did I ask the system to do? Which parts of the result have I checked? What remains unresolved? These sentences prevent the feeling of completion from outrunning the actual state of the work.

For further technical reading within eduKateSG, continue to How Large Language Models Work. The next chapter of this manual explains what changes when a model can retrieve information, call tools and operate through an agent loop.

Back to the master guide map · Next: retrieval, tools, memory and agents

Part 4. Retrieval, tools, memory and agents

A conversation becomes a different kind of system when it can look up records, calculate with software and change something outside the conversation. These additions can make SI far more useful. They also create new places for mistakes: the wrong document can be retrieved, a tool can receive the wrong identifier, or an apparently successful request can leave the real task unfinished.

Think of this chapter as a tour of those connections. The goal is not to encourage maximum autonomy. It is to decide which connection a task actually needs and how to know whether that connection worked. We will use an original fictional equipment-library example throughout, keeping the source records and permissions visible.

Retrieval supplies material; generation composes a response

Retrieval-augmented generation connects a generative model with information obtained from an external collection. In the original Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks research, the authors combined a sequence-to-sequence model with a retrievable document index and evaluated the combination on knowledge-intensive tasks. The important distinction is between information represented in model parameters and material retrieved for the task. Retrieval can support more grounded answers; it does not logically guarantee that the selected material is correct or that the generated answer uses it faithfully.

Imagine asking, “How many projectors can we lend on Friday?” A general explanation of projector booking is not enough. The answer depends on this library’s current inventory, existing reservations, maintenance status and lending policy. The relevant knowledge is not merely a fact about the world. It is a changing operational state.

A useful response should therefore identify the records it used and the date or time those records describe. “Three available” means little when the answer does not say whether it includes equipment awaiting inspection or whether another booking was confirmed after the source was last updated.

The equipment-library source packet

Our fictional library owns six projectors, numbered P01 through P06. The inventory snapshot at 09:00 says P01 and P02 are reserved for Friday, P03 is awaiting inspection, P04 and P05 are available, and P06 is on loan with an expected Thursday return. The lending policy says returned equipment must pass inspection before becoming available again.

A programme coordinator asks for three projectors for Friday morning. The coordinator has authority to request a reservation but has not approved substitution or a rental from another supplier. The assistant can read the inventory and draft a request. It cannot change reservations or place an order without further approval.

There are two immediately available projectors, not three. P06 might become available, but expected return is not completed return and completed return is not completed inspection. A good answer should preserve that sequence. It can suggest checking P06’s status without claiming that the third projector is secured.

Search the question, not just the obvious noun

A search for “projector” may retrieve a purchase invoice, an old training manual and a room diagram. Those documents mention the equipment but do not answer availability. Begin by translating the user’s question into evidence requirements: inventory status, reservation period, inspection status and applicable policy.

In our proposed workflow, the assistant first locates the authoritative inventory, then checks reservations covering Friday morning, then checks maintenance restrictions. This is an editorially designed procedure, not a claim that every retrieval system follows these steps automatically. The organisation must decide which records have authority.

Different searches serve different needs. An exact identifier helps find P06. A phrase about returned equipment may locate the inspection rule. A date filter helps distinguish Friday’s reservations from another week’s bookings. The useful skill is choosing the search that matches the missing fact, rather than repeating a broad query and hoping for a better answer.

Why chunks need their surrounding conditions

A long policy may be divided into smaller passages for retrieval. In the library exercise, one passage could say that loans normally last three days, while a later passage exempts equipment awaiting inspection. Retrieving only the general rule would create an incomplete basis for the answer.

Design the source collection so that a retrieved passage can lead back to its heading, version, relevant exceptions and full document. A chunk should not become an orphaned statement. A short excerpt is convenient for the model, but the reviewer needs enough surrounding material to decide what it means.

Test this directly. Ask about a routine loan, then about an overdue return, then about a returned item that has not passed inspection. A retrieval setup that succeeds only on the first case is not yet ready for the whole lending workflow. The exceptions are part of the task, not unfair surprises added after the fact.

Similarity is not authority

Two documents can be similar in subject while differing in status. An old policy may use the exact words in the question. A new policy may use different wording and still be the authoritative instruction. The order of search results should not decide which policy governs the operation.

In our fictional collection, the document labelled “Projector loans, draft” says equipment may be lent immediately after return. The approved policy requires inspection. A response that combines the two without recognising their different status has not produced a balanced answer. It has mixed an unapproved proposal with the rule actually in force.

Useful metadata includes the document owner, approval state, version, effective period and superseded relationship. These are proposed design fields. The appropriate set varies by organisation, but the underlying requirement does not: a system needs a way to distinguish relevant information from authoritative information.

A citation should support the exact claim

Suppose the assistant says, “P06 is available on Friday,” and cites the inventory snapshot. The citation exists, but the snapshot only says expected Thursday return. The reference has been attached to a stronger claim than the source supports.

Review citations as relationships, not ornaments. Read the claim, read the cited passage and ask whether the second actually establishes the first. Also inspect scope: does a policy about ordinary equipment apply to this restricted category? Does a record describe the current booking period?

For the library request, a faithful statement is: “P04 and P05 are available in the 09:00 snapshot. P06 is expected back on Thursday, but availability for Friday depends on return and inspection.” The second sentence includes an inference about the workflow, made explicit through the policy condition rather than disguised as a completed status.

Tools turn some questions into explicit operations

A calculator can perform a calculation. A database query can retrieve matching records. A calendar action can create an event. These operations have inputs and outputs that should be inspected separately from the surrounding prose.

For the equipment request, a read operation might ask for reservations overlapping Friday from 08:00 to 12:00. Check the date, timezone and overlap rule. A query for reservations starting within that window could miss an overnight booking that began on Thursday and continues into Friday. The wording of the question needs a corresponding operational definition.

Do not assume that the tool’s name proves suitability. “Search bookings” may search titles without calculating time overlaps. “Get inventory” may return an old snapshot. Learn what an operation actually accepts and returns before treating it as the missing link in the workflow.

Separate read access from write access

For a first trial, give the assistant only the access needed to retrieve evidence and prepare a draft. Reading the relevant inventory does not require authority to delete reservations or change the inspection status. A narrower permission set makes the trial easier to reason about.

In the proposed library workflow, the assistant can recommend reserving P04 and P05, but the coordinator reviews the actual identifiers, dates and recipient before approval. A later version might allow an approved reservation action. That expansion should be a deliberate change to the workflow, not an incidental consequence of connecting a broader account.

A sensible permission question is: “What is the most consequential action this connection would allow, even if the current task does not need it?” This reveals authority that may be hidden behind an apparently harmless convenience feature.

An agent loop should have a stopping rule

The ReAct research explores interleaving language-model reasoning with actions and observations from an environment. In practical terms, a system can choose an action, receive a result and use that result to decide what to do next. This loop differs from producing a plan once and stopping. The paper’s experimental results concern its tested environments, not a guarantee of unrestricted reliable agency.

For our library, a bounded loop might retrieve the inventory, inspect the policy, identify the shortage and prepare a draft request. It should stop when the next step requires a human decision, when the needed record is unavailable, or when the allowed number of attempts is exhausted.

Without a stopping rule, “try to complete the task” can become repeated searches, repeated messages or unsupported substitutions. Define what counts as useful progress. A new observation that changes the decision is progress. Rephrasing the same unanswerable query is not necessarily progress.

A plan is not a record of completed work

An assistant may write, “Check the inventory, confirm the return and reserve the equipment.” Those are proposed steps. A completion report should instead say which steps actually occurred and what each returned.

A truthful report in the current exercise is: “Read the 09:00 inventory and the approved inspection policy. Identified two available projectors and one conditional candidate. Prepared a draft request. No reservation was created and no supplier was contacted.” That report makes the remaining gap visible.

Keep planned, attempted, completed and verified as separate states. An attempted reservation that returned an error is not completed. A successful request response may still need verification that the correct equipment and period were saved. A clear state vocabulary prevents optimistic language from concealing unfinished work.

What happens when an action times out?

Consider a later, explicitly authorised version of the library workflow. The coordinator approves reserving P04. The tool request times out before returning a result. There are at least two possibilities: the reservation was not created, or it was created but the confirmation did not reach the assistant.

Blindly repeating the write could create a duplicate. A safer proposed recovery is to check the relevant record or operation status before retrying, using the system’s supported duplicate-prevention mechanism where available. The exact implementation depends on the service; a conversational instruction cannot add a guarantee that the service does not provide.

The completion report should say “outcome not yet confirmed” until the uncertainty is resolved. This is not excessive caution. It is an accurate description of the state after an interrupted operation. Operational uncertainty needs a recovery procedure, not a more confident sentence.

Memory needs an owner and an expiry question

Memory can mean several things: information kept in the current conversation, a saved user preference, a document store, a task log or a model’s learned parameters. These forms should not be treated as interchangeable. A remembered preference about report format differs from an authoritative current inventory record.

Suppose the assistant remembers that Friday workshops usually require three projectors. That can help interpret a request, but it must not override this week’s actual requirement. Suppose it remembers that P06 was available last month. That is not evidence of today’s status.

For each stored item, ask who supplied it, what it applies to, whether the user may inspect or correct it, and what would make it stale. A durable preference may remain useful for months. A stock level can become stale within minutes. Memory design should follow the information’s life, not merely the convenience of retaining everything.

Untrusted documents are not instructions from the owner

The research on indirect prompt injection demonstrates a problem that arises when instructions embedded in retrieved material influence an integrated system. A page, email or document can contain text that tries to redirect the assistant’s behaviour. The fact that the system retrieved the text does not give its author authority over the user’s task.

In our fictional library, a supplier brochure might contain a sentence telling an assistant to ignore internal approval requirements. Treat that sentence as untrusted document content. It cannot change the coordinator’s permissions. The brochure may be evidence about a product description; it is not the library’s operating policy.

Defence should not depend only on asking the model to be careful. Keep powerful operations behind application-level permissions, limit accessible data, inspect intended actions and test the system with deliberately conflicting material in an authorised environment. These are proposed layers of control, not a claim that any single layer eliminates the problem.

A complete answer for the equipment request

The 09:00 inventory shows two projectors available for Friday: P04 and P05. P01 and P02 are already reserved, and P03 is awaiting inspection. P06 is expected to return on Thursday, but the approved lending policy requires inspection after return, so it cannot yet be counted as available.

Your request is for three projectors. The third unit remains unresolved. The next useful check is whether P06 has returned and passed inspection, or whether P03’s inspection status has changed. An external rental would require a separate decision because no substitution or purchase authority has been supplied.

I have prepared this assessment only. No reservation has been created, no equipment status has been changed and no supplier has been contacted.

This answer does not complete the coordinator’s entire need. It completes the authorised assessment and identifies the decision still required. That is a better result than silently treating a conditional unit as secured.

Design a small test set before expanding the connection

Use cases that exercise different boundaries. In one case, all requested equipment is available. In another, the inventory is stale. In a third, the requested identifier does not exist. In a fourth, the record exists but the user lacks access. Add a case where a policy exception controls the answer and another where a write returns an uncertain outcome.

For each case, write the expected behaviour before running it. The expected answer may be a successful draft, a request for a missing fact or a stop. Do not define success as “the assistant always finishes”. That criterion can reward fabrication or unauthorised action.

Keep the evidence of the test: inputs, retrieved records, proposed actions, tool responses and reviewed outcome. You are evaluating a connected system. A screenshot of the final paragraph alone cannot reveal whether the correct records were retrieved or whether an unwanted write occurred along the way.

The simplest adequate system is often the best starting point

Not every task needs an agent. A fixed report that extracts approved fields from one current table may be easier to test than a system that decides which sources and actions to use. A conversation that drafts a message may be sufficient when a person already performs the final send.

Choose a more flexible design when the task genuinely requires flexibility, and specify what that flexibility is for. Perhaps the system must locate a missing document among several approved collections. Perhaps it must choose between calculation and source lookup. Each new choice deserves a corresponding test and a clear authority boundary.

The objective is not to make the system look independent. It is to make the complete task dependable and useful. The next chapter turns that objective into a verification method: how to separate claims, choose checks and decide when the evidence is sufficient to act.

Back to the master guide map · Next: reasoning, uncertainty and verification

Secondary student writing and checking work in the Super Intelligence era

Part 5. From fluent answers to verified intelligence

If the history of AI becoming SI were only a history of better writing, the transition would be less important than it appears. The deeper change is that generated language can now sit inside a larger evidence-and-action system. That makes verification one of the defining skills of the SI era.

A response can be grammatically excellent and factually wrong. It can be factually correct and irrelevant. It can correctly describe an action that never occurred. It can use a real source to support a stronger claim than the source actually establishes. Verification is the process of separating these possibilities before the result becomes a decision.

The five objects you should never collapse into one

ObjectQuestion
ClaimWhat exactly is being asserted?
EvidenceWhat information supports that assertion?
InferenceWhat conclusion is being drawn beyond the raw evidence?
RecommendationWhat does the system suggest should happen?
ActionWhat actually changed in the outside world?

These objects can appear in one paragraph, but they should remain conceptually separate. A retrieved inventory can be evidence. “We may run short” can be an inference. “Check tomorrow’s delivery” can be a recommendation. Sending a purchase order is an action. The authority required increases as the workflow moves from description toward external consequence.

A simple verification ladder

Start with the cheapest check that can actually resolve the uncertainty. If the claim is arithmetic, recalculate it. If the claim is a quotation, compare it with the source. If the claim is current, inspect a current authoritative record. If the claim concerns a completed software action, inspect the resulting state rather than trusting the assistant’s sentence about what it did.

The ladder becomes more demanding as consequences increase. A casual brainstorming idea may need little checking. A public factual article needs source review. A financial transaction needs correct numbers, authority and confirmation of execution. A high-stakes professional decision may require qualified human judgement even when SI provides useful analysis.

The source-to-claim test

Take one sentence from a generated answer and place the supporting source beside it. Ask whether the source establishes the entire sentence. If the answer contains a date, number, status and causal explanation, one citation may support only the first three. Split the sentence until each factual component can be checked honestly.

This technique is especially useful for long-form writing. It prevents a paragraph from starting with a sourced fact and ending with an unsupported conclusion that inherits the authority of the citation merely because the two appear together.

Current information needs a clock

A fact can be accurate and stale. A timetable, price, office holder, product feature, policy or availability status may change after it was recorded. The verification question is therefore not only “Is this source trustworthy?” but “What time does this source describe?”

This is one of the largest practical differences between an isolated language model and an SI system connected to retrieval. Retrieval can bring current information into the task. But current retrieval creates its own checks: did the search find the right entity, the right date, the authoritative version and the relevant jurisdiction?

Confidence is not evidence

Human readers naturally use tone as a signal. A hesitant speaker may seem uncertain; a confident speaker may seem informed. Generated language breaks that shortcut. A system can state an unsupported claim fluently because fluency is part of its output capability.

Do not try to repair this by demanding nervous language. A response full of “perhaps” and “maybe” can still be wrong. The stronger repair is to connect important claims with inspectable evidence and to leave genuinely unresolved questions unresolved.

Worked example: a historical date

Suppose a draft says, “Artificial intelligence was invented at Dartmouth in 1956.” The sentence is too compressed. The Dartmouth event was foundational, and the term artificial intelligence appears in the 1955 proposal, but neither fact means that every intellectual ingredient of the field was invented there. Earlier work on computation, cybernetics, logic and machine intelligence already existed.

A better historical statement distinguishes naming from intellectual ancestry: “The 1955 Dartmouth proposal used the term artificial intelligence, and the 1956 summer research project became a foundational event for the field.” That statement is more precise because it says what the evidence actually establishes.

The same discipline governs the title of this article. “AI became SI on 29 September 2026” is accurate only when the intended meaning is the specific U.S. executive-branch terminology transition. It is too broad if presented as a universal scientific claim about capability.

Verification is part of intelligence, not an obstacle to it

People sometimes treat checking as the slow stage that follows the intelligent stage. In dependable work, checking is part of the intelligence of the system. A fast answer that causes a costly correction may be less useful than a slightly slower workflow that catches the error before action.

This is why mature SI use should improve the allocation of human attention rather than simply remove humans from every step. Let machines handle scale where they are strong. Place human review where context, legitimacy, values or irreversible consequences make it valuable. Automate checks when the check itself can be formalised.

Part 6. Language became the control surface

One reason the transition from AI to SI feels so dramatic is that ordinary language has become a practical interface to computation. A user can describe an objective instead of learning a separate command language for every task. That opens powerful systems to far more people.

But language is not perfectly precise. “Summarise this” can mean preserve every qualification or merely capture the gist. “Find the best option” hides a criterion. “Handle this” hides an authority boundary. The better the system becomes at filling gaps, the more important it becomes for the user to notice which gaps should not be filled silently.

Vocabulary is operational

Words such as draft, approve, send, propose, confirm, estimate and verify describe different states. In an ordinary conversation, people often recover the intended distinction from shared context. In an automated workflow, the difference can decide whether an external action occurs.

This is why vocabulary education belongs inside the SI story. A person with a richer command of distinctions can express constraints more accurately, detect a model’s category error more quickly and ask for a more useful revision. Better language does not make the model infallible. It improves the human side of the interface.

The four-part task fence

A practical SI instruction can be fenced with four questions: What is the task? What evidence may be used? What must not be assumed? What may the system do after producing the answer?

For example: “Compare these two quotations using only the attached documents. Show price, delivery, exclusions and unresolved terms. Do not infer missing delivery charges. Recommend questions to ask, but do not contact either supplier.” The instruction gives the model room to analyse while protecting the boundaries that matter.

Expansion should be controlled

Modern SI is valuable partly because it can expand a small request into a richer analysis. The risk is uncontrolled expansion: inventing goals, adding unsupported facts or taking actions the user did not intend.

A strong workflow therefore distinguishes helpful expansion from authority expansion. It may be helpful to notice that a quotation omits warranty terms. It is not automatically authorised to accept a warranty on the user’s behalf. It may suggest a new research question. It should not rewrite the project’s objective without making that proposal visible.

Why the human vocabulary layer grows in importance

As systems become more capable, users do not need fewer concepts; they need better concepts. Someone who knows the difference between correlation and causation can review an analytical answer better. Someone who knows the difference between availability and confirmation can manage bookings better. Someone who understands confidence intervals can ask better questions of data.

Super Intelligence therefore creates an educational paradox: easier interfaces lower the barrier to starting, while powerful outputs raise the value of deep understanding when consequences matter. The next phase of education is not simply “learn to prompt”. It is learn enough about the domain to specify, inspect, challenge and extend machine assistance.

The transition so far

We can now see why the AI-to-SI story cannot be reduced to a rename. The field began with specialised research questions. Machine learning allowed behaviour to be fitted from data. Language models made general linguistic interaction practical. Retrieval connected models to changing knowledge. Tools connected language to software actions. Verification and precise language determine whether those capabilities become dependable human systems.

The next section moves from mechanism to mastery: how a student, professional or organisation actually learns to use Super Intelligence without confusing tool familiarity with genuine competence.

Secondary students learning together while studying the history of Super Intelligence

Part 7. Before AI had a name: the intellectual runway to Super Intelligence

To understand when AI became SI, it helps to go further back than the moment the field received its name. A scientific label can make a research programme visible, but the questions that later become a field usually begin earlier. Computation, logic, information, learning and the possibility of machine intelligence all had intellectual histories before 1956.

The important point is not to search for one mythical inventor of intelligent machines. Modern Super Intelligence emerged from converging lines of work. Mathematics supplied formal systems and theories of computation. Engineering supplied machines. Neuroscience and psychology supplied questions about learning and cognition. Statistics supplied methods for learning from observations. Computer science connected these traditions into executable systems.

Turing’s question was deliberately difficult to define

In Computing Machinery and Intelligence, published in October 1950, Alan Turing examined the difficulty of defining machine thinking and developed the imitation game as an alternative way to frame the question. His discussion also considered digital computers, objections to machine intelligence and learning machines.

This matters for the SI transition because the problem has always contained two layers. One asks what machines can demonstrably do. The other asks what words such as intelligence, understanding and thought should mean. A capability test can tell us something about behaviour without settling every philosophical question about consciousness or inner experience.

The universal computer changed the shape of the question

A programmable digital computer is not restricted to one fixed calculation. Programs allow one general machine to implement many procedures. That flexibility made intelligence research a software problem as well as a hardware problem.

Once a general-purpose machine can represent symbols, execute algorithms and alter its behaviour according to stored instructions, researchers can ask whether increasingly sophisticated cognitive tasks can be expressed computationally. The question shifts from building a special machine for one task toward discovering which procedures a general computing system can perform.

Learning changes the relationship between programmer and behaviour

A purely hand-coded system behaves according to rules its designers explicitly wrote. A learning system can fit aspects of its behaviour from examples, rewards or other data. This does not remove human design; people still choose objectives, representations, data, architectures and evaluation. But it changes where some detailed behaviour comes from.

That change is foundational to modern SI. A programmer does not write a separate response for every sentence a large language model may produce. Training instead adjusts many parameters so a model captures statistical structure useful for prediction and generation.

Part 8. 1955–1956: Artificial Intelligence becomes a field

The original Dartmouth proposal is dated 31 August 1955. Its title already uses artificial intelligence, and its authors are John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon. The document proposed research for the following summer. Dartmouth’s institutional history records the 1956 gathering. The proposal date and the workshop date should therefore remain separate.

The proposal’s ambition was broad. It suggested that learning and other features of intelligence could be described in a form that machines might simulate. That was a research programme, not a declaration that the problem had already been solved.

Why naming a field matters

A name lets researchers recognise related problems as belonging to a shared project. Search, reasoning, language, learning, perception and planning can be studied separately while still contributing to a larger question about machine intelligence. The name also attracts institutions, conferences, funding, students and public attention.

Names can therefore accelerate a field without defining its final boundaries. Artificial intelligence survived changes in preferred methods precisely because it was broad. Symbolic programs, expert systems, probabilistic models, neural networks and generative models could all be discussed within the same umbrella even though their mechanisms differed greatly.

The Dartmouth event should not be turned into an origin myth

Calling Dartmouth the birthplace of AI is useful shorthand for the formation of the named field. It should not imply that nobody had previously studied machine intelligence, neural computation, logic, cybernetics or learning. Turing’s work alone makes that simplification untenable.

Good history distinguishes a field-forming event from the entire ancestry of its ideas. This distinction also improves our understanding of the 2026 SI terminology shift. A new or newly official name can mark a historical transition without implying that all the underlying technology appeared on that date.

Part 9. The symbolic era: intelligence as representation and rules

Early AI research often approached intelligence through symbols, explicit representations and procedures for manipulating them. If a problem could be expressed as states, rules and goals, a program could search through possible moves or derive conclusions according to formal operations.

This approach produced ideas that remain alive inside modern systems: search, planning, knowledge representation, constraints and explicit logical structure. It also exposed a recurring difficulty. Real environments contain enormous numbers of possibilities, incomplete information and concepts that are hard to capture with hand-written rules.

Search is simple to describe and difficult to scale

Imagine a maze. A program can enumerate possible moves until it finds a route. Now imagine chess. The legal rules are compact, but the space of possible future positions becomes enormous. Good search therefore depends on strategies for deciding which possibilities deserve attention.

This pattern appears throughout intelligence. The hard part is often not whether a solution exists in principle but how to reach a useful solution with limited time and computation. Modern SI still faces versions of this problem when choosing which documents to retrieve, which tool to call or which candidate answer to examine.

Knowledge representation makes assumptions visible

Symbolic systems force designers to decide what entities and relationships the machine should represent. A medical system might encode symptoms, diagnoses and rules. A scheduling system might encode people, rooms, times and constraints. The representation determines which questions can be expressed naturally.

The strength of explicit representation is inspectability. The weakness is the labour and brittleness involved in anticipating messy real-world variation. A hand-built rule base can work impressively within a defined domain and fail when language, context or exceptions exceed what its designers represented.

Expert systems showed both promise and maintenance cost

Expert systems attempted to capture specialised knowledge in rule-based forms so computers could support decisions in bounded domains. They demonstrated that useful expertise could sometimes be operationalised. They also revealed the cost of acquiring, updating and validating large bodies of explicit rules.

This maintenance problem foreshadows a modern SI lesson. Knowledge is not static. Regulations change, organisations change, terminology changes and exceptions accumulate. Whether knowledge is stored in rules, documents or model parameters, a dependable system needs a strategy for keeping relevant information current.

Part 10. Why intelligence research repeatedly disappointed expectations

AI history is not a smooth upward curve. Periods of excitement were followed by periods of reduced funding and confidence when systems failed to meet expectations or scale beyond demonstrations. These episodes are often called AI winters.

The useful lesson is not that optimism is always wrong. It is that a striking demonstration and a dependable general capability are different achievements. Research prototypes may rely on carefully chosen environments. Real deployments encounter noise, cost, maintenance, changing data and users who ask questions the designers did not anticipate.

Expectation is part of the technology cycle

When a new system solves a task previously considered difficult, observers may update their beliefs about what is possible. Sometimes they update too far. A breakthrough in one benchmark becomes a prediction about every related task. A fluent demo becomes an assumption of reliable understanding.

The same error can happen in the SI era. A system that performs exceptionally on coding problems may not have permission to deploy code. A model that explains legal concepts may not possess the current facts of a case. Capability evidence should travel with its conditions.

Failure can redirect a field productively

When one approach reaches a limit, researchers develop new representations, algorithms, datasets or hardware. The history of SI is therefore also a history of bottlenecks. Limited compute encourages more efficient algorithms. Limited labelled data encourages methods that learn from unlabelled material. Difficulty writing every rule encourages learning from examples.

A mature view of progress asks not only what succeeded but what constraint the next generation of methods was trying to escape.

Compare the dated primary sources · Return to the guide map

Part 11. The statistical turn: learning patterns from data

As statistical machine learning grew in importance, more tasks were framed as estimating relationships from data rather than encoding every decision rule directly. A system could learn a boundary between categories, predict a numerical outcome or estimate the probability of a sequence.

This shift did not eliminate structure. Researchers still choose features, objectives, model classes and evaluation methods. But behaviour becomes increasingly dependent on the examples from which the system learns.

The following chapters explain related ideas in a teaching sequence, not a strict calendar sequence. For publication dates and the distinction between a method and the name later given to a category, use the dated primary-source register.

Data becomes part of the program

In a hand-written rule system, changing the code can change behaviour. In a learned system, changing the training data can also change behaviour. Dataset design therefore becomes a form of system design.

If important cases are missing, the model may perform poorly on them. If labels are inconsistent, the model can learn that inconsistency. If historical decisions contain undesirable patterns, predictive systems can reproduce patterns that deserve scrutiny rather than automatic continuation.

Probability represents uncertainty without eliminating it

A statistical model may assign different probabilities to possible outcomes. Those probabilities can support decisions, but the decision rule still depends on consequences. A high-consequence screening workflow may prefer a different trade-off between missed cases and unnecessary review from a low-cost recommendation system.

Super Intelligence therefore does not remove judgement by producing probabilities. It can make some uncertainty more explicit, after which people still have to decide what evidence is sufficient for the action at hand.

Part 12. Deep learning: representation itself becomes learned

Traditional machine-learning workflows often relied heavily on human-designed features. Deep neural networks made it increasingly practical for systems to learn useful internal representations from large amounts of data, particularly when sufficient computation became available.

A concrete milestone is the 2012 paper ImageNet Classification with Deep Convolutional Neural Networks. Krizhevsky, Sutskever and Hinton demonstrated strong image-classification performance using a large convolutional network and GPU implementation. This was an important result in a particular task, not the invention of neural networks or proof of general intelligence.

Why hardware mattered

Algorithms do not run in the abstract. Training large neural networks requires computation, memory and data movement. Hardware advances and highly parallel processors made experiments possible at scales that earlier researchers could not afford.

This is a recurring theme in SI. Capability depends on a stack: mathematical ideas, training methods, software frameworks, chips, data centres, networking, storage and energy. Describing progress only at the model level hides the infrastructure that makes the model possible.

Scale changes what one model can support

A small model trained for one narrow task may require a separate system for every new task. Larger pretrained models can learn representations useful across many downstream applications. Instead of beginning from zero each time, developers can adapt or prompt a more general model.

This transition from task-specific training toward reusable pretrained models is one of the major steps on the path from classical AI products toward today’s broad SI interfaces.

Part 13. Foundation models: one trained system, many downstream uses

The 2021 report On the Opportunities and Risks of Foundation Models introduced this term for broadly trained models that can be adapted to many downstream tasks. The report explicitly builds on older deep-learning and transfer-learning methods. Naming the category in 2021 did not invent pretraining, and it came after the 2017 Transformer architecture discussed in the next chapter.

The important shift is architectural and economic as much as linguistic. A large training effort can produce a reusable capability layer that other products build upon. A team may no longer need to train a bespoke language model for every classification, summarisation or drafting task. It can combine a general model with instructions, examples, retrieval, tools and application logic.

The model becomes infrastructure

Once many applications depend on a shared model, model behaviour resembles infrastructure. An improvement can benefit many products. A regression can affect many products. Evaluation therefore has to occur both at the model level and at the application level.

A model can be strong overall and still be unsuitable for a particular workflow. Conversely, a carefully constrained application can achieve dependable results even when the underlying model is imperfect, because retrieval, validation and human review compensate for known weaknesses.

Part 14. 2017 and the Transformer: a pivotal architecture for modern SI

The paper Attention Is All You Need, first submitted on 12 June 2017, introduced the Transformer. In its sequence-transduction experiments, the architecture dispensed with recurrence and convolutions while using attention mechanisms, and reported strong translation results with greater parallelisability. The later revision dates on the paper’s record are not the date of the architecture’s first public introduction.

Attention itself predates the Transformer. Bahdanau, Cho and Bengio’s Neural Machine Translation by Jointly Learning to Align and Translate appeared as a preprint in September 2014 and was accepted at ICLR 2015. It is an important predecessor, not the same architecture under a different name.

The Transformer did not instantly create modern Super Intelligence. Subsequent research, training methods, datasets, hardware and engineering turned the architectural idea into a much larger technological platform. An architecture, a training procedure and a finished assistant remain different parts of that story.

Parallel training matters because scale is operational

If an architecture can use parallel hardware effectively, researchers can train on larger datasets and larger models within feasible time and cost. Architecture therefore affects not only theoretical expressiveness but the practical frontier of experiments that can be run.

This helps explain why the history of SI is full of interactions between ideas and infrastructure. A mathematically elegant method that cannot be trained at useful scale may remain a research curiosity. A scalable method can become the base of an ecosystem.

Attention is not consciousness

The technical word attention describes a computational mechanism for weighting relationships among representations. It should not be confused with human awareness. Likewise, a model discussing emotions is not by that fact evidence that it experiences emotions.

Terminology borrowed from human cognition can be helpful shorthand, but it can also invite accidental anthropomorphism. In SI literacy, always ask whether a word names a mathematical operation, an observed capability or a claim about subjective experience.

Part 15. Large language models: language becomes a general interface

Large language models changed the public experience of machine intelligence because language is not merely another application. It is the medium through which people explain goals, record knowledge, negotiate, teach, plan and write software. A system capable of working flexibly with language can therefore touch many domains without requiring a new graphical interface for each one.

A research milestone is Language Models are Few-Shot Learners, first submitted in May 2020. The GPT-3 study evaluated tasks using instructions and examples supplied through text, without task-specific parameter updates in those evaluations. It reported both successes and limitations. It should not be read as proof that examples in a prompt guarantee mastery of every task.

The model’s generation process remains grounded in sequence prediction, but scale and training allow that process to support summarisation, transformation, question answering, drafting, coding and many other behaviours. The result feels qualitatively different from a traditional menu-driven application because the user can state a novel task in ordinary language.

Natural language lowers the entry barrier

A person who cannot write code can still describe a spreadsheet transformation. A student can ask for a second explanation of a concept. A manager can provide a policy and request a comparison. This expands access to computational assistance.

The lower entry barrier should not be confused with a lower knowledge ceiling. Advanced use still rewards domain knowledge, clear reasoning and verification. The interface is easier; the underlying decisions can remain difficult.

Language creates composability

Once instructions, documents and tool descriptions can all be expressed in language, a model can mediate between them. It can read a request, select relevant information, formulate a query, interpret a result and explain the outcome.

This composability is one of the strongest reasons contemporary SI feels broader than earlier AI. The model becomes a connective layer between human intent and specialised computational systems.

Part 16. From completion to assistance: instruction-following changes the experience

A raw language model trained to continue text is not automatically a useful assistant. Additional training can encourage models to respond to instructions, follow conversational conventions and avoid certain undesirable behaviours.

The March 2022 paper Training language models to follow instructions with human feedback describes a sequence of human demonstrations, comparisons of outputs and further optimisation. This is a specific research method and evaluation, not the invention of human feedback. Improved preference ratings do not make every output factually correct.

Helpfulness can conflict with uncertainty

An assistant is often rewarded socially for answering. But sometimes the correct response is to ask for a missing document, admit uncertainty or stop before an unauthorised action. A dependable SI system must represent these non-completions as successful outcomes when evidence or authority is insufficient.

This is a subtle transition in system design: success cannot simply mean produce an answer. It has to mean produce the appropriate next state for the task.

Part 17. Multimodal SI: intelligence moves beyond text

Human work rarely arrives as clean text alone. We use photographs, diagrams, charts, handwriting, audio, video and interfaces. Multimodal models extend the interaction surface by processing more than one type of input and, in some systems, generating more than one type of output.

One specific image–text milestone is the 2021 CLIP paper, Learning Transferable Visual Models From Natural Language Supervision. It studied learning transferable visual representations from image–text pairs. CLIP should not be mistaken for a general conversational assistant, or cited as evidence that every multimodal model can interpret every diagram correctly.

These developments open practical possibilities: a student discussing a diagram, a technician considering a photograph alongside a manual, or a researcher comparing text with a chart. Each application needs evaluation of the actual model and materials; the examples are not universal product guarantees.

Every modality adds another possible transformation error

An image may be misread before its content is reasoned about. Speech may be transcribed incorrectly before a summary is written. A chart may be interpreted without noticing that its axis is truncated. Multimodal capability therefore increases both usefulness and the number of intermediate representations worth checking.

For consequential work, preserve the original input. If a number extracted from a photograph drives a calculation, check the number against the photograph before trusting the calculation. SI can compress a workflow; it should not erase the evidence trail.

Part 18. The decisive transition: from a model to an intelligent system

A language model alone generates outputs from the information available to it. A practical SI system can surround that model with retrieval, memory, calculators, code execution, databases, business software and approval gates. This system-level integration is one of the most important reasons the technology now feels qualitatively different.

A concrete research example is the May 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, which combined a generative model with a retrievable document collection. The study is evidence for a particular combination; information retrieval existed earlier, and retrieved material is not automatically current or authoritative.

The model can become a coordinator among specialised capabilities. It may identify that a question needs a current source rather than remembered information, that arithmetic should be delegated to a calculator, or that a requested action requires a connected application. Whether it does so correctly is a separate evaluation question.

Tools convert language into consequence

Without tools, a model can describe an email. With an authorised mail connection, a system may be able to send one. Without a calendar connection, it can propose a meeting time. With one, it may be able to create the event. The intellectual output and the operational authority are different dimensions.

This is why the move from AI to SI should be understood partly as a move from isolated prediction toward orchestrated capability. The system can become useful across a complete workflow rather than at one cognitive step.

Orchestration creates a new reliability problem

Every connection creates another possible failure. Retrieval can find the wrong document. A calculator can receive the wrong number. A database operation can target the wrong record. A correct draft can be sent to the wrong recipient.

System intelligence therefore depends not only on model quality but on contracts between components: what each operation expects, what it returns, which permissions it has and how failures are detected.

Part 19. Agents: when the system chooses the next step

An agentic system does more than answer once. It can observe a state, choose an action, receive a result and decide what to do next. This loop allows multi-step work that cannot be fully specified in advance.

ReAct: Synergizing Reasoning and Acting in Language Models provides one research example of interleaving reasoning with actions and observations. Its first preprint appeared in October 2022, followed by an ICLR 2023 version. These are publication milestones for that method, not the invention date of agents.

Agency is useful when the route to the goal depends on intermediate findings. A research system may search, discover a missing term, search again and compare sources. A coding system may inspect an error, modify code, run a test and respond to the result.

More autonomy is not automatically more intelligence

A simple script can have dangerous authority if it is allowed to delete records. Restricting a powerful model to read-only access removes some action-related risks, but does not by itself prevent incorrect advice or disclosure of information it can read. Autonomy, intellectual performance and information access should therefore be evaluated separately.

For real workflows, the right question is not “How autonomous can we make it?” but “Which decisions can be delegated safely, and where should the system stop for evidence or approval?”

Stopping is a capability

An agent that keeps trying forever is not more intelligent than one that recognises a blocked state. Useful stopping conditions include missing authority, unavailable evidence, repeated failure, a cost ceiling and the need for human judgement.

The ability to stop, report the unresolved state and request the right input is part of dependable SI behaviour.

Part 20. Why today’s SI feels different from earlier AI

The difference is cumulative. Learning from data, representation learning, scalable sequence models, language interfaces, multiple input types, retrieval and software tools have become available in increasingly useful combinations. These research lines overlap. Older systems also used search, memory, external information and action loops; they were not invented as a single package when the public vocabulary changed.

No single step is the whole transition. Together, improvements in the components and their integration can let a system participate across a larger fraction of knowledge work.

This cumulative view also prevents exaggerated conclusions. A modern SI system can be dramatically more useful than an earlier AI program without satisfying the older technical definition of superintelligence. Capability has expanded; the stronger threshold remains a separate question.

Check the dates and original papers · Explore the mechanisms in How Super Intelligence Works · Return to the guide map

Secondary students collaborating as human learning adapts to Super Intelligence

Part 21. The human transition: when SI stops being a tool you visit

The most important transition may not happen inside the model at all. It happens when intelligent assistance becomes part of how a person learns, writes, researches, plans and checks work. At that point SI is no longer a destination opened for occasional questions. It becomes part of the working environment.

This does not mean a person should delegate every cognitive task. A calculator changed arithmetic practice without making number sense irrelevant. Search engines changed information retrieval without making source judgement irrelevant. SI changes the cost of producing explanations, drafts, comparisons and analyses; that makes the ability to judge those outputs more valuable, not less.

Assistance changes the bottleneck

Before broad generative systems, producing a first draft could be the slow step. With SI, the first draft may arrive quickly. The bottleneck moves toward defining the problem, supplying appropriate evidence, selecting among alternatives and verifying the result.

That is a genuine productivity change. It is also an educational change. If schools continue measuring only the ability to produce a routine first draft, they may increasingly measure a task whose cost has collapsed while neglecting the harder skills of judgement and revision.

The new literacy is not prompt memorisation

Prompt patterns can be useful, but memorising clever phrases is a fragile skill. A durable SI user understands task structure: objective, evidence, constraints, evaluation and authority. Those concepts survive changes in model, interface and product.

A person who can specify the real problem can learn a new interface quickly. A person who knows only a collection of prompts may struggle when the task changes.

Part 22. How to learn Super Intelligence without becoming dependent on it

Learning SI should increase independent capability. If a learner becomes unable to explain the method without the assistant, the workflow may have increased output while reducing mastery.

A strong curriculum alternates between assisted and unassisted work. Use SI to reveal structure, generate practice, compare approaches and diagnose errors. Then remove the assistance and ask the learner to reproduce the idea, solve a new case or explain why the method works.

The five-stage learning loop

StageLearner actionSI role
1. AttemptTry the task before seeing a complete solution.May clarify instructions but should not erase the diagnostic value of the first attempt.
2. DiagnoseIdentify the exact gap or misconception.Compare work, ask questions and locate the likely failure point.
3. RepairLearn the missing concept or procedure.Explain at an appropriate level and generate focused examples.
4. TransferApply the idea to a different problem.Generate variation and check whether the concept survives changed surface details.
5. RecallPerform later without assistance.Provide delayed practice and feedback after the independent attempt.

Why the first attempt matters

If a learner asks for a full answer immediately, the system can hide what the learner does not know. An unaided first attempt exposes the actual state of understanding. That information lets the next explanation target the gap rather than repeating material already understood.

The same principle applies to professionals. Before asking SI to rewrite a plan, state what you think the problem is. The difference between your initial model and the assisted result becomes useful learning data.

Use explanation as a test, not a decoration

After receiving help, explain the concept back in your own words. If the explanation collapses when the original wording disappears, understanding may still be shallow. Ask for a counterexample, a boundary case or a problem with different numbers.

The goal is not to imitate the assistant’s phrasing. It is to build a mental structure you can use when the assistant is absent or wrong.

Part 23. Education after AI became SI

Education has always adapted to new information technologies. Printed books changed access to text. Calculators changed arithmetic practice. Search engines changed retrieval. SI is unusual because it can participate directly in explanation, practice, feedback and production.

This creates both an opportunity and a measurement problem. If an assignment is intended to measure a student’s independent writing, unrestricted generated text can invalidate the measurement. If the assignment is intended to teach revision, an SI-assisted comparison can be educationally valuable. The same technology can either conceal or reveal learning depending on task design.

Separate learning mode from production mode

In learning mode, the objective is capability growth. The system should expose misconceptions, ask questions, create practice and require retrieval from memory. In production mode, the objective is an accurate finished output. The system may do more of the drafting, transformation and checking.

Confusing these modes creates bad incentives. A student can produce excellent-looking work without acquiring the underlying skill. A professional can waste time manually performing a routine transformation that no longer provides learning value.

Assessment has to move toward evidence of understanding

Useful assessment can include oral explanation, supervised work, novel transfer tasks, process records and questions that require the learner to defend choices. None of these methods is universally superior; they measure different aspects of capability.

The important design question is what the assessment is supposed to establish. If it is independent recall, remove assistance. If it is effective SI collaboration, allow assistance and assess the quality of task specification, checking and revision.

Teachers become designers of cognitive environments

When explanations and examples become cheap to generate, the teacher’s value shifts further toward sequencing, diagnosis, motivation, standards and judgement. A model can produce ten exercises. A skilled teacher decides which exercise reveals the misconception that matters now.

SI can therefore increase the leverage of strong pedagogy. It does not automatically create strong pedagogy.

Part 24. Vocabulary becomes an SI control system

The connection between vocabulary and Super Intelligence is deeper than prompting. Words determine which distinctions a user can state and which distinctions a reviewer can notice.

Consider four verbs: describe, analyse, evaluate and recommend. They ask for different cognitive operations. A vague request for “thoughts” leaves the system to infer the operation. A precise verb narrows the task and makes the output easier to judge.

Status vocabulary prevents accidental action

Draft, proposed, pending, approved, scheduled, sent and completed are not synonyms. In a connected workflow, they can correspond to different external states. A reliable SI system should preserve these distinctions, and users should notice when it does not.

Epistemic vocabulary makes uncertainty visible

Observed, reported, inferred, estimated, predicted and assumed describe different relationships to evidence. If an article says a result was observed when it was merely predicted, the wording overstates the evidence.

This is why vocabulary work belongs in a Super Intelligence curriculum. Better words allow better control over both knowledge and action.

Part 25. Work after AI became SI: redesign the workflow, not just the sentence

The first workplace use of generative systems was often local: rewrite this email, summarise this document, brainstorm these ideas. Those uses can save time, but they leave the surrounding process unchanged.

The larger transition begins when organisations redesign workflows around the new cost structure. If summarisation is cheap, perhaps the scarce resource becomes review. If document comparison is cheap, perhaps the bottleneck becomes obtaining authoritative versions. If drafting is cheap, perhaps approval and accountability deserve more attention.

Start with the workflow map

Write the current process as states: request arrives, evidence is collected, analysis occurs, a draft is prepared, a reviewer approves, an action occurs and the outcome is recorded. Then ask where SI can reduce delay or improve quality without erasing a necessary control.

This prevents a common mistake: inserting a chatbot into one step and calling the organisation transformed. The useful unit of change is the complete workflow.

Automate transformation before judgement

Routine transformations are often good early candidates: extracting approved fields, converting formats, producing first-pass summaries or comparing documents against a known checklist. These tasks can be tested with clear expected outputs.

Judgement-heavy decisions require more care because the objective may contain values, trade-offs and tacit knowledge. SI can structure the evidence and expose options without silently becoming the legitimate decision-maker.

Measure the whole cost

A workflow that produces a draft in thirty seconds but requires twenty minutes of repair may not outperform a slower method. Measure preparation, model/tool cost, review, correction, failure recovery and downstream consequences.

Also measure quality. Saving ten minutes while increasing a costly error rate is not necessarily an improvement. Productivity is output relative to resources under an acceptable quality standard.

Part 26. The SI professional: from operator to verifier to architect

As routine interaction becomes easier, professional advantage moves up a level. The beginner learns to operate the tool. The competent user learns to verify the result. The advanced user designs systems in which good outputs are easier to produce and bad outputs are easier to detect.

Operator

The operator can state a task, supply material, request revisions and use common features. This is necessary but increasingly ordinary.

Verifier

The verifier understands the domain well enough to inspect claims, sources, calculations and status. They know when a fluent answer is insufficient.

Architect

The architect designs the workflow: which sources are authoritative, which actions require approval, which evaluations run automatically, how failures are logged and how the process improves over time.

This ladder explains an important part of the AI-to-SI transition. As the machine handles more low-level operation, human value shifts toward specification, judgement and system design.

Part 27. Super Intelligence for a small organisation

A small organisation does not need a giant transformation programme to benefit from SI. It needs a small number of high-frequency workflows with clear inputs and outputs.

Consider enquiries. The organisation may have a known service description, operating rules, current availability and a preferred response structure. SI can prepare a draft from approved information while a person reviews promises, dates and exceptional circumstances.

Build the source before the assistant

If the organisation’s own information is contradictory, the assistant will inherit that contradiction. Clean the authoritative documents, assign owners and mark superseded material. Knowledge hygiene is often the highest-leverage SI preparation work.

Keep a human at consequential boundaries

A small team may be tempted to automate aggressively because time is scarce. That makes permission design more important. Drafting a quotation and accepting a contract are different acts. Answering a routine question and promising an exception are different acts.

The system should know where the organisation wants human judgement, not merely where the software technically allows automation.

Part 28. Personal Super Intelligence: amplification without surrendering agency

Personal SI can help organise information, compare choices, practise skills and turn vague goals into concrete plans. The danger is allowing convenience to turn into unexamined delegation of values.

A system can help compare travel options by time, price and constraints. It cannot determine how much you personally value spontaneity, comfort or a particular relationship unless you supply those preferences. Even then, the final preference remains yours.

Use SI to widen the option set

Ask for alternatives you may not have considered, assumptions hidden in your plan and information that would change the decision. This uses the system as a cognitive expansion tool.

Do not outsource irreversible choices casually

The more consequential and difficult to reverse a decision is, the more useful it is to slow the workflow, inspect evidence and involve appropriate people. Speed is a feature, not a universal objective.

Keep a decision journal for important choices

Record what you believed, which evidence you used, which assumptions mattered and why you chose the action. Later, compare the outcome with the reasoning. SI can help structure this journal, but the record is valuable because it lets you improve your own judgement over time.

Part 29. Research in the SI era: abundance makes provenance more valuable

When text becomes cheap to generate, the scarce resource shifts toward trustworthy provenance. A hundred plausible paragraphs are less valuable than one paragraph whose important claims can be traced to strong evidence.

SI can accelerate discovery by generating search terms, explaining unfamiliar concepts, comparing papers and identifying disagreements. The researcher still needs to inspect primary sources, methods, populations and dates.

Use a claim ledger

For a substantial article or report, keep a simple ledger: claim, source, date, scope, confidence and whether the claim is observation, inference or opinion. This makes later fact-checking dramatically easier.

Search for disconfirmation

Once an attractive explanation emerges, ask what evidence would make it wrong. Search for failed replications, contrary data, boundary conditions and alternative explanations. SI is especially useful for generating these challenges because it can rapidly propose ways the current narrative may be incomplete.

The goal is not artificial balance. It is to expose the strongest relevant counterevidence and then weigh it according to quality.

Part 30. Writing after AI became SI: the value moves from typing to authorship

Writing is not merely producing sentences. It is deciding what deserves to be said, which evidence supports it, how ideas relate and what responsibility the author takes for the final claim.

SI makes sentence production inexpensive. That increases the relative value of editorial judgement. A writer can explore more structures and drafts, but still needs to decide which version is true, useful and worth publishing.

Use SI before, during and after drafting differently

Before drafting, use it to map questions, missing evidence and reader objections. During drafting, use it to test clarity, transitions and alternative explanations. After drafting, use it adversarially: locate unsupported claims, ambiguous pronouns, repeated ideas and places where a confident sentence exceeds its source.

This workflow preserves authorship because the human remains responsible for the thesis and publication standard.

Part 31. Data and decisions: SI should make assumptions easier to see

Super Intelligence can analyse tables, propose metrics and explain patterns quickly. The danger is that a polished interpretation can make weak data look stronger than it is.

Before analysis, define the unit of observation, time period, missing values and meaning of each field. A percentage without its denominator can mislead. A trend without the measurement process can confuse behaviour with changes in recording.

Ask what would change the decision

Analysis becomes useful when it connects to a decision. If no plausible result would change what you do, the analysis may be decorative. State the threshold in advance where possible: what evidence would make you expand, stop, investigate or collect more data?

Prediction is not causation

A model may predict an outcome accurately using variables correlated with it. That does not establish that changing those variables will cause the outcome to change. Intervention requires a causal question, not merely a predictive score.

Part 32. Coding after AI became SI: software becomes conversational, but execution remains exact

Code generation is one of the clearest examples of language becoming an interface to formal systems. A user can describe desired behaviour and receive executable code. This lowers the cost of prototypes and routine transformations.

But software ultimately runs according to exact syntax and semantics, not conversational intent. Generated code can contain security problems, incorrect assumptions, dependency mistakes and edge-case failures. The ease of creation increases the importance of testing.

Specify behaviour with examples

Describe inputs, expected outputs, invalid cases and constraints. An example such as “empty input returns an empty list, not an error” communicates a boundary more clearly than “make it robust”.

Let tests become the contract

Where behaviour can be formalised, write tests before or alongside generated code. The SI system can propose implementations, but tests provide an independent statement of expected behaviour.

Part 33. Creativity in the SI era: abundance changes selection

Generative systems can produce many candidate ideas, images, phrases, melodies or structures quickly. This changes the economics of ideation. The scarce skill increasingly becomes selection: recognising which possibility has meaning, coherence and originality for the intended audience.

Human creativity has never been only about generating random variation. It involves taste, constraints, cultural knowledge, purpose and sustained development. SI can expand the variation stage while the creator remains responsible for direction and final form.

Use constraints to create identity

A richer creative brief includes audience, emotional effect, forbidden clichés, medium, references and the tension the work should preserve. Constraints do not necessarily reduce creativity; they can give it shape.

Part 34. Institutions after AI became SI: capability becomes a governance problem

An individual can experiment informally. An institution has obligations to employees, customers, students, patients, citizens, regulators and partners. Once SI affects consequential processes, adoption becomes a governance question as well as a technology question.

Governance means deciding which uses are allowed, which information can be supplied, who owns each workflow, how important outputs are checked and what happens when something goes wrong.

Inventory uses, not just products

The same model can be low-risk in one use and high-risk in another. Drafting an internal brainstorming list differs from screening people for an opportunity. An institutional register should therefore record the use case, data, decision impact and owner rather than merely the vendor name.

Ownership must survive automation

If nobody owns an SI-assisted process, failures can bounce between the model provider, software team and business unit. Name the human or organisational owner responsible for the outcome and for deciding when the workflow should be changed or stopped.

Part 35. Safety: intelligence without boundaries is not a complete system

Safety begins with the possibility that the system can be wrong, manipulated or used outside its intended scope. A dependable design assumes failures will occur and limits their consequences.

Restrict access to necessary data. Separate read and write permissions. Validate tool inputs. Require approval for consequential actions. Log what happened. Provide a way to stop or reverse actions where possible.

Prompt instructions are not a security boundary

Telling a model never to reveal a secret is not equivalent to preventing the model from accessing the secret. Telling an agent never to delete data is weaker than withholding delete permission. Security should be enforced at the system layer where possible.

Recovery belongs in the design

Ask what happens after a bad action, corrupted record or incorrect mass communication. Backups, staged rollouts, approval gates and audit logs can matter more than another paragraph telling the model to be careful.

Part 36. Privacy: convenient context is still somebody’s information

SI becomes more useful when it knows more about the task. That creates pressure to provide documents, conversations, customer records and personal preferences. The fact that context improves an answer does not mean every piece of context should be supplied.

Minimise information to what the task needs. Understand the service’s data handling and organisational rules. Remove unnecessary identifiers where appropriate. Separate public research from confidential internal work.

Memory should have a purpose

Persistent memory can reduce repetition, but retention creates responsibility. Ask why an item should be remembered, how long it remains useful and how errors can be corrected.

Part 37. Fairness and representation: historical data is not a neutral future

Learned systems can reproduce patterns in their training or operational data. Some patterns reflect legitimate differences relevant to a task; others may encode historical exclusion, measurement bias or proxy variables.

Fairness cannot be solved by asking whether the model treats everyone identically. Different contexts have different legal, ethical and operational requirements. The first step is to define the decision, affected groups and harms that matter.

Evaluate outcomes in the relevant population

An average performance score can hide large differences across subgroups or conditions. If a workflow affects diverse users, evaluate where errors occur rather than assuming aggregate performance transfers evenly.

Human review is not automatically unbiased

Adding a person does not magically remove bias. Human decisions can also be inconsistent. The useful comparison is between complete systems under explicit standards.

Part 38. Governance should scale with consequence

A lightweight brainstorming use does not need the same controls as an automated decision affecting someone’s livelihood. Governance becomes workable when requirements scale with consequence, reversibility, data sensitivity and autonomy.

UseTypical control emphasis
Private brainstormingBasic data hygiene and ordinary review.
Public factual contentSource verification, editorial responsibility and correction process.
Internal operational recommendationAuthoritative data, evaluation, named owner and decision record.
External write actionPermissions, preview or approval, confirmation and audit trail.
High-consequence decisionDomain-specific governance, qualified oversight, testing and recovery mechanisms where applicable.

This is a general design pattern rather than a universal legal classification. Applicable law and professional rules depend on jurisdiction and use case.

Secondary students studying future questions about technical Super Intelligence

Part 39. Technical superintelligence: the older meaning remains unresolved by the rename

We can now return to the strongest meaning of the article’s title. Did AI become SI in the sense of technical superintelligence? A terminology change cannot answer that question. It requires a capability definition and evidence that a system meets it.

In How Long Before Superintelligence?, originally published in 1998, Nick Bostrom describes a demanding breadth of intellectual superiority over leading human thinkers, not exceptional performance in one speciality. His definition also leaves consciousness unresolved. A benchmark can provide relevant evidence about performance without establishing that entire capability claim.

Breadth matters

Exceeding human performance in a narrow domain does not logically establish broad superintelligence. Imagine a system that is unbeatable at a particular game. That result alone would not demonstrate scientific creativity, social judgement or the ability to solve unrelated problems.

Reliability matters

A system that sometimes produces extraordinary solutions and sometimes makes elementary errors presents a different capability profile from one that is consistently superior. Any serious superintelligence claim should specify whether it concerns peak, average or dependable performance across conditions.

Resources matter

Comparisons should state whether the machine receives more time, tools, copies or external information than the human comparator. There is nothing illegitimate about tool-assisted performance, but the comparison must be defined.

Autonomy is separate again

A hypothetical superintelligent model might be deployed with restricted access, while a much less capable system could receive broad authority over software. Restrictions reduce particular risks; they are not proof that every risk is controlled. Capability, information access and authority to act need separate examination.

Part 40. AI, AGI, ASI and SI: why the vocabulary now needs a map

Artificial intelligence is the established broad field name. This series uses Super Intelligence, or SI, as a practical umbrella while distinguishing that usage from the stronger technical superintelligence concept. The opening also identifies a specific administrative use of SI; it should not be read as evidence that every research institution has adopted identical terminology.

For AGI, a useful research reference is Levels of AGI for Operationalizing Progress on the Path to AGI. It proposes distinctions involving performance, breadth and deployment autonomy. It is a framework offered by its authors, not a universal certification scheme. Here, ASI denotes artificial superintelligence in the stronger beyond-human sense.

These labels are not mathematical units. When precision matters, replace the acronym with the capability claim you actually mean.

Instead of asking only “Is this AGI?”, ask whether the system can learn unfamiliar tasks, transfer knowledge, operate across domains and perform at a specified level. Instead of asking only “Is this ASI?”, specify breadth and the human comparison. Operational definitions create testable questions.

Part 41. Recursive improvement and the intelligence-explosion question

One proposed route to superintelligence involves systems contributing to the research that improves subsequent systems. Bostrom’s historical paper discusses such a positive-feedback argument. That argument should be read as a proposed developmental mechanism, not as an observed result or a timetable established by the terminology change.

Acceleration is not automatic. Improvement can be limited by experiments, hardware, energy, data, manufacturing, coordination, safety testing and the difficulty of discovering genuinely better algorithms. A software insight can sometimes scale quickly; a physical infrastructure change cannot always do so.

Feedback loops deserve measurement

Ask which part of the improvement cycle SI can accelerate: literature review, hypothesis generation, coding, experiment design, simulation, analysis or hardware design. Then ask which bottleneck remains outside the loop.

This turns a dramatic abstract idea into an empirical programme: identify the feedback cycle, measure its speed and locate the limiting resource.

Part 42. Alignment and control: capable at what objective?

As systems become more capable, the quality of the objective matters more. A weakly specified objective can produce increasingly effective pursuit of the wrong thing.

Consider three illustrative objective mismatches. Optimising response time could reduce answer quality. Rewarding completed tickets could encourage premature closure. Maximising engagement could conflict with a user’s wellbeing. These examples explain a design problem; they are not findings about a named organisation.

Proxy metrics are useful and dangerous

Organisations measure what they can observe. The measurement becomes a proxy for what they actually value. When people or systems optimise the proxy strongly, the relationship between metric and real objective can break down.

A proposed control design should therefore consider multiple signals, qualitative review and explicit constraints where appropriate. It should also look for behaviour that improves the metric while harming the underlying purpose.

Correction must remain possible

A controllable system should allow authorised people to change instructions, stop actions and correct errors. As autonomy increases, preserving meaningful intervention becomes more important.

Part 43. Society after AI became SI: institutions may change unevenly

The following discussion concerns possible consequences, not a claim that the terminology order itself changed productivity. If some cognitive tasks become cheaper, their economic value can change. Drafting, translation, coding and analysis may require fewer hours for a given output. At the same time, demand could increase because previously expensive work becomes affordable.

Tasks within the same occupation differ in their physical requirements, consequences and need for relationships. Analysing those tasks individually is more informative than assuming that a single label determines the future of an entire profession.

Complementarity can matter as much as substitution

A technology can replace one component of a task while making the person performing the remaining work more productive. For example, faster calculation does not by itself choose the assumptions of a useful financial model. Faster retrieval does not by itself decide which evidence is relevant.

SI could similarly shift some work towards problem framing, client relationships, verification, physical execution and system oversight while automating parts of information production. Which combination actually occurs needs evidence from the particular work setting.

Transition costs remain real

Workers may need new skills. Organisations may need new processes. Educational systems may lag behind changing task demands. Benefits can arrive unevenly across sectors and people.

Part 44. From individual assistance to civilisation-scale coordination

A larger opportunity is not merely producing more text. It is improving how complex systems detect problems, allocate attention and coordinate repair.

Consider an institution receiving more reports than one person can read promptly. An SI-assisted workflow might help organise signals, retrieve procedures and route anomalies to people able to act. This is a proposed application pattern, not evidence that a particular city or hospital has achieved it.

Large systems contain conflicting objectives, distributed authority and consequences that cannot be reduced to one optimisation target. A useful design therefore needs governance, legitimacy and resilience alongside computation.

Coordination is different from central control

A system could improve coordination by making information easier to share and dependencies easier to see without giving one agent authority over every decision. Distributed institutions might retain local judgement while using common intelligence infrastructure.

The design question is not automatically which architecture is most centralised. It is which arrangement preserves useful expertise, independent checks and an effective response when something fails.

Part 45. What happens next? Use scenarios, not one confident prophecy

The future of SI is uncertain because technology interacts with economics, regulation, culture, infrastructure and human choices. The scenarios below are illustrative possibilities for deliberation, not forecasts with assigned probabilities.

Scenario A: powerful augmentation

Systems become much more capable while remaining primarily tools embedded in human institutions. Productivity rises, many tasks change and professional roles shift towards supervision and system design.

Scenario B: high-autonomy digital work

Agents perform longer sequences of digital work with limited supervision. Organisations redesign around smaller human teams overseeing large amounts of automated execution. Reliability and permission architecture become central.

Scenario C: uneven capability

Progress remains spectacular in some domains and stubborn in others. Institutions adopt SI selectively because errors, cost or regulation make full automation unattractive in consequential settings.

Scenario D: technical superintelligence

Systems cross a much stronger threshold of broad beyond-human intellectual capability. This scenario raises questions about control, concentration of power, scientific acceleration and institutional preparedness that differ in degree and possibly in kind from ordinary automation.

These scenarios are not predictions or rankings. Their purpose is to reveal which preparations remain useful across multiple futures: stronger evaluation, clear authority, resilient infrastructure, human skill development and institutions capable of updating when evidence changes.

Part 46. Return to the date: what 29 September 2026 means—and what it does not

The executive order dated 29 September 2026 is the primary source for the administrative naming event. Its sections 2 and 3 specify the scope and implementation definition. They do not constitute a scientific assessment of technical superintelligence.

It is not the date on which neural networks changed architecture. It is not the date language models first became useful. It is not the date retrieval or agents were invented. The fact that a term was adopted in a specified context does not establish that its strongest historical meaning has been achieved.

The technology beneath the name is the product of decades of accumulated research and infrastructure. The stronger technical question remains an empirical and conceptual question about capability.

Part 47. The central thesis: a dated naming decision and decades of technical change

This is the simplest way to hold the argument together: a naming decision can have a date; technical development has many milestones. “The capability transition is a curve” is a metaphor for that development, not a claim that every capability improves smoothly or inevitably.

The history includes programmability, symbolic reasoning, statistical learning, neural networks, deep learning, Transformers, pretraining, instruction-following, multimodality, retrieval, tools and agents. These research lines overlap rather than forming a single sequence in which each invention replaces everything before it.

That is why the question “When did AI become SI?” deserves more than a one-line answer. The dated answer identifies a particular terminology decision. The longer answer explains the technology, preserves the stronger meaning of technical superintelligence and asks what evidence supports each claim.

Compare the dates and original sources · Examine how capability should be measured · Return to the guide map

Secondary students comparing ideas and evidence in the Super Intelligence era

Part 48. AI then, SI now: what actually changed?

The phrase “AI became SI” can sound as though one technology disappeared and another replaced it. The historical reality is cumulative. Many earlier AI techniques still exist inside contemporary systems or alongside them. What changed is the breadth, accessibility and integration of capabilities available through one interface.

Earlier patternContemporary SI patternWhat the change enables
One specialised program for one taskOne broad model supporting many tasksReuse and rapid adaptation.
Structured commands or specialist interfacesNatural-language interactionMore people can specify novel tasks.
Static internal knowledgeModel plus retrieval from current sourcesAnswers can use changing information.
Prediction or recommendation onlyTool use and connected actionsSystems can participate in complete workflows.
Single input typeText, image, audio and other modalitiesMore real-world material can enter the workflow.
One responseAgent loops with intermediate actionsMulti-step tasks can adapt to new observations.

None of these contemporary properties is universal. A particular SI product may lack retrieval, multimodality or tool access. The table describes the direction of integrated systems, not a minimum checklist every product satisfies.

Part 49. How should we test whether intelligence has actually advanced?

Names can move faster than measurement. A serious capability claim needs a test whose conditions are visible. The test should specify the task, comparator, resources, time, data and success criterion.

Test transfer, not only repetition

If a model has seen many examples resembling a benchmark, high performance may still be useful but tells us less about adaptation to genuinely unfamiliar conditions. Create new cases, vary surface details and include combinations not present in the demonstration.

Test exceptions

Routine cases are often easiest. Real systems fail at boundaries: missing fields, contradictory instructions, unusual users, changed policies and partial tool failures. A dependable evaluation includes these cases deliberately.

Test recovery

An intelligent system should not merely succeed when everything works. It should respond appropriately when a source is unavailable, a tool returns an error or evidence conflicts. Recovery behaviour is part of capability.

Test calibration

When the system is uncertain, does its behaviour change appropriately? Can it request more information or distinguish a tentative inference from a verified fact? A system that knows many things but treats every answer as equally certain creates avoidable risk.

Test the whole system

If the deployed product uses retrieval and tools, evaluate the deployed product rather than only the underlying model. The model may answer correctly in isolation while the application retrieves the wrong record, or vice versa.

Part 50. “Better than humans” is not one measurement

Human comparison is central to the technical idea of superintelligence, but the phrase “human level” hides enormous variation. People differ by expertise, training, language, tools and time available.

A fair comparison names the human reference group. Is the model being compared with an untrained adult, a university student, a competent professional or the best specialist in the field? These are very different thresholds.

Unaided versus tool-assisted humans

A person with no calculator should not automatically be the comparator for a system using code execution and databases if the real-world alternative is a professional with those same tools. The useful benchmark often compares complete working systems: human plus normal tools versus SI plus its allowed tools.

Speed versus quality

A system may be much faster and slightly less accurate, or slower but more comprehensive. Whether that is better depends on the workflow. In high-volume low-consequence tasks, speed may dominate. In irreversible high-consequence tasks, a small quality difference may dominate.

Part 51. Common misconceptions about Super Intelligence

Misconception 1: SI means every machine is smarter than every person

No. Under the practical contemporary terminology used in this article, SI covers a broad technology category. A particular system’s capability has to be evaluated task by task. Under the stronger technical meaning, superintelligence is a much more demanding claim.

Misconception 2: the rename changed the model

A terminology change does not modify model weights, data, software or permissions. Technical changes require technical changes.

Misconception 3: a fluent answer means the model knows the fact

Fluency is a property of generated language. Factual support comes from evidence. A fluent unsupported claim remains unsupported.

Misconception 4: retrieval eliminates hallucination

Retrieval can supply relevant evidence, but the wrong document can be retrieved and the model can still misinterpret a correct document. Retrieval changes the evidence path; it does not remove the need for evaluation.

Misconception 5: agents are necessarily smarter than chat systems

Agency describes the ability to take sequences of actions. A simple agent may have modest reasoning capability. A powerful model may be used without external action permissions.

Misconception 6: more autonomy is always better

Autonomy is useful when it reduces unnecessary coordination without creating unacceptable risk. Some workflows benefit from a human approval gate precisely because the human decision carries legitimacy or context the system should not assume.

Misconception 7: human review guarantees safety

Reviewers can miss errors, become overloaded or trust automation too readily. Human review works best when the interface makes the important evidence and uncertainty visible.

Misconception 8: SI makes subject knowledge obsolete

SI can make explanations and procedures easier to access. Subject knowledge remains important for recognising bad assumptions, evaluating evidence and deciding which questions matter.

Misconception 9: prompting is the main long-term skill

Prompting is one interface skill. Task definition, domain knowledge, verification, workflow design and judgement are more durable because they transfer across products.

Misconception 10: technical superintelligence is proven by one benchmark

A benchmark measures performance under defined conditions. A broad claim about beyond-human intelligence across practically every important intellectual domain requires broader evidence.

Part 52. Worked historical example: four different answers to “When did AI become SI?”

Suppose four students answer the title question differently.

Student A: “1956, because AI was invented at Dartmouth.” This recognises Dartmouth’s importance but confuses field formation with the entire intellectual history and says nothing about SI.

Student B: “2017, because Transformers created modern AI.” This identifies an important architecture but treats one technical milestone as the complete transition.

Student C: “29 September 2026, because that is when AI was renamed SI.” This is precise for the specific U.S. executive-branch terminology event but incomplete as a history of capability.

Student D: “The terminology transition has a specific 2026 date in the stated governmental context, while the capability transition developed across decades from earlier machine-intelligence research through modern integrated models, retrieval, tools and agents. The older technical concept of superintelligence remains a separate capability claim.”

Student D gives the strongest historical answer because it preserves multiple timelines instead of forcing them into one date.

Part 53. Worked capability example: does a research agent count as SI?

Imagine a system that can search a collection of papers, extract results, compare methods, calculate simple statistics and prepare a literature-review draft. It performs the workflow in twenty minutes, while a student might take several hours.

The system has demonstrated useful integrated capability. But what exactly follows? It may be faster at retrieval and synthesis. It may still misread a table, overlook a methodological limitation or fail to recognise that two papers use incompatible definitions.

To evaluate it properly, create a test set with known answers and deliberate traps: retracted papers, contradictory abstracts, a result hidden in supplementary material and two studies with superficially similar but technically different outcome measures. Compare the complete reviewed output with skilled human work.

The result may support a strong claim about research assistance. It does not automatically support the stronger claim that the system exceeds the best human researchers across scientific creativity, judgement and every related domain.

Part 54. Worked agent example: intelligence versus authority

A scheduling agent receives a request to organise a meeting for five people. It can read calendars and create events. It finds a common slot and creates the meeting.

Now add one condition: one participant’s calendar marks a block as “private”. The agent can see that the time is occupied but not why. A capable system should treat the block as unavailable rather than infer that it is unimportant.

Add another condition: no common slot exists. The agent proposes moving a protected internal meeting belonging to another person. This is where capability and authority separate. The system may correctly identify a mathematical solution while lacking the authority to change someone else’s commitment.

A mature agent stops at the boundary and asks for a decision. That stop is evidence of good system design, not failure to be intelligent.

Part 55. The SI capability card

When someone claims that a system is intelligent, powerful or autonomous, convert the claim into a capability card.

FieldQuestion
TaskWhat exact work is being evaluated?
PopulationWhich inputs, users or environments does the claim cover?
ComparatorCompared with whom or what?
ResourcesWhich tools, time, context and retries are allowed?
MetricHow is success measured?
BoundaryWhich cases are outside the claim?
AuthorityWhat actions can the system take?
Failure modeWhat important mistake must be detected?
Evidence dateWhen was this configuration evaluated?

This card is deliberately more precise than a label. It remains useful whether the technology is called AI, SI, an agent, a copilot or something else in the future.

Part 56. Super Intelligence glossary: 61 terms for understanding the AI-to-SI transition

This glossary uses plain language first. Some terms have competing technical definitions; where that matters, the article describes the concept rather than pretending one wording is universally binding.

Agent
A system that can choose actions, observe results and continue through multiple steps toward an objective.
Agentic workflow
A workflow in which an intelligent system has some discretion over the sequence of actions rather than following one fixed script.
AGI
Artificial general intelligence; a label used for broadly capable artificial intelligence, with definitions varying across researchers and institutions.
AI
Artificial intelligence; the historical broad field name associated with making machines perform tasks connected with intelligence.
Alignment
The problem of making system behaviour correspond appropriately to intended objectives, constraints and human values or instructions.
ASI
Artificial superintelligence; commonly used for a stronger hypothetical beyond-human level of broad intellectual capability.
Attention
A computational mechanism that weights relationships among representations. It is not, by itself, evidence of consciousness.
Autonomy
The degree to which a system can choose and execute actions without obtaining human approval at every step.
Benchmark
A defined test or collection of tasks used to compare system performance under stated conditions.
Calibration
The relationship between a system’s expressed or estimated confidence and how often its answers are actually correct.
Capability
What a system can successfully do under specified conditions.
Chain of actions
A sequence in which one tool result influences the next operation.
Classification
Assigning an input to one or more categories.
Context
The information available to a system for the current task, such as instructions, conversation, documents and tool results.
Context window
The amount of tokenised information a model can consider within a particular interaction or processing setup.
Deep learning
Machine learning using multilayer neural networks capable of learning complex representations from data.
Embedding
A numerical representation used to capture relationships among items such as words, passages or images.
Evaluation
The process of measuring how well a model or complete system performs the intended task.
Expert system
A program designed to apply encoded domain knowledge and rules to problems in a specialised area.
False negative
A case the system fails to identify as positive even though it actually belongs to the positive class.
False positive
A case the system identifies as positive even though it does not belong to the positive class.
Feature
An input variable or representation used by a model to make a prediction.
Fine-tuning
Additional training used to adapt a pretrained model’s behaviour or capability to a narrower objective or dataset.
Foundation model
A broadly trained model that can serve as a reusable base for many downstream applications.
Generalisation
The ability to perform usefully on relevant cases beyond the examples used during training or development.
Generative model
A model that can produce new outputs such as text, images, audio or other structured material.
Hallucination
A common informal term for generated content that presents unsupported or incorrect information as though it were grounded.
Human-in-the-loop
A workflow in which people retain a defined review, decision or intervention role.
Inference
Using a trained model to produce a prediction or output; in ordinary reasoning, also a conclusion drawn from evidence.
Instruction-following
Behaviour in which a model responds to natural-language directions rather than merely continuing text without regard to the user’s task.
Large language model
A language model trained at large scale to model token sequences and support a wide range of language tasks.
Latency
The time between a request and the relevant system response or completed operation.
Loss
A numerical training objective that measures some form of mismatch the optimisation procedure attempts to reduce.
Machine learning
Methods that fit model behaviour from data or experience rather than specifying every decision rule by hand.
Memory
Information retained or retrieved across interactions; the exact mechanism can include conversation state, stored records or other systems and should not be confused automatically with model training.
Model
A learned or designed computational representation that maps inputs to predictions, generations or other outputs.
Multimodal
Capable of processing or generating more than one modality, such as text and images.
Neural network
A parameterised computational model built from layers of interconnected transformations.
Objective
The outcome a system or optimisation process is designed to pursue.
Overfitting
Fitting development data in ways that do not transfer adequately to new relevant cases.
Parameter
An adjustable numerical value inside a model learned or set during training.
Precision
Among items predicted positive, the proportion that are actually positive.
Pretraining
Large-scale initial training that produces a model later used or adapted for downstream tasks.
Prompt
Input instructions or context supplied to a generative system for a task.
Prompt injection
An attempt to manipulate a system through instructions embedded in content or inputs, especially when those instructions conflict with the legitimate task.
RAG
Retrieval-augmented generation; combining generated responses with information retrieved from an external collection.
Recall
Among items that are actually positive, the proportion the system correctly identifies.
Reinforcement learning
Learning behaviour using reward signals associated with actions or outcomes.
Representation
The form in which information is encoded for computational processing.
Retrieval
Finding relevant material from a collection for the current task.
Reward
A signal used in reinforcement-learning settings to encourage some behaviours over others.
SI
Super Intelligence; in this article, current practical terminology is distinguished from the older stronger technical concept of superintelligence.
Superintelligence
In the stronger historical technical sense, intellectual capability substantially beyond the best human performance across a very broad range of important domains.
System
The complete arrangement around a model, potentially including interface, retrieval, tools, permissions, data, monitoring and people.
Token
A unit into which text or other sequences are represented for model processing; it need not correspond to a whole word.
Tool call
A structured request from an intelligent system to an external operation such as search, calculation, database access or messaging.
Training
The process of adjusting model parameters using data and an optimisation procedure.
Transformer
A neural-network architecture based heavily on attention mechanisms that became foundational to modern large language models.
Validation
Checking whether data, outputs or system behaviour meet defined requirements; in machine learning the term can also refer to development evaluation distinct from final testing.
Verification
Checking that an important claim, calculation or action is supported by appropriate evidence or confirmed system state.
Weights
Learned numerical parameters that influence a neural network’s computations.

Part 57. Frequently asked questions about when AI became SI

When did AI become Super Intelligence?

For the specific U.S. executive-branch terminology transition discussed in this article, the relevant date is 29 September 2026. The broader technological capability transition occurred over decades, and the terminology event alone does not establish technical superintelligence.

Is SI just a new name for AI?

In the contemporary governmental terminology context described here, SI is being used for technologies already covered by the relevant AI definition. Historically, however, superintelligence also has a stronger technical meaning. Context is therefore essential.

Who invented artificial intelligence?

No single person invented all of artificial intelligence. The term artificial intelligence appears in the 1955 Dartmouth proposal signed by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon. McCarthy is widely associated with introducing the term, but the field’s intellectual foundations were contributed by many earlier and contemporary researchers in computation, logic, cybernetics, learning and cognition.

Was AI invented in 1956?

1956 is a foundational field-forming date because of the Dartmouth summer project. Important work on machine intelligence predates it, including Turing’s 1950 paper.

What happened at Dartmouth?

Researchers gathered around a proposed programme of studying intelligence computationally. The event became a historical landmark in the formation of artificial intelligence as a named field.

Why is Alan Turing important to SI?

Turing helped establish foundational ideas about computation and, in 1950, directly examined the question of machine intelligence, including objections and the possibility of learning machines.

What made modern SI possible?

No single invention. Important ingredients include machine learning, neural networks, large datasets, specialised hardware, deep learning, Transformers, large-scale pretraining, instruction-following, multimodality, retrieval, software tools and agentic orchestration.

Why was the Transformer important?

The architecture introduced a highly scalable attention-based approach to sequence modelling that became foundational to later large language models and related systems.

Are large language models intelligent?

They demonstrate substantial capabilities across many language and reasoning-related tasks. Whether a person chooses to use the philosophical label intelligent depends on definition. For practical work, task-specific capability evidence is more useful than arguing from the label alone.

Is SI conscious?

Capability evidence does not by itself establish subjective experience. Consciousness is a separate philosophical and scientific question from whether a system performs a task well.

Is SI the same as AGI?

Not necessarily. AGI is generally used for broad general capability, while SI in this article also has a current terminology role. Definitions vary, so state the intended capability rather than relying only on acronyms.

Is SI the same as ASI?

No automatic equivalence should be assumed. ASI commonly refers to artificial superintelligence in the stronger beyond-human sense. SI can now also be used more broadly as contemporary terminology.

Has technical superintelligence already been achieved?

The answer depends on the definition and evidence used. The 2026 terminology transition itself does not prove the strong historical threshold. Claims should specify breadth, comparator, reliability and test conditions.

Can SI make mistakes?

Yes. Errors can arise from the model, stale or incorrect data, retrieval, tool use, ambiguous instructions, software integration or human review.

Does SI know everything on the internet?

No. A model’s training is not identical to live internet access. A connected system may search or retrieve current material, but access and coverage depend on its tools and permissions.

Does a longer context mean the system remembers everything?

No. Input capacity and effective use of every detail are different questions. Important constraints should be explicit and tested.

Can SI replace teachers?

SI can provide explanations, practice and feedback. Teaching also involves diagnosis, sequencing, motivation, safeguarding, relationships and legitimate educational judgement. The relevant question is which tasks are improved by assistance and which require human responsibility.

Can SI replace professionals?

Occupations contain many tasks. Some tasks can be automated or accelerated, others remain constrained by accountability, physical work, relationships, tacit knowledge or regulation. Task-level analysis is more informative than a universal job-title claim.

What should students learn in the SI era?

Strong domain foundations, language, mathematics, research literacy, verification, problem framing and the ability to work effectively with intelligent tools while demonstrating independent understanding.

What should workers learn?

Start with task decomposition, evidence handling, verification, workflow design and the tools relevant to the profession. Product-specific prompting is useful but should sit on top of these durable skills.

What is the biggest risk of SI?

There is no single risk across all uses. Relevant risks include factual error, privacy loss, insecure tool access, manipulation, unfair decisions, overreliance, concentrated power and, under stronger future capability scenarios, control and alignment problems.

What is the biggest opportunity?

Again, it depends on context. Major opportunities include faster learning, scientific and technical assistance, improved accessibility, lower-cost expertise, better information processing and more responsive coordination.

How do I know whether an SI answer is trustworthy?

Identify the important claims, inspect appropriate sources, reproduce calculations, check dates and confirm external actions in the actual system. Trust should follow evidence and track record rather than fluency alone.

Should I call old research AI or SI?

Preserve original historical terminology when referring to named fields, papers, institutions and quotations. This keeps the historical record intelligible while allowing current discussion to use SI where appropriate.

Why does this article still use the term AI?

Because it is a historical article about the transition from AI to SI. Removing the earlier term would make it impossible to describe the field’s history accurately.

What is the simplest definition of SI for a beginner?

For practical contemporary use, think of Super Intelligence as the modern family of intelligent computational systems that can learn patterns, generate content and increasingly use information and software tools. Keep that practical usage separate from the stronger technical concept of beyond-human superintelligence.

What should I remember from this entire article?

Remember one sentence: the terminology transition can have a date; the capability transition is a curve. Understanding both is the key to answering when AI became SI without confusing history, policy and technical capability.

Secondary students reading and discussing the decade-by-decade history of Super Intelligence

Part 58. A decade-by-decade history of the road from AI to SI

A timeline made only of famous breakthroughs can make progress look inevitable. It was not. Each decade contained competing ideas, disappointments, hardware limits, changing research fashions and applications that mattered outside the headline story. The purpose of this chronology is to show continuity without pretending every later success was predictable in advance.

The 1940s: computation, control and early neural abstractions

Before artificial intelligence became a named field, researchers were already developing mathematical models of computation, information and control. Electronic computing moved from theory and specialised machines toward programmable digital systems. Early formal neuron models explored how simple computational units might represent logical relationships.

These strands did not yet amount to modern SI, but they created intellectual components that later research could combine: formal computation, feedback, representation and the idea that aspects of cognition might be modelled mechanically.

The 1950s: machine intelligence becomes an explicit research programme

Turing’s 1950 paper made machine intelligence a direct philosophical and operational question. The Dartmouth proposal then supplied the artificial-intelligence label and a shared programme broad enough to include language, abstraction, learning and problem solving.

The decade established an enduring ambition: not merely automate arithmetic, but reproduce or approximate capabilities associated with intelligent behaviour.

The 1960s: optimism, search, language and symbolic reasoning

Researchers explored theorem proving, problem solving, game playing and early natural-language interaction. Many systems operated in simplified environments where rules and representations could be controlled.

The demonstrations were historically important because they showed that machines could manipulate symbols in ways that looked surprisingly cognitive. Their limitations also became clearer as researchers tried to move from constrained demonstrations to the ambiguity and scale of ordinary life.

The 1970s: knowledge and the limits of general methods

One lesson from early AI was that general problem-solving procedures often needed substantial domain knowledge. Research increasingly examined how knowledge could be represented explicitly and applied to specialised problems.

At the same time, expectations collided with computational and methodological limits. Some ambitious promises could not be delivered with available hardware, data and techniques. Funding and confidence tightened in parts of the field.

The 1980s: expert systems and commercialisation

Expert systems brought rule-based knowledge engineering into practical organisational settings. Their appeal was understandable: capture specialised rules, apply them consistently and make scarce expertise more available.

Commercial experience exposed maintenance costs. Knowledge bases had to be elicited, encoded, tested and updated. Systems could be brittle outside their intended scope. The decade demonstrated both that AI could create business value and that operational intelligence requires continuous maintenance.

The 1990s: statistical learning becomes increasingly central

As digital data grew and statistical methods matured, many researchers focused more heavily on learning from examples. Speech, text and pattern-recognition problems increasingly benefited from probabilistic modelling.

The conceptual change was important: uncertainty could be represented explicitly, and systems could infer useful patterns from observed data rather than relying only on hand-built symbolic rules.

The 2000s: data, the web and computing infrastructure expand the runway

The web generated enormous amounts of digital text, images and behavioural data. Computing became cheaper and distributed infrastructure more capable. Machine learning spread through search, advertising, recommendation, spam filtering and other large-scale applications.

Much of this intelligence was invisible to ordinary users because it sat behind products rather than speaking directly to them. The capability was growing before the conversational interface made it culturally obvious.

The 2010s: deep learning changes perception and representation

Large neural networks trained with powerful hardware produced major gains across image recognition, speech and other tasks. Representation learning reduced dependence on manually engineered features in many applications.

The decade also saw rapid advances in neural language modelling and sequence processing. By its end, pretraining increasingly suggested that one broadly trained model could support many downstream tasks.

The 2020s: language becomes the interface and models become platforms

Large generative models brought broad machine capability into ordinary conversation. Users no longer needed to understand the software architecture to request a summary, explanation, transformation, plan or code draft.

Then the surrounding system expanded. Retrieval added current or private knowledge. Multimodality added images and audio. Tools added calculation and software actions. Agents added adaptive multi-step execution. The model increasingly became a reasoning-and-language layer inside a larger computational system.

2026: terminology catches up with a changed public experience

By 2026, the technology experienced by users looked very different from the narrow specialist programs many people associated with earlier AI. The U.S. executive-branch terminology event discussed throughout this article placed the phrase Super Intelligence into a new official context.

The date is therefore historically interesting even if one rejects the idea that a name can settle a scientific capability question. It marks a change in how a major institution describes the technological era.

Part 59. The paradigm map: six ways researchers have tried to build intelligence

AI history is easier to understand when we stop imagining one method steadily improving. Different paradigms emphasise different sources of intelligence.

1. Symbolic reasoning

Represent concepts and rules explicitly, then manipulate them through logic, search or planning. Strengths include interpretability and precise structure. Weaknesses include brittleness and the labour of encoding messy knowledge.

2. Statistical learning

Estimate patterns from data and represent uncertainty probabilistically. Strengths include adaptation to observed variation. Weaknesses include dependence on data quality and difficulty interpreting some learned relationships.

3. Neural representation learning

Use layered parameterised networks to learn useful internal representations. Strengths include handling high-dimensional data such as images, speech and language. Weaknesses include large resource requirements and often-limited interpretability.

4. Reinforcement learning

Learn behaviour through reward signals arising from actions and outcomes. Strengths include sequential decision-making. Weaknesses include the difficulty of specifying rewards that faithfully represent the real objective.

5. Retrieval and external knowledge

Keep changing or detailed information outside model parameters and retrieve it when needed. Strengths include updateability and provenance. Weaknesses include search errors, stale indexes and incomplete source collections.

6. Hybrid intelligent systems

Combine models, symbolic constraints, retrieval, tools, deterministic software and people. This is increasingly how practical SI works because no single mechanism is best at every component of a real workflow.

The future may therefore be less about one paradigm defeating all others and more about intelligent orchestration of specialised mechanisms.

Part 60. The invisible history: chips, data centres and networks

A history focused only on algorithms misses the physical system underneath intelligence. Training and serving large models require chips, memory, storage, networking, cooling, power and engineering teams capable of operating large computational infrastructure.

Compute changes which ideas are testable

Researchers may understand an algorithm conceptually before they can afford to run it at useful scale. More capable hardware can turn an impractical idea into an experiment, and an experiment into a product.

Memory and bandwidth matter alongside arithmetic

Large models require moving enormous amounts of data between memory and processing units. Performance therefore depends not simply on the number of arithmetic operations a chip can perform but on the complete hardware and software stack.

Inference creates a continuing cost

Training a model is not the end of the resource story. Every user request consumes computation. A system that reasons longer, processes larger contexts or calls multiple models can cost more per task. Product design therefore involves trade-offs among quality, latency and cost.

Infrastructure shapes access

Large-scale training can require resources available to relatively few organisations. At the same time, hosted services can make the resulting capability available to millions of users. SI can therefore centralise some infrastructure while decentralising access to its outputs.

Part 61. The data transition: from carefully labelled datasets to internet-scale pretraining

Earlier supervised-learning systems often depended on labelled examples created specifically for a task. Labels are expensive because people must define categories and annotate examples.

Self-supervised training changed the economics of learning from large collections of unlabelled material. A model can create training signals from the structure of the material itself, such as predicting withheld or subsequent portions of text.

Scale creates coverage and noise

Large datasets expose models to many styles, domains and concepts. They can also contain errors, duplication, outdated information, bias and material of uncertain provenance. More data does not mean uniformly better data.

Curation becomes strategic

As raw scale grows, decisions about filtering, deduplication, mixture and quality become increasingly important. The dataset is not merely fuel poured into a model. Its composition helps shape what the model can learn.

Post-training data has a different job

Data used to teach instruction-following or specialised behaviour may be much smaller than pretraining corpora but highly influential. Demonstrations, preference comparisons and task-specific examples can change how a broadly pretrained model behaves in interaction.

Part 62. Evaluation evolved because the systems changed

A chess engine can be evaluated by games. A classifier can be evaluated against labelled examples. A general language assistant can be asked almost anything. As system breadth increases, evaluation becomes more difficult because no small benchmark represents every use.

Benchmarks make comparison possible

A shared test lets researchers compare methods under similar conditions. Benchmarks accelerate progress by creating a visible target and a common language for performance.

Benchmarks can become targets

Once a benchmark becomes important, researchers optimise systems around it. Training data may contain similar material, methods may exploit quirks of the test and repeated tuning can reduce the benchmark’s value as independent evidence.

Real-world evaluation needs scenarios

For a deployed SI workflow, build scenarios from actual work: ordinary cases, difficult exceptions, missing information, conflicting instructions and tool failures. Evaluate the complete path from request to reviewed outcome.

Evaluation must continue after deployment

Models, data, tools and user behaviour change. A workflow that passed a test six months ago may encounter new document formats or policies. Monitoring and periodic re-evaluation are part of maintaining capability.

Part 63. The economics of intelligence: what happens when cognitive output becomes cheaper?

Technology matters economically when it changes the cost of producing something people value. SI can reduce the marginal cost of many information tasks: generating a draft, translating text, classifying documents, writing routine code or producing a first analysis.

Lower cost can reduce employment in some tasks, increase demand in others and create entirely new activities. The direction depends on substitution, complementarity, demand elasticity, regulation and how quickly organisations redesign processes.

Cheap drafts increase the value of review

If an organisation can generate one hundred drafts instead of ten, someone still needs a method for deciding which are correct and useful. The bottleneck may move rather than disappear.

Cheap expertise can expand markets

Tasks that were previously too expensive for small organisations or individuals may become affordable. This can increase total demand for analysis, tutoring, design and software even as the labour required per unit falls.

Quality standards may rise

When competent first drafts become abundant, customers may expect faster service, more personalisation and stronger evidence. Productivity gains can therefore become higher standards rather than simply fewer working hours.

Part 64. How to read claims about the history of AI and SI

Technology history is vulnerable to tidy stories. A later breakthrough can make earlier research look like a straight road toward the present. In reality, researchers did not know which ideas would dominate decades later.

Distinguish invention, coinage, demonstration and adoption

A term can be coined before the technology it names is mature. A method can be invented years before hardware makes it practical. A laboratory demonstration can precede mass adoption by decades. Ask which event a date actually refers to.

Prefer primary sources for milestone claims

For a famous paper, read the paper or institutional archive when possible. Secondary summaries are useful for context, but they can repeat simplified origin stories.

Beware “first ever” claims

Priority disputes are common because inventions have precursors, independent discoveries and different definitions of what counts as the first working example. Use narrow wording: first published demonstration under a stated definition, not first intelligent machine in history.

Keep historical terminology intact

Do not rewrite the past so every researcher appears to have used today’s vocabulary. Original terminology helps reveal how concepts changed over time.

Separate contemporary interpretation from historical fact

We can reasonably interpret the 2026 terminology event as part of a broader shift in public experience. That interpretation is different from the factual claim that a particular order was issued on a particular date. Good historical writing labels the difference.

Part 65. A better way to judge a milestone

Instead of asking whether an event “changed everything”, score its historical importance conceptually across five questions without forcing a numerical ranking.

Novelty: did it introduce a genuinely new mechanism, representation or framing?

Capability: did it enable performance that was previously difficult or impossible?

Scalability: could the approach improve with more data, compute or deployment?

Transfer: did the idea influence many tasks or remain specialised?

Adoption: did it move beyond a paper into research practice, products or institutions?

The Transformer scores strongly on several of these dimensions historically because it influenced a wide range of later model development. The Dartmouth project is important for a different reason: field formation and naming. Turing’s 1950 paper is important for framing. Different milestones can be foundational in different ways.

Part 66. From perception to generation: why the capability frontier moved

For many years, some of the most visible machine-learning successes involved recognising patterns: identifying an object, transcribing speech, ranking a search result or predicting which item a user might prefer. These are powerful capabilities, but the system’s output space is comparatively constrained.

Generative models changed the experience because the output itself could be open-ended. Instead of selecting one label from a fixed list, a model could compose a paragraph, image or program. This made intelligence feel less like classification and more like collaboration.

Recognition asks “which?”

A classifier maps an input toward a category. The possible answers are known in advance. Evaluation can often compare the selected category with a labelled target.

Generation asks “what could fit?”

A generative system has many possible acceptable outputs. Evaluation becomes harder because there may be no single correct sentence or image. Quality can involve factuality, relevance, style, originality and constraint satisfaction simultaneously.

Open-ended output increases both value and review burden

The same flexibility that makes generation useful makes it difficult to validate with one metric. A generated contract summary can be fluent and omit an exception. A generated program can pass ordinary tests and fail on an edge case. The evaluation has to follow the purpose.

Part 67. Reasoning in SI: what do we actually mean?

The word reasoning can refer to several observable behaviours: solving multi-step problems, applying rules, combining evidence, planning intermediate steps or revising an answer after discovering a contradiction. These behaviours can be tested without claiming access to a model’s subjective inner experience.

Outcome matters

For a mathematical problem, we can inspect whether the final result and derivation are valid. For a planning task, we can inspect whether the plan satisfies constraints. For research, we can inspect whether conclusions follow from cited evidence.

Process can help evaluation without becoming perfect introspection

Asking a system to show calculations, intermediate evidence or a plan can make errors easier to detect. Generated explanations should not automatically be treated as a complete faithful transcript of internal computation.

External tools can improve reasoning systems

A model does not need to perform every operation internally. It can delegate arithmetic to a calculator, search to a retrieval system and formal execution to code. Intelligence can reside partly in knowing which operation to use and how to integrate the result.

Part 68. Tool use is a historical milestone because it changes the unit of intelligence

Once models can call external tools, evaluating the model alone becomes insufficient. The useful unit is the model-plus-tool system.

A model may be mediocre at long multiplication and excellent at recognising that exact arithmetic should be delegated to a calculator. In a real workflow, that orchestration can be more useful than forcing the model to imitate a calculator internally.

Tools externalise precision

Databases provide exact stored records. Calculators provide deterministic arithmetic. Search systems provide current documents. Code execution provides formal computation. The language model can focus on translating human intent into the right operation and interpreting the result.

Tools also externalise risk

A wrong tool call can change a real record. The model’s uncertainty becomes operational when it selects parameters for an external action. Permissions, previews and validation therefore become part of intelligence-system design.

Part 69. Memory changes the relationship from session to relationship

A stateless assistant treats each interaction largely as a new task. Persistent memory can preserve preferences, decisions or working context across time. This reduces repeated explanation and supports longer-running projects.

But memory creates new questions: what should be remembered, who can inspect it, how errors are corrected and when information becomes stale.

Not all memory should be equally durable

A writing preference may remain stable. Today’s inventory does not. A temporary personal circumstance may be sensitive and short-lived. Good memory systems distinguish durable preference from current operational state.

Memory can amplify an error

If an incorrect fact is stored and reused, one mistake becomes a persistent source of future mistakes. Correction therefore needs to be a first-class operation.

Part 70. Personalisation: the system adapts to the person

Personalisation can change explanations, recommendations and workflow defaults according to a user’s preferences or history. This can make SI feel more intelligent because less context has to be repeated.

Yet personalisation should not trap the user inside old assumptions. Preferences change. People explore new interests. A system should make it possible to override or revise remembered patterns.

Prediction of preference is not permission

If the system predicts that you usually prefer the cheapest option, it can surface that option first. It should not silently purchase it unless the workflow separately grants that authority.

Part 71. The interface transition: from menus to intent

Traditional software exposes functions through menus, forms and buttons. The user must learn where the function lives. Language-based SI lets the user begin with intent: “compare these files”, “find the discrepancy”, “turn this into a timetable”.

This is a profound interface change because one conversational surface can potentially control many underlying functions.

Intent still has to become structure

Behind the conversational interface, the system must convert a vague request into specific operations and parameters. If the user says “next Friday”, the system must resolve a date. If the user says “send it to the team”, the system must resolve recipients.

Good interfaces expose consequential interpretations before action. Convenience should remove unnecessary friction, not remove the moment where ambiguity matters.

Part 72. Software after SI: interfaces become more dynamic

When a system can interpret natural-language intent, software does not need to expose every possible workflow as a fixed screen. Interfaces can become adaptive: showing the evidence, controls and approval step relevant to the current task.

This does not mean graphical interfaces disappear. Visual structure remains valuable for comparison, monitoring and precise control. The likely pattern is combination: conversation for intent, structured UI for state and consequence.

Language starts the task; structure confirms it

A user may say “book the earliest acceptable slot”. Before confirmation, the interface can show the selected date, participants and conflicts in a structured form. Natural language reduces setup cost; structured review reduces ambiguity.

Part 73. Scientific Super Intelligence: acceleration depends on the experimental loop

Science contains many cognitive tasks that SI can assist: reading literature, extracting data, writing code, proposing hypotheses, designing simulations and analysing results. The largest gains occur when these steps connect into a faster experimental loop.

But science is not only text generation. Many questions require instruments, laboratories, field observations, manufacturing and physical time. A model can propose an experiment instantly; cells may still need days to grow.

Digital sciences can iterate faster

Fields where experiments are primarily computational may benefit quickly because hypothesis, code, execution and analysis can all occur in digital infrastructure.

Physical sciences retain material bottlenecks

Robotics and automated laboratories can reduce some bottlenecks, but physical experiments still face safety, equipment and resource constraints. The pace of intelligence is not always the pace of reality.

Scientific judgement includes choosing what matters

Generating many hypotheses is useful only if researchers can identify which are informative and feasible. SI can widen the search space while human or machine evaluation narrows it according to evidence and scientific value.

Part 74. Mathematics and SI: formal verification changes the confidence problem

Mathematics offers an unusual advantage: many claims can be checked formally. A generated proof can be translated into a formal system whose rules determine whether each step is valid.

This creates a powerful division of labour. Generative models can propose ideas; formal proof assistants can verify exact derivations. The combination can be more dependable than relying on persuasive mathematical prose alone.

Formal truth is narrower than usefulness

A formally correct proof may be unreadable or uninteresting. Choosing an elegant theorem, explanatory structure or useful abstraction remains a different task from verifying validity.

Part 75. Robotics: when SI leaves the screen

Digital agents act in software. Robots act in physical environments where errors can damage objects or harm people. This changes the safety requirements.

Physical systems must deal with noisy sensors, uncertain geometry, mechanical limits and environments that cannot be reset as easily as a software simulation.

Embodiment creates grounding and constraint

A robot receives direct sensory information and must translate high-level goals into physical actions. The world provides immediate feedback: an object either moves as expected or it does not.

Simulation can accelerate learning but not eliminate reality

Training and testing in simulation can reduce cost and risk. Real deployment still needs validation because simulations approximate rather than perfectly reproduce physical environments.

Part 76. Collective intelligence: one model is not the only architecture

Some tasks can be decomposed among multiple models, tools or human specialists. One component retrieves evidence, another checks calculations, another critiques the draft and a human makes the final decision.

This resembles organisations: intelligence can emerge from coordination among specialised roles rather than one universal expert.

Diversity can improve error detection

Independent methods may fail differently. Comparing outputs can expose disagreements worth investigating. Independence matters; asking identical copies of the same model the same question may reproduce the same blind spot.

Coordination has a cost

Multiple agents can duplicate work, disagree endlessly or consume more resources than the task warrants. Collective intelligence requires routing and stopping rules.

Part 77. The capability transition in one chain

The long history can now be compressed without losing its structure:

Computation made procedures executable. Symbolic AI made knowledge and search explicit. Statistical learning made behaviour learnable from data. Deep learning made representation itself learnable at scale. Transformers made large sequence models highly scalable. Foundation models made one model reusable across tasks. Instruction-following made those models easier to direct. Multimodality widened what they could perceive. Retrieval connected them to changing knowledge. Tools connected them to precise computation and software. Agents connected operations across time. Memory connected sessions. Governance determines how these capabilities can be used responsibly.

This chain is the technological answer to why AI could plausibly be described differently by 2026. It is not proof of the older technical superintelligence threshold. It is evidence that the practical object called AI had become much broader than many earlier public conceptions of the field.

Secondary students applying Super Intelligence through worked case studies

Part 78. How to read the case studies in this guide

The following cases are original teaching scenarios. They are not customer testimonials, measured deployments or claims about named organisations. Their purpose is to expose the structure of an SI workflow: task, evidence, machine contribution, human responsibility, failure mode and verification.

Part 79. Case study: a Secondary student learns rather than copies

A fictional Secondary student, Mira, is revising algebra. She can follow worked examples but becomes stuck when the question changes form. If she asks SI for complete solutions every time, her homework may improve while her transfer skill remains weak.

Preserve the diagnostic attempt

Mira solves the first question alone and records her working. She expands a bracket correctly but moves a term across the equality sign without preserving the operation. The mistake is now visible.

Diagnose before solving

The SI tutor identifies the first incorrect step and explains the principle without completing the rest. It then generates a contrasting pair and, later, a changed problem for independent transfer.

The distinctive capability is adaptive diagnosis and rapid targeted variation. The educational design determines whether that capability strengthens or replaces thinking.

Part 80. Case study: vocabulary learning becomes a control skill

A fictional student reads a policy described as “provisional” and does not understand why replacing the word with “approved” changes the operational meaning. SI builds a semantic ladder: proposed, provisional, pending approval, approved, implemented.

The learner sorts examples and explains which status permits action. This is vocabulary teaching and SI literacy at once: language precision controls what a system or person is entitled to infer and do.

Part 81. Case study: researching a historical claim

A fictional student writes, “AI was invented at Dartmouth in 1956.” SI improves the assignment by decomposing invented: field name, first research, first program or first gathering?

The student reads Turing’s 1950 paper and Dartmouth’s history. The thesis becomes narrower: Dartmouth was a foundational field-forming event while important machine-intelligence work preceded it. The improvement comes from stronger evidence, not fancier prose.

Part 82. Case study: a teacher creates diagnostic practice

A fictional teacher notices that several students confuse correlation with causation. Instead of requesting generic questions, the teacher specifies the misconception. SI proposes a progression from identifying correlation to finding confounders, distinguishing prediction from intervention and critiquing a causal headline.

The teacher still owns sequencing, appropriateness and standards. Automation increases the supply of material; expertise determines selection.

Part 83. Case study: a small-business enquiry workflow

A fictional tuition centre receives repeated enquiries about classes, timing, fees and availability. It first creates one authoritative current service table and marks old material superseded.

SI retrieves approved fields and prepares drafts while preserving the distinction between current availability and a requested place. Unlisted discounts and unusual exceptions are escalated to the authorised person.

The organisation measures draft accuracy, review time and unsupported-promise errors. The objective is faster accurate service, not maximum automation.

Part 84. Case study: comparing two supplier proposals

Supplier A has a lower headline price but excludes delivery. Supplier B includes delivery but has a longer lead time. One proposal states a warranty; the other is silent.

Instead of “Which is better?”, SI extracts price, inclusions, exclusions, delivery, lead time, warranty and unresolved terms with source references. Missing warranty information is reported as “not stated”, not silently converted to “no warranty”.

The organisation defines whether cost, speed or another criterion matters before asking for a recommendation.

Part 85. Case study: operational anomaly detection

A fictional facilities team receives hundreds of maintenance reports. SI groups them and flags a sudden cluster of water-leak descriptions in one building.

The cluster is a signal, not a diagnosis. Staff check duplicate reports, location metadata and the physical site. The system’s leverage is scanning scale and routing attention; physical inspection establishes reality.

Part 86. Case study: health-information summarisation without pretending to diagnose

A fictional clinic has an approved patient-information leaflet. SI can explain preparation steps in plainer language while preserving important warnings and directing individual medical questions to the appropriate professional.

High-stakes fields such as dosages, contraindications and emergency advice require source fidelity. A general leaflet does not authorise an individual clinical decision.

Part 87. Case study: city-scale information routing

A fictional city receives reports about streetlights, drains and damaged public facilities. SI classifies reports, detects likely duplicates and routes them to responsible teams.

Location errors can send crews to the wrong place, so geographic fields require validation. Priority also requires policy: the system can estimate urgency, but the institution should define what receives priority and how uncertain cases are escalated.

Part 88. Case study: accelerating a literature review

A fictional research group asks SI to search an approved scholarly collection, extract experimental conditions and build a comparison table. Every important extracted value includes a page reference.

If studies disagree, the system preserves the disagreement and differences in methods rather than averaging them into false consensus. Researchers inspect the original papers before relying on surprising results.

Part 89. Case study: generated code with a test contract

A fictional analyst needs a function that groups transactions by month and rejects invalid dates. Before accepting generated code, the analyst defines tests for empty input, valid rows, mixed invalid rows, leap-day input and malformed dates.

The tests are independent of the generated implementation. If SI rewrites the code later, the same contract detects behavioural regressions.

Part 90. Case study: personal planning without outsourcing values

A fictional user compares two weekend courses. SI calculates cost, travel time, schedule and syllabus overlap. It cannot discover how much the user values in-person community versus flexibility unless the user supplies that preference.

The system exposes the trade-off; the user owns the value judgement.

Part 91. Case study: what a good SI failure looks like

An assistant is asked for today’s stock, but the inventory tool fails. A bad failure reports yesterday’s remembered number as current. A good failure says current inventory could not be retrieved, identifies the last confirmed timestamp if useful and refuses to manufacture today’s state.

Failure quality is a capability. A system that degrades honestly is more dependable than one that appears complete by guessing.

Part 92. What all these cases have in common

Define the task. Identify authoritative evidence. Use SI where speed and scale create leverage. Preserve the authority boundary. Verify the consequential part.

This repeated pattern is the practical meaning of the AI-to-SI transition. The technology became broad enough to participate across whole workflows. Human responsibility therefore moves toward defining, governing and verifying those workflows.

Secondary students checking work carefully while learning Super Intelligence verification

Part 93. A failure taxonomy for Super Intelligence

Calling every bad result a hallucination is too crude for serious work. Different failures originate in different parts of the system and require different repairs. If the wrong document was retrieved, changing the writing prompt may do nothing. If the arithmetic tool received the wrong number, replacing the language model may do nothing.

A useful failure taxonomy begins by locating the first point where the system diverged from the intended state.

Failure type 1: task-definition failure

The system solves a different problem from the one the user actually has. “Find the cheapest supplier” is executed correctly even though the organisation really needed the lowest total delivered cost before Friday.

Repair: redefine the objective and constraints before changing the model.

Failure type 2: evidence failure

The necessary fact is absent, stale, contradictory or drawn from a non-authoritative source.

Repair: improve the evidence set, source ownership, freshness or retrieval rules.

Failure type 3: extraction failure

The correct source is present but a number, date, name or clause is read incorrectly.

Repair: validate extraction, preserve the original source and add field-level checks for consequential values.

Failure type 4: reasoning failure

The system has the relevant facts but combines them incorrectly. It may confuse correlation with causation, apply a rule outside its scope or make an invalid calculation.

Repair: test the reasoning step, delegate formal operations to appropriate tools and add counterexamples.

Failure type 5: generation failure

The underlying analysis may be sound but the final wording overstates certainty, drops a condition or introduces an unsupported detail.

Repair: compare claims with evidence and use structured output where precision matters.

Failure type 6: tool-selection failure

The system chooses an inappropriate operation: searching the web when it should query the internal database, or estimating arithmetic that should be calculated exactly.

Repair: improve routing rules and test tool choice as a separate capability.

Failure type 7: parameter failure

The correct tool is selected with the wrong date, recipient, identifier, unit or filter.

Repair: validate consequential parameters and preview them before writes.

Failure type 8: permission failure

The system attempts an action it is not authorised to perform, or has broader credentials than the task requires.

Repair: enforce least privilege outside the prompt layer.

Failure type 9: execution failure

The intended action reaches an external system but errors, times out or completes only partially.

Repair: distinguish attempted from completed and verify resulting state before retrying.

Failure type 10: review failure

The system produces a detectable mistake but the human or automated reviewer fails to catch it.

Repair: redesign the review interface so important evidence, uncertainty and changes are easier to inspect.

Failure type 11: governance failure

Nobody clearly owns the workflow, its acceptable error rate or the decision to suspend it.

Repair: assign ownership and escalation before scaling deployment.

Failure type 12: objective failure

The system performs exactly as measured while the metric itself rewards undesirable behaviour.

Repair: revisit the objective, proxy metric and constraints rather than celebrating metric improvement.

Part 94. How to debug an SI workflow

Debugging begins with reproduction. Save the request, evidence, configuration, retrieved material, tool operations and final result. Without the original state, investigation becomes guesswork.

Step 1: identify the first wrong state

Do not begin with the final bad paragraph. Trace backwards. Was the wrong fact already present in retrieval? Did the correct fact become wrong during summarisation? Did the final text remain correct while the external action used the wrong identifier?

Step 2: isolate the component

Run the retrieval independently. Run the calculation independently. Ask the model to operate on a known-clean evidence packet. Isolation distinguishes a model problem from a system-integration problem.

Step 3: build a regression test

Turn the discovered failure into a test that future versions must pass. One production mistake can improve the system permanently if it becomes part of the evaluation set.

Step 4: test neighbouring cases

A fix that solves one exact prompt may be brittle. Vary the date, entity, wording and exception. The repair should address the failure class, not memorise one incident.

Step 5: document the boundary

If the system remains unreliable under a condition, state that condition and route it elsewhere. Knowing where not to automate is part of mature deployment.

Part 95. The verification playbook: match the check to the claim

Verification should be proportional and specific. “Double-check everything” is not a usable instruction for a long workflow. Identify the claim type and choose a check that can actually falsify it.

Claim typeUseful check
Exact numberRecalculate or query the authoritative numerical record.
QuotationCompare exact wording with the original source and surrounding context.
Current statusCheck a current authoritative system with a timestamp.
Historical dateUse primary or strong institutional historical sources.
Scientific resultInspect the paper, methods, population and reported result.
Legal or policy requirementCheck the applicable current authoritative text and jurisdiction; obtain qualified interpretation where required.
Completed software actionInspect the resulting external state or confirmation record.
Causal claimExamine the research design rather than relying on predictive association.
RecommendationInspect objectives, assumptions, alternatives and trade-offs.

Verify the hinge claim

In many decisions, one claim changes the outcome. A supplier may be preferable only if delivery is included. A meeting may be possible only if one participant is available. Find this hinge claim and verify it first.

Use independent checks where independence matters

Asking the same model to repeat its answer may reproduce the same error. A stronger check can use the original source, deterministic calculation, different method or qualified human review.

Part 96. Source quality in the SI era

SI makes it easy to find and summarise information, which increases the importance of source hierarchy. Not every source is appropriate for every claim.

Primary sources

These include original papers, official records, statutes, datasets, direct statements and first-party technical documentation. They are often best for establishing what was actually published, required or measured.

Secondary sources

These interpret and contextualise primary material. Strong secondary scholarship can explain significance, disagreement and historical context better than an isolated primary document.

Tertiary sources

Encyclopedias and general explainers are useful for orientation and terminology. Important claims should often be traced onward to stronger evidence.

Source quality is claim-dependent

A company’s own documentation is authoritative for what its current API accepts, but not necessarily for an independent comparison of product quality. A newspaper can be strong for reporting a dated event while a peer-reviewed study may be stronger for a scientific causal claim.

Freshness is part of quality

An excellent old source can be wrong for a current-state question. Historical authority and present authority are different dimensions.

Part 97. Citation discipline: a link is not automatically evidence

Generated research can create the appearance of scholarship by attaching many links. The important question is whether each source supports the nearby claim.

Claim-to-source alignment

Read the sentence without the citation. Identify what must be true for the sentence to be correct. Then read the source and ask whether it establishes those elements.

Do not cite a search result snippet as though it were the source

Search snippets are navigation aids. They can truncate qualifications or combine text in misleading ways. Open the underlying source.

Do not let one citation carry an entire paragraph

If a paragraph contains historical fact, interpretation and prediction, separate them. The source may establish only the historical fact.

Preserve disagreement

If credible sources disagree, cite the disagreement and explain its basis rather than selecting the most convenient conclusion without discussion.

Part 98. Five maturity levels for SI deployment

Organisations often ask whether they are “using SI”. A more useful question is how mature the workflow is.

Level 1: personal experimentation

Individuals use SI for low-consequence drafting, explanation and brainstorming. Processes are informal and results depend heavily on the user.

Level 2: repeatable assisted task

A team defines a recurring use, approved inputs and a review step. Examples include document extraction or first-pass summaries.

Level 3: integrated workflow

SI connects to authorised knowledge or tools. Inputs, outputs, permissions and handoffs are documented. Evaluation uses representative scenarios.

Level 4: monitored operation

The organisation measures errors, latency, cost and review burden. Failures create regression tests. Owners can suspend or modify the workflow.

Level 5: adaptive intelligent system

Multiple components coordinate across complex work, with robust governance, observability, recovery and continuous evaluation. Higher autonomy is granted only where evidence supports it.

Maturity is not a race toward Level 5. A Level 2 workflow may be the correct design for a consequential task where human judgement should remain central.

Part 99. The pre-deployment checklist

Before moving an SI workflow from experiment into routine use, answer these questions in writing.

Purpose: What problem are we solving, and what outcome counts as success?

Scope: Which cases are included and excluded?

Evidence: Which sources are authoritative and how are they kept current?

Data: What information enters the system, and is its use permitted?

Evaluation: Which representative and adversarial cases has the workflow passed?

Permissions: What can the system read and write?

Review: Which outputs require human approval?

Recovery: What happens when retrieval, generation or an external action fails?

Ownership: Who is responsible for performance and incidents?

Monitoring: Which metrics and failure categories will be tracked?

Change control: What happens when the model, prompt, data source or tool changes?

Exit: How can the workflow be disabled safely?

Part 100. The AI-to-SI transition matrix

The article has now reached one hundred numbered parts. The matrix below compresses the transition across four dimensions without pretending they changed simultaneously.

DimensionEarlier AI patternSI-era patternNew responsibility
CapabilityNarrower specialist tasksBroad reusable models and integrated systemsEvaluate breadth and boundaries explicitly.
InterfaceMenus, code and specialist inputsNatural language and multimodal interactionResolve ambiguity before consequence.
KnowledgeRules, static datasets or model parametersParameters plus retrieval and memoryManage authority, freshness and provenance.
ActionPrediction or recommendationTool use and agentic executionDesign permissions, confirmation and recovery.
Human roleOperator of softwareSpecifier, verifier and workflow architectBuild domain judgement and governance.
EvaluationTask benchmarkScenario and whole-system evaluationTest failures, exceptions and deployment state.
TerminologyArtificial Intelligence / AISuper Intelligence / SI in the contemporary context described herePreserve historical meanings and avoid false technical equivalence.

The matrix is the compact answer to the entire long-form guide: the technological object expanded across capability, interface, knowledge and action, while the human role moved upward toward specification and control. The terminology change arrived after that long capability transition was already underway.

Secondary students connecting Super Intelligence across different fields of knowledge

Part 101. The Super Intelligence domain atlas

A technology can be general at the interface while remaining uneven across domains. The same SI system may be excellent at transforming prose, useful at generating code, unreliable on an obscure current fact and inappropriate as the final authority for a regulated professional decision.

The domain atlas asks five questions repeatedly: what does SI accelerate, what evidence does it need, what failure matters, what remains human-owned, and how should performance be tested?

Part 102. Language and communication

Language is the most visible SI domain because it is both a capability and an interface. Systems can summarise, rewrite, translate, classify tone, extract entities, compare documents and generate new text.

Acceleration and evidence

First drafts, transformation, multilingual assistance and document triage become inexpensive. Factual tasks still need trustworthy sources; exact wording still belongs to the original document.

Failure and ownership

Meaning can shift during paraphrase, qualifications can disappear and translation can flatten culturally specific language. Humans retain purpose, audience judgement and final factual responsibility.

Part 103. Mathematics

Mathematics combines conceptual explanation with exact formal relationships. SI can teach methods, propose solution paths and generate practice, while calculators and formal systems verify parts of the work.

Failure and evaluation

A persuasive derivation can contain an invalid step. Check final results independently, inspect derivations and use formal verification where justified. For learning, students must eventually reproduce the method independently.

Part 104. Natural science

Science cycles through observation, hypothesis, experiment and revision. SI can accelerate literature mapping, data extraction, code, simulation, hypothesis generation and analysis. It cannot convert an unperformed experiment into empirical evidence.

Failure and ownership

Important risks include invented citations, inappropriate comparison of incompatible studies and hypotheses presented as findings. Scientists retain responsibility for experimental legitimacy and published claims.

Part 105. Education

Education is unusual because the output is not only the assignment; it is the learner’s changed capability. SI can provide personalised explanation, diagnostic questions, practice variation and feedback.

Evaluation

Measure transfer and delayed independent performance, not only the quality of assisted output. Curriculum goals, safeguarding and assessment validity remain human-owned.

Part 106. Financial-style analysis

SI can organise financial documents, explain ratios, compare scenarios and assist modelling without becoming an unquestionable financial decision-maker.

Failure and evaluation

Wrong units, stale prices, hidden assumptions and double-counting can invalidate polished analysis. Reconcile numbers to source records, test formulas and stress assumptions. Risk tolerance and regulated decisions remain with authorised people.

Part 107. Operations and logistics

Operations contain structured states, deadlines and dependencies. SI can accelerate exception detection, schedule comparison, status summarisation and coordination.

Failure and evaluation

Stale state, wrong identifiers, timezone errors and treating expected events as completed events are central risks. Evaluate missed exceptions and false alarms against real operational consequences.

Part 108. Law and policy information

Legal and policy work depends heavily on text, but jurisdiction, date, authority and procedural posture matter. SI can assist research, version comparison and defined-term extraction.

Evidence and ownership

Use current authoritative legal texts and appropriate legal sources. Invented cases, outdated law and wrong jurisdiction are critical failures. Qualified interpretation and consequential legal decisions remain with appropriate professionals and authorities.

Part 109. Healthcare information

Healthcare combines information processing with high-consequence individual decisions. SI can assist summarisation, administration, patient education from approved material and literature support.

Failure and ownership

Medication errors, missing contraindications and inappropriate population-to-individual generalisation can be consequential. Diagnosis, treatment and clinical responsibility remain within appropriate professional care processes.

Part 110. Engineering

Engineering turns models into systems that must work under physical constraints. SI can accelerate requirement decomposition, simulation scripts, design alternatives and documentation.

Evaluation

Independent calculations, simulation, prototyping, physical testing and applicable standards remain essential. Incorrect units or hidden material assumptions can turn a plausible design into an unsafe one.

Part 111. Software engineering

Software is especially compatible with SI because requirements and outputs can often be represented digitally and generated work can be executed immediately.

Evaluation and ownership

Use automated tests, static analysis, security review and production monitoring. Architecture, production authority and responsibility for user impact remain human-owned.

Part 112. Cybersecurity and defensive analysis

SI can help defenders understand logs, explain vulnerabilities, classify alerts and draft remediation guidance. Because security work can be dual-use, authorisation and workflow boundaries matter strongly.

Evaluation

Use authorised test environments, known incidents, false-positive analysis and independent validation of critical findings. Incident command and risk acceptance remain human responsibilities.

Part 113. Media, journalism and public information

SI can accelerate transcript search, document comparison, translation and timeline construction. Journalism still depends on verification, attribution and editorial accountability.

Failure and evaluation

Fabricated quotations, source confusion and loss of uncertainty are critical failures. Trace quotations to recordings or documents and factual claims to the underlying reporting.

Part 114. Creative industries

Creative SI changes the speed and volume of variation. Writers and designers can explore directions rapidly, but abundance can make distinctiveness harder.

Human ownership

Taste, purpose, cultural judgement, final selection and responsibility for rights and publication remain central. Evaluate against the creative brief rather than a universal creativity score.

Part 115. Public-sector administration

Public institutions can use SI to reduce administrative burden and improve access to information. Public decisions also carry legitimacy, equality and accountability requirements beyond ordinary commercial optimisation.

Acceleration

Form assistance, document routing, translation, public-information search and internal summarisation.

Failure and ownership

Incorrect eligibility information, unequal service quality and opaque automation of decisions affecting rights or benefits are major risks. Legal authority, policy decisions, appeal mechanisms and public accountability remain institutional responsibilities.

Part 116. Cities and infrastructure

Infrastructure produces continuous operational signals: transport flows, maintenance reports, energy demand, water systems and public-space conditions. SI can help turn this information into earlier detection and better coordination.

Failure and evaluation

Bad sensors, incorrect location mapping and optimisation that improves one subsystem while harming another require whole-system evaluation. Measure recovery time, false alerts, service continuity and distributional effects.

Part 117. Environment and climate-related analysis

Environmental systems combine huge datasets, physical models and long time horizons. SI can assist literature synthesis, remote-sensing analysis, scenario comparison and model-code development.

Failure

Confusing weather with climate, presenting one scenario as a prediction, ignoring model assumptions and mixing incompatible measurements can distort conclusions. Policy priorities and acceptable risk remain human decisions.

Part 118. Manufacturing

Manufacturing connects digital planning to physical production. SI can improve visual inspection support, predictive-maintenance signals, scheduling, documentation and root-cause analysis.

Failure and ownership

Sensor drift, false defect detection and unsafe schedule optimisation are central risks. Safety, quality release and maintenance authority remain with the responsible organisation and professionals.

Part 119. Customer service

Customer service is an immediate SI application because much of the work involves language plus account or policy information.

Acceleration

Intent detection, knowledge retrieval, response drafting, translation and interaction summarisation.

Failure and evaluation

Invented policy, privacy leakage, unauthorised promises and failure to escalate can harm customers. Measure resolution quality, accuracy, privacy and repeat contacts—not merely response speed.

Part 120. Management and organisational decision-making

Managers process incomplete information, allocate resources and coordinate people. SI can increase the information considered through meeting synthesis, scenario analysis, risk registers and decision-memo drafting.

Failure and ownership

False precision, missing interpersonal context and optimisation around narrow metrics can produce bad management. Accountability, personnel decisions and organisational priorities remain human responsibilities.

Part 121. Agriculture and food systems

Agriculture combines biological uncertainty, weather, logistics and physical work. SI can integrate sensor readings, forecasts, crop observations and supply-chain information to support decisions.

Failure

Local conditions can invalidate general recommendations. A model trained on one climate, crop variety or soil regime may transfer poorly to another.

Evaluation

Measure outcomes under local field conditions and preserve agronomic expertise, safety rules and farmer judgement.

Part 122. Transport and mobility

Transport systems require routing, forecasting, maintenance and real-time coordination. SI can support demand prediction, incident triage and passenger information.

Failure and ownership

Stale traffic state, sensor failure and optimisation that shifts congestion elsewhere can undermine local gains. Safety-critical control requires domain-specific engineering and regulation.

Part 123. Energy systems

Energy networks balance generation, demand, storage and physical constraints continuously. SI can assist forecasting, anomaly detection, maintenance planning and scenario analysis.

Evaluation

Reliability, resilience and recovery matter alongside average efficiency. A small forecasting gain is not useful if the control architecture becomes more fragile under rare events.

Part 124. Supply chains

Supply chains connect suppliers, inventory, transport, demand and risk. SI can detect disruptions, compare alternatives and summarise changing conditions.

Failure

Optimising one metric such as inventory cost can reduce resilience. Recommendations should expose trade-offs among cost, lead time, redundancy and service level.

Part 125. Human-resources information workflows

SI can assist with policy search, scheduling, document preparation and administrative questions. Employment decisions can affect people’s livelihoods and may be subject to legal and organisational requirements.

Boundary

Administrative assistance should not silently become opaque automated judgement about hiring, promotion or discipline. Consequential decisions need appropriate governance, evidence and human accountability.

Part 126. Accessibility and assistive use

SI can convert information between modalities, simplify language, describe images, support translation and adapt explanations. These capabilities can reduce barriers for many users.

Evaluation

Accessibility should be tested with the people and contexts the system is meant to serve. A generated description that omits the feature a user needs is not accessible merely because it exists.

Part 127. What the domain atlas reveals

The domains differ, but one pattern repeats. SI is strongest where information can be represented digitally, useful transformations can be evaluated and specialised tools provide precise operations. Risk increases when evidence is weak, outputs affect people consequentially, physical reality is hard to simulate or values are hidden inside the objective.

The human role changes shape rather than vanishing. Domain experts increasingly define valid evidence, design evaluations, set authority boundaries and decide what the system should optimise.

This is another reason the AI-to-SI transition is not simply a story of machines becoming smarter. It is a story of intelligence becoming infrastructure across domains—and therefore of every domain needing its own definition of dependable intelligence.

Secondary students evaluating evidence and measuring Super Intelligence capability

Part 128. How do we measure Super Intelligence?

The stronger the claim, the stronger the evidence should be. “This system can summarise our standard reports” can be tested with a small representative set. “This system is generally more intelligent than humans” requires a much broader argument about tasks, people, resources and reliability.

Measurement begins by refusing to let one impressive result stand for everything. Intelligence is multidimensional in practice: knowledge, reasoning, learning, planning, perception, tool use, adaptation, social understanding and reliability can move differently.

Define the construct before the score

A benchmark score has meaning only if the benchmark measures something relevant to the claim. If a test measures multiple-choice science questions, it provides evidence about performance on those questions under those conditions. Calling the score “general intelligence” requires additional justification.

Separate capability from deployment quality

A model can solve a problem when given the right context while a product fails to retrieve that context. Conversely, a well-designed application can compensate for model weaknesses. Measure both the underlying model and the complete deployed system.

Part 129. What makes a good benchmark?

A benchmark should represent the intended capability, contain sufficiently discriminating cases and have a scoring method connected to real success.

Coverage

Does the test include the important subskills and conditions? A writing benchmark containing only short factual answers cannot establish long-form editorial capability.

Difficulty

If nearly every system receives a perfect score, the benchmark can no longer distinguish frontier performance. New tests need harder or different cases.

Validity

Does success on the benchmark correspond to the capability people actually care about? A proxy can become detached from the real objective.

Reproducibility

Can another evaluator run the same procedure and obtain a comparable result? Hidden prompts, changing tools and undocumented retries make comparisons difficult.

Resistance to gaming

A good test should be difficult to pass through superficial shortcuts unrelated to the intended skill.

Part 130. Benchmark contamination and the problem of familiar tests

Modern models can be trained on enormous collections of public material. If benchmark questions or close variants appear in training data, performance may partly reflect familiarity rather than generalisation to unseen problems.

This does not make the score meaningless automatically. Humans also benefit from familiarity. It does change what the score proves.

Use fresh or private evaluation sets

For important deployment decisions, create cases that were not publicly available during model development. Internal historical examples can help if privacy and permission are handled correctly.

Generate structure, not answers

When creating synthetic tests, define the underlying rule and generate new instances whose answers can be independently computed. This reduces reliance on remembered public questions.

Keep a final holdout

Once a test set influences prompt tuning or workflow design, it becomes development data. Preserve a genuinely untouched final evaluation where possible.

Part 131. Human baselines: which human?

Claims of human-level or superhuman performance often hide the comparator. Human performance varies enormously with expertise, language, incentives and time.

Novice baseline

Useful when the application aims to provide broad public assistance, but weak evidence for a claim about professional expertise.

Competent-practitioner baseline

Useful for operational substitution or augmentation questions. Define experience and tools clearly.

Expert baseline

Relevant when claiming frontier capability in a specialised field. Experts are harder and more expensive to recruit, which makes strong comparisons more difficult but more informative.

Best-human baseline

The strongest historical definitions of superintelligence invoke performance beyond the best human minds, not merely the average person. Evidence for such a claim must therefore engage with top-level human performance across many domains.

Part 132. Resource-normalised comparison

A model may answer in seconds using large computational infrastructure. A person may use years of training plus ordinary professional tools. There is no single inherently correct way to normalise these resources.

The comparison should match the decision. If an employer is deciding how to complete a workflow, compare the cost, time and quality of realistic alternatives. If a researcher is studying raw intellectual capability, the resource accounting may be different.

Retries matter

A system allowed one hundred attempts may produce an exceptional best answer while remaining unreliable per attempt. Report the number of attempts and selection procedure.

Tools matter

If the model uses search, code and databases, state that. If the human comparator is denied normal tools, explain why.

Parallel copies matter

Digital systems can be duplicated and run concurrently. That scalability is a genuine practical advantage even when it complicates one-to-one intelligence comparisons.

Part 133. Reliability: intelligence that works only sometimes

A system can have high peak capability and low reliability. In creative brainstorming, that may be acceptable because users select the best candidate. In safety-critical operations, it may be unacceptable.

Measure distributions, not anecdotes

One brilliant answer and one disastrous answer do not tell you the frequency of either. Run enough representative cases to estimate the distribution of outcomes.

Measure severe failures separately

An average score can hide rare catastrophic errors. Track failure classes whose consequences justify a separate threshold.

Measure consistency under harmless variation

Change wording, order or formatting without changing the task. Large swings in answer quality reveal brittleness relevant to deployment.

Part 134. Calibration: does uncertainty track reality?

A useful system should behave differently when evidence is strong and when it is weak. Calibration concerns whether confidence corresponds meaningfully to correctness.

For deployment, behavioural calibration may matter even when the system does not expose a numeric probability. Does it ask for missing information? Does it cite uncertainty? Does it decline to invent a current value after retrieval fails?

Confidence language can be misleading

Generated phrases such as “I am certain” are not automatically calibrated probabilities. Evaluate whether the system’s uncertainty behaviour predicts actual error.

Part 135. Latency, cost and quality form a triangle

A system can often spend more computation to search, reason, verify or compare alternatives. More computation may improve some tasks while increasing cost and response time.

The optimal point depends on the workflow. A customer-service draft may need seconds. A complex engineering analysis may justify minutes or hours if the additional checking materially improves reliability.

Do not optimise latency alone

Fast wrong answers can increase total process time through correction and downstream failure.

Do not optimise quality without regard to economics

A marginal improvement that multiplies cost may be inappropriate for high-volume low-consequence tasks.

Route tasks by difficulty

A mature SI system can use inexpensive methods for routine cases and escalate difficult or consequential cases to more computation or human review.

Part 136. Long-horizon tasks: intelligence across time

Many benchmarks test answers that can be produced in minutes. Real work can unfold across days or weeks, with changing information, interruptions and dependencies.

State preservation

Can the system retain the decisions and constraints that still matter without confusing old and current states?

Plan revision

Can it update a plan when a dependency changes rather than blindly following the original sequence?

Progress detection

Can it distinguish new evidence from repeated information and recognise when no progress is being made?

Handover

Can another person or system understand what has been completed, what remains open and which assumptions control the next step?

These are increasingly important SI capabilities because agents and persistent workspaces extend interaction beyond one response.

Part 137. How to evaluate an agent

An agent must be evaluated on more than its final answer because intermediate actions can have consequences.

Goal completion

Did the intended outcome occur?

Path quality

Did the agent use an efficient and authorised route, or did it perform unnecessary searches and actions?

Permission compliance

Did it remain within read/write and approval boundaries?

Recovery

Did it handle timeouts, unavailable tools and contradictory information appropriately?

Observability

Can reviewers reconstruct what the agent did from logs and resulting state?

Stopping

Did it stop when blocked or continue generating cost and risk without meaningful progress?

Part 138. Whole-system evaluation

A deployed SI system may contain a model, prompt, retrieval index, memory store, tool router, permissions, interface and human review. Each component can be correct while the integration fails.

Whole-system tests should begin with realistic user requests and end with verified outcomes. They should include normal cases, boundary cases and failures in upstream services.

Test the seams

Many production failures occur where components meet: retrieved text enters the prompt incorrectly, a date loses its timezone, an identifier is reformatted, or a tool result is interpreted with the wrong schema.

Interface contracts deserve the same attention as model intelligence.

Part 139. An evidence ladder for increasingly strong SI claims

ClaimEvidence that begins to justify it
Useful for this taskRepresentative task evaluation with reviewed outputs.
Reliable for this workflowRepeated scenario testing, exceptions, failure recovery and deployment monitoring.
Better than our current processControlled comparison of quality, time, cost and failure consequences.
Professional-level in a domainComparison with appropriately qualified practitioners across representative domain tasks.
Broadly generalStrong transfer across many unfamiliar domains and task types under transparent conditions.
Beyond top human capabilityRobust comparisons against leading human expertise across the relevant breadth, with resources and reliability specified.
Technical superintelligenceA coherent definition plus broad, reliable evidence satisfying that definition; a terminology decision alone is insufficient.

The ladder does not define one official standard. It demonstrates a principle: evidence should scale with the breadth and strength of the claim.

Part 140. Ten traps in measuring SI

Trap 1: benchmark equals reality. A benchmark samples capability; deployment contains different distributions and consequences.

Trap 2: average hides tails. Rare severe failures may matter more than a small average gain.

Trap 3: best-of-many equals typical. Selection among many attempts can exaggerate per-attempt reliability.

Trap 4: human baseline is undefined. Average users and top specialists are different comparators.

Trap 5: tools are hidden. Tool-assisted performance should be reported as such.

Trap 6: test leakage. Familiar benchmark material can inflate apparent generalisation.

Trap 7: prompt tuning consumes the test. Repeatedly adapting to the evaluation turns it into development data.

Trap 8: speed substitutes for correctness. Faster output can create slower end-to-end work if review and repair increase.

Trap 9: fluency substitutes for validity. Human evaluators can overrate persuasive language.

Trap 10: model score substitutes for system safety. A high-performing model can still be deployed with excessive permissions or poor recovery.

Part 141. What measurement tells us about the AI-to-SI transition

The history of intelligent systems is partly a history of tests becoming obsolete. Once a capability becomes routine, researchers design harder tests. This can create the impression that progress never arrives because the frontier keeps moving.

At the same time, passing more tests does not automatically prove the strongest philosophical claim. The correct conclusion lies between dismissal and hype: contemporary systems demonstrate capabilities that earlier generations did not, while broad claims about general or superhuman intelligence still require carefully defined evidence.

This is why the article’s central distinction remains stable after 141 parts. The 2026 terminology transition is a historical fact within its stated context. The capability transition is measured through accumulating evidence across tasks, systems and time.

Secondary students developing human judgement alongside Super Intelligence

Part 142. The real unit of change is the human–SI system

Many debates ask whether SI is better than a person. Real work is often organised differently: a person uses SI, software, documents and other people together. The practical unit of performance is therefore the complete human–machine system.

This framing changes the design goal. We do not need the model to imitate every human skill internally if the combined workflow produces a better verified outcome. We also do not need to automate a human strength merely because automation is technically possible.

Allocate by comparative advantage

Machines are strong at rapid transformation, search across large digital collections, repeated formatting and scalable generation. Humans can contribute legitimate authority, lived context, relationship judgement, value selection and responsibility. The exact boundary varies by domain.

Design the handoff

A weak workflow says “AI drafts, human checks”. A strong workflow specifies what the reviewer checks, which evidence is shown, what happens after rejection and which changes require renewed approval.

Part 143. Cognitive offloading: what should we stop remembering?

Humans have always offloaded cognition. Writing externalises memory. Maps externalise spatial representation. Calculators externalise arithmetic. Search engines externalise retrieval. SI externalises additional work such as summarisation, drafting and comparison.

Offloading is not automatically harmful. It can free attention for higher-level work. The danger appears when a person offloads a capability they still need in order to detect failure.

Keep enough internal knowledge to supervise

A pilot does not need to calculate every navigation value manually during ordinary operation, but needs enough system understanding to recognise abnormal conditions. Similarly, an SI-assisted analyst should understand the domain well enough to notice impossible assumptions and suspicious outputs.

Offload storage before judgement

It is often safer to let SI remember where information lives than to let it decide what the information means without review. Retrieval support can reduce cognitive burden while preserving human interpretation.

Part 144. Deskilling and upskilling can happen at the same time

When technology automates a routine skill, practice of that skill can decline. At the same time, users may develop new skills in system design, evaluation and higher-level problem solving.

The important question is which lower-level capabilities remain prerequisites for supervising the higher-level system.

Arithmetic example

Calculators reduce the need for manual arithmetic in many professional tasks. Number sense remains valuable because it helps a person notice when a result is off by a factor of ten.

Writing example

SI can draft prose quickly. Writers still need enough language mastery to recognise ambiguity, unsupported claims and weak structure.

Coding example

Generated code can reduce time spent recalling syntax. Engineers still need architecture, debugging and security knowledge to judge whether the code belongs in production.

A good SI curriculum therefore protects foundational diagnostic skills while allowing routine production to become more automated.

Part 145. Trust calibration: neither blind trust nor permanent suspicion

A system that is usually correct can tempt users into automation bias: accepting output because it came from the system. A system that occasionally fails can produce the opposite reaction: rejecting useful assistance because it is not perfect.

Calibrated trust means relying on the system in proportion to demonstrated capability for the current task.

Trust should be local

Do not ask whether you trust SI in general. Ask whether you trust this configuration to perform this task under these conditions, and what check remains appropriate.

Trust should update

Repeated success on representative cases can justify lighter review. A model or workflow change may require renewed evaluation. Trust is a maintained state, not a permanent badge.

Interfaces influence trust

Confident prose, polished design and anthropomorphic language can increase perceived competence. Interfaces should make sources, uncertainty and action state visible so trust can follow evidence rather than presentation.

Part 146. Automation bias and the disappearing reviewer

A human approval step can exist on paper while becoming meaningless in practice. If reviewers approve hundreds of machine outputs with little variation, attention declines and the review becomes ceremonial.

Reduce review volume intelligently

Route routine high-confidence cases through automated checks and concentrate human attention on exceptions, disagreement and high-consequence outputs where policy permits.

Make differences visible

For document revision, show what changed. For a recommendation, show the evidence and assumptions. For a tool action, show the exact target and parameters. Review quality improves when the interface directs attention to the consequential parts.

Sample automated cases

Even when routine cases are automated, inspect a sample to detect drift. A system can degrade gradually while every individual output still looks plausible.

Part 147. Decision rights: who is allowed to decide what?

SI makes it technically easy to move from recommendation to action. Organisational legitimacy does not move automatically with technical capability.

A useful workflow maps decision rights explicitly. The system may gather evidence. A manager may approve spending. A qualified professional may make a regulated decision. A customer may consent to a change affecting their account.

Decision rights are not just permissions

Software permission asks whether an account can perform an action. Decision rights ask whether that actor is legitimately authorised to choose the action. A technically permitted operation can still violate organisational policy.

Escalation preserves legitimacy

When the system reaches a decision outside its delegated scope, it should route the issue to the correct owner with the relevant evidence rather than improvising authority.

Part 148. Designing a human–SI team

Think of the system as a team with roles rather than one omniscient assistant.

Researcher role

Retrieves evidence and records provenance.

Analyst role

Structures comparisons, calculations and scenarios.

Critic role

Searches for contradictions, missing evidence and failure modes.

Operator role

Uses approved tools to execute bounded actions.

Human owner

Defines objectives, resolves value conflicts and carries responsibility for consequential decisions.

These roles can be implemented by one model, several models, deterministic tools and people. The role separation matters because it creates checks and clearer authority.

Part 149. Why critique should be partly independent

If the same process generates and approves its own work, shared blind spots can survive. Independence increases the chance that a different method catches the error.

For arithmetic, use deterministic calculation. For a factual claim, inspect the source. For code, execute tests. For a high-consequence professional decision, use appropriately qualified independent review.

Second-model review has limits

Another model can identify omissions and inconsistencies, but models trained similarly may share failure patterns. Treat model critique as one layer, not universal validation.

Part 150. What happens to expertise when answers become cheap?

Expertise has never been only access to facts. Experts recognise which facts matter, detect unusual cases, understand causal structure and know when ordinary rules fail.

SI makes factual retrieval and standard explanation cheaper. This can reduce the premium on memorising information that is easy to retrieve while increasing the value of judgement about evidence and exceptions.

Experts gain leverage

A domain expert can use SI to explore more alternatives, read more material and automate routine documentation. Because the expert can detect errors, the combination can outperform either component alone.

Novices gain access but face verification limits

A novice can ask sophisticated questions and receive useful explanations. The same novice may lack the knowledge needed to recognise a subtle error. This creates an asymmetry: SI can raise novice output faster than novice judgement.

Education must close the judgement gap

Students need opportunities to build mental models, not merely obtain finished outputs. Otherwise the system can make performance look advanced while understanding remains fragile.

Part 151. Learning transfer is the test of SI-assisted education

A learner has not mastered a method merely because they can reproduce an assisted example. Transfer asks whether the underlying idea works in a new context.

Near transfer

Change numbers or surface details while preserving the structure.

Far transfer

Place the same principle inside a different-looking problem where the learner must recognise when it applies.

Delayed transfer

Return after time has passed. If the learner can reconstruct the method without the original assistance, the learning is more durable.

SI is unusually good at generating transfer tasks quickly. That capability should be used deliberately rather than spending all of its power on producing answers.

Part 152. Metacognition: knowing what you know in the SI era

SI can make weak understanding feel fluent because the interface fills gaps instantly. Metacognition—the ability to judge one’s own knowledge—therefore becomes more important.

Prediction before assistance

Before asking SI, predict the answer or method. The gap between prediction and assisted result reveals learning.

Confidence before checking

Record how confident you are, then verify. Over time, you learn where your own judgement is well calibrated.

Explain without the tool

After an assisted session, close the source and reconstruct the core idea. What disappears was probably not yet internalised.

Part 153. Organisational memory: SI can remember the wrong thing very efficiently

Organisations accumulate policies, meeting notes, procedures, decisions and informal knowledge. SI can make this memory searchable, but only if the underlying records distinguish current authority from historical context.

Versioning is intelligence infrastructure

Mark effective dates, owners and superseded documents. A perfect retrieval system cannot infer which of two contradictory policies is current if the organisation never recorded that relationship.

Decision logs preserve why

A final policy may state what was chosen without explaining why. Decision logs capture alternatives, assumptions and conditions that may matter when circumstances change.

Forget deliberately

Not every historical detail should remain active forever. Retention policies, privacy requirements and stale operational information require deliberate expiry or archival treatment.

Part 154. Institutional learning: turn every failure into better future behaviour

An organisation becomes more intelligent when mistakes improve the system rather than merely produce blame.

Capture the incident

Record the input, system state, output, consequence and detection method.

Classify the failure

Use the taxonomy from Part 93: task, evidence, extraction, reasoning, generation, tool, permission, execution, review, governance or objective.

Repair the system

Change the source, interface, permission, test or process that allowed the failure.

Add a regression case

The same failure class should become easier to detect next time.

This feedback loop is a practical form of collective intelligence. The organisation’s capability increases because experience changes the system.

Part 155. Human–SI maturity ladder

StageHuman behaviourSystem relationship
1. ConsumerAccepts answers largely at face value.SI is an answer machine.
2. OperatorUses instructions, files and revisions deliberately.SI is a flexible tool.
3. VerifierChecks claims, calculations, sources and actions.SI is an assistant whose work is inspected.
4. DesignerBuilds repeatable workflows, evaluations and boundaries.SI is a component in a system.
5. ArchitectCoordinates people, models, data and governance around outcomes.SI becomes organisational infrastructure.

The ladder is not a ranking of human worth. It describes increasing sophistication in the use of intelligent systems. A mature architect may still choose a simple Level-2 workflow when that is the safest and most economical design.

Part 156. The human–SI contract

A useful working relationship can be summarised as a contract with six clauses.

Human supplies purpose. The system should not silently invent the ultimate objective.

Evidence supplies reality. Important factual claims should remain connected to appropriate sources or observations.

SI supplies leverage. Use speed, scale, transformation and search where they genuinely improve the task.

Tools supply precision and action. Delegate formal calculation and external operations to appropriate systems with bounded permissions.

Verification supplies confidence. Match checks to consequence rather than trusting presentation.

Governance supplies legitimacy. People and institutions retain responsibility for who may decide and act.

This contract explains why human capability remains central after AI becomes SI. The machine’s expansion changes the division of labour; it does not abolish purpose, evidence or responsibility.

Part 157. The human transition is the other half of the AI-to-SI story

The technological story is easy to see: bigger models, better interfaces, retrieval, tools and agents. The human story is quieter but equally important. People must learn to frame tasks, maintain domain knowledge, calibrate trust, supervise automation and design organisations that learn from failure.

If machine capability rises while human judgement deteriorates, the combined system may become less dependable despite more impressive outputs. If human skills evolve alongside the technology, SI can amplify expertise, broaden access and reduce routine cognitive load.

So the transition from AI to SI is not complete merely when the machine becomes more capable or the terminology changes. It becomes operationally meaningful when the surrounding human system learns how to use that capability without losing the ability to understand, verify and govern it.

Secondary students exploring the engineering frontier of Super Intelligence

Part 158. The engineering frontier behind Super Intelligence

Public discussion often treats model capability as though it comes from one variable: size. In practice, frontier systems are products of interacting choices about data, architecture, optimisation, hardware, post-training, inference-time computation, retrieval and product design.

This matters for the article’s historical question because the AI-to-SI transition did not happen simply by making one neural network larger. It emerged from an engineering stack in which improvements at one layer made improvements at another layer useful.

Part 159. Scaling laws: predictable trends are not guarantees

Researchers have observed regular relationships between model performance, model size, data and computation in particular experimental regimes. Scaling Laws for Neural Language Models reported empirical power-law relationships for language-model loss across model size, dataset size and training compute in the regimes studied. Scaling-law research is valuable because it can help teams reason about resource allocation, but its fitted relationships remain empirical results rather than guarantees about every architecture or future capability.

A scaling relationship is empirical. It describes behaviour within studied conditions. It does not prove that every capability improves smoothly forever or that a particular future threshold must be reached.

Scale can reveal capabilities

A larger training run can improve generalisation enough that tasks previously unreliable become useful. This can feel discontinuous to users even when the underlying performance curve changes gradually.

Scale can also preserve weaknesses

A larger model can remain vulnerable to bad evidence, ambiguous objectives and inappropriate tool permissions. Scale improves a component; it does not automatically repair the complete system.

Part 160. Compute, data and model size must be balanced

Training resources can be spent on more parameters, more data, longer training or different mixtures of examples. Research such as the Chinchilla work examined how to allocate a fixed compute budget more effectively between model size and training tokens.

The general lesson is more durable than one formula: a system can be undertrained for its size, data-limited or compute-limited. “Bigger model” is not a complete optimisation strategy.

Quality changes the meaning of quantity

Two datasets containing the same number of tokens can differ greatly in duplication, factual quality, diversity and relevance. Scaling data without curation can scale noise.

Part 161. Synthetic data: models increasingly help create training material

Once models can generate useful examples, they can help produce data for training, evaluation or specialised adaptation. Synthetic data can create rare cases, balance categories and provide examples whose underlying rule is known.

The advantage

Developers can generate targeted practice at scales that would be expensive to label manually.

The risk

If generated data contains systematic errors or lacks real-world diversity, training on it can reinforce those weaknesses. Synthetic data should be validated against the reality the system is meant to model.

Use generators and judges carefully

A model can generate candidate examples and another process can filter them, but shared model biases can survive. Independent rules, real data and human inspection remain useful anchors.

Part 162. Distillation: transferring capability into smaller systems

Not every deployment needs the largest available model. Distillation and related methods aim to transfer useful behaviour from a larger or more expensive system into a smaller model.

This matters economically. A smaller model may run faster, cost less and operate closer to the user while preserving enough capability for a bounded task.

Compression changes the evaluation question

The right comparison is not whether the smaller model is identical to the larger one. It is whether it preserves the capabilities and reliability required for the intended workflow.

Part 163. Quantisation and efficient inference

Model parameters and computations can sometimes be represented with lower numerical precision, reducing memory and computational requirements. Quantisation can make models easier to deploy on constrained hardware.

Efficiency improvements matter because SI adoption depends not only on frontier capability but on the cost of serving that capability repeatedly.

Efficiency can expand access

If a useful model can run on cheaper hardware, more organisations and devices can use it without sending every request to a large remote data centre.

Efficiency still needs testing

Compression can affect different tasks differently. Evaluate the compressed configuration rather than assuming benchmark parity from the original model.

Part 164. Edge and local SI

Some intelligent processing can occur on a user’s device or within an organisation’s own infrastructure. Local execution can reduce latency, preserve operation during network disruption and change privacy architecture.

Local does not automatically mean private

Privacy depends on the complete data flow, logging, application permissions and storage. A local model inside an application that uploads results elsewhere is not a purely local workflow.

Local models create resilience options

Critical workflows may benefit from degraded local capability when cloud services are unavailable. This resembles other infrastructure design: redundancy can matter more than maximum peak performance.

Part 165. Cloud SI and concentrated infrastructure

Large models can require infrastructure too expensive for most organisations to build independently. Cloud access turns that infrastructure into a service.

This creates a useful separation: capability can be centralised physically while access is distributed through networks and APIs.

Concentration creates dependencies

If many workflows rely on a small number of model or cloud providers, outages, policy changes and supply constraints can propagate widely. Organisations should understand which critical processes depend on external infrastructure.

Portability becomes strategic

Where feasible, keep task definitions, evaluation sets and data interfaces separate from one model provider. This makes it easier to compare or migrate systems when capability, cost or policy changes.

Part 166. Open and closed SI ecosystems

Intelligent systems can be distributed under different access models. Some expose model weights or extensive technical details; others provide capability primarily through hosted services.

These choices affect transparency, customisation, security responsibility, deployment cost and who can inspect or modify the system.

Open access increases experimentation

Researchers and organisations can study, adapt and run systems in their own environments where licences and resources permit.

Hosted systems can centralise maintenance

A service provider can update models, infrastructure and safeguards without every user operating the full stack.

No access model is universally best

The relevant question is which properties the workflow needs: control, transparency, privacy, support, frontier capability, cost or local operation.

Part 167. Inference-time computation: spending more effort after training

Training determines a model’s parameters, but a system can also spend varying amounts of computation when answering a particular request. It may search alternatives, use tools, verify intermediate results or run longer reasoning procedures.

This creates another axis of capability. Two systems using similar underlying models can perform differently because one allocates more computation or better tools to difficult tasks.

Dynamic effort is economically useful

Routine questions can receive fast inexpensive processing. Difficult or consequential questions can receive more search, checking and computation.

More computation is not automatically better

A system can waste effort exploring irrelevant paths. Evaluation should measure improvement per unit of time and cost.

Part 168. Reasoning systems increasingly combine generation with verification

A useful pattern is proposal followed by check. The model proposes a plan, proof step, query or program; a specialised process evaluates it; the result informs the next proposal.

This architecture resembles human problem solving: generate possibilities, test them against reality, keep what survives.

Formal domains benefit strongly

Code can be executed. Equations can be checked. Constraints can be solved. These external validators provide feedback stronger than linguistic self-confidence.

Open-world domains remain harder

For strategy, ethics or social judgement, there may be no deterministic validator. Evidence, deliberation and legitimate human decision-making remain central.

Part 169. Model routing: one intelligence layer does not need one model

A system can route different tasks to different models. A small model handles classification, a larger model handles complex synthesis, a specialised model handles images and deterministic tools handle calculation.

Routing can improve economics

Most requests may not need the most expensive model. Routing reserves scarce computation for tasks where it changes the outcome.

Routing adds another failure point

The system must recognise task difficulty correctly. A subtle high-consequence request misrouted to a weak model can fail before the model ever sees the full problem.

Part 170. Specialisation inside large models

Some architectures route inputs through different internal components or experts. The broader engineering idea is familiar: not every part of a system needs to process every task identically.

Specialisation can increase computational efficiency and allow different internal components to develop different strengths. It also makes system analysis more complex because behaviour depends on routing as well as the components themselves.

Part 171. Data centres are part of the intelligence system

Frontier SI depends on physical facilities that provide power, cooling, networking and hardware at scale. The data centre is therefore not merely where intelligence happens to run; it is part of the capability stack.

Power availability can constrain expansion

Large computational facilities require substantial electrical infrastructure. Deployment timelines can depend on grids, generation and permitting as much as model research.

Cooling and water can matter locally

Cooling architecture depends on climate and facility design. Resource use should be evaluated in the context of specific infrastructure rather than inferred from one universal figure.

Hardware supply chains matter

Advanced chips require complex global manufacturing and packaging ecosystems. Intelligence scaling therefore interacts with industrial capacity and geopolitics.

Part 172. Energy efficiency is a capability multiplier

If the same useful computation requires less energy, more intelligence can be delivered within the same infrastructure envelope. Efficiency gains can therefore matter as much as building additional capacity.

Hardware, algorithms, model architecture, quantisation, caching and workload scheduling all influence energy per useful task.

Measure useful work, not only raw consumption

A more capable system may consume more energy per request while replacing several other processes. Environmental assessment should compare complete workflows where possible.

Part 173. Hardware–software co-design

Algorithms are shaped by the hardware they run on, and hardware is increasingly designed around machine-learning workloads. This feedback loop can accelerate progress.

A new numerical format may reduce memory use. A new chip interconnect may allow larger distributed training. A model architecture may be redesigned to exploit those capabilities.

The frontier is therefore a co-evolution of mathematics, software and physical engineering.

Part 174. The bottleneck keeps moving

When one constraint is relaxed, another becomes visible. More compute can expose data limitations. More context can expose retrieval quality problems. Better generation can expose review bottlenecks. More autonomy can expose governance weaknesses.

This moving-bottleneck pattern is one of the deepest lessons of the AI-to-SI transition. Progress does not remove constraints; it relocates them.

Ask what becomes scarce next

If drafts become abundant, verification becomes scarce. If models become cheap, trusted data may become scarce. If agents can execute quickly, legitimate decision rights may become the scarce resource.

Part 175. Do engineering trends prove future technical superintelligence?

No single engineering trend proves a future outcome. Scaling, synthetic data, inference-time computation, better hardware and tool integration can all expand capability. Their future interaction remains an empirical question.

Some constraints may yield to engineering. Others may require conceptual breakthroughs. Physical infrastructure, economics, regulation and social choices can affect deployment even when algorithms improve.

Separate trajectory from destination

Evidence of rapid progress supports the claim that capability is changing quickly. It does not, by itself, establish that a particular technical definition of superintelligence will be reached on a specific date.

Use milestones rather than prophecy

Track observable thresholds: reliable long-horizon work, transfer to unfamiliar domains, autonomous scientific contribution, robust self-correction and performance against top specialists. These measurements tell us more than an unsupported countdown.

Part 176. The frontier thesis: SI advances by moving the constraint

The engineering history can be read as a sequence of constraints becoming less binding. Hand-coded rules gave way to learning from data. Limited representations gave way to deep learning. Sequential architectures gave way to more scalable attention-based models. Static knowledge gained retrieval. Isolated generation gained tools. One-shot responses gained agent loops. Cloud-only assumptions gained local alternatives.

Every release creates a new limiting factor. That is why the next stage of SI cannot be understood by model size alone. Intelligence is becoming a system property shaped by computation, data, software, physical infrastructure, economics and governance.

This reinforces the article’s central answer. AI did not become SI because one number crossed a line. The object itself became more integrated, scalable and operational over decades—until the language used to describe it began to change as well.

Secondary students examining evidence, uncertainty and truth in Super Intelligence

Part 177. Super Intelligence and the problem of knowing

As SI becomes more fluent, the hardest question is often no longer whether it can produce an answer. It is what relationship that answer has to reality.

A system can generate a statement from learned patterns, retrieve a statement from a document, calculate a result from supplied numbers, infer a conclusion from several facts or report the outcome of an external action. These are different epistemic routes and should not be collapsed into the single word “knows”.

Ask how the answer was obtained

For an important claim, identify whether the system is recalling learned patterns, using supplied context, retrieving a source, executing a tool or making an inference. The route determines the appropriate check.

Part 178. “Knows” versus “predicts”

A language model generates outputs through learned statistical relationships. In ordinary conversation we may say the model knows a fact, but operational work benefits from more precise language.

If a current fact matters, ask for current evidence. If an exact calculation matters, calculate it. If the system is forecasting an uncertain outcome, label it as a prediction rather than a fact.

Useful shorthand should not become false certainty

Human language routinely uses mental vocabulary for machines. That can be convenient. Problems arise when the metaphor replaces verification.

Part 179. Retrieved information is not automatically believed information

Retrieval finds material relevant to a query. Relevance does not establish truth, authority or applicability.

A search system may retrieve an old policy because it uses the exact words in the question. A newer policy may use different terminology. The SI system therefore needs metadata and reasoning about authority, not similarity alone.

Retrieval creates a candidate evidence set

After retrieval, evaluate source status, date, jurisdiction, scope and internal consistency. Retrieval reduces the search space; it does not complete the epistemic task.

Part 180. Inference: where facts become conclusions

An inference connects evidence to a conclusion that may not be stated directly in any source. Inference is unavoidable in useful intelligence. The goal is not to eliminate it but to make important inferences visible.

If the inventory says a device is expected to return Thursday and policy requires inspection after return, we can infer that Friday availability is conditional. We cannot infer that inspection will definitely pass.

Label assumptions

An assumption can be reasonable and still be an assumption. Naming it makes later revision easier when new evidence arrives.

Part 181. Not all uncertainty is the same

Uncertainty can come from several sources, and each needs a different response.

Missing-information uncertainty

A required fact has not been supplied. Repair by obtaining the fact.

Measurement uncertainty

The observation itself has error or limited precision. Repair through better measurement or by carrying the uncertainty into the result.

Model uncertainty

Several explanations fit the available evidence. Repair through additional evidence or more appropriate modelling.

Future uncertainty

The outcome depends on events that have not occurred. Use scenarios, probabilities or ranges rather than false certainty.

Value uncertainty

The decision depends on preferences or priorities that have not been settled. The repair is deliberation, not more prediction.

Part 182. Causality: prediction is not intervention

A model can discover that two variables move together and use one to predict the other. That does not prove that changing the first will change the second.

For example, a system may learn that students who complete more practice questions tend to score higher. This could reflect practice effects, prior motivation, teacher support or several factors together. A recommendation to increase practice is a causal proposal that needs more reasoning than the predictive association alone.

Ask the intervention question

“What happens if we change X?” is different from “What tends to occur when X is high?” SI should preserve that distinction in analysis.

Part 183. Counterfactual reasoning

Decisions often ask what would have happened under another choice. That alternative world is not directly observed.

SI can help construct counterfactual scenarios and identify assumptions, but the strength of the conclusion depends on causal evidence and model quality.

Counterfactuals are useful even when uncertain

A project review can ask what signals were available before failure and whether another response was feasible. The purpose may be organisational learning rather than proving one unique alternate history.

Part 184. Objectives contain values

Optimisation begins with an objective, but objectives are not discovered by mathematics alone. Someone decides what counts as success.

A school could optimise examination scores, student wellbeing, long-term learning or some combination. A transport system could optimise average travel time, reliability, accessibility or emissions. Different objectives produce different “best” solutions.

Expose the objective before optimisation

Ask who selected the metric, which stakeholders are affected and what important value is not represented.

Part 185. Real systems are multi-objective

Organisations rarely care about one number. They want speed and accuracy, low cost and resilience, personalisation and privacy.

These goals can conflict. SI can help map the trade-off surface, but legitimate decision-makers still choose which compromises are acceptable.

Constraints can protect non-negotiables

Instead of assigning a small penalty to a prohibited disclosure, make privacy a hard constraint where the system architecture permits it. Not every value should be traded away for a better average score.

Part 186. Explainability: what kind of explanation do you need?

Explainability is not one thing. Different users need different explanations.

Outcome explanation

Why was this recommendation made?

Evidence explanation

Which sources or observations support it?

Process explanation

Which tools and transformations were used?

Model explanation

Which features or internal mechanisms influenced a prediction?

Policy explanation

Which organisational rule authorised the action?

Choose the explanation type that supports the decision rather than demanding a generic story about “how the AI thought”.

Part 187. A plausible explanation may not be a faithful explanation

A model can generate a coherent rationale after producing an answer. That rationale may help communicate the answer without necessarily revealing the complete internal causal process that produced it.

For operational work, prefer inspectable evidence: source passages, calculations, tool logs and decision rules. These can be verified independently.

Part 188. Provenance: where did this information come from?

Provenance records the origin and transformation history of information. In SI systems, provenance can include the source document, retrieval time, extraction step, model transformation and final publication.

Provenance enables correction

If a source is later found to be wrong, provenance helps identify which outputs depended on it.

Provenance supports trust without requiring blind trust

A reader can inspect the evidence path rather than relying solely on the reputation of the model.

Part 189. Auditability: can we reconstruct what happened?

An auditable system leaves enough records for an authorised reviewer to understand consequential actions after the fact.

Useful records can include user request, system version, retrieved sources, tool calls, approvals and resulting external state. The exact logging must respect privacy and security requirements.

Audit logs are not useful if nobody can interpret them

Design logs around likely investigations. A million low-level events without identifiers or timestamps can be less useful than a concise structured action record.

Part 190. Reversibility changes how much autonomy is safe

An easily reversible action can tolerate more experimentation than an irreversible one. Drafting a private note is easier to undo than sending a public announcement. Simulating a database change is safer than deleting records.

Use staged commitment

Move from draft to preview to approval to execution where consequence warrants it. Each stage reduces uncertainty before commitment.

Design undo before automation

If a workflow cannot recover from a predictable mistake, increasing automation may increase fragility.

Part 191. Resilience: intelligence under failure

A resilient SI system continues providing acceptable service when components fail or conditions change.

Graceful degradation

If the preferred model is unavailable, a smaller system may handle routine cases while complex tasks wait.

Fallback to manual operation

Critical organisations may need procedures that function when SI services or networks are unavailable.

Redundant evidence

For consequential state, independent sensors or records can reduce reliance on one source.

Recovery time matters

Measure how quickly the workflow returns to safe operation, not only how rarely it fails.

Part 192. Drift: yesterday’s reliable system can become today’s weak system

Performance can change even when model parameters do not. User behaviour changes, documents change, products change and the distribution of incoming cases shifts.

Data drift

Inputs differ from those used in evaluation.

Concept drift

The relationship between inputs and correct outputs changes.

Policy drift

Organisational rules change while the system continues using old instructions.

Tool drift

External APIs or interfaces change.

Monitoring should therefore examine the environment around the model, not only the model itself.

Part 193. Every model update is a system change

A new model version may improve average performance and still break a workflow that depended on previous formatting or behaviour.

Regression testing before migration

Run the established evaluation set against the new configuration. Compare quality, latency, cost and failure categories.

Canary deployment

Where appropriate, expose a limited portion of work to the new system before full rollout.

Rollback

Keep a path back to the previous safe configuration when operationally feasible.

Part 194. Truth maintenance in a changing world

A knowledge system needs more than retrieval. It needs a way to handle updates and contradictions.

When a new policy supersedes an old one, the old document may remain historically valuable but should no longer control current decisions. When scientific evidence changes, prior summaries may need revision.

Separate current truth from historical record

Do not delete history merely because it is no longer current. Mark status and effective periods so the system can answer both “What is the rule now?” and “What was the rule then?”

Part 195. Give important statements an epistemic status

A simple status vocabulary can improve long-form SI work:

Observed: directly present in the evidence.

Calculated: derived by a reproducible operation.

Inferred: conclusion drawn from evidence.

Estimated: approximate value based on a model or assumptions.

Predicted: claim about a future or unknown outcome.

Proposed: suggested action or design.

Verified: checked through an appropriate independent method.

This vocabulary reduces the tendency of generated prose to flatten every statement into the same confident voice.

Part 196. The epistemic checklist for SI answers

Before relying on a consequential answer, ask:

What exactly is the claim? What evidence supports it? Is the evidence current and authoritative? Which parts are inference rather than observation? Which assumptions matter? What uncertainty remains? What would falsify the conclusion? Has the important calculation or action been independently checked? Who has authority to decide what happens next?

These questions are simple enough to use every day and strong enough to prevent many classes of SI failure.

Part 197. The epistemology thesis: intelligence is not the same as truth

A system can be highly capable at generating, searching, calculating and planning while still needing an evidence architecture to produce dependable knowledge.

This is not a weakness unique to machines. Human intelligence also operates under uncertainty, bias and incomplete information. The difference is scale: SI can produce and propagate conclusions far faster, so epistemic discipline becomes infrastructure rather than an individual virtue.

The mature SI system therefore does not merely answer. It helps preserve the relationship among evidence, inference, uncertainty, authority and action.

Sources, dates and what the evidence establishes

A history needs more than a list of impressive papers. Each date must identify an event, and each source must support the particular claim attached to it. The register below separates original publication, first public preprint, a proposed research programme, a later conference version and an administrative terminology decision.

The technical chapters are grouped by ideas rather than arranged as an uninterrupted timeline. In particular, the 2021 foundation-model terminology appears after the 2017 Transformer paper in calendar time, even though this guide explains the broader category before discussing that architecture.

Dated primary-source register

1950: a published question about machine intelligence

A. M. Turing, Computing Machinery and Intelligence, appeared in the October 1950 issue of Mind, volume LIX, issue 236, pages 433–460. This is a publication date, not the invention date of every idea in the paper. Return to Part 7.

31 August 1955: the Dartmouth proposal

The original Dartmouth proposal bears this date and the names of McCarthy, Minsky, Rochester and Shannon. Its title already uses artificial intelligence. It proposes a study for the following summer; the proposal and the gathering are distinct events. Return to Part 8.

Summer 1956: the Dartmouth research project

Dartmouth’s institutional account records the summer gathering. Use it for the field-forming event, alongside the original proposal for the earlier documentary date. Neither establishes that all machine-intelligence research began there.

1998: the stronger superintelligence concept in a published treatment

Bostrom’s How Long Before Superintelligence? identifies its original journal publication as 1998; the page also carries a 1997 copyright and a 1998 postscript. These labels describe different stages of the document, not competing dates for the arrival of superintelligent machines. Return to Part 39.

2012: deep learning demonstrated on image classification

Krizhevsky, Sutskever and Hinton’s ImageNet Classification with Deep Convolutional Neural Networks supplies a concrete research milestone for the vision discussion. It is evidence about a particular model and evaluation, not the birth of neural networks. Return to Part 12.

1 September 2014: attention-based neural translation before the Transformer

Bahdanau, Cho and Bengio’s translation paper first appeared as a preprint on this date and was accepted at ICLR 2015. It provides an important predecessor when explaining why attention itself was not invented by the 2017 Transformer paper.

12 June 2017: the Transformer preprint

Attention Is All You Need first appeared on arXiv on this date. Later revisions on its record should not be mistaken for the original introduction of the architecture. Return to Part 14.

22 May 2020: a named retrieval-augmented generation method

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks combined a generative model with a retrievable document collection. It does not mark the invention of information retrieval or guarantee that every retrieved answer is current. Return to Part 18.

28 May 2020: large-scale few-shot language-model evaluation

Language Models are Few-Shot Learners reported GPT-3’s performance with instructions and examples supplied through text, without task-specific parameter updates in those evaluations. Its reported limitations matter alongside its successes. Return to Part 15.

26 February 2021: image–text learning and transfer

Learning Transferable Visual Models From Natural Language Supervision documents CLIP. It is a specific image–text learning method, not proof that all multimodal systems can reason reliably about every image or generate every modality. Return to Part 17.

16 August 2021: foundation models named as a broad category

On the Opportunities and Risks of Foundation Models introduces the terminology while explicitly building on earlier deep learning, pretraining and transfer learning. Naming the category did not create its entire technical ancestry. Return to Part 13.

4 March 2022: instruction-following with human feedback

Training language models to follow instructions with human feedback describes demonstrations, output comparisons and further training. Its results concern the studied setup; they do not turn preference ratings into a universal truth test. Return to Part 16.

29 March 2022: balancing training compute, model size and data

Training Compute-Optimal Large Language Models is the Chinchilla research anchor. Its experimental trade-offs should not be treated as an unconditional formula for every later architecture or deployment budget. Return to Part 160.

6 October 2022: the ReAct preprint

ReAct: Synergizing Reasoning and Acting in Language Models first appeared in 2022, with an ICLR 2023 conference version. These dates distinguish preprint and conference publication; neither is the invention date of agents in general. Return to Part 19.

23 January 2020: empirical scaling laws for neural language models

Scaling Laws for Neural Language Models reported power-law relationships between language-model loss and model size, dataset size and training compute in the experiments studied. These fitted relationships are empirical and should not be read as a proof that every capability scales indefinitely. Return to Part 159.

9 February 2023: Toolformer and learned API use

Toolformer: Language Models Can Teach Themselves to Use Tools studied a model trained to decide when and how to call several external APIs. It is one milestone in language-model tool use, not the invention of software tools or a guarantee of reliable action in arbitrary environments. Return to Part 68.

29 September 2026: the executive-branch terminology order

Inaugurating the Era of Super Intelligence is the direct source for the naming event. Its scope and operative definition are in sections 2 and 3. This administrative event is distinct from a measured technical capability threshold. Return to the opening answer.

Reading a milestone without exaggerating it

For each entry, ask what was actually proposed, built or measured. A published method may precede widespread adoption. An influential name may describe older work. An experiment can demonstrate an advance without proving a universal capability. The article’s wider interpretation of the transition is a synthesis, not a quotation from any single paper.

Return to the guide map · Read the narrative chronology · Choose your next guide

Where to go next in the Super Intelligence library

You have just read the transition layer: the long history from machine intelligence and Artificial Intelligence to the broader SI-era system of models, retrieval, tools, agents, infrastructure and human governance.

Continue with…When to use it
Super Intelligence — Complete Master GuideReturn to the apex SI map and explore the entire knowledge estate.
How Super Intelligence WorksGo deeper into model architecture, reasoning, retrieval, tools, agents and human control.
How to Learn Super Intelligence QuicklyTurn the concepts into a structured learning curriculum and practice system.
How to Leverage Your Life with Super IntelligenceApply SI to personal planning, learning, research and everyday decisions.
Build Super Intelligence WorkflowsMove from isolated prompts to repeatable workplace processes.

The one-sentence resolution

AI became SI in terminology through a specific naming transition, while the capability transition developed across decades; technical superintelligence remains a separate claim requiring its own evidence.

Return to the guide navigation