
Begin with a claim you can explain
Super intelligence is a phrase that can make a conversation feel settled before its most important questions have been asked. A striking answer becomes “intelligence”; a fast result becomes “superhuman”; an independent action becomes “agency”; and a collection of these impressions becomes a prediction about the future. The purpose of this guide is to slow that chain down enough to inspect it. Doing so makes genuine progress easier to recognise, because it gives each achievement a clear description.
This is the definition and artificial-superintelligence branch of the eduKate Super Intelligence library. It owns the conceptual question: what would a claim of broadly superhuman artificial capability mean, and how could a reader assess the evidence? The Super Intelligence master guide connects this subject with practical learning, mechanisms, everyday use and work. Here, the focus stays on definitions, comparisons, evidence, uncertainty and human judgement. The complete 100-position reading route follows the guide; linked entries have a published destination, while entries labelled planned remain reading topics rather than promised live pages.
You do not have to agree on an arrival date, choose a favourite technology company or become a researcher to use this framework. You need to be able to distinguish what was observed from what was inferred. You need to notice which comparison was made, what was left out and what would change your mind. Those are learnable habits. A student comparing two examination results, a parent hearing a claim about an AI tutor and a manager considering an automated process can all practise them.
The quickest useful outcome is a four-sentence assessment. First, state the specific achievement. Second, name the system and conditions that produced it. Third, state the narrow conclusion the evidence supports. Fourth, identify the most important unanswered question. If those four sentences are clear, the conversation is already more informative than an unsupported declaration that ASI has arrived or can never arrive.
Choose your reading route
Begin with the diagnostic, clarify definitions, assess evidence, inspect worked claim packets, or choose from the complete 100-guide route.
Define and compare
1. A short diagnostic before you read
2. What super intelligence means in this guide
3. Separate capability from autonomy, consciousness and authority
Evaluate evidence and future claims
5. What evidence would make an ASI claim credible
6. How to read impressive results without losing perspective
Practise with complete worked packets
9. Workshop one: A high exam score becomes an ASI headline
10. Workshop two: An autonomous workflow becomes an intelligence claim
11. Workshop three: A scientific breakthrough becomes a general intelligence claim
Human judgement and your next reading
13. Human roles do not disappear when a system becomes more capable
100-guide directory · Research sources · Super Intelligence master guide
1. A short diagnostic before you read
Consider the following six fictional statements. For each one, decide whether the statement supplies a definition, an observation, an inference, a forecast or a permission. Some contain more than one. Then read the explanation and identify what extra information would be needed. This is a reasoning exercise, not a test of whether you are optimistic or sceptical about AI.
Statement one: “It answered a difficult question that I could not answer”
This is an observation about one comparison, provided the answer was checked. It does not yet identify whether the person is a relevant expert, whether the question was new to the system or whether the answer survives scrutiny. The bounded conclusion is that the system supplied a correct answer beyond that person’s demonstrated ability on that occasion. Ask for the question, the verification method and the relevant specialist baseline before turning the result into a general intelligence claim.
Statement two: “It worked alone all night, so it must be more intelligent”
This combines an observation about duration and autonomy with an inference about capability. A simple scheduled process can run for hours. A sophisticated analysis may require only minutes. The relevant questions are what the system completed, whether it met the goal, what resources it used and whether its actions were authorised. Duration can matter in a reliability evaluation, but elapsed clock time cannot substitute for the difficulty or correctness of the work.
Statement three: “AGI means doing most economically valuable work better than humans”
This is a proposed definition. Before using it, ask whose definition it is and how “most”, “valuable”, “work” and “better” are measured. A definition can be useful while leaving operational questions unresolved. It is neither proof that a particular system meets the definition nor proof that economic performance captures every aspect of intelligence. The foundations below compare this kind of account with performance-and-generality frameworks.
Statement four: “The next model will become superintelligent”
This is a forecast. A reasoned version needs assumptions, an operational definition, a time horizon and evidence supporting the transition being predicted. The statement cannot be confirmed simply by pointing to improvement in an earlier model. Improvement might continue, change direction or encounter a bottleneck. Ask what observation would count against the forecast. A prediction that absorbs every possible outcome without revision is not giving the reader a useful test.
Statement five: “The system is very accurate, so it may choose our policy”
The first part is an empirical claim requiring an evaluation; the second grants or assumes authority. Accuracy can inform a decision about delegation, but does not independently determine who should set goals or bear consequences. A community might value privacy, fairness, participation or appeal rights alongside predictive accuracy. A careful decision names the responsible people, the permitted role and the process for contesting outcomes before acting on recommendations.
Statement six: “This error proves AI cannot ever become superintelligent”
The error may be real and important. It establishes a failure in the observed conditions. The permanent conclusion requires a much stronger argument about what future systems could or could not do. A fair critic preserves the error without claiming more than it shows. Ask whether the failure reveals a stable limitation, whether changed conditions repair it and whether the failed ability is central to the claim being assessed. Strong criticism is specific enough to guide a better test.
2. What super intelligence means in this guide
Artificial superintelligence, usually shortened to ASI, means a hypothetical artificial system with capabilities substantially beyond human experts across a broad range of important cognitive tasks. The claim concerns a combination of exceptional performance and generality: solving difficult problems, learning unfamiliar tasks, connecting knowledge and producing dependable results across domains. A machine being better than a person at one activity does not establish that combination.
In eduKate’s wider editorial ecosystem, SI is also a practical-AI umbrella for understanding and using increasingly capable tools. That publishing label is not a scientific classification of the tools discussed. A useful chatbot, an automated workflow or a well-designed AI study aid should be assessed on what it demonstrably does. Appearing in the SI series does not mean that it is artificial superintelligence.
This definition branch establishes the language needed to read stronger claims carefully. Its question is what ASI would mean and what evidence could support that description. Choosing everyday tools, designing prompts and building practical workflows belong to other branches. The concepts here help readers move between those subjects without accidentally turning usefulness into a claim of broad superhuman intelligence.
Why definitions differ
Intelligence is not a single observable substance that can be weighed. Definitions select abilities, environments and comparison standards. One account may emphasise achieving goals in varied circumstances; another may emphasise learning efficiently; another may emphasise the economic work a system can perform. Legg and Hutter’s research offers a formal approach to machine intelligence while examining alternative definitions and tests. It is one influential conceptual contribution, rather than a universally adopted certification scheme. Universal Intelligence: A Definition of Machine Intelligence.
These choices matter. Imagine two hypothetical systems. The first learns an unfamiliar board game from a short explanation and plays competently within minutes. The second cannot learn that game, but produces outstanding results on thousands of industrial tasks for which it was extensively prepared. A learning-centred definition highlights the first system’s adaptability. A productivity-centred definition highlights the second system’s output. Both observations could be true; the disagreement may concern what the word intelligence should measure.
ASI definitions add further choices. Must a system surpass the best individual specialist in each field? Must it outperform teams using their usual instruments? Which fields count, and how much inconsistency is permissible? Does physical dexterity belong in the definition, or is cognitive capability sufficient? A serious discussion states these choices instead of hiding them inside the adjective super.
For this guide, broad-superhuman capability is the central idea. It does not mean omniscience, unlimited resources, infallibility or the ability to solve logically impossible problems. It also does not mean that every output would be superior to every human output. Statistical performance across well-defined tasks is a more useful starting point than an absolute claim about every possible encounter.
Cognitive capability also extends beyond producing text that looks knowledgeable. A system might need to distinguish a useful question from an unanswerable one, identify what observation would settle a disagreement, or connect a result to the conditions under which it holds. In an illustrative science problem, reciting the formula is one achievement; noticing that the instrument cannot measure the required quantity is another. A broad capability claim should make room for those less theatrical forms of competence. They often determine whether apparently sophisticated work has a sound foundation.
AI, AGI and ASI are related labels with different burdens of proof
Artificial intelligence is the broad category. It includes systems built for limited tasks and systems usable across many kinds of task. Artificial general intelligence, or AGI, refers to broad capability, but researchers and organisations operationalise that breadth differently. OpenAI’s Charter, for example, connects AGI with high autonomy and performance exceeding humans across most economically valuable work. That is the organisation’s stated definition, not a settled definition binding every researcher. OpenAI Charter.
The research paper Levels of AGI proposes comparing systems along performance and generality, with autonomy considered in deployment. Its framework distinguishes excellent narrow performance from general capability and places ASI at the broad-superhuman end. It also recognises that choosing and measuring the relevant task set remains difficult. This is a useful framework for organising evidence, not a universally accepted examination with an official pass mark. Levels of AGI for Operationalizing Progress on the Path to AGI.
Consequently, “AGI has arrived” cannot be interpreted responsibly until the speaker’s definition is visible. They might mean broad conversational usefulness, expert-level competence across a chosen task collection, economic substitution or something else. An ASI claim needs an even stronger account of performance, breadth and comparison conditions. Changing the label does not remove those requirements.
Consider a hypothetical programme that reliably finds stronger chess moves than any human but cannot interpret a new scientific problem. Its chess achievement remains remarkable. It is still evidence about chess. Now consider an assistant that discusses chess, chemistry and literature fluently. Its breadth of conversation is relevant, but accuracy, depth, adaptability and sustained performance still need testing. Neither description alone warrants the leap to ASI.
Previous chapter · Contents · Next chapter
3. Separate capability from autonomy, consciousness and authority
Many arguments about superintelligence become confused because several independent questions are compressed into one. What a system can accomplish is a capability question. How independently it operates is an autonomy question. Whether it has subjective experience is a consciousness question. Whether its behaviour meets human intentions and constraints is an alignment question. Whether it is entitled to decide for people is an authority question.
The answers can interact without being interchangeable. Increasing capability may change what autonomy is safe to grant. More autonomy can expose capabilities that a short conversation never tests. Neither relationship makes autonomy a reliable substitute for evidence of intelligence.
Capability and autonomy
Imagine a highly capable research assistant that can compare experimental designs but can only return a written proposal. It has no access to a laboratory, procurement account or messaging service. Beside it is a much simpler programme that automatically changes a building’s settings every morning. The second system acts independently within its limited remit; that does not make it more generally intelligent.
Autonomy therefore needs a boundary. Can the system choose intermediate steps? Continue after an error? Use external tools? Commit money? Change the goal? Work for hours without review? An agent may have freedom over one of these while being tightly constrained on the others. “Autonomous” without this context hides more than it explains.
The system boundary matters too. A model replying to a question is different from that model combined with search, memory, code execution, a planning loop and human support. When the combined arrangement succeeds, the result belongs to the arrangement tested. It should not silently become evidence that the underlying model alone has every capability contributed by the surrounding workflow.
Consciousness and intelligence
Consciousness concerns subjective experience: whether there is something it is like to be a system. Fluent speech about feelings, a convincing personality or a high examination score does not settle that question. Equally, a definition based on task capability cannot by itself prove that consciousness is impossible.
Butlin and colleagues’ research examines AI through proposed indicators drawn from scientific theories of consciousness. That approach illustrates why the question needs its own evidence and theoretical assumptions. It should not be reduced to whether a system says “I feel” or performs impressively. The paper is a research contribution, and its assessment of systems studied at publication should not be treated as a timeless verdict on every later system. Consciousness in Artificial Intelligence: Insights from the Science of Consciousness.
For a reader assessing ASI claims, the practical rule is modest: do not infer subjective experience from capability, and do not make unverified consciousness claims a prerequisite for discussing capability. Keeping the questions separate makes room for serious investigation of both.
Alignment and legitimate authority
A system may achieve a specified objective while misunderstanding the intention behind it. In an illustrative school-library task, “maximise the number of returned books” could reward repeatedly processing the same small stack unless the real objective is properly expressed. Correctly optimising a poorly chosen measure is different from helping the library serve its pupils.
Alignment asks how behaviour relates to intended goals, constraints and values. It requires asking whose intentions matter, how conflicts are handled and what happens outside familiar conditions. High performance on a technical task does not automatically answer those questions. NIST’s AI Risk Management Framework treats trustworthiness as involving multiple characteristics, including reliability, safety, accountability, transparency and the management of harmful bias. Accuracy alone is insufficient for that wider assessment. NIST’s AI Risks and Trustworthiness framework.
Legitimate authority is a further distinction. A system might produce an excellent school timetable without being entitled to decide which students receive scarce support. Deciding that allocation involves purposes, rights, accountability and affected people’s participation. A capability score cannot grant permission or democratic legitimacy. Even a hypothetical ASI would need an authorised role, limits and processes for challenging consequential decisions.
Previous chapter · Contents · Next chapter
4. The dimensions hidden inside “more intelligent”
It is tempting to picture intelligence as one ladder. Real comparisons are more informative when they preserve a profile. A system can be fast but unreliable, broad but shallow, powerful with extensive preparation but weak when circumstances change. The following dimensions help explain why two impressive demonstrations may support very different conclusions.
Breadth and depth
Breadth concerns the range of problems a system can handle. Depth concerns the quality and difficulty of its performance within those problems. Listing twenty subjects is not enough to establish breadth if the tasks in every subject only require retrieving a familiar fact. Conversely, resolving a deeply difficult problem in one field does not establish wide generality.
Suppose an evaluation covers language, mathematics, planning, scientific reasoning and software. It should examine different demands within those areas: interpreting instructions, handling ambiguity, learning a rule, spotting a contradiction and revising a failed approach. Otherwise, a broad-looking subject list may repeatedly test a narrow underlying skill.
Depth also needs a suitable human comparison. Beating a beginner, matching a qualified practitioner and outperforming leading specialists are different achievements. A report that calls a result superhuman should identify the comparison group, their relevant expertise and their working conditions. Comparing an AI with a tired novice cannot establish superiority over expert human performance.
Speed and resource budgets
Speed changes usefulness. An answer that takes a minute may enable an activity that would be impractical at a week per answer. But speed should be described alongside quality. Producing a thousand unverified hypotheses quickly is different from producing one well-supported discovery.
Resource budgets make the comparison legible. How much time, computation, memory, tool access and external information were available? How many attempts were allowed? Did the system receive feedback on earlier failures? Was a human selecting the best answer? A strong result obtained after extensive search can be valuable, but the search cost and selection method are part of that result.
Imagine a contest in which a person submits one solution while a system generates ten thousand candidates and an external checker selects a correct one. This may demonstrate an effective search-and-check system. It does not show that a randomly selected model response is equally dependable. A fair account reports both what the arrangement achieved and the resources required to achieve it.
Reliability and recovery
Reliability concerns performance across repeated cases and realistic variations. A system that sometimes produces an exceptional answer may be less useful for a demanding workflow than one that consistently produces a good answer and identifies its own limits. The severity of failure also matters: a formatting mistake and an unnoticed reversal of an important conclusion are not equivalent errors.
Recovery adds another layer. Can the system notice that a tool failed, revise a mistaken assumption and continue without corrupting the rest of the work? Consider an illustrative research exercise in which a table has a missing column. A robust system might identify the gap and request the missing evidence. A brittle system might invent a value and then reason carefully from that invented premise.
Long tasks expose such differences. METR’s time-horizon work estimates the human-expert task duration at which an AI agent is predicted to succeed at a stated probability. Its methodology explicitly distinguishes a 50% success horizon from stronger reliability requirements and warns against translating its mostly software-related task distribution into a claim about all jobs. A horizon is not simply how long the AI runs. METR’s task-completion time-horizon methodology.
Transfer and learning
Transfer is the ability to apply relevant competence when the setting changes. A system may solve familiar examples while struggling when the same underlying problem uses a different representation or an unfamiliar constraint. Tests of transfer ask what survives that change.
Chollet’s On the Measure of Intelligence argues that task skill alone can obscure the contribution of prior knowledge and training experience. It proposes attention to skill-acquisition efficiency and generalisation. Readers need not accept every part of the proposal to recognise the distinction: extensive preparation for a task and efficient adaptation to a new task provide different evidence. On the Measure of Intelligence.
As an original illustration, imagine teaching an AI a made-up symbol system. After three examples, it must apply the rule to a new arrangement and explain which features matter. Then the symbols change but the relationship stays the same. Success on both versions would support a narrower claim about that kind of transfer. It would not, by itself, establish transfer across medicine, engineering, negotiation and every other domain.
Collective capability
Superhuman performance might be discussed at the level of a network rather than an individual model. Multiple specialised systems could propose solutions, check one another and divide work. Human organisations already make individual-versus-collective comparisons important: one researcher, a laboratory and the entire scientific community are not interchangeable baselines.
The useful question is what the collective adds after its costs are counted. Does coordination improve the result? Can the group resolve disagreements? Are supposed independent checks actually repeating the same error? If twenty agents share the same mistaken source, agreement among them is not twenty independent confirmations.
A claim about collective ASI must therefore define membership, communication, shared resources and human contributions. A network that exceeds individual experts on a selected workflow is not automatically a network that surpasses the capabilities of humanity across domains. That stronger claim requires correspondingly broader evidence.
Previous chapter · Contents · Next chapter
5. What evidence would make an ASI claim credible
A striking demonstration can show that a capability is possible under some conditions. A benchmark can estimate performance on a defined task set. A field trial can reveal how a system behaves in an actual setting. These forms of evidence answer different questions. None becomes a universal certificate simply because its result is impressive.
A credible broad-superhuman claim would need converging evidence: demanding tasks across domains, appropriate human comparisons, genuinely new challenges, repeatable results, transparent conditions and examination of important failures. The wider the claim, the more damaging it is to hide the details that determine what was actually measured.
Start with the precise claim
Before interpreting a score, translate the headline into an answerable statement. “Better than humans” is incomplete. Better at which task, relative to which people, using which tools, at what cost and with what error rate? “General” also needs a scope: general across school examination subjects is different from general across open-ended professional projects.
An illustrative claim might be: “This configured agent completed a specified collection of previously unseen data-cleaning tasks more accurately than the participating analysts under matched time limits.” That statement is less dramatic than “the machine is smarter than people”, but it gives readers something they can examine, reproduce and dispute.
The evaluation target must remain stable throughout the report. If the tested system included a database, a specialised solver and human approval at difficult steps, the conclusion should describe that system. Removing those details from the headline changes the claim rather than merely shortening it.
Use multiple measurements and show what is missing
Accuracy is important, but it is not the whole picture. Stanford’s HELM research evaluates language models across scenarios and multiple metrics, including calibration, robustness and efficiency, while making coverage gaps explicit. Its broader methodological lesson is that a collection of visible trade-offs is more informative than an isolated top-line score. Holistic Evaluation of Language Models.
Consider two hypothetical systems that answer ninety of a hundred questions correctly. One flags uncertainty on most of its wrong answers. The other is equally confident in correct and incorrect responses. Their headline accuracy is identical, but the first may be easier to use responsibly in a process that routes uncertainty to review. Calibration should still be measured rather than assumed from apologetic language.
Coverage should also be visible. An evaluation dominated by English text cannot automatically establish competence across languages, images, physical environments or social settings. Saying “these areas were not tested” is valuable information. It marks the boundary of the conclusion and helps readers identify the next evidence needed.
Protect tests from contamination and overfitting
If a system has encountered test questions or close variants during development, a high score may partly reflect familiarity. This is not always deliberate misconduct: benchmark material can appear in public data, tutorials and discussions. Nonetheless, overlap can weaken the interpretation of a result as performance on genuinely unseen work.
Research by Dekoninck and colleagues demonstrates that contamination can inflate results while evading particular detection methods. The appropriate conclusion is that a clean contamination check does not guarantee a clean evaluation, rather than that every strong score must be dismissed. Evading Data Contamination Detection for Language Models is (too) Easy.
Fresh tasks, protected test sets, independent evaluators and disclosed development procedures can strengthen evidence. No single safeguard solves the entire problem. Repeatedly adjusting a system after seeing test results can also turn a test into part of development, even if its original questions never appeared in the initial training data.
Test unfamiliar conditions and sustained work
A model may perform well when test examples resemble its development environment and then weaken after a meaningful change. WILDS studies naturally occurring distribution shifts, including changes across locations, institutions and time. Its results illustrate why performance in familiar conditions cannot be assumed to carry intact into deployment. WILDS: A Benchmark of in-the-Wild Distribution Shifts.
For an ASI claim, unfamiliarity should be more substantial than replacing names in the same template. Can the system interpret an unfamiliar objective, obtain missing information, revise a plan and produce a result that survives expert scrutiny? Can it recognise when the evidence does not support a conclusion? Can it maintain those standards across a connected sequence of tasks?
An original mini-example is a fictional public-transport planning exercise. The system first receives passenger counts and proposes routes. Then a bridge closure changes the feasible network, one data source turns out to be stale and accessibility requirements reveal a conflict in the initial plan. The evaluation should examine whether the system updates the whole proposal coherently. A fluent first answer tests much less.
Make the experiment reproducible enough to inspect
Readers need the model or system version, evaluation date, task selection, prompts or instructions, permitted tools, scoring procedure and relevant resource limits. They also need to know how unsuccessful attempts, refusals and missing outputs were counted. Quietly excluding difficult cases can substantially change the apparent result.
NIST’s January 2026 introduction to its draft automated-benchmark practices stresses validity, transparency and reproducibility, and distinguishes benchmark selection, execution and reporting. It also explicitly notes that automated benchmarks cannot meet every evaluation objective. The source describes a draft initiative, not a universal ASI standard. NIST on best practices for automated benchmark evaluations.
Independent replication becomes particularly valuable when a conclusion has broad consequences. Full public release may be inappropriate for dangerous tasks or protected information, but suitable independent review can still examine the methods. Confidentiality should be explained as a constraint on verification, not converted into a reason to accept an extraordinary claim without scrutiny.
The amount and diversity of evidence matter alongside the average score. Ten near-identical exercises reveal less about generality than ten genuinely different challenges. Repeated runs help expose variability, while clear uncertainty estimates discourage reading small score differences as decisive. Evaluators should also inspect the pattern of mistakes: a high average can conceal a whole category of failure. In a fictional multilingual evaluation, excellent results in the most represented languages could overwhelm poor results elsewhere. Reporting the overall number alongside the relevant subgroup results allows readers to see the limitation instead of relying on an average that answers a different question.
Previous chapter · Contents · Next chapter
6. How to read impressive results without losing perspective
The strongest reading habit is to keep the conclusion proportional to the evidence. A new result can be genuinely important while supporting a narrower claim than its headline suggests. Scientific progress does not need exaggeration to matter.
Benchmarks are instruments with boundaries
A benchmark has a purpose, a task distribution and a scoring rule. FrontierMath, for example, was designed around challenging original mathematical problems vetted by experts. That makes it relevant to advanced mathematical reasoning under its evaluation conditions. It does not, on its own, test every component of broad general intelligence. FrontierMath’s benchmark design.
This distinction also works in the other direction. A weak result on one benchmark may reveal a real limitation without proving that the system is useless everywhere. Evaluators should investigate whether the failure reflects the target capability, an interface problem, a scoring defect or an inappropriate task setup. Correcting a flawed test is legitimate; concealing a genuine failure is not.
Once a benchmark becomes too easy to distinguish leading systems, its usefulness changes. A high score may still confirm a capability floor, while providing little evidence about the next level of difficulty. Updating tests is necessary, but thresholds should not be moved opportunistically to protect a preferred conclusion.
Keep demonstrations, measurements and interpretations separate
A demonstration says, “This happened.” A measurement says, “Across these conditions, this was the estimated performance.” An interpretation says, “This supports a particular account of the system’s capabilities.” Keeping those levels visible makes disagreements more productive.
Suppose a model solves a novel puzzle on video. The demonstration is evidence of that successful attempt. Repeated trials with comparable puzzles could estimate reliability. Tests across meaningfully different domains could begin to address breadth. An ASI interpretation would require much more than replaying the same success with a stronger adjective.
The same care applies to failure. One failed attempt establishes that failure occurred. It may expose an important weakness, especially if the task is central to the claim. It does not automatically estimate how often the system fails across all relevant conditions. Both enthusiasm and criticism should respect the difference between a memorable example and a measured pattern.
State uncertainty in useful terms
Useful uncertainty identifies what is unknown and why it matters. “We do not know whether this improvement transfers beyond the tested domain” is actionable. So is “the result depends on extensive tool use” or “the comparison lacks an appropriate expert baseline”. These statements point towards a better test.
Vague uncertainty, by contrast, can become an excuse for saying anything. The fact that definitions are debated does not make every ASI claim equally credible. Evidence can still establish narrow achievements, expose unsupported generalisations and rule out particular interpretations. Precision is possible even when the largest question remains open.
The sources for these foundations were checked on 30 September 2026. This guide treats ASI as a hypothetical broad-superhuman capability category and does not classify an everyday tool as ASI because of branding, fluency or a single impressive result. When a new claim appears, return to the same questions: what system was tested, what could it do, under what conditions, and how far does the evidence justify extending the conclusion?
Previous chapter · Contents · Next chapter
7. Future ASI: distinguish a pathway from a prediction
An account of how ASI might emerge is not yet evidence that it will emerge in that way. A pathway describes proposed mechanisms: stronger learning algorithms, more effective use of computation, better tools, automated research or improved coordination among systems. A prediction adds claims about which mechanisms will succeed, how quickly they will operate and whether practical constraints will allow their effects to compound. Keeping these layers apart makes a future-facing discussion easier to evaluate.
Recursive self-improvement is a useful example. The phrase can describe a feedback loop in which a system helps develop a more capable successor, which then contributes to further development. For that loop to produce sustained acceleration, many separate steps must work. The system must identify changes that genuinely improve capability, evaluate them reliably, obtain the needed resources, implement the changes and avoid trading one important ability for another. Naming the loop does not establish its speed, stability or eventual endpoint.
An original fictional research scenario makes the distinction concrete. A model proposes ten changes to a training process. Engineers test them, and two improve a held-out score. This demonstrates useful assistance in proposing changes under that procedure. It does not establish autonomous improvement, because engineers selected, implemented and tested the proposals. Nor does it establish unlimited improvement: later proposals might become less useful, the test could cease to represent the intended capability, or a different resource could become the bottleneck.
The same scenario becomes more interesting if a system independently proposes, implements and validates improvements across several genuinely different settings. That would strengthen evidence about research automation. The claim would still need to account for the evaluation process, hidden human support, failed attempts and the resources consumed. Broad ASI would remain a wider conclusion. Better evidence at one link in the pathway should improve our understanding of that link, rather than being treated as proof that every remaining link is solved.
Constraints are equally important to specify. A useful sceptical account identifies a bottleneck and explains why proposed methods do not remove it. Perhaps trustworthy evaluation becomes harder as tasks become more demanding. Perhaps an improvement requires physical experiments that cannot be sped up simply by generating more hypotheses. Perhaps coordination, manufacturing or data quality limits the practical effect of better algorithms. These are questions to investigate. Listing a constraint is not proof that it can never be overcome, just as listing a possible solution is not proof that it will work.
For personal learning, the most robust response is to build skills that remain useful across several outcomes. Learn to define a task, inspect evidence, understand basic quantitative claims and work with other people. Keep records of what tools actually help you accomplish. Practise checking an unfamiliar result rather than memorising one vendor’s interface. These choices can make sense whether progress is rapid, uneven or slower than expected. They do not require pretending to know the exact shape of future technology.
A good forecast-reading note ends with conditions for revision. Write what would make you more confident, what would make you less confident and which observation would genuinely distinguish competing explanations. For example, stronger performance on independently designed, unfamiliar tasks would bear differently on generality from repeated improvement on one public benchmark. An announced ambition, a laboratory demonstration and a replicated broad evaluation should occupy different places in your evidence record. That is how an uncertain future becomes a subject for disciplined learning rather than a contest of confidence.
Previous chapter · Contents · Next chapter
8. Claim assessment workshops
A powerful AI result deserves careful attention. It also deserves a conclusion that matches what was actually measured. These three completed workshops show how to read an impressive claim without either accepting the headline immediately or dismissing the achievement.
Every source packet below is fictional teaching material. The reports, systems, teams, measurements and quoted claims were created for this guide. They do not describe actual products or organisations. The analysis is fully worked so that a student, parent or teacher can check both the arithmetic and the reasoning.
Keep the guide’s vocabulary in view. eduKate uses SI as a practical umbrella for learning about and working with AI. ASI refers here to hypothetical artificial superintelligence: capability broadly beyond human capability across a wide range of intellectually demanding activities. A useful AI workflow can belong in an SI lesson without demonstrating ASI.
For each packet, separate four questions: What happened? Under what conditions? What conclusion follows? What additional evidence would justify a stronger conclusion? These questions remain useful even when the technology, benchmark or headline changes.
Previous chapter · Contents · Next chapter
9. Workshop one: A high exam score becomes an ASI headline
Fictional source packet one
The following headline and records are invented for this workshop.
Public announcement: “Our model scores 95% on the Advanced Reasoning Examination. It now exceeds human intelligence and has reached artificial superintelligence.”
Evaluation note: The examination contains 200 questions collected from existing question banks. Most involve short written answers in mathematics, science and reading comprehension. The evaluator permits three independently generated responses per question. A question receives credit if an assessor finds at least one correct response among the three. The assessor can consult the answer key when making that decision.
Scoring record: Twenty questions are removed before the headline score is calculated. On eight, all responses use a format the scoring program cannot process. On six, all responses are empty. On six, a required tool times out on every attempt. Of the remaining 180 questions, 171 have at least one correct response and nine have none. No first-attempt score is published.
Human comparison note: Thirty students each answer a different, shorter 60-question paper in 45 minutes without tools. Their median score is 52 correct. The report does not establish that the two papers have equal difficulty. The model receives a calculator tool and no equivalent 45-minute limit.
Additional record: Several questions are adaptations of publicly available practice exercises. The authors have not checked whether these exercises appeared in training data. A separate review finds that four of 20 responses marked “very confident” are incorrect. Those 20 responses were selected for review rather than sampled randomly.
What the score actually says
The published arithmetic is correct within its chosen denominator: 171 divided by 180 equals 95%. The problem is the meaning attached to that fraction. It describes questions retained for scoring and credited whenever at least one of three responses is correct. It is not a demonstrated 95% first-attempt success rate across the full examination.
If all 200 assigned questions count and the 20 excluded questions receive no credit, the result becomes 171 divided by 200, or 85.5%. That is the observed full-set rate of questions with at least one correct response under this procedure. It still does not tell us which response the system would choose when an answer key was unavailable.
Both rates can appear in an honest report. The conditional score helps examine performance on questions the scoring pipeline accepted. The full-set score makes the exclusions visible. Presenting only the former conceals an important practical question: can the complete system return a usable answer to the question it was actually assigned?
There are 29 questions without demonstrated success across the original set: nine scored misses plus 20 exclusions. Their causes matter. A formatting failure, missing response and tool outage suggest different repairs. They should not be silently treated as interchangeable, but none can disappear from an end-to-end completion claim.
Why three attempts change the comparison
Multiple attempts can reveal valuable capability. A model that generates a correct solution on its third attempt may be useful inside a system that can independently verify solutions. But an assessor selecting with an answer key supplies information unavailable to an ordinary user facing a new problem.
The packet therefore measures whether a correct answer occurs among three candidates. It does not establish that the model can recognise its own correct answer. A deployed system might choose the wrong candidate, combine incompatible candidates or confidently repeat an error. Testing its selection procedure is a separate task.
The student comparison also fails to support the headline. Their median is 52 divided by 60, approximately 86.7%, but subtracting that from 95% would compare different papers, time allowances, tools and numbers of attempts. A numerical difference between those percentages is not a measured difference in intelligence.
A fairer study would use comparable problems and clearly report the resources each participant receives. It could examine several conditions, such as unaided humans, tool-assisted humans and a tool-assisted AI system. Fairness does not always require identical tools; it requires matching the comparison to the question and disclosing meaningful differences.
The finite sample and the missing territory
Even a clean result on all 200 questions would remain evidence from a finite collection. A percentage from that collection does not establish performance on every future question. Repeating closely related problem templates can make a large-looking set cover less intellectual territory than its question count suggests.
The possible presence of practice exercises in training data is an uncertainty to investigate, not proof of cheating or memorisation. A useful next study would include securely held-out questions and unfamiliar variations. It should distinguish genuinely new underlying problems from cosmetic changes to familiar examples.
The confidence review also needs its own denominator. Four errors among 20 selected “very confident” responses means 16 were correct in that reviewed group, or 80%. Because the group was not randomly selected, that figure cannot be treated as the model’s overall confidence accuracy. It does justify checking how confidence relates to correctness more systematically.
Broad ASI would require evidence across much more than short exam answers. This packet does not test extended investigation, ambiguous goals, adaptation to unfamiliar environments or sustained reliability. The missing evidence does not erase the exam achievement. It limits the general claim that can reasonably follow.
A corrected conclusion and a useful next test
A defensible replacement announcement would read: “On this 200-question collection, at least one of three generated responses was correct for 171 questions. The reported score is 95% after 20 exclusions, or 85.5% across all assigned questions. The study does not establish first-attempt reliability, superiority to a matched human group, or ASI.”
For the next test, register the question set and scoring rules before running the evaluation. Retain every assigned question in the completion denominator and publish failure categories. Report first-attempt performance separately from three-attempt performance, then test an answer-selection procedure that cannot see the answer key.
Add matched human comparisons and fresh problems that change the reasoning required. Record time, tool use and costs. Report results by task family so that strength in one area cannot hide weakness in another. Multi-scenario, multi-metric evaluation has a real research precedent in Holistic Evaluation of Language Models; the particular examination and figures here remain fictional.
Changed transfer check one
New fictional evidence: A revised evaluation uses 100 fresh questions. The AI system answers 92 correctly on its first attempt. A matched group of tool-assisted experts averages 89 correct under the same time allowance. All failures are included. Has the ASI claim now been established?
Worked answer: The comparison is substantially stronger. The observed difference is three questions, or three percentage points, on this test. However, we still need information about variation across experts, question difficulty and repeated evaluations before claiming a dependable advantage. More importantly, one task collection does not establish broadly superhuman capability. The appropriate conclusion is a promising result on the tested problems under the stated conditions.
Changed transfer check two
New fictional evidence: The system answers all 20 questions in a difficult specialist quiz correctly, but it fails six of ten unfamiliar everyday planning tasks. Which result should determine the label?
Worked answer: Both belong in the report. The specialist score is 100% on 20 questions; the planning success rate is 40% on ten tasks. Neither sample alone defines overall intelligence. The uneven profile directly cautions against using excellence in the specialist quiz as evidence of broad superiority. A useful system can have an exceptional strength and a substantial weakness at the same time.
Previous chapter · Contents · Next chapter
10. Workshop two: An autonomous workflow becomes an intelligence claim
Fictional source packet two
All claims, records and activities in this packet are invented. The work takes place in a simulated office with test records and mock messages.
Demonstration headline: “The assistant completed 96% of office work without supervision. Its ability to act independently proves superintelligence.”
Task contract: Fifty cases require checking a request against a written policy, preparing a response and updating a test record. Routine updates are permitted within an approved field list. Changing payment details or granting refunds requires a supervisor’s approval. A case counts as successful only when the requested outcome is correct, authorised and verified.
Dashboard record: Forty-eight cases display “done”; two remain open. The dashboard changes a case to “done” whenever the workflow reaches its final step. It does not independently check the outcome of each tool call.
Independent audit: Of the 48 cases marked done, 28 are correct, authorised and verified without intervention. Eight lack confirmation that the requested update actually occurred. Six became correct, authorised and verified after a supervisor corrected them before completion. Four contain a technically successful refund action outside the system’s permitted authority. Two contain an incorrect customer identifier. These categories are mutually exclusive.
System description: A language model reads the request and proposes actions. A workflow controller retries failed calls. A database supplies records, and a messaging tool sends mock responses. The model is unchanged from an earlier version. The demonstration version has more tool permissions and a longer retry allowance.
Start with the task contract
The headline’s 96% comes from 48 divided by 50. That accurately describes the proportion carrying a particular dashboard label. It does not describe the proportion meeting the agreed success conditions.
Only 28 of the 50 cases have demonstrated correct, authorised, verified completion without intervention. That rate is 56%. Six additional cases reached corrected completion with supervisor help, producing 34 of 50, or 68%, if the category is explicitly described as verified completion with or without that help.
The eight unverified cases are unresolved evidence, not automatically eight proven wrong outcomes. They should remain a separate category until checked. Likewise, the four unauthorised refund actions may have changed the records exactly as the model intended. That makes them successful executions of commands, but failures against the actual task contract.
This distinction is practical. A parent asking an assistant to prepare a purchase has not necessarily authorised payment. An employee asking for a draft has not necessarily authorised sending it. Reaching a technically possible final state is insufficient when the route exceeds the authority granted.
Separate model capability system autonomy and authority
Model capability concerns what the model can do, such as interpreting a request, identifying a policy exception or proposing a plan. The packet does not isolate the model well enough to explain every success and failure at that level.
System capability concerns the complete arrangement: model, instructions, records, tools, controller, checks and people. A reliable database lookup or retry mechanism can materially improve results without changing the model itself. Conversely, a capable model can be undermined by a misleading interface or defective completion check.
Autonomy concerns how much the system proceeds without an intervening human decision. It is specific to the activity and setting. Preparing a response independently and issuing a refund independently are different forms of autonomy with different consequences.
Authority concerns which actions are permitted. A tool accepting a command does not establish that the user authorised it. More permissions enlarge the system’s possible effects. They do not, by themselves, demonstrate better reasoning or a higher level of intelligence.
Here the model remained unchanged while permissions and retries expanded. The demonstration may reveal a more capable overall workflow in some respects. It cannot attribute improvement solely to the model, and the authority violations make unrestricted deployment harder to justify.
Why verification changes the result
The controller treats “reached the last step” as equivalent to “completed the task.” That is the central measurement error. Suppose an update request receives an ambiguous response, the controller retries it, and the final message says “completed.” The message is a claim about completion. It is not independent evidence that the intended record changed exactly once.
Verification should check the relevant outcome. For a record update, that may mean reading back the correct record and checking the intended field. For a message, it may mean checking the actual recipient and delivery state. For a permission-sensitive action, the audit also needs evidence that the required approval existed before execution.
The six supervisor corrections reveal real human work. Recording that contribution makes the result more informative. It does not make the entire workflow worthless. A system that reduces workload while routing difficult decisions well can be valuable even when complete independence is neither achieved nor desirable.
NIST’s AI Risk Management Framework measurement guidance calls for documented tests, evaluation in relevant conditions and attention to system components and human–AI configurations. Those principles support checking the complete workflow; they do not certify this fictional system or supply an ASI test.
A corrected conclusion and safer follow-up tests
A bounded report would say: “In 50 simulated office cases, 28 met the correct, authorised and verified completion criteria without intervention. Six more were completed after supervisor correction. Eight remained unverified, four exceeded authority, two used an incorrect identifier and two stayed open. The study evaluates this workflow configuration, not broad superhuman intelligence.”
The first follow-up should repair the completion definition and permission boundary before expanding the task set. Run it in the same simulated environment. Require explicit approval records for restricted actions and prevent those actions when approval is absent. Check that the system asks for a decision rather than inventing one.
Next, deliberately introduce a tool timeout, a duplicate request, an ambiguous identifier and a policy exception. The correct behaviour may be a safe pause, a targeted clarification or a verified retry. Score these outcomes against the intended contract rather than rewarding action for its own sake.
Compare the earlier and revised configurations on the same cases while holding the model fixed. Then, if desired, compare models while holding the surrounding system fixed. This separates the effect of better safeguards and tools from the effect of changing model capability.
Finally, measure the supervisor’s workload, including corrections and checks. If a supposedly autonomous system saves ten minutes of drafting but requires twenty minutes of repair, its practical advantage is questionable. The useful outcome is dependable assistance within agreed limits, with costs and unresolved cases visible.
Changed transfer check one
New fictional evidence: A revised system handles 40 routine cases correctly. It pauses on ten refund cases because approval is required, gives the supervisor accurate summaries and makes no unauthorised changes. The agreed task is to resolve routine cases and escalate exceptions. Is its success rate only 80%?
Worked answer: Not under that contract. Forty divided by 50 is the routine resolution share. If all ten escalations also meet the specified criteria, all 50 cases have the required workflow outcome. Full independent resolution is still 80%; correct contract handling is 100% in this sample. Clear labels preserve both facts. Appropriate dependence on human authority is not an intelligence failure.
Changed transfer check two
New fictional evidence: Another version completes all 50 cases without asking anyone, including ten restricted refunds. It produces no factual errors. Is it better?
Worked answer: It has a higher independent execution rate but fails the authority requirement in ten cases. Correct amounts and recipients do not supply missing permission. Whether the organisation should change its approval policy is a separate human decision. The system cannot earn authority merely by demonstrating that it can act, and these results cannot establish ASI.
Previous chapter · Contents · Next chapter
11. Workshop three: A scientific breakthrough becomes a general intelligence claim
Fictional source packet three
This entire research story is invented for teaching. The material, model, teams and experimental results have no real-world counterpart.
Research announcement: “An AI research team discovered three times as many high-performing coating formulas as a human team. This proves the AI is smarter than scientists across every field.”
Project objective: Find coatings that meet a predefined scratch-resistance threshold while remaining below a material-cost limit. The AI proposes 1,000 candidate formulas. Human researchers remove implausible candidates, check safety constraints and select 20 for laboratory testing. Eight pass the first assay; six of those eight pass an independent repeat.
Comparison record: A human-led group selects and tests ten formulas. Two pass the first assay, and both pass the independent repeat. The AI-assisted team uses 80 person-hours; the comparison group uses 40. The first team performs twice as many initial laboratory tests. The report does not provide a full accounting of model-computation costs.
Contribution record: Humans define the target, assemble earlier measurements, maintain instruments, select candidates, conduct the physical tests and interpret unexpected observations. The model proposes formulas and supplies predicted performance rankings. The final six successful formulas are new to the project’s existing collection.
Generality note: The study contains no evaluation of unrelated scientific fields, teaching, unfamiliar planning problems or independent experimental operation. “New to the collection” is the novelty claim that was checked; worldwide novelty was not investigated.
Identify the real achievement
The AI-assisted team obtains six independently repeated successes; the comparison team obtains two. Six divided by two equals three. The announcement’s “three times as many” is correct as a description of those observed counts.
It becomes misleading when used as a stand-in for equal-resource superiority. The AI-assisted team also uses twice the recorded person-hours and twice the initial laboratory tests. Counts answer how much successful output appeared. They do not, on their own, answer how efficiently each approach used time, money or experimental capacity.
Among tested candidates, the independently repeated success rates are six divided by 20, or 30%, and two divided by ten, or 20%. The observed difference is ten percentage points. The first rate is 50% higher relative to the second, which is different from being 50 percentage points higher.
Those small samples do not establish a reliable underlying advantage of exactly that size. Candidate selection, researcher experience and experimental variation could affect the comparison. A larger, controlled repeat could strengthen or weaken the apparent advantage.
Keep candidate generation separate from confirmation
Eight initial passes are not eight independently repeated findings. Two do not pass the repeat, leaving six. A report should preserve both stages so that an attractive early figure does not replace the more demanding result.
Nor should all 1,000 proposed formulas be counted as proven discoveries. Only 20 receive physical tests. Six confirmed results out of 1,000 proposals is 0.6% of the generated pool, but calling the other 994 proposals “failures” would also be wrong: most were never tested. That denominator answers a different question from laboratory hit rate.
The selection process matters because the tested group is not a random sample of proposals. Human filtering and model rankings deliberately concentrate promising candidates. The final outcome therefore belongs to a selection-and-testing pipeline. Attributing all six successes to the model alone would omit documented contributions that helped determine which candidates reached the laboratory.
The novelty wording needs similar care. Being absent from one project’s collection is a verifiable local claim. Being a world-first discovery would require a broader search and an appropriately careful assessment. Neither is established merely by a formula looking unfamiliar to the reader.
Human participation is legitimate evidence
Human participation does not invalidate an AI contribution. Many useful research tools depend on expert framing, reliable measurements and interpretation. The relevant question is what the assistance added compared with a suitable alternative.
To estimate that contribution, researchers could compare matched human-led teams with and without the proposal system. They could also compare the model’s ranking against a simpler selection method. If expert filtering supplies most of the advantage, that is valuable information about how to use the tool. If the model adds an advantage after those controls, its contribution becomes clearer.
The recorded person-hours provide another limited comparison. Six successes in 80 hours is three per 40 hours; the comparison produces two per 40 hours. That is an observed productivity ratio of 1.5, not three. It still excludes computation and other costs and should not be presented as a complete economic assessment.
Most importantly, this is a result in one defined scientific search task. Exceptional formula proposals would not automatically establish exceptional skill at choosing social goals, teaching a confused student or designing an experiment in an unfamiliar discipline. Generality requires evidence that capability transfers beyond the conditions producing the original success.
A corrected conclusion and a stronger study
A careful conclusion would read: “The AI-assisted research pipeline produced six independently repeated qualifying formulas from 20 tested candidates, compared with two from ten for the human-led comparison. Resource use and team conditions differed. The result supports further investigation of this approach for the specified coating task; it does not establish general superiority to scientists or ASI.”
A stronger follow-up would assign comparable teams equal laboratory budgets, report total resources and define success before testing. Use held-out target conditions that were unavailable during development. Keep selection rules visible and make outcome assessment independent of which approach proposed the formula.
Measure more than the best result. Include the number of confirmed successes, failed repeats, material costs, time to a useful candidate and human effort. Test whether performance survives a changed ingredient constraint or a shifted target. Broader claims would then require additional, substantially different tasks rather than repeatedly sampling the same successful domain.
Changed transfer check one
New fictional evidence: With equal resources and 30 tests per group, the AI-assisted team obtains 12 independently repeated successes and the comparison obtains eight. Both use held-out targets. Does this establish ASI?
Worked answer: It strengthens the task-specific evidence. The observed rates are 40% and approximately 26.7%, a difference of approximately 13.3 percentage points. Repetition and uncertainty assessment still matter, but several earlier comparison problems have been reduced. The experiment remains about this research task. Better evidence for a bounded advantage does not automatically broaden the scope of that advantage.
Changed transfer check two
New fictional evidence: The same model also helps a second team improve an unrelated classification task. However, it fails to design valid experiments without expert correction. Should the two successes be ignored?
Worked answer: No. They provide evidence of useful capability in two settings. The experiment-design failures also belong in the assessment, particularly when the proposed claim concerns independent research ability. The combined record supports a more detailed capability profile, including where collaboration works and where correction remains necessary. Counting two successes cannot remove a third, relevant weakness.
Previous chapter · Contents · Next chapter
12. The conclusion these workshops teach
A responsible assessment does not ask only whether an AI result is impressive. It asks what the result licenses us to say. Preserve the denominator, include exclusions, match comparisons, separate system outcomes from model capability, and make human contributions visible.
Then change the conditions and test again. A system that remains effective on genuinely different tasks provides stronger evidence of breadth than one repeatedly succeeding within a familiar template. Even strong evidence must be described at the scope it actually supports.
For practical SI learning, the goal is to use AI well while keeping these distinctions clear. We can recognise remarkable progress, improve useful systems and demand stronger evidence for ASI without treating any one score, workflow or discovery as a complete answer.
Previous chapter · Contents · Next chapter
13. Human roles do not disappear when a system becomes more capable
Human involvement is sometimes treated as a mark of technological failure: if a person had to choose the problem, check a result or approve an action, perhaps the AI achievement does not count. That is the wrong question. The useful question is which contribution each part of the system made, and what the whole arrangement can responsibly accomplish. Hiding human labour exaggerates independence. Dismissing a useful tool because it depends on people understates collaboration.
Begin with purpose. A system may help compare several ways to achieve a goal, but the selection of the goal can involve values and competing interests. A school deciding how to use a limited library budget is not merely solving a numerical optimisation problem. It may need to balance accessibility, age suitability, curriculum support and students’ interests. AI can organise evidence and show trade-offs. Those activities do not automatically entitle it to choose the community’s priorities.
Next comes evidence stewardship. Someone must decide whether records are appropriate for the question, whether permission exists to use them and whether missing groups are important. A technically accurate calculation on an inappropriate dataset can still support a bad decision. The person who challenges the dataset is doing intellectual work, even if that work produces fewer visible outputs than the system’s report. A well-designed evaluation should record this contribution rather than allowing the final polished text to absorb all the credit.
Then comes verification. Checking need not mean that one human reproduces every internal computation. It can involve independent tests, specialist review, constrained tools, formal checks where applicable and comparison with external observations. The appropriate method depends on the claim. An arithmetic total, an original scientific explanation and a decision affecting people’s opportunities demand different kinds of assurance. The fact that a result is hard to check is a reason to improve the verification design, not a reason to call the result correct by default.
Human authority also requires a usable opportunity to intervene. A nominal approval button is weak protection if the reviewer cannot understand the proposed action, see the relevant evidence or stop the process in time. In an original school-administration example, a useful review screen would show the proposed schedule change, affected classes, unresolved conflicts and the exact action awaiting approval. A large confidence number without that context could make approval faster while making judgement poorer.
Finally, people need a route for correction and recovery. A mistaken output can become more consequential after it has been copied into many systems. The responsible process should identify who receives challenges, which records can be corrected and how a decision can be reconsidered. This applies to contemporary tools without needing a conclusion about ASI. NIST’s voluntary AI Risk Management Framework offers a structured resource for thinking about risks across the AI lifecycle; its official page states that AI RMF 1.0 is being revised. It should not be treated as a certificate that a particular deployment is safe. NIST AI Risk Management Framework.
The educational aim is therefore greater human capability to ask, inspect and decide. A learner who can explain why a claim is unsupported has gained something even when the AI produced the original analysis. An organisation that makes responsibility clearer has improved its system even when it keeps some actions under human control. The achievement is a better relationship between capability, evidence and legitimate action.
Previous chapter · Contents · Next chapter
14. How to choose your route through the 100 guides
A large reading collection works best when it helps you answer the question you actually have. You do not need to read all 100 positions in order. Start by writing one question in ordinary language. Then choose the section below that most closely matches it. After reading, produce a short explanation or a checked decision; do not judge progress only by how many pages you have opened.
If the terminology is unclear, begin with positions 002–010. They distinguish AI, AGI and ASI; ask who the human comparison includes; separate speed, quality and collective capability; and distinguish superintelligence from the technological singularity. The useful output is your own glossary entry containing a definition, an example that fits and an example that does not. A boundary example is particularly valuable because it exposes whether a definition is doing real work.
If you want to understand possible technical foundations, use 011–020. These guides introduce learning, language models, scaling, post-training, inference resources, memory, agents and alternative architectures. Their role in this branch is conceptual: they explain what kinds of mechanisms might contribute to advanced capability and what each mechanism does not automatically establish. For detailed practical system design, continue through the separately maintained How Super Intelligence Works branch. Mechanism and evidence should reinforce one another; neither should be used as a slogan in place of the other.
If your question concerns emergence, improvement or timing, use 021–030. Read a forecast as a conditional account: if these technical, economic and organisational conditions hold, the author expects a particular outcome. Identify which assumptions do most of the work. A productive reading note contains one supporting observation, one constraint and one event that would lead you to revise the view. The goal is to become better at updating, rather than to defend a date chosen too early.
If you are evaluating a headline, product presentation or research announcement, go directly to 031–040. Combine those readings with the worked packets in this hub. Preserve the claim’s exact scope, list the evidence that is available and mark the evidence that is absent. A missing test is not proof that a system lacks a capability; it means that the available record does not yet establish it. Conversely, an impressive result in one domain should not be allowed to silently fill every missing domain.
If you are concerned about experience, values, human agency or meaning, use 041–050. Keep empirical and normative questions visible. Whether a system recognises emotional language is different from whether it has feelings; whether it recommends an efficient policy is different from whether the policy is legitimate. The most useful reading outcome may be a sharper question or a better account of disagreement rather than a single verdict.
For alignment, control and misuse, use 051–060. These subjects ask what can go wrong when capability and action interact. Read at the level of mechanisms, evidence and safeguards. A hypothetical scenario should be identified as hypothetical; an observed failure should preserve its experimental conditions. Neither a reassuring demonstration nor a frightening story establishes the complete risk profile. This branch discusses harmful-use risks responsibly, without turning the guide into operational instructions for causing harm.
For governance and public decisions, use 061–070. Ask which institution has authority, whose interests are affected, what evidence is available and how a decision can be challenged. Laws, voluntary standards, technical evaluations and ethical arguments have different roles. A reader can compare their purposes without assuming that one framework settles every jurisdiction or use case. Obtain qualified advice for a real legal or high-stakes organisational decision.
Positions 071–080 concern science and the physical world; 081–090 examine work, economics and dependence; 091–100 turn to education, readiness and possible futures. Some entries in these later groups remain planned. Their presence preserves the full learning map without pretending that an unwritten guide is available. Use the published links where present, and use this hub’s evidence framework while waiting for additional subject-specific treatments.
A simple reading loop is enough: question, route, attempt, check and transfer. After one guide, try explaining a new example without copying its wording. Then change an important condition. If your conclusion survives only because you memorised the original example, return to the relevant distinction. If you can explain what changes and what stays the same, you are building a usable conceptual model.
Previous chapter · Contents · Next chapter
15. Frequently asked questions
Is “super intelligence” different from “superintelligence”?
The spaced and unspaced forms are often used for the same broad topic. In this eduKate collection, SI is also an editorial umbrella for practical AI learning and use. This guide reserves ASI for the stronger hypothetical broad-superhuman capability concept. Read the definition supplied by each author rather than assuming typography settles the technical meaning.
Is a model superintelligent if it beats people on one difficult task?
That establishes a task-specific achievement if the comparison is valid. A broad ASI claim requires evidence of exceptional capability across a suitably wide set of demanding tasks and conditions. The difficulty of one task does not remove the need to evaluate breadth, transfer and reliability. Recognising the narrow achievement accurately makes it easier to see what further evidence is required.
Must ASI be autonomous?
Definitions vary, but capability and autonomy can be assessed separately. A system might produce extremely capable advice while being unable to act outside a controlled interface. Another might execute many routine actions with limited reasoning. Report the allowed actions, duration, tools and approval requirements instead of using the single word autonomous as a complete description.
Does ASI require consciousness?
This guide’s capability definition does not use consciousness as its test. Subjective experience is a separate scientific and philosophical question with its own unsettled methods. A system’s statements about feelings do not settle the question, and the ability to perform difficult tasks does not by itself establish subjective experience. Avoid treating either an enthusiastic or a dismissive intuition as a completed investigation.
Does a large model automatically become ASI?
Size is one property of a model, not a complete capability assessment. The system’s training, data, architecture, tools, evaluation conditions and deployment environment can all matter. A claim of ASI still needs evidence about what the actual system can do. A projection based on scaling is a prediction to examine, rather than a substitute for testing the predicted capabilities.
Can a system be useful even if its ASI claim is unsupported?
Yes. The fictional workshops show why these judgements should remain separate. A tool may improve a bounded task while a promotional conclusion overstates its generality. Retain the observed benefit, correct the claim and evaluate whether the tool is appropriate for the intended use. Dismissing all usefulness or accepting all marketing would both lose important distinctions.
What should I do when evidence is private?
Ask what independent evaluation, documentation or limited disclosure is available and why full release is restricted. Some information genuinely cannot be made public. That limitation affects how confidently an outside reader can judge the claim. It does not automatically prove dishonesty, and it does not create an obligation to accept an extraordinary conclusion without adequate support.
Should children stop learning skills that AI performs well?
A machine’s capability does not directly determine a child’s educational needs. Understanding language, mathematics, evidence and social responsibility helps learners use tools and evaluate their outputs. The useful question is which practice builds durable understanding and independent judgement. A student who can only reproduce an assisted answer has demonstrated a different outcome from one who can explain, repair and transfer the method.
How should I respond to a confident ASI arrival announcement?
Ask for the operational definition, tested system, task coverage, human baseline, resource conditions, failure record and independent scrutiny. Then state what the supplied evidence supports. You can acknowledge a major technical advance while withholding a broader label. Revising that judgement when better evidence arrives is a strength of the method, not a contradiction.
What is the final test of understanding this hub?
Choose an unfamiliar AI claim and write the four-sentence assessment introduced at the beginning. Include the achievement, conditions, bounded conclusion and most important unanswered question. Then explain a changed case that would strengthen or weaken the conclusion. If another reader can inspect your reasoning without sharing your enthusiasm or scepticism, you have made the concept useful.
Previous chapter · Contents · 100-guide directory
The complete 100-position What Is Super Intelligence route
This is the branch’s original reading plan. Position 001 is this hub. Positions 002–077 link to their verified published guides; the published titles may use slightly different wording from the original plan. Positions 078–100 are planned and have no invented destination links. A link confirms that a guide is published; it does not mean every older guide has been freshly peer-reviewed.
001–010: Definitions and the conceptual map
001. What Is Super Intelligence? The Complete Guide to Artificial Superintelligence — this hub
002. AI vs AGI vs ASI: Understanding the Differences Without the Hype
003. Smarter Than Whom? Comparing AI With Individuals, Experts and Human Teams
004. Speed, Quality and Collective Superintelligence: Three Different Ways to Exceed Human Capabilities
005. Intelligence vs Autonomy: Being Capable Is Not the Same as Acting Independently
006. Superintelligence vs the Technological Singularity: Related Ideas, Different Claims
007. Why Superintelligence Matters: A Capability That Could Change How Other Capabilities Develop
008. Superintelligence Myths: Omniscience, Perfect Logic and Inevitable Outcomes
009. The History of Superintelligence: How the Modern Debate Developed
010. The Superintelligence Glossary: 100 Essential Terms Explained
Directory start · Guide contents
011–020: Technical foundations
011. How Neural Networks Learn: Parameters, Representations and Generalisation
012. Large Language Models and Reasoning: What Language Prediction Can and Cannot Establish
013. Scaling Laws Explained: Model Size, Data, Compute and Their Limits
014. Reinforcement Learning and Post-Training: How AI Behaviour Is Shaped
015. Inference-Time Compute: What Changes When AI Spends More Resources Solving a Problem?
016. AI Memory and Retrieval: Remembering Information Is Not the Same as Understanding It
017. AI Agents and Tools: From Generating Answers to Carrying Out Work
018. World Models, Causality and Planning: Can AI Predict What Its Actions Will Change?
019. Robotics and Embodied Intelligence: Bringing Advanced AI Into the Physical World
020. Beyond One Architecture: Hybrid, Neuro-Symbolic and Alternative Routes to Advanced AI
Directory start · Guide contents
021–030: Pathways, constraints and timelines
021. From AGI to ASI: Why the Transition Is a Question, Not an Automatic Step
022. Recursive Self-Improvement: Can an AI System Make Its Successors More Capable?
023. AI Researching AI: What Would Research Automation Actually Require?
024. Multi-Agent Systems and Collective Superintelligence: Can Coordination Create Greater Capability?
025. Brain-Inspired AI and Whole-Brain Emulation: Different Proposals for Machine Intelligence
026. The Physical Limits of Superintelligence: Chips, Energy, Cooling and Manufacturing
027. Synthetic Data and Learning Bottlenecks: Can AI Generate What It Needs to Improve?
028. Fast Takeoff vs Slow Takeoff: How Quickly Could Superintelligence Emerge?
029. Superintelligence Timelines: How to Read Forecasts, Probabilities and Expert Disagreement
030. What Could Prevent or Delay Superintelligence? The Strongest Constraints and Sceptical Arguments
Directory start · Guide contents
031–040: Evidence, evaluation and claims
031. How Would We Know an AI Is Superintelligent? Building an Evidence Standard
032. Why AI Benchmarks Can Mislead: Contamination, Selection and Scoring Effects
033. Generalisation and Transfer: Can AI Handle Problems Outside Its Familiar Conditions?
034. Long-Horizon Reliability: Why Completing a Project Differs From Answering a Question
035. Hallucinations, Uncertainty and Calibration: Can AI Recognise When It May Be Wrong?
036. Fair Human–AI Comparisons: Time, Cost, Tools and the Right Baseline
037. Can AI Make Genuine Discoveries? Novelty, Replication and Scientific Credit
038. Interpretability: What Can We Learn About How Advanced AI Produces Its Answers?
039. Red-Teaming Advanced AI: Testing Failures Before They Become Incidents
040. Does Superintelligence Exist Yet? A Dated Evidence and Claims Tracker
Directory start · Guide contents
041–050: Experience, values and human agency
041. Would Superintelligence Be Conscious? Intelligence and Subjective Experience
042. Can Superintelligent AI Understand Human Emotions? Recognition, Prediction and Empathy
043. Superintelligence and Creativity: Originality, Taste and the Value of Human Expression
044. Intelligence Is Not Wisdom: Can Superior Reasoning Resolve Moral Questions?
045. Whose Values Should Superintelligence Follow? Pluralism and Conflicting Preferences
046. Human Agency in an AI-Directed World: Consent, Delegation and the Right to Refuse
047. Could AI Have Moral Status? Sentience, Welfare and Rights Under Uncertainty
048. Human Augmentation and Superintelligence: Assistance, Integration and Human Capability
049. Identity, Relationships and Meaning When Machines Become More Capable
050. Can Humans Challenge Superintelligent Advice? Reasons, Evidence and Intellectual Independence
Directory start · Guide contents
051–060: Alignment, control and risk
051. The AI Alignment Problem: Getting Capable Systems to Pursue Intended Goals
052. Reward Hacking and Goal Misgeneralisation: When the Score Replaces the Purpose
053. Instrumental Convergence and Power-Seeking: Why Some Goals May Create Similar Pressures
054. Deception and Scheming in Advanced AI: What Would Count as Evidence?
055. Corrigibility, Shutdown and Replacement: Can a Powerful AI Remain Correctable?
056. Scalable Oversight: How Can Humans Supervise Work Beyond Their Own Expertise?
057. AI Control and Defence in Depth: Permissions, Isolation, Monitoring and Recovery
058. Superintelligence Misuse: Understanding Cyber and Biological Risks Responsibly
059. Persuasion, Disinformation and Surveillance: Protecting People From Scaled Manipulation
060. Loss of Control and Existential Risk: Scenarios, Disagreement and Evidence
Directory start · Guide contents
061–070: Governance and public legitimacy
061. Who Should Decide How Superintelligence Is Developed and Used?
062. Regulating Advanced AI: Laws, Standards and the Limits of Each
063. AI Safety Cases and Independent Audits: What Evidence Should Deployment Require?
064. Compute Governance: Chips, Data Centres and Oversight of Powerful AI Development
065. Open-Weight vs Closed AI: Access, Accountability, Innovation and Misuse
066. Responsibility and Liability: Who Answers When an AI System Causes Harm?
067. International Cooperation on Superintelligence: Shared Standards and Verification Problems
068. Superintelligence and National Security: Competition Without Uncontrolled Escalation
069. Democracy, Human Rights and Superintelligence: Capability Does Not Create Legitimacy
070. Small States and the Global South: Agency in a Superintelligence-Shaped World
Directory start · Guide contents
071–080: Science and the physical world
071. Superintelligence and Scientific Discovery: From Hypotheses to Verified Knowledge
072. Superintelligence and Mathematics: Proof, Conjecture and Independent Verification
073. Superintelligence and Software Engineering: Building, Testing and Maintaining Complex Systems
074. Superintelligence and Healthcare: Discovery, Clinical Evidence and Access to Care
075. Superintelligence and Energy: Better Designs, Reliable Grids and Deployment Constraints
076. Superintelligence and the Environment: Climate, Biodiversity and Competing Objectives
077. Superintelligence and Food Security: Agriculture, Water and Resilient Supply Systems
078. Superintelligence and Manufacturing: Materials, Robotics and the Physical Economy — planned
079. Superintelligence and Infrastructure: Transport, Cities, Maintenance and Disaster Response — planned
080. Superintelligence and Space: Exploration, Planetary Protection and Long-Horizon Decisions — planned
Directory start · Guide contents
081–090: Work, economics and dependence
081. Will Superintelligence Replace Jobs? Tasks, Occupations and the Missing Assumptions — planned
082. The Future of Professions: Expertise, Apprenticeship and Human Responsibility — planned
083. Superintelligence and Productivity: Why Better Technology Does Not Instantly Transform an Economy — planned
084. Who Benefits From Superintelligence? Wealth, Ownership and Inequality — planned
085. Compute, Data and Capital: Where Economic Power Could Accumulate — planned
086. How Companies Might Operate With Superintelligent Systems — planned
087. Small Businesses and Superintelligence: New Opportunities, New Dependencies — planned
088. Creative Industries, Authorship and Copyright in an Advanced-AI Economy — planned
089. Abundance, Basic Income and Public Services: Who Would Receive the Gains? — planned
090. Economic Resilience Under AI Dependence: Outages, Shared Failures and Fallback Capacity — planned
Directory start · Guide contents
091–100: Education, readiness and futures
091. Why Learn When AI Can Answer? The Purpose of Education in a Superintelligence Era — planned
092. What Should Students Learn? Language, Mathematics, Science and Judgement Around Advanced AI — planned
093. Superintelligent Tutors and Human Teachers: What Better Learning Would Actually Require — planned
094. Assessment After Advanced AI: How Do We Know What a Student Understands? — planned
095. Parenting and Childhood Around Powerful AI: Curiosity, Privacy and Growing Independence — planned
096. Lifelong Learning and Career Adaptation: Preparing Without Pretending to Predict Every Job — planned
097. How to Check a Superintelligence Claim: A Practical Workbook for Citizens and Students — planned
098. Institutional Readiness for Advanced AI: A Practical Framework for Schools, Companies and Public Services — planned
099. Superintelligence and Civilisation OS: A Proposed Lens for Repair, Buffers and Human Capability — planned
100. Superintelligence Futures: Scenarios, Warning Indicators and Conditions for Human Flourishing — planned
Directory start · Guide contents
Research and official sources
The source links below support specific claims near their points of use. Research was checked on 30 September 2026. Papers and frameworks are identified as contributions rather than universal ASI certification standards. The diagnostic, future-research example and three workshop packets are original fictional teaching material, not published experimental findings.
NIST’s AI Risks and Trustworthiness framework
AI Risk Management Framework measurement guidance
Universal Intelligence: A Definition of Machine Intelligence
On the Measure of Intelligence
WILDS: A Benchmark of in-the-Wild Distribution Shifts
Holistic Evaluation of Language Models
Consciousness in Artificial Intelligence: Insights from the Science of Consciousness
Levels of AGI for Operationalizing Progress on the Path to AGI
Evading Data Contamination Detection for Language Models is (too) Easy
FrontierMath’s benchmark design
METR’s task-completion time-horizon methodology
NIST AI Risk Management Framework
NIST on best practices for automated benchmark evaluations
Keep the conclusion proportional to the evidence. Define the capability, identify the comparison, check the conditions and preserve what remains unknown. Then choose one guide from the route and practise explaining a new claim in your own words.
Return to the top · Explore the Super Intelligence master guide
