Super Intelligence (SI) becomes more consequential when intelligence can contribute to the production of further intelligence. This article examines AI researching AI: breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. It preserves the locked Clementi-depth floor with mechanisms, worked cases, failure analysis, verification, transfer, diagnostics and RFE closure.
Search Intent and Direct Answer
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
First principles begin with a loop: propose a change, implement it, test it, interpret the result and decide what to do next. Improvement exists only when the tested system performs better on the intended objective without unacceptable regressions elsewhere. Generating a plausible proposal is not yet improvement. Generating more proposals faster is useful only if evaluation can distinguish real gains from noise.
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Definition and Scope
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Failed improvements are valuable evidence. A system may optimise the metric it can see while degrading robustness, cost, interpretability or another hidden requirement. Record the failure instead of discarding it. The distribution of failed proposals tells us how much reliable research judgement the system possesses and how much human repair remains in the loop.
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
First Principles
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Coordination can add capability through specialisation. One system can search literature, another implement, another test and another criticise. But coordination also introduces communication overhead, duplicated work and the risk that every component inherits the same false premise. Collective SI should therefore be measured against both the strongest individual component and an appropriate human team.
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
The Core Feedback Loop
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Generalisation asks whether the improvement survives a changed context. Evaluate on held-out tasks, different distributions and adversarial cases. A system that optimises a familiar benchmark without transferring has learned something narrower than the headline suggests. Broad SI requires repeated transfer, not merely repeated optimisation of known tests.
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
What Must Be Measured
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
Independent replication protects against shared blind spots. If the same model generates the hypothesis, writes the code, designs the evaluation and judges the result, correlated errors can survive every internal check. Independent tools, separate models, human reviewers or external experiments can provide diversity in the verification path.
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
What the Idea Does Not Assume
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Economics matters because an improvement that costs far more than the value it creates may not scale. Measure performance together with training cost, inference cost, human supervision and experimental throughput. A research system can be scientifically impressive before it becomes economically transformative.
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Worked Example: A Small Improvement
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Safety analysis asks whether the improvement loop can be interrupted, audited and reversed. Faster iteration can compress the time available for review. Shared models can propagate the same error across agents. More capable research systems can also increase dual-use potential. Controls should therefore scale with the consequence and autonomy of the workflow.
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Worked Example: A Failed Improvement
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Students should be able to draw the pipeline and label where evidence enters. Then they should invent one failure at each stage and explain how it would be detected. This transforms an abstract SI concept into systems reasoning and matches the Clementi progression from definition to independent diagnosis.
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
Worked Example: A Research Pipeline
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
RFE closes the article. Receiver: who benefits from the improved capability? Function: what exact research or coordination job must close? Evidence: what verified outcome demonstrates improvement? Exit: when should the loop be paused, redesigned or retired? Applied to AI researching AI, RFE prevents acceleration from becoming the objective when reliable progress is the actual goal.
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
Worked Example: A Team of Systems
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
First principles begin with a loop: propose a change, implement it, test it, interpret the result and decide what to do next. Improvement exists only when the tested system performs better on the intended objective without unacceptable regressions elsewhere. Generating a plausible proposal is not yet improvement. Generating more proposals faster is useful only if evaluation can distinguish real gains from noise.
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Worked Example: Human Comparison
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Failed improvements are valuable evidence. A system may optimise the metric it can see while degrading robustness, cost, interpretability or another hidden requirement. Record the failure instead of discarding it. The distribution of failed proposals tells us how much reliable research judgement the system possesses and how much human repair remains in the loop.
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Verification Before Credit
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Coordination can add capability through specialisation. One system can search literature, another implement, another test and another criticise. But coordination also introduces communication overhead, duplicated work and the risk that every component inherits the same false premise. Collective SI should therefore be measured against both the strongest individual component and an appropriate human team.
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Generalisation Beyond the Test
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Generalisation asks whether the improvement survives a changed context. Evaluate on held-out tasks, different distributions and adversarial cases. A system that optimises a familiar benchmark without transferring has learned something narrower than the headline suggests. Broad SI requires repeated transfer, not merely repeated optimisation of known tests.
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
Long-Horizon Reliability
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
Independent replication protects against shared blind spots. If the same model generates the hypothesis, writes the code, designs the evaluation and judges the result, correlated errors can survive every internal check. Independent tools, separate models, human reviewers or external experiments can provide diversity in the verification path.
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Coordination Costs
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Economics matters because an improvement that costs far more than the value it creates may not scale. Measure performance together with training cost, inference cost, human supervision and experimental throughput. A research system can be scientifically impressive before it becomes economically transformative.
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Shared Failure Modes
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Safety analysis asks whether the improvement loop can be interrupted, audited and reversed. Faster iteration can compress the time available for review. Shared models can propagate the same error across agents. More capable research systems can also increase dual-use potential. Controls should therefore scale with the consequence and autonomy of the workflow.
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Human Oversight
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Students should be able to draw the pipeline and label where evidence enters. Then they should invent one failure at each stage and explain how it would be detected. This transforms an abstract SI concept into systems reasoning and matches the Clementi progression from definition to independent diagnosis.
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
Independent Replication
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
RFE closes the article. Receiver: who benefits from the improved capability? Function: what exact research or coordination job must close? Evidence: what verified outcome demonstrates improvement? Exit: when should the loop be paused, redesigned or retired? Applied to AI researching AI, RFE prevents acceleration from becoming the objective when reliable progress is the actual goal.
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
Compute and Hardware
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
First principles begin with a loop: propose a change, implement it, test it, interpret the result and decide what to do next. Improvement exists only when the tested system performs better on the intended objective without unacceptable regressions elsewhere. Generating a plausible proposal is not yet improvement. Generating more proposals faster is useful only if evaluation can distinguish real gains from noise.
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Data and Experimental Bottlenecks
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Failed improvements are valuable evidence. A system may optimise the metric it can see while degrading robustness, cost, interpretability or another hidden requirement. Record the failure instead of discarding it. The distribution of failed proposals tells us how much reliable research judgement the system possesses and how much human repair remains in the loop.
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Physical-World Latency
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Coordination can add capability through specialisation. One system can search literature, another implement, another test and another criticise. But coordination also introduces communication overhead, duplicated work and the risk that every component inherits the same false premise. Collective SI should therefore be measured against both the strongest individual component and an appropriate human team.
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Economics and Scale
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Generalisation asks whether the improvement survives a changed context. Evaluate on held-out tasks, different distributions and adversarial cases. A system that optimises a familiar benchmark without transferring has learned something narrower than the headline suggests. Broad SI requires repeated transfer, not merely repeated optimisation of known tests.
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
What Current Evidence Supports
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
Independent replication protects against shared blind spots. If the same model generates the hypothesis, writes the code, designs the evaluation and judges the result, correlated errors can survive every internal check. Independent tools, separate models, human reviewers or external experiments can provide diversity in the verification path.
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
What Current Evidence Does Not Establish
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Economics matters because an improvement that costs far more than the value it creates may not scale. Measure performance together with training cost, inference cost, human supervision and experimental throughput. A research system can be scientifically impressive before it becomes economically transformative.
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Connection to Super Intelligence (SI)
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Safety analysis asks whether the improvement loop can be interrupted, audited and reversed. Faster iteration can compress the time available for review. Shared models can propagate the same error across agents. More capable research systems can also increase dual-use potential. Controls should therefore scale with the consequence and autonomy of the workflow.
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Safety Implications
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Students should be able to draw the pipeline and label where evidence enters. Then they should invent one failure at each stage and explain how it would be detected. This transforms an abstract SI concept into systems reasoning and matches the Clementi progression from definition to independent diagnosis.
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
Governance and Responsibility
Progress has four levels. Level 1: describe the idea. Level 2: explain the mechanism. Level 3: predict which bottleneck dominates under changed conditions. Level 4: design an evaluation that can falsify the claim. The locked article floor aims for Level 4 rather than vocabulary-only familiarity.
RFE closes the article. Receiver: who benefits from the improved capability? Function: what exact research or coordination job must close? Evidence: what verified outcome demonstrates improvement? Exit: when should the loop be paused, redesigned or retired? Applied to AI researching AI, RFE prevents acceleration from becoming the objective when reliable progress is the actual goal.
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
Student Learning Route
The core search question in AI researching AI is breaking AI research into stages so automation claims can be measured from hypothesis to verified contribution. The cleanest analysis separates hypothesis, implementation, experiment, evaluation, interpretation and replication. Each stage can succeed or fail independently, which means the final outcome cannot be inferred from one impressive intermediate result. Super Intelligence (SI) requires end-to-end evidence because a feedback loop is only as strong as its weakest verification step.
First principles begin with a loop: propose a change, implement it, test it, interpret the result and decide what to do next. Improvement exists only when the tested system performs better on the intended objective without unacceptable regressions elsewhere. Generating a plausible proposal is not yet improvement. Generating more proposals faster is useful only if evaluation can distinguish real gains from noise.
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Organisation Diagnostic Route
A small worked example makes this concrete. Suppose a system proposes a code optimisation that raises one benchmark by five percent. Before crediting intelligence, repeat the test, inspect side effects, run unrelated benchmarks and compare against a simple baseline. If the gain survives, the evidence strengthens. If it disappears under a changed workload, the original conclusion was too broad.
Failed improvements are valuable evidence. A system may optimise the metric it can see while degrading robustness, cost, interpretability or another hidden requirement. Record the failure instead of discarding it. The distribution of failed proposals tells us how much reliable research judgement the system possesses and how much human repair remains in the loop.
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Progress Ladder
A research pipeline contains different cognitive jobs: selecting a question, reviewing prior work, forming hypotheses, implementing experiments, choosing controls, analysing data, interpreting anomalies and communicating results. Automation can be high in one stage and low in another. The phrase “AI automates research” is meaningful only after the stages and success criteria are specified.
Coordination can add capability through specialisation. One system can search literature, another implement, another test and another criticise. But coordination also introduces communication overhead, duplicated work and the risk that every component inherits the same false premise. Collective SI should therefore be measured against both the strongest individual component and an appropriate human team.
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Counterexamples and Falsification
Verification is the gate between activity and knowledge. A generated theorem needs proof, a software improvement needs tests, a scientific claim needs evidence and a model change needs evaluation across representative tasks. As systems become more capable, verification may itself become the bottleneck because humans can struggle to judge outputs beyond their expertise.
Generalisation asks whether the improvement survives a changed context. Evaluate on held-out tasks, different distributions and adversarial cases. A system that optimises a familiar benchmark without transferring has learned something narrower than the headline suggests. Broad SI requires repeated transfer, not merely repeated optimisation of known tests.
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
RFE Closure
Long-horizon reliability matters because research and coordination involve many dependent steps. Context can drift, experiments can fail and early assumptions can become obsolete. Measure whether the system notices contradiction, updates plans and recovers. A process that works only when every previous step is correct is fragile.
Independent replication protects against shared blind spots. If the same model generates the hypothesis, writes the code, designs the evaluation and judges the result, correlated errors can survive every internal check. Independent tools, separate models, human reviewers or external experiments can provide diversity in the verification path.
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Frequently Asked Questions
Compute and hardware set a material boundary. Recursive improvement may discover better algorithms while still depending on chips, fabrication, electricity and cooling. Some improvements are software-fast; others require new hardware generations or physical experiments. Takeoff forecasts should specify which loop they mean.
Economics matters because an improvement that costs far more than the value it creates may not scale. Measure performance together with training cost, inference cost, human supervision and experimental throughput. A research system can be scientifically impressive before it becomes economically transformative.
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Continue the Super Intelligence (SI) Series
Current frontier systems already assist with coding, literature synthesis, experiment design and evaluation. These are meaningful pieces of research automation. They do not by themselves establish a self-sustaining recursive improvement loop, collective superintelligence or whole-brain emulation. The correct conclusion preserves the scope of the demonstrated capability.
Safety analysis asks whether the improvement loop can be interrupted, audited and reversed. Faster iteration can compress the time available for review. Shared models can propagate the same error across agents. More capable research systems can also increase dual-use potential. Controls should therefore scale with the consequence and autonomy of the workflow.
Governance asks who sets the objective and who bears responsibility. A system that can optimise its successors does not acquire authority to choose social goals. Research institutions still need decision rights, documentation, access controls and accountability for consequential deployment.
Bottleneck Map: What Can Slow the Loop?
Map the bottlenecks before forecasting speed. List idea generation, implementation, compute availability, experiment duration, evaluation quality, hardware access and decision approval. Estimate which stage has the longest latency and which has the least spare capacity. The fastest cognitive component cannot make the whole loop faster than a hard downstream bottleneck indefinitely. This map turns vague takeoff language into a systems model.
A falsification case is an apparent improvement that disappears when the task changes. Perhaps a new method raises a benchmark because it exploits a regularity specific to that evaluation. Test it on held-out domains, altered distributions and unrelated objectives. If the gain vanishes, the result is still useful: it tells researchers the improvement was local rather than general. Super Intelligence (SI) evidence should reward honest narrowing.
Experiments fail for ordinary reasons: code bugs, noisy measurements, missing data, unstable infrastructure and incorrect assumptions. A capable research system should distinguish these categories rather than repeatedly retrying the same plan. Measure whether it can localise the failure, choose a diagnostic test and update its hypothesis. Recovery quality is part of research intelligence.
Self-grading creates correlated risk. If one system proposes a change, implements it and declares it successful, a shared blind spot can survive every stage. Use independent evaluation where practical: separate models, formal checks, held-out benchmarks, human reviewers or physical replication. Independence is not perfect, but diversity in the verification path reduces the chance that one mistaken premise controls the entire conclusion.
Falsification Case: An Improvement That Does Not Transfer
A falsification case is an apparent improvement that disappears when the task changes. Perhaps a new method raises a benchmark because it exploits a regularity specific to that evaluation. Test it on held-out domains, altered distributions and unrelated objectives. If the gain vanishes, the result is still useful: it tells researchers the improvement was local rather than general. Super Intelligence (SI) evidence should reward honest narrowing.
Experiments fail for ordinary reasons: code bugs, noisy measurements, missing data, unstable infrastructure and incorrect assumptions. A capable research system should distinguish these categories rather than repeatedly retrying the same plan. Measure whether it can localise the failure, choose a diagnostic test and update its hypothesis. Recovery quality is part of research intelligence.
Self-grading creates correlated risk. If one system proposes a change, implements it and declares it successful, a shared blind spot can survive every stage. Use independent evaluation where practical: separate models, formal checks, held-out benchmarks, human reviewers or physical replication. Independence is not perfect, but diversity in the verification path reduces the chance that one mistaken premise controls the entire conclusion.
The workbook follows one complete cycle. Write the baseline capability. State the proposed change. Predict the measurable improvement before running it. Implement the change. Test on the target and at least one transfer condition. Record regressions. Decide whether the evidence justifies keeping, revising or rejecting the change. Then ask how much human intervention was required. This produces an auditable improvement claim.
Recovery Case: When the Experiment Fails
Experiments fail for ordinary reasons: code bugs, noisy measurements, missing data, unstable infrastructure and incorrect assumptions. A capable research system should distinguish these categories rather than repeatedly retrying the same plan. Measure whether it can localise the failure, choose a diagnostic test and update its hypothesis. Recovery quality is part of research intelligence.
Self-grading creates correlated risk. If one system proposes a change, implements it and declares it successful, a shared blind spot can survive every stage. Use independent evaluation where practical: separate models, formal checks, held-out benchmarks, human reviewers or physical replication. Independence is not perfect, but diversity in the verification path reduces the chance that one mistaken premise controls the entire conclusion.
The workbook follows one complete cycle. Write the baseline capability. State the proposed change. Predict the measurable improvement before running it. Implement the change. Test on the target and at least one transfer condition. Record regressions. Decide whether the evidence justifies keeping, revising or rejecting the change. Then ask how much human intervention was required. This produces an auditable improvement claim.
Acceleration and progress are different quantities. More experiments per day can increase the chance of discovery, but only if evaluation keeps pace. A loop that produces results faster than they can be checked accumulates uncertainty rather than knowledge. The strongest route toward SI is therefore not merely faster iteration; it is faster reliable iteration with preserved verification, transfer and recovery.
Independent Evaluation: Avoiding Self-Grading
Self-grading creates correlated risk. If one system proposes a change, implements it and declares it successful, a shared blind spot can survive every stage. Use independent evaluation where practical: separate models, formal checks, held-out benchmarks, human reviewers or physical replication. Independence is not perfect, but diversity in the verification path reduces the chance that one mistaken premise controls the entire conclusion.
The workbook follows one complete cycle. Write the baseline capability. State the proposed change. Predict the measurable improvement before running it. Implement the change. Test on the target and at least one transfer condition. Record regressions. Decide whether the evidence justifies keeping, revising or rejecting the change. Then ask how much human intervention was required. This produces an auditable improvement claim.
Acceleration and progress are different quantities. More experiments per day can increase the chance of discovery, but only if evaluation keeps pace. A loop that produces results faster than they can be checked accumulates uncertainty rather than knowledge. The strongest route toward SI is therefore not merely faster iteration; it is faster reliable iteration with preserved verification, transfer and recovery.
Map the bottlenecks before forecasting speed. List idea generation, implementation, compute availability, experiment duration, evaluation quality, hardware access and decision approval. Estimate which stage has the longest latency and which has the least spare capacity. The fastest cognitive component cannot make the whole loop faster than a hard downstream bottleneck indefinitely. This map turns vague takeoff language into a systems model.
Practical Workbook: Trace One Improvement Cycle
The workbook follows one complete cycle. Write the baseline capability. State the proposed change. Predict the measurable improvement before running it. Implement the change. Test on the target and at least one transfer condition. Record regressions. Decide whether the evidence justifies keeping, revising or rejecting the change. Then ask how much human intervention was required. This produces an auditable improvement claim.
Acceleration and progress are different quantities. More experiments per day can increase the chance of discovery, but only if evaluation keeps pace. A loop that produces results faster than they can be checked accumulates uncertainty rather than knowledge. The strongest route toward SI is therefore not merely faster iteration; it is faster reliable iteration with preserved verification, transfer and recovery.
Map the bottlenecks before forecasting speed. List idea generation, implementation, compute availability, experiment duration, evaluation quality, hardware access and decision approval. Estimate which stage has the longest latency and which has the least spare capacity. The fastest cognitive component cannot make the whole loop faster than a hard downstream bottleneck indefinitely. This map turns vague takeoff language into a systems model.
A falsification case is an apparent improvement that disappears when the task changes. Perhaps a new method raises a benchmark because it exploits a regularity specific to that evaluation. Test it on held-out domains, altered distributions and unrelated objectives. If the gain vanishes, the result is still useful: it tells researchers the improvement was local rather than general. Super Intelligence (SI) evidence should reward honest narrowing.
Final Synthesis: Acceleration Is Not the Same as Verified Progress
Acceleration and progress are different quantities. More experiments per day can increase the chance of discovery, but only if evaluation keeps pace. A loop that produces results faster than they can be checked accumulates uncertainty rather than knowledge. The strongest route toward SI is therefore not merely faster iteration; it is faster reliable iteration with preserved verification, transfer and recovery.
Map the bottlenecks before forecasting speed. List idea generation, implementation, compute availability, experiment duration, evaluation quality, hardware access and decision approval. Estimate which stage has the longest latency and which has the least spare capacity. The fastest cognitive component cannot make the whole loop faster than a hard downstream bottleneck indefinitely. This map turns vague takeoff language into a systems model.
A falsification case is an apparent improvement that disappears when the task changes. Perhaps a new method raises a benchmark because it exploits a regularity specific to that evaluation. Test it on held-out domains, altered distributions and unrelated objectives. If the gain vanishes, the result is still useful: it tells researchers the improvement was local rather than general. Super Intelligence (SI) evidence should reward honest narrowing.
Experiments fail for ordinary reasons: code bugs, noisy measurements, missing data, unstable infrastructure and incorrect assumptions. A capable research system should distinguish these categories rather than repeatedly retrying the same plan. Measure whether it can localise the failure, choose a diagnostic test and update its hypothesis. Recovery quality is part of research intelligence.
Research Automation Is a Pipeline, Not One Task
To ask whether AI can research AI, break research into stages: identify a worthwhile problem, review prior work, formulate hypotheses, design experiments, implement them, monitor runs, analyse results, distinguish real gains from noise, write conclusions, and decide what question should come next. A system that automates only code generation has not automated research. A system that closes most of this pipeline is much closer.
For Super Intelligence (SI), the most consequential possibility is not merely faster coding. It is shortening the complete loop from research question to verified capability improvement.
Problem Selection Is the Highest-Leverage Research Decision
Research time can be wasted brilliantly on the wrong question. A capable research system therefore needs more than experimental skill: it must estimate which bottleneck is limiting progress and which intervention has the highest expected value. This requires synthesising evidence, recognising uncertainty and resisting fashionable but low-yield directions.
Current systems remain much easier to evaluate when humans choose the research objective. Removing that human role increases both the ambition and the evaluation difficulty.
Hypothesis Generation Must Produce Testable Differences
A research hypothesis is useful when it predicts something that distinguishes it from alternatives. “Try a larger model” is a suggestion. “This memory architecture should improve long-context task completion specifically because it reduces context compression loss, and should therefore outperform the baseline most strongly on tasks with delayed dependencies” is closer to a scientific hypothesis.
Automated research improves when systems learn to formulate experiments that discriminate among mechanisms rather than merely generate many configurations.
Experiment Design Is an Information-Gain Problem
A good experiment does not only seek a high score. It reduces uncertainty. The system must choose controls, baselines, ablations and evaluation metrics that reveal why a change worked. Without these, an AI research agent can find empirical improvements while learning very little about the mechanism.
Mechanistic understanding matters because it improves transfer. If the system knows why an intervention helps, it can predict where else the intervention should work.
Sakana’s AI Scientist Shows How Much of the Research Loop Can Be Automated
Sakana AI’s AI Scientist project, published in Nature in 2026, demonstrates an agentic system aimed at automating large parts of the machine-learning research lifecycle. Earlier versions generated research ideas, ran experiments and produced papers; the later AI Scientist-v2 produced a fully AI-generated paper that passed a human peer-review process according to the project’s published account.
This is a substantial milestone in research automation, but peer review is not identical to scientific truth. The stronger long-term test is whether automated work produces reproducible findings that other researchers use, extend and validate.
Google DeepMind’s Co-Scientist Shows a Different Research Architecture
DeepMind’s Co-Scientist, published in Nature in May 2026, uses multiple agents to generate, critique, rank and refine scientific hypotheses. The design focuses on helping human scientists develop experimentally testable ideas rather than replacing the whole scientific institution.
This illustrates an important architectural choice: research automation can be designed as an autonomous scientist, a collaborative thought partner, or a modular system inside a human research team. The best design may differ by domain.
Implementation Is Becoming the Easy Part in Some AI Research
Frontier coding agents can increasingly implement experiments, manage repositories, run training jobs and inspect logs. Anthropic’s 2026 reporting on AI-assisted research describes a shift toward humans providing ideas while models execute and test them much faster than before. This can move the research bottleneck upward from implementation toward problem selection and interpretation.
When one bottleneck falls, another becomes visible. Research automation should therefore be measured stage by stage rather than declared complete because one stage became fast.
Evaluation Is Hardest When the System Is Exploring Beyond Known Benchmarks
If an AI discovers an optimisation that improves a standard benchmark, verification is straightforward. If it proposes a new training paradigm or interpretability result, evaluation may require judgement, replication and theoretical understanding. The more novel the research, the less likely a pre-existing automatic grader can settle it.
This creates a paradox for SI research automation: the most valuable discoveries may be the least amenable to fully automated verification.
Multi-Agent Research Can Create Internal Peer Review
One agent can propose a hypothesis, another can challenge assumptions, another can design experiments and another can inspect statistical validity. This role separation can reduce some single-agent failure modes by introducing structured disagreement.
It can also create group failure. If every agent shares the same pretrained blind spot, debate may produce consensus around the same mistake. Diversity of prompts, models, data and verification tools matters.
Research Infrastructure Becomes Part of the Intelligence Stack
An AI researcher needs access to code, compute, experiment tracking, datasets, version control, literature and evaluation systems. Capability therefore depends on the surrounding research infrastructure. A weaker model with excellent tools and automation can outperform a stronger isolated model on end-to-end research throughput.
This is why “AI researching AI” should be analysed as a system capability rather than a model-only property.
Worked Example: Improving an Inference Algorithm
An AI research system notices that a reasoning model spends excessive tokens on easy questions. It hypothesises that a learned difficulty estimator could allocate test-time compute adaptively. It implements the controller, evaluates it on held-out problems, measures accuracy and cost, performs ablations and checks whether the gains transfer to a second benchmark.
If the system then recognises that the controller improves its own research efficiency and deploys it into the next round of experiments, research automation begins to overlap with recursive self-improvement.
Worked Example: A False Discovery Caused by Benchmark Leakage
An agent proposes a training method that produces a large gain. Later analysis reveals that examples from the evaluation benchmark entered the training data. The experimental pipeline was fast but the research conclusion was invalid.
A mature research agent must therefore track provenance, data leakage, multiple comparisons and replication. Scientific speed without methodological control produces faster error.
What Would Full Research Automation Actually Require?
It would require competent direction-setting, literature synthesis, hypothesis formation, experiment design, implementation, debugging, evaluation, interpretation, replication, documentation and prioritisation of the next question. It would also require governance over compute and access, because an autonomous research system can consume substantial resources and modify consequential systems.
That standard is much higher than “can write a paper” or “can run experiments.”
RFE Closure: Automated Research Must Produce Verified New Knowledge
The problem is equating research activity with research progress. The function of AI research automation is to reduce the time and cost required to produce verified improvements or discoveries. The receiver is the scientific or engineering process. Evidence includes reproducibility, transfer, independent validation, research efficiency and downstream use.
The exit condition is to narrow autonomy when direction-setting becomes unreliable, validation cannot keep pace, or the system generates more experiments than the organisation can responsibly verify. The scarce resource may become judgement rather than generation.
Continue the Super Intelligence (SI) Series
Next: Super Intelligence | Multi-Agent Systems and Collective SI.
