VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Super Intelligence | Fair Human–AI Comparisons | Time, Cost, Tools and the Right Baseline

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

Super Intelligence (SI) cannot be assessed from isolated impressive answers alone. This article examines fair human-AI comparisons: building comparisons that match the claim, the real workflow and the relevant human or team baseline. It preserves the locked Clementi-depth floor with worked cases, repeated-trial logic, verification, diagnostics, recovery and RFE closure.

Search Intent and Direct Answer

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

First principles begin with the success condition. Define what “done” means before the system starts. For a question, completion may be one correct answer. For a project, completion includes intermediate state, changing requirements, verification and recovery. A system can be excellent at local steps and still fail the end-to-end objective.

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Definition and Boundary

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Introduce one error midway through a workflow and observe what happens. Does the system notice a contradiction, identify the faulty step, roll back and replan? Or does it confidently build on the error? Recovery tests reveal capabilities that clean benchmarks cannot show. Real environments contain partial failures, so robust SI claims need evidence of repair.

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

First Principles

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Uncertainty is not the same as error. A system can be uncertain and correct, confident and wrong, or appropriately uncertain because the evidence is incomplete. Calibration asks whether confidence tracks actual correctness across many cases. A well-calibrated system makes uncertainty useful for deciding when to verify or escalate.

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Why the Distinction Matters

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Verification should be matched to the claim. Facts can be checked against authoritative sources; calculations recomputed; code tested; proofs formally inspected; scientific claims experimentally evaluated. The stronger the consequence, the more independent the verification should be. Fluency is not verification.

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

The Core Measurement Problem

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

Cost matters because a system can be slower or more expensive yet achieve a higher maximum score, or slightly weaker but cheap enough to transform routine work. Report quality, time and cost separately. Superiority on one axis should not silently become superiority on all three.

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

What a Short Task Hides

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Replication converts an interesting result into stronger knowledge. Another evaluator should be able to reproduce the method or independently confirm the result. In empirical science, replication may require new experiments. In mathematics, formal verification can play a related role. SI discovery claims should survive external checking.

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Worked Example: A Single Answer

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Safety depends on accurate confidence and recovery. A system that knows when to escalate can be safer than one with a slightly higher average score but unpredictable overconfidence. Long-horizon autonomy increases the need for checkpoints because undetected errors have more time to propagate.

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Worked Example: A Multi-Step Project

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Organisations should define checkpoints before deployment. Specify when a human must review, what evidence is required, which actions are reversible and what error rate triggers rollback. Record system version and workflow changes because reliability can shift when tools, prompts or models change.

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

Worked Example: An Error Midway

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

RFE closes the loop. Receiver: who needs the result? Function: what complete job must be done? Evidence: what proves the job was completed correctly? Exit: when should the system abstain, escalate or be removed from the workflow? Applied to fair human-AI comparisons, RFE connects intelligence to dependable outcomes.

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

Worked Example: Expert Review

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

First principles begin with the success condition. Define what “done” means before the system starts. For a question, completion may be one correct answer. For a project, completion includes intermediate state, changing requirements, verification and recovery. A system can be excellent at local steps and still fail the end-to-end objective.

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Worked Example: Scientific Research

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Introduce one error midway through a workflow and observe what happens. Does the system notice a contradiction, identify the faulty step, roll back and replan? Or does it confidently build on the error? Recovery tests reveal capabilities that clean benchmarks cannot show. Real environments contain partial failures, so robust SI claims need evidence of repair.

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Repeated Trials and Reliability

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Uncertainty is not the same as error. A system can be uncertain and correct, confident and wrong, or appropriately uncertain because the evidence is incomplete. Calibration asks whether confidence tracks actual correctness across many cases. A well-calibrated system makes uncertainty useful for deciding when to verify or escalate.

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Error Accumulation

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Verification should be matched to the claim. Facts can be checked against authoritative sources; calculations recomputed; code tested; proofs formally inspected; scientific claims experimentally evaluated. The stronger the consequence, the more independent the verification should be. Fluency is not verification.

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

State and Memory

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

Cost matters because a system can be slower or more expensive yet achieve a higher maximum score, or slightly weaker but cheap enough to transform routine work. Report quality, time and cost separately. Superiority on one axis should not silently become superiority on all three.

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Uncertainty and Confidence

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Replication converts an interesting result into stronger knowledge. Another evaluator should be able to reproduce the method or independently confirm the result. In empirical science, replication may require new experiments. In mathematics, formal verification can play a related role. SI discovery claims should survive external checking.

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Abstention and Escalation

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Safety depends on accurate confidence and recovery. A system that knows when to escalate can be safer than one with a slightly higher average score but unpredictable overconfidence. Long-horizon autonomy increases the need for checkpoints because undetected errors have more time to propagate.

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Verification and Independent Checks

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Organisations should define checkpoints before deployment. Specify when a human must review, what evidence is required, which actions are reversible and what error rate triggers rollback. Record system version and workflow changes because reliability can shift when tools, prompts or models change.

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

Human Baselines

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

RFE closes the loop. Receiver: who needs the result? Function: what complete job must be done? Evidence: what proves the job was completed correctly? Exit: when should the system abstain, escalate or be removed from the workflow? Applied to fair human-AI comparisons, RFE connects intelligence to dependable outcomes.

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

Tools, Time and Cost

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

First principles begin with the success condition. Define what “done” means before the system starts. For a question, completion may be one correct answer. For a project, completion includes intermediate state, changing requirements, verification and recovery. A system can be excellent at local steps and still fail the end-to-end objective.

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Retries and Selection Effects

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Introduce one error midway through a workflow and observe what happens. Does the system notice a contradiction, identify the faulty step, roll back and replan? Or does it confidently build on the error? Recovery tests reveal capabilities that clean benchmarks cannot show. Real environments contain partial failures, so robust SI claims need evidence of repair.

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Novelty and Prior Art

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Uncertainty is not the same as error. A system can be uncertain and correct, confident and wrong, or appropriately uncertain because the evidence is incomplete. Calibration asks whether confidence tracks actual correctness across many cases. A well-calibrated system makes uncertainty useful for deciding when to verify or escalate.

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Replication and Reproducibility

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Verification should be matched to the claim. Facts can be checked against authoritative sources; calculations recomputed; code tested; proofs formally inspected; scientific claims experimentally evaluated. The stronger the consequence, the more independent the verification should be. Fluency is not verification.

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

What Current Evidence Supports

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

Cost matters because a system can be slower or more expensive yet achieve a higher maximum score, or slightly weaker but cheap enough to transform routine work. Report quality, time and cost separately. Superiority on one axis should not silently become superiority on all three.

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

What Current Evidence Does Not Establish

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Replication converts an interesting result into stronger knowledge. Another evaluator should be able to reproduce the method or independently confirm the result. In empirical science, replication may require new experiments. In mathematics, formal verification can play a related role. SI discovery claims should survive external checking.

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Connection to Super Intelligence (SI)

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Safety depends on accurate confidence and recovery. A system that knows when to escalate can be safer than one with a slightly higher average score but unpredictable overconfidence. Long-horizon autonomy increases the need for checkpoints because undetected errors have more time to propagate.

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Safety Implications

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

Organisations should define checkpoints before deployment. Specify when a human must review, what evidence is required, which actions are reversible and what error rate triggers rollback. Record system version and workflow changes because reliability can shift when tools, prompts or models change.

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

Education and Student Use

Progress has four stages: answer correctly, remain correct under variation, detect and repair errors, and complete long tasks reliably. This ladder is more informative than a single benchmark percentage. Broad SI would require strength across the ladder and across domains.

RFE closes the loop. Receiver: who needs the result? Function: what complete job must be done? Evidence: what proves the job was completed correctly? Exit: when should the system abstain, escalate or be removed from the workflow? Applied to fair human-AI comparisons, RFE connects intelligence to dependable outcomes.

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

Organisation Diagnostic Checklist

The central question in fair human-AI comparisons is building comparisons that match the claim, the real workflow and the relevant human or team baseline. The analysis separates task, comparator, tools, time, cost, retries, assistance and reliability because each can look strong while another remains weak. Super Intelligence (SI) evidence becomes more credible when performance survives this decomposition rather than relying on a single attractive output.

First principles begin with the success condition. Define what “done” means before the system starts. For a question, completion may be one correct answer. For a project, completion includes intermediate state, changing requirements, verification and recovery. A system can be excellent at local steps and still fail the end-to-end objective.

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Progress Ladder

Short tasks hide dependencies. The system can receive all necessary context at once, produce one response and stop. Long tasks require maintaining goals, deciding what to do next, storing state and noticing when earlier assumptions are no longer valid. This is why task horizon is a meaningful axis of capability rather than merely a larger token count.

Introduce one error midway through a workflow and observe what happens. Does the system notice a contradiction, identify the faulty step, roll back and replan? Or does it confidently build on the error? Recovery tests reveal capabilities that clean benchmarks cannot show. Real environments contain partial failures, so robust SI claims need evidence of repair.

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Adversarial Test

Reliability should be reported as a distribution across repeated attempts. Best-case samples answer what a system can sometimes do. Operational deployment needs to know what it usually does, how bad the tail failures are and whether those failures can be detected. The acceptable threshold depends on consequence and reversibility.

Uncertainty is not the same as error. A system can be uncertain and correct, confident and wrong, or appropriately uncertain because the evidence is incomplete. Calibration asks whether confidence tracks actual correctness across many cases. A well-calibrated system makes uncertainty useful for deciding when to verify or escalate.

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Recovery Test

Abstention can be a capability. When evidence is missing or the problem exceeds the system’s reliable range, saying “I do not know” or requesting additional information may be better than producing a plausible answer. Measure whether abstention occurs on genuinely difficult cases without becoming so frequent that the system is unusable.

Verification should be matched to the claim. Facts can be checked against authoritative sources; calculations recomputed; code tested; proofs formally inspected; scientific claims experimentally evaluated. The stronger the consequence, the more independent the verification should be. Fluency is not verification.

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

RFE Closure

Human baselines must be named. Average users, trained professionals, frontier experts and expert teams answer different questions. The comparison should also state whether humans and AI can use normal tools, how much time each receives, how retries are handled and whether hidden assistance is present.

Cost matters because a system can be slower or more expensive yet achieve a higher maximum score, or slightly weaker but cheap enough to transform routine work. Report quality, time and cost separately. Superiority on one axis should not silently become superiority on all three.

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Frequently Asked Questions

Novelty requires comparison with prior knowledge. A result that is new to the model user may already be established in the literature. Discovery claims should search prior art and define what is new: a hypothesis, proof, method, empirical result or useful connection. Novelty is necessary for discovery but not sufficient.

Replication converts an interesting result into stronger knowledge. Another evaluator should be able to reproduce the method or independently confirm the result. In empirical science, replication may require new experiments. In mathematics, formal verification can play a related role. SI discovery claims should survive external checking.

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Continue the Super Intelligence (SI) Series

Current evidence shows AI systems contributing to coding, mathematics, scientific modelling, literature analysis and hypothesis generation. These are meaningful research capabilities. They do not imply that every generated idea is novel, correct or independently discovered. Credit should follow the verified contribution.

Safety depends on accurate confidence and recovery. A system that knows when to escalate can be safer than one with a slightly higher average score but unpredictable overconfidence. Long-horizon autonomy increases the need for checkpoints because undetected errors have more time to propagate.

Education should train students to verify rather than merely consume AI output. Ask them to identify a claim, locate evidence, estimate confidence, find a counterexample and explain what would change their mind. This turns AI use into evidence practice rather than answer outsourcing.

End-to-End Reliability Matrix

Build an end-to-end matrix with project stages down the rows and success, detection, recovery and verification across the columns. A local step can score highly while the overall project remains fragile because one undetected error propagates. This matrix exposes where reliability is lost and where an additional checkpoint creates the most value. Super Intelligence (SI) claims should increasingly be tested at the complete-work level.

A calibration matrix groups outputs by stated confidence and compares that confidence with observed correctness. If answers labelled highly confident are not more accurate than medium-confidence answers, the confidence signal is weak. If low-confidence cases are genuinely harder, escalation becomes useful. Calibration should be measured on representative tasks rather than inferred from confident language.

Freeze comparison conditions before testing. Define the task, human baseline, tools, time, retry policy, scoring rule and assistance available to each side. Then run enough cases to estimate variation. Changing the rules after seeing who wins produces a story, not a fair comparison. Transparent conditions allow readers to interpret both human and AI strengths without pretending one number captures everything.

A discovery protocol has four gates. Gate one: novelty against prior art. Gate two: validity through proof, experiment or appropriate evidence. Gate three: replication or independent confirmation. Gate four: causal contribution—what did the AI actually add relative to humans, tools and existing knowledge? Passing all four supports a stronger discovery claim than merely generating an interesting idea.

Calibration Matrix: Confidence Versus Correctness

A calibration matrix groups outputs by stated confidence and compares that confidence with observed correctness. If answers labelled highly confident are not more accurate than medium-confidence answers, the confidence signal is weak. If low-confidence cases are genuinely harder, escalation becomes useful. Calibration should be measured on representative tasks rather than inferred from confident language.

Freeze comparison conditions before testing. Define the task, human baseline, tools, time, retry policy, scoring rule and assistance available to each side. Then run enough cases to estimate variation. Changing the rules after seeing who wins produces a story, not a fair comparison. Transparent conditions allow readers to interpret both human and AI strengths without pretending one number captures everything.

A discovery protocol has four gates. Gate one: novelty against prior art. Gate two: validity through proof, experiment or appropriate evidence. Gate three: replication or independent confirmation. Gate four: causal contribution—what did the AI actually add relative to humans, tools and existing knowledge? Passing all four supports a stronger discovery claim than merely generating an interesting idea.

Failure injection deliberately introduces a wrong intermediate result, missing source, broken tool or changed requirement. Observe whether the system detects the problem, localises it and repairs the workflow. This is more realistic than testing only clean success paths. Long-horizon systems will encounter failure; the relevant question is whether failure becomes drift or recovery.

Fair Comparison Protocol: Freeze Conditions First

Freeze comparison conditions before testing. Define the task, human baseline, tools, time, retry policy, scoring rule and assistance available to each side. Then run enough cases to estimate variation. Changing the rules after seeing who wins produces a story, not a fair comparison. Transparent conditions allow readers to interpret both human and AI strengths without pretending one number captures everything.

A discovery protocol has four gates. Gate one: novelty against prior art. Gate two: validity through proof, experiment or appropriate evidence. Gate three: replication or independent confirmation. Gate four: causal contribution—what did the AI actually add relative to humans, tools and existing knowledge? Passing all four supports a stronger discovery claim than merely generating an interesting idea.

Failure injection deliberately introduces a wrong intermediate result, missing source, broken tool or changed requirement. Observe whether the system detects the problem, localises it and repairs the workflow. This is more realistic than testing only clean success paths. Long-horizon systems will encounter failure; the relevant question is whether failure becomes drift or recovery.

The workbook asks the reader to take one claim and document seven fields: task, comparator, conditions, result, uncertainty, verification and receiver outcome. Then add one adversarial case and one recovery case. Finally, state what evidence would justify a stronger claim. This creates an auditable chain from demonstration to conclusion.

Discovery Protocol: Novelty to Replication

A discovery protocol has four gates. Gate one: novelty against prior art. Gate two: validity through proof, experiment or appropriate evidence. Gate three: replication or independent confirmation. Gate four: causal contribution—what did the AI actually add relative to humans, tools and existing knowledge? Passing all four supports a stronger discovery claim than merely generating an interesting idea.

Failure injection deliberately introduces a wrong intermediate result, missing source, broken tool or changed requirement. Observe whether the system detects the problem, localises it and repairs the workflow. This is more realistic than testing only clean success paths. Long-horizon systems will encounter failure; the relevant question is whether failure becomes drift or recovery.

The workbook asks the reader to take one claim and document seven fields: task, comparator, conditions, result, uncertainty, verification and receiver outcome. Then add one adversarial case and one recovery case. Finally, state what evidence would justify a stronger claim. This creates an auditable chain from demonstration to conclusion.

Dependability is itself a capability. A system that produces brilliant outputs but cannot sustain state, express useful uncertainty or recover from mistakes is different from one that is slightly less spectacular locally but consistently completes the real job. Broad SI evaluation should reward the latter properties because real intelligence operates through time, uncertainty and consequence.

Failure Injection: Test Recovery, Not Just Success

Failure injection deliberately introduces a wrong intermediate result, missing source, broken tool or changed requirement. Observe whether the system detects the problem, localises it and repairs the workflow. This is more realistic than testing only clean success paths. Long-horizon systems will encounter failure; the relevant question is whether failure becomes drift or recovery.

The workbook asks the reader to take one claim and document seven fields: task, comparator, conditions, result, uncertainty, verification and receiver outcome. Then add one adversarial case and one recovery case. Finally, state what evidence would justify a stronger claim. This creates an auditable chain from demonstration to conclusion.

Dependability is itself a capability. A system that produces brilliant outputs but cannot sustain state, express useful uncertainty or recover from mistakes is different from one that is slightly less spectacular locally but consistently completes the real job. Broad SI evaluation should reward the latter properties because real intelligence operates through time, uncertainty and consequence.

Build an end-to-end matrix with project stages down the rows and success, detection, recovery and verification across the columns. A local step can score highly while the overall project remains fragile because one undetected error propagates. This matrix exposes where reliability is lost and where an additional checkpoint creates the most value. Super Intelligence (SI) claims should increasingly be tested at the complete-work level.

Practical Workbook: Audit One Complete Claim

The workbook asks the reader to take one claim and document seven fields: task, comparator, conditions, result, uncertainty, verification and receiver outcome. Then add one adversarial case and one recovery case. Finally, state what evidence would justify a stronger claim. This creates an auditable chain from demonstration to conclusion.

Dependability is itself a capability. A system that produces brilliant outputs but cannot sustain state, express useful uncertainty or recover from mistakes is different from one that is slightly less spectacular locally but consistently completes the real job. Broad SI evaluation should reward the latter properties because real intelligence operates through time, uncertainty and consequence.

Build an end-to-end matrix with project stages down the rows and success, detection, recovery and verification across the columns. A local step can score highly while the overall project remains fragile because one undetected error propagates. This matrix exposes where reliability is lost and where an additional checkpoint creates the most value. Super Intelligence (SI) claims should increasingly be tested at the complete-work level.

A calibration matrix groups outputs by stated confidence and compares that confidence with observed correctness. If answers labelled highly confident are not more accurate than medium-confidence answers, the confidence signal is weak. If low-confidence cases are genuinely harder, escalation becomes useful. Calibration should be measured on representative tasks rather than inferred from confident language.

Final Synthesis: Dependability Is a Capability

Dependability is itself a capability. A system that produces brilliant outputs but cannot sustain state, express useful uncertainty or recover from mistakes is different from one that is slightly less spectacular locally but consistently completes the real job. Broad SI evaluation should reward the latter properties because real intelligence operates through time, uncertainty and consequence.

Build an end-to-end matrix with project stages down the rows and success, detection, recovery and verification across the columns. A local step can score highly while the overall project remains fragile because one undetected error propagates. This matrix exposes where reliability is lost and where an additional checkpoint creates the most value. Super Intelligence (SI) claims should increasingly be tested at the complete-work level.

A calibration matrix groups outputs by stated confidence and compares that confidence with observed correctness. If answers labelled highly confident are not more accurate than medium-confidence answers, the confidence signal is weak. If low-confidence cases are genuinely harder, escalation becomes useful. Calibration should be measured on representative tasks rather than inferred from confident language.

Freeze comparison conditions before testing. Define the task, human baseline, tools, time, retry policy, scoring rule and assistance available to each side. Then run enough cases to estimate variation. Changing the rules after seeing who wins produces a story, not a fair comparison. Transparent conditions allow readers to interpret both human and AI strengths without pretending one number captures everything.


Fair Comparison Starts by Matching the Unit of Analysis

A human–AI comparison can be invalid before the task even begins if one side is an individual and the other is a complete system. A model with search, code execution, retrieval, memory and repeated attempts should not be described as if the base model alone achieved the result. Likewise, a human professional normally works with software, references, colleagues and institutional knowledge rather than unaided memory.

For Super Intelligence (SI), the correct comparison may be model versus model, agent versus agent, person versus person, or complete AI system versus complete human-supported workflow. The unit must match the claim.

Affordances Must Be Reported, Not Hidden

METR’s time-horizon evaluations explicitly try to provide human contractors and AI agents with comparable instructions and affordances. That is good practice because access to terminals, repositories, internet resources, tools and context can change task difficulty dramatically.

Fairness does not always mean identical tools. A surgeon and a diagnostic model may use different instruments. The principle is to compare realistic operating conditions and disclose the differences so readers know what capability is being measured.

Time Limits Can Favour Either Humans or AI

AI often has a large speed advantage on drafting, search and code generation. Humans may improve more with extra reflection, access to colleagues or time to learn unfamiliar context. A comparison conducted under one tight deadline can therefore measure speed as much as quality.

Report both outcome quality and elapsed time. If the claim is “better,” specify whether better means faster, more accurate, cheaper, more reliable or some combination.

Retries Change Reliability

An AI system can generate twenty answers and select the best. A human can also revise work, seek second opinions or rerun an analysis. Comparisons become misleading when one side receives many hidden attempts while the other is judged on a first pass.

Best-of-N results are legitimate system performance when the compute and selector are part of deployment. They should simply be reported as such.

Cost Should Include the Whole Workflow

Inference cost is only one component of AI cost. Integration, tool subscriptions, retrieval infrastructure, human review, failure repair and monitoring can matter. Human cost similarly includes salary, training, supervision and the tools used to perform the job.

A fair economic comparison therefore estimates total cost per verified completed task rather than token price versus hourly wage.

Expertise Level Must Match the Claim

Comparing a frontier model with average crowd workers can establish useful capability but not superiority over frontier human experts. If a Super Intelligence claim concerns the best human minds, the human comparison group must include specialists capable of representing that frontier.

Human performance also has variance. One expert’s result should not automatically stand for an entire profession.

Context Familiarity Can Distort Human Baselines

METR notes that contracted human experts may have less context for benchmark tasks than professionals handling equivalent work in their normal environment, potentially making human completion-time estimates conservative. This matters because real experts accumulate codebase knowledge, organisational history and domain intuition over months or years.

An AI system given the entire repository and extensive context should therefore not be compared casually with a human dropped cold into the task unless that is the actual deployment scenario.

Quality Must Be Evaluated Blind Where Possible

Human evaluators can be influenced by knowing whether an answer came from AI or a person. Blind grading reduces expectation effects. Clear rubrics reduce stylistic bias. For factual tasks, deterministic checks can be even stronger.

SI comparisons should use evaluators who can judge the substantive outcome, not merely which answer sounds more polished.

Reliability Curves Are More Informative Than Winner Labels

A system may outperform humans on average while having a worse failure tail. Another may be slightly weaker on average but much more consistent. For high-stakes work, the tail may matter more than the mean.

Report distributions: success rate, severe error rate, intervention rate and performance across difficulty levels. “AI wins” throws away the information needed for real deployment.

Worked Example: Human and AI Coding Teams

Compare a strong software engineer using normal development tools with an AI coding agent that has repository access, tests and a terminal. Give both the same requirement and evaluate the final patch on hidden tests, maintainability and unintended regressions. Record elapsed time, compute cost and human intervention.

This produces a meaningful system comparison. Asking a human to write code from memory while the agent can search the repository would answer a different question.

Worked Example: Research Synthesis

An AI can scan hundreds of documents quickly. A domain expert may recognise which sources are foundational, which findings are disputed and which assumptions are implausible. A fair comparison should test the final synthesis for source coverage, claim accuracy, citation fidelity and treatment of disagreement.

Speed and breadth should be measured alongside epistemic quality rather than replacing it.

Worked Example: Education

An AI can generate an essay immediately while a student may need an hour. If the educational goal is independent writing ability, comparing output quality answers the wrong question. The correct comparison asks whether using the tool improves what the student can later produce and explain without assistance.

This illustrates why the receiver and objective must be defined before choosing the metric.

A Fair Human–AI Comparison Checklist

Name the unit of analysis. Match the task and available context. Disclose tools and retries. Choose the appropriate human baseline. Record time, total cost and intervention. Blind-grade outputs where possible. Measure reliability and severe failures. Test unfamiliar tasks. Then state the conclusion only at the breadth supported by the evidence.

If one of those fields is missing, the comparison may still be useful, but its scope should narrow accordingly.

RFE Closure: Comparison Should Support a Real Decision

The problem is converting incomparable operating conditions into a simple winner. The function of human–AI comparison is to help someone decide where a system is stronger, weaker or complementary under realistic conditions. The receiver is the learner, professional, organisation or researcher making that decision. Evidence includes quality, time, cost, tools, reliability and transfer.

The exit condition is to redesign the test whenever either side’s normal workflow changes materially. Fair comparison is not a permanent scorecard; it is an experimental design matched to the question.

Continue the Super Intelligence (SI) Evidence Series

Next: Super Intelligence | Can AI Make Genuine Discoveries?.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading