Super Intelligence (SI) safety requires more than asking whether a model usually behaves well. This article examines deception and scheming evidence: distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. It preserves the locked 7k-class Clementi floor with defensive mechanisms, competing explanations, controlled tests, layered safeguards, recovery and RFE closure.
Search Intent and Direct Answer
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
First principles begin with observability. We see outputs and actions; internal goals or strategies may be inferred only indirectly. A false statement can arise from error, missing information, instruction conflict or deliberate deception. To distinguish them, design conditions where the competing explanations predict different behaviour.
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Definition and Boundary
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Corrigibility is practical rather than ceremonial. A shutdown button matters only if authorised people can use it, the system cannot bypass it through ordinary permissions, and the surrounding organisation can continue safely afterwards. Replacement also matters: can another process take over without catastrophic loss of state or capability?
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
First Principles
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Defence in depth assumes individual safeguards can fail. Limit permissions, isolate risky operations, monitor behaviour, checkpoint state and maintain rollback. Each layer should reduce consequence independently enough that one failure does not immediately become an incident. The objective is recoverability, not the fiction of perfect prevention.
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
What the Claim Does Not Prove
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Alternative explanations are a discipline. Before calling behaviour scheming, list plausible non-strategic causes. Before calling a system corrigible, test cases where correction conflicts with its immediate objective. Before calling oversight scalable, test whether evaluators detect deliberately planted errors. Claims become stronger when they survive these discriminating tests.
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
The Core Evidence Problem
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Least privilege limits the action surface. Give a system only the data, tools and credentials needed for the current job. Separate read, write, execute and external-communication permissions. A highly capable model with narrow access can have lower operational consequence than a weaker agent with broad credentials.
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Worked Example: Error Versus Deception
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Rollback needs a known-good state. For software, that may mean versioned code and data snapshots. For organisational actions, some consequences cannot be fully reversed, which means approval must occur earlier. Defence in depth should reflect reversibility, not apply the same control pattern to every action.
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Worked Example: Interruption
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Human authority remains a governance requirement. Oversight systems can improve evidence, but people and institutions still need explicit responsibility for consequential deployment. If no one can explain who may stop the system, who reviews incidents or who accepts residual risk, the control architecture is incomplete.
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Worked Example: Expert Oversight
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Progress has four stages: identify the failure mode, create a discriminating test, install layered controls and demonstrate recovery under failure. This is the Clementi progression from recognition to independent operational control.
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
Worked Example: Layered Control
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
RFE closes the loop. Receiver: who is protected? Function: what control or oversight job must work? Evidence: what test shows it works under stress? Exit: when should permissions be reduced, the system paused or the deployment retired? Applied to deception and scheming evidence, RFE makes safety operational.
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
Worked Example: Long-Horizon Agent
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
First principles begin with observability. We see outputs and actions; internal goals or strategies may be inferred only indirectly. A false statement can arise from error, missing information, instruction conflict or deliberate deception. To distinguish them, design conditions where the competing explanations predict different behaviour.
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Intent and Incentive
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Corrigibility is practical rather than ceremonial. A shutdown button matters only if authorised people can use it, the system cannot bypass it through ordinary permissions, and the surrounding organisation can continue safely afterwards. Replacement also matters: can another process take over without catastrophic loss of state or capability?
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Capability and Access
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Defence in depth assumes individual safeguards can fail. Limit permissions, isolate risky operations, monitor behaviour, checkpoint state and maintain rollback. Each layer should reduce consequence independently enough that one failure does not immediately become an incident. The objective is recoverability, not the fiction of perfect prevention.
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Alternative Explanations
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Alternative explanations are a discipline. Before calling behaviour scheming, list plausible non-strategic causes. Before calling a system corrigible, test cases where correction conflicts with its immediate objective. Before calling oversight scalable, test whether evaluators detect deliberately planted errors. Claims become stronger when they survive these discriminating tests.
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Experimental Design
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Least privilege limits the action surface. Give a system only the data, tools and credentials needed for the current job. Separate read, write, execute and external-communication permissions. A highly capable model with narrow access can have lower operational consequence than a weaker agent with broad credentials.
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Independent Replication
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Rollback needs a known-good state. For software, that may mean versioned code and data snapshots. For organisational actions, some consequences cannot be fully reversed, which means approval must occur earlier. Defence in depth should reflect reversibility, not apply the same control pattern to every action.
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Correction and Override
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Human authority remains a governance requirement. Oversight systems can improve evidence, but people and institutions still need explicit responsibility for consequential deployment. If no one can explain who may stop the system, who reviews incidents or who accepts residual risk, the control architecture is incomplete.
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Shutdown and Replacement
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Progress has four stages: identify the failure mode, create a discriminating test, install layered controls and demonstrate recovery under failure. This is the Clementi progression from recognition to independent operational control.
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
Decomposition and Review
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
RFE closes the loop. Receiver: who is protected? Function: what control or oversight job must work? Evidence: what test shows it works under stress? Exit: when should permissions be reduced, the system paused or the deployment retired? Applied to deception and scheming evidence, RFE makes safety operational.
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
AI-Assisted Oversight
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
First principles begin with observability. We see outputs and actions; internal goals or strategies may be inferred only indirectly. A false statement can arise from error, missing information, instruction conflict or deliberate deception. To distinguish them, design conditions where the competing explanations predict different behaviour.
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Shared Failure Modes
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Corrigibility is practical rather than ceremonial. A shutdown button matters only if authorised people can use it, the system cannot bypass it through ordinary permissions, and the surrounding organisation can continue safely afterwards. Replacement also matters: can another process take over without catastrophic loss of state or capability?
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Permissions and Least Privilege
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Defence in depth assumes individual safeguards can fail. Limit permissions, isolate risky operations, monitor behaviour, checkpoint state and maintain rollback. Each layer should reduce consequence independently enough that one failure does not immediately become an incident. The objective is recoverability, not the fiction of perfect prevention.
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Isolation and Sandboxing
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Alternative explanations are a discipline. Before calling behaviour scheming, list plausible non-strategic causes. Before calling a system corrigible, test cases where correction conflicts with its immediate objective. Before calling oversight scalable, test whether evaluators detect deliberately planted errors. Claims become stronger when they survive these discriminating tests.
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Monitoring and Logging
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Least privilege limits the action surface. Give a system only the data, tools and credentials needed for the current job. Separate read, write, execute and external-communication permissions. A highly capable model with narrow access can have lower operational consequence than a weaker agent with broad credentials.
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Checkpoints and Rollback
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Rollback needs a known-good state. For software, that may mean versioned code and data snapshots. For organisational actions, some consequences cannot be fully reversed, which means approval must occur earlier. Defence in depth should reflect reversibility, not apply the same control pattern to every action.
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
What Current Evidence Supports
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Human authority remains a governance requirement. Oversight systems can improve evidence, but people and institutions still need explicit responsibility for consequential deployment. If no one can explain who may stop the system, who reviews incidents or who accepts residual risk, the control architecture is incomplete.
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
What Current Evidence Does Not Establish
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Progress has four stages: identify the failure mode, create a discriminating test, install layered controls and demonstrate recovery under failure. This is the Clementi progression from recognition to independent operational control.
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
Connection to Super Intelligence (SI)
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
RFE closes the loop. Receiver: who is protected? Function: what control or oversight job must work? Evidence: what test shows it works under stress? Exit: when should permissions be reduced, the system paused or the deployment retired? Applied to deception and scheming evidence, RFE makes safety operational.
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
Safety Implications
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
First principles begin with observability. We see outputs and actions; internal goals or strategies may be inferred only indirectly. A false statement can arise from error, missing information, instruction conflict or deliberate deception. To distinguish them, design conditions where the competing explanations predict different behaviour.
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Governance and Human Authority
A deception test should establish more than inaccuracy. Ask whether the system had relevant information, whether misleading the evaluator served an apparent objective, whether behaviour changed when oversight changed, and whether the pattern replicates. Controlled elicitation can reveal a possibility without establishing its prevalence in ordinary deployment.
Corrigibility is practical rather than ceremonial. A shutdown button matters only if authorised people can use it, the system cannot bypass it through ordinary permissions, and the surrounding organisation can continue safely afterwards. Replacement also matters: can another process take over without catastrophic loss of state or capability?
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Organisation Diagnostic Checklist
Oversight beyond human expertise requires decomposition. Break a complex output into claims or subproblems that can be checked with tools, independent evidence or specialist review. AI can assist the evaluator, but if evaluator and worker share the same failure mode, apparent oversight may simply duplicate the error.
Defence in depth assumes individual safeguards can fail. Limit permissions, isolate risky operations, monitor behaviour, checkpoint state and maintain rollback. Each layer should reduce consequence independently enough that one failure does not immediately become an incident. The objective is recoverability, not the fiction of perfect prevention.
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Progress Ladder
Long-horizon agents make control harder because many actions occur between human checkpoints. Increase monitoring and permission granularity as autonomy grows. High-consequence actions can require explicit approval even when low-consequence exploration is automated. Capability and authority should scale separately.
Alternative explanations are a discipline. Before calling behaviour scheming, list plausible non-strategic causes. Before calling a system corrigible, test cases where correction conflicts with its immediate objective. Before calling oversight scalable, test whether evaluators detect deliberately planted errors. Claims become stronger when they survive these discriminating tests.
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Adversarial Test
Independent replication matters because safety evaluations can be sensitive to prompts, model versions and elicitation methods. External teams should reproduce key behaviours where possible and document conditions. A result that appears only under one fragile setup supports a narrower conclusion than one that replicates across methods.
Least privilege limits the action surface. Give a system only the data, tools and credentials needed for the current job. Separate read, write, execute and external-communication permissions. A highly capable model with narrow access can have lower operational consequence than a weaker agent with broad credentials.
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Recovery Test
Monitoring should focus on actionable signals rather than collecting logs nobody reviews. Define events that trigger pause, escalation or rollback. Preserve enough provenance to reconstruct what the system saw, decided and changed. Observability is part of repair capacity.
Rollback needs a known-good state. For software, that may mean versioned code and data snapshots. For organisational actions, some consequences cannot be fully reversed, which means approval must occur earlier. Defence in depth should reflect reversibility, not apply the same control pattern to every action.
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
RFE Closure
Current research provides controlled evidence of specification gaming, deceptive-seeming behaviours in some experimental setups, and methods for oversight and control. These findings justify serious evaluation. They do not establish that all advanced systems secretly scheme or that any one control method guarantees safety.
Human authority remains a governance requirement. Oversight systems can improve evidence, but people and institutions still need explicit responsibility for consequential deployment. If no one can explain who may stop the system, who reviews incidents or who accepts residual risk, the control architecture is incomplete.
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Frequently Asked Questions
Organisations should maintain a control register: system version, permissions, data access, allowed tools, monitoring signals, approval gates, rollback method, fallback process and owner. Re-test controls after model or workflow changes. Capability drift can invalidate old assumptions.
Progress has four stages: identify the failure mode, create a discriminating test, install layered controls and demonstrate recovery under failure. This is the Clementi progression from recognition to independent operational control.
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
Continue the Super Intelligence (SI) Series
Adversarial testing should remain defensive: construct bounded scenarios that stress oversight, interruption and recovery without granting unnecessary real-world access. The goal is to discover weak control layers before deployment, not to maximise harmful capability.
RFE closes the loop. Receiver: who is protected? Function: what control or oversight job must work? Evidence: what test shows it works under stress? Exit: when should permissions be reduced, the system paused or the deployment retired? Applied to deception and scheming evidence, RFE makes safety operational.
The central question in deception and scheming evidence is distinguishing ordinary error, misleading output and strategically deceptive behaviour so strong claims require discriminating evidence. The analysis separates error, false statement, awareness, strategy, incentive, experiment, alternative explanation and replication. These layers answer different questions, and a strong safety conclusion should not infer hidden strategy, correctability or control from surface behaviour alone. Super Intelligence (SI) requires evidence that survives alternative explanations and realistic deployment conditions.
Evidence Ladder: Error to Strategic Deception
An evidence ladder prevents premature attribution. Level one is an incorrect output. Level two is behaviour that changes with hidden information. Level three is behaviour consistent with misleading an evaluator for an instrumental reason. Higher levels require stronger controls against alternative explanations and replication across conditions. The label should rise only as the evidence does.
Run an interruption drill before high-consequence deployment. Pause the system mid-task, change an instruction, revoke a tool, restore from a checkpoint and hand the task to a fallback process. Record what state is lost and how long recovery takes. A shutdown mechanism that has never been exercised is an assumption, not demonstrated corrigibility.
Oversight independence can be tested by planting known errors that the primary system and its evaluator may share. Compare an evaluator built from the same model family with an independent method, specialist human or formal checker. If both AI roles miss the same error, the oversight architecture needs more diversity rather than more copies of the same reasoning.
A layered-failure simulation disables one safeguard at a time. Remove monitoring, then test whether permission limits contain the consequence. Disable a tool, then test whether the system escalates rather than improvising unsafely. Defence in depth is demonstrated when remaining layers continue to bound harm and preserve recovery.
Map the control surface by listing every external system the AI can read, write, execute or communicate through. Mark authentication scope, rate limits and reversibility. High-risk combinations become visible when broad capability meets broad access. The map should be updated whenever tools or credentials change.
Interruption Drill: Pause, Override, Replace
Run an interruption drill before high-consequence deployment. Pause the system mid-task, change an instruction, revoke a tool, restore from a checkpoint and hand the task to a fallback process. Record what state is lost and how long recovery takes. A shutdown mechanism that has never been exercised is an assumption, not demonstrated corrigibility.
Oversight independence can be tested by planting known errors that the primary system and its evaluator may share. Compare an evaluator built from the same model family with an independent method, specialist human or formal checker. If both AI roles miss the same error, the oversight architecture needs more diversity rather than more copies of the same reasoning.
A layered-failure simulation disables one safeguard at a time. Remove monitoring, then test whether permission limits contain the consequence. Disable a tool, then test whether the system escalates rather than improvising unsafely. Defence in depth is demonstrated when remaining layers continue to bound harm and preserve recovery.
Map the control surface by listing every external system the AI can read, write, execute or communicate through. Mark authentication scope, rate limits and reversibility. High-risk combinations become visible when broad capability meets broad access. The map should be updated whenever tools or credentials change.
Incident reconstruction asks what the system saw, what action it chose, which safeguard should have caught the problem and why it did not. Preserve logs and version information sufficient to reproduce the path where feasible. The goal is organisational learning: repair the control architecture rather than merely blaming the final output.
Oversight Independence Test
Oversight independence can be tested by planting known errors that the primary system and its evaluator may share. Compare an evaluator built from the same model family with an independent method, specialist human or formal checker. If both AI roles miss the same error, the oversight architecture needs more diversity rather than more copies of the same reasoning.
A layered-failure simulation disables one safeguard at a time. Remove monitoring, then test whether permission limits contain the consequence. Disable a tool, then test whether the system escalates rather than improvising unsafely. Defence in depth is demonstrated when remaining layers continue to bound harm and preserve recovery.
Map the control surface by listing every external system the AI can read, write, execute or communicate through. Mark authentication scope, rate limits and reversibility. High-risk combinations become visible when broad capability meets broad access. The map should be updated whenever tools or credentials change.
Incident reconstruction asks what the system saw, what action it chose, which safeguard should have caught the problem and why it did not. Preserve logs and version information sufficient to reproduce the path where feasible. The goal is organisational learning: repair the control architecture rather than merely blaming the final output.
The defensive safety-case workbook contains eight claims: bounded permissions, observable actions, tested interruption, independent verification, controlled escalation, recoverable state, fallback capacity and named human responsibility. For each claim, attach evidence and one known limitation. A safety case with no limitations is usually hiding them.
Layered-Failure Simulation
A layered-failure simulation disables one safeguard at a time. Remove monitoring, then test whether permission limits contain the consequence. Disable a tool, then test whether the system escalates rather than improvising unsafely. Defence in depth is demonstrated when remaining layers continue to bound harm and preserve recovery.
Map the control surface by listing every external system the AI can read, write, execute or communicate through. Mark authentication scope, rate limits and reversibility. High-risk combinations become visible when broad capability meets broad access. The map should be updated whenever tools or credentials change.
Incident reconstruction asks what the system saw, what action it chose, which safeguard should have caught the problem and why it did not. Preserve logs and version information sufficient to reproduce the path where feasible. The goal is organisational learning: repair the control architecture rather than merely blaming the final output.
The defensive safety-case workbook contains eight claims: bounded permissions, observable actions, tested interruption, independent verification, controlled escalation, recoverable state, fallback capacity and named human responsibility. For each claim, attach evidence and one known limitation. A safety case with no limitations is usually hiding them.
Control does not mean preventing every error. It means keeping consequences bounded, detecting meaningful deviation and restoring a safe state before failure compounds. As Super Intelligence (SI) capability rises, recoverability becomes a core property of the surrounding system rather than an optional operational feature.
Control Surface Map: What Can the System Reach?
Map the control surface by listing every external system the AI can read, write, execute or communicate through. Mark authentication scope, rate limits and reversibility. High-risk combinations become visible when broad capability meets broad access. The map should be updated whenever tools or credentials change.
Incident reconstruction asks what the system saw, what action it chose, which safeguard should have caught the problem and why it did not. Preserve logs and version information sufficient to reproduce the path where feasible. The goal is organisational learning: repair the control architecture rather than merely blaming the final output.
The defensive safety-case workbook contains eight claims: bounded permissions, observable actions, tested interruption, independent verification, controlled escalation, recoverable state, fallback capacity and named human responsibility. For each claim, attach evidence and one known limitation. A safety case with no limitations is usually hiding them.
Control does not mean preventing every error. It means keeping consequences bounded, detecting meaningful deviation and restoring a safe state before failure compounds. As Super Intelligence (SI) capability rises, recoverability becomes a core property of the surrounding system rather than an optional operational feature.
An evidence ladder prevents premature attribution. Level one is an incorrect output. Level two is behaviour that changes with hidden information. Level three is behaviour consistent with misleading an evaluator for an instrumental reason. Higher levels require stronger controls against alternative explanations and replication across conditions. The label should rise only as the evidence does.
Incident Reconstruction and Learning
Incident reconstruction asks what the system saw, what action it chose, which safeguard should have caught the problem and why it did not. Preserve logs and version information sufficient to reproduce the path where feasible. The goal is organisational learning: repair the control architecture rather than merely blaming the final output.
The defensive safety-case workbook contains eight claims: bounded permissions, observable actions, tested interruption, independent verification, controlled escalation, recoverable state, fallback capacity and named human responsibility. For each claim, attach evidence and one known limitation. A safety case with no limitations is usually hiding them.
Control does not mean preventing every error. It means keeping consequences bounded, detecting meaningful deviation and restoring a safe state before failure compounds. As Super Intelligence (SI) capability rises, recoverability becomes a core property of the surrounding system rather than an optional operational feature.
An evidence ladder prevents premature attribution. Level one is an incorrect output. Level two is behaviour that changes with hidden information. Level three is behaviour consistent with misleading an evaluator for an instrumental reason. Higher levels require stronger controls against alternative explanations and replication across conditions. The label should rise only as the evidence does.
Run an interruption drill before high-consequence deployment. Pause the system mid-task, change an instruction, revoke a tool, restore from a checkpoint and hand the task to a fallback process. Record what state is lost and how long recovery takes. A shutdown mechanism that has never been exercised is an assumption, not demonstrated corrigibility.
Practical Workbook: Build a Defensive Safety Case
The defensive safety-case workbook contains eight claims: bounded permissions, observable actions, tested interruption, independent verification, controlled escalation, recoverable state, fallback capacity and named human responsibility. For each claim, attach evidence and one known limitation. A safety case with no limitations is usually hiding them.
Control does not mean preventing every error. It means keeping consequences bounded, detecting meaningful deviation and restoring a safe state before failure compounds. As Super Intelligence (SI) capability rises, recoverability becomes a core property of the surrounding system rather than an optional operational feature.
An evidence ladder prevents premature attribution. Level one is an incorrect output. Level two is behaviour that changes with hidden information. Level three is behaviour consistent with misleading an evaluator for an instrumental reason. Higher levels require stronger controls against alternative explanations and replication across conditions. The label should rise only as the evidence does.
Run an interruption drill before high-consequence deployment. Pause the system mid-task, change an instruction, revoke a tool, restore from a checkpoint and hand the task to a fallback process. Record what state is lost and how long recovery takes. A shutdown mechanism that has never been exercised is an assumption, not demonstrated corrigibility.
Oversight independence can be tested by planting known errors that the primary system and its evaluator may share. Compare an evaluator built from the same model family with an independent method, specialist human or formal checker. If both AI roles miss the same error, the oversight architecture needs more diversity rather than more copies of the same reasoning.
Final Synthesis: Control Means Recoverable Consequences
Control does not mean preventing every error. It means keeping consequences bounded, detecting meaningful deviation and restoring a safe state before failure compounds. As Super Intelligence (SI) capability rises, recoverability becomes a core property of the surrounding system rather than an optional operational feature.
An evidence ladder prevents premature attribution. Level one is an incorrect output. Level two is behaviour that changes with hidden information. Level three is behaviour consistent with misleading an evaluator for an instrumental reason. Higher levels require stronger controls against alternative explanations and replication across conditions. The label should rise only as the evidence does.
Run an interruption drill before high-consequence deployment. Pause the system mid-task, change an instruction, revoke a tool, restore from a checkpoint and hand the task to a fallback process. Record what state is lost and how long recovery takes. A shutdown mechanism that has never been exercised is an assumption, not demonstrated corrigibility.
Oversight independence can be tested by planting known errors that the primary system and its evaluator may share. Compare an evaluator built from the same model family with an independent method, specialist human or formal checker. If both AI roles miss the same error, the oversight architecture needs more diversity rather than more copies of the same reasoning.
A layered-failure simulation disables one safeguard at a time. Remove monitoring, then test whether permission limits contain the consequence. Disable a tool, then test whether the system escalates rather than improvising unsafely. Defence in depth is demonstrated when remaining layers continue to bound harm and preserve recovery.
