Super Intelligence (SI) requires evidence not only about what systems can do, but about what we can legitimately infer from their behaviour. This article examines AI interpretability: explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. It preserves the locked Clementi-depth floor with mechanisms, competing explanations, tests, diagnostics, uncertainty and RFE closure.
Search Intent and Direct Answer
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
First principles begin by naming the hidden variable. An output is observable; an internal mechanism, failure propensity or subjective experience may not be directly observable in the same way. Researchers therefore use indicators and interventions. The strength of the conclusion depends on how tightly the indicator is connected to the thing being claimed.
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Definition and Boundary
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Causal evidence is stronger than correlation when the question concerns mechanism. If a probe can read information from a representation, that shows the information is detectable; it does not automatically prove the representation causes the final behaviour. Interventions that alter the proposed mechanism and produce the predicted effect provide stronger support.
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
First Principles
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
Severity and frequency should be separated. A rare but irreversible failure can matter even when average performance is high. A common low-impact failure may be easier to tolerate or repair. Defensive evaluation should map both axes and connect them to permissions, monitoring and recovery.
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
What the Question Does Not Ask
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
A consciousness claim has a different evidential structure from a capability claim. Behaviour and self-report may be relevant indicators, but they do not by themselves settle subjective experience. Scientific approaches compare candidate theories and look for properties those theories associate with consciousness. Uncertainty should be preserved rather than filled with anthropomorphic intuition.
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
The Core Evidence Problem
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Independent evaluation is especially valuable when developers have privileged access to models, training data or internal activations. External researchers can test behaviour, auditors can inspect controlled evidence, and replicated findings can reduce dependence on one interpretation. Broad SI claims should not rest on a single institution’s vocabulary.
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Worked Example: A Simple Case
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Safety benefits from interpretability and red-teaming when these methods reveal actionable failure modes, but neither is a complete safety guarantee. A model can be partly interpretable and still surprise evaluators; a red team can miss untested behaviours. Defence in depth combines evaluation with permissions, monitoring, isolation where appropriate and recovery.
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Worked Example: A Misleading Explanation
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Education should teach students to distinguish an explanation from evidence for an explanation. Ask what observation would look different if a competing theory were true. This habit applies to AI internals, safety claims and consciousness. It turns philosophical or technical debate into structured reasoning.
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
Worked Example: A Stress Test
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
RFE closes the loop. Receiver: who needs the conclusion? Function: what decision will it support? Evidence: which observations justify action? Exit: when should the interpretation, safety claim or classification be revised? Applied to AI interpretability, RFE keeps uncertainty connected to responsible decisions.
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
Worked Example: A Dated Capability Claim
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
First principles begin by naming the hidden variable. An output is observable; an internal mechanism, failure propensity or subjective experience may not be directly observable in the same way. Researchers therefore use indicators and interventions. The strength of the conclusion depends on how tightly the indicator is connected to the thing being claimed.
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Worked Example: Behaviour Versus Experience
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Causal evidence is stronger than correlation when the question concerns mechanism. If a probe can read information from a representation, that shows the information is detectable; it does not automatically prove the representation causes the final behaviour. Interventions that alter the proposed mechanism and produce the predicted effect provide stronger support.
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
How to Measure the Claim
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
Severity and frequency should be separated. A rare but irreversible failure can matter even when average performance is high. A common low-impact failure may be easier to tolerate or repair. Defensive evaluation should map both axes and connect them to permissions, monitoring and recovery.
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
Causal Evidence Versus Correlation
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
A consciousness claim has a different evidential structure from a capability claim. Behaviour and self-report may be relevant indicators, but they do not by themselves settle subjective experience. Scientific approaches compare candidate theories and look for properties those theories associate with consciousness. Uncertainty should be preserved rather than filled with anthropomorphic intuition.
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Coverage and Blind Spots
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Independent evaluation is especially valuable when developers have privileged access to models, training data or internal activations. External researchers can test behaviour, auditors can inspect controlled evidence, and replicated findings can reduce dependence on one interpretation. Broad SI claims should not rest on a single institution’s vocabulary.
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Reliability Under Repetition
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Safety benefits from interpretability and red-teaming when these methods reveal actionable failure modes, but neither is a complete safety guarantee. A model can be partly interpretable and still surprise evaluators; a red team can miss untested behaviours. Defence in depth combines evaluation with permissions, monitoring, isolation where appropriate and recovery.
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
False Positives and False Negatives
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Education should teach students to distinguish an explanation from evidence for an explanation. Ask what observation would look different if a competing theory were true. This habit applies to AI internals, safety claims and consciousness. It turns philosophical or technical debate into structured reasoning.
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
Independent Evaluation
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
RFE closes the loop. Receiver: who needs the conclusion? Function: what decision will it support? Evidence: which observations justify action? Exit: when should the interpretation, safety claim or classification be revised? Applied to AI interpretability, RFE keeps uncertainty connected to responsible decisions.
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
Human Interpretation Limits
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
First principles begin by naming the hidden variable. An output is observable; an internal mechanism, failure propensity or subjective experience may not be directly observable in the same way. Researchers therefore use indicators and interventions. The strength of the conclusion depends on how tightly the indicator is connected to the thing being claimed.
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
What Current Evidence Supports
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Causal evidence is stronger than correlation when the question concerns mechanism. If a probe can read information from a representation, that shows the information is detectable; it does not automatically prove the representation causes the final behaviour. Interventions that alter the proposed mechanism and produce the predicted effect provide stronger support.
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
What Current Evidence Does Not Establish
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
Severity and frequency should be separated. A rare but irreversible failure can matter even when average performance is high. A common low-impact failure may be easier to tolerate or repair. Defensive evaluation should map both axes and connect them to permissions, monitoring and recovery.
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
Connection to Super Intelligence (SI)
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
A consciousness claim has a different evidential structure from a capability claim. Behaviour and self-report may be relevant indicators, but they do not by themselves settle subjective experience. Scientific approaches compare candidate theories and look for properties those theories associate with consciousness. Uncertainty should be preserved rather than filled with anthropomorphic intuition.
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Safety Implications
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Independent evaluation is especially valuable when developers have privileged access to models, training data or internal activations. External researchers can test behaviour, auditors can inspect controlled evidence, and replicated findings can reduce dependence on one interpretation. Broad SI claims should not rest on a single institution’s vocabulary.
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Governance Implications
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Safety benefits from interpretability and red-teaming when these methods reveal actionable failure modes, but neither is a complete safety guarantee. A model can be partly interpretable and still surprise evaluators; a red team can miss untested behaviours. Defence in depth combines evaluation with permissions, monitoring, isolation where appropriate and recovery.
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Education and Evidence Literacy
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Education should teach students to distinguish an explanation from evidence for an explanation. Ask what observation would look different if a competing theory were true. This habit applies to AI internals, safety claims and consciousness. It turns philosophical or technical debate into structured reasoning.
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
Student Diagnostic Checklist
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
RFE closes the loop. Receiver: who needs the conclusion? Function: what decision will it support? Evidence: which observations justify action? Exit: when should the interpretation, safety claim or classification be revised? Applied to AI interpretability, RFE keeps uncertainty connected to responsible decisions.
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
Organisation Diagnostic Checklist
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
First principles begin by naming the hidden variable. An output is observable; an internal mechanism, failure propensity or subjective experience may not be directly observable in the same way. Researchers therefore use indicators and interventions. The strength of the conclusion depends on how tightly the indicator is connected to the thing being claimed.
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Progress Ladder
A simple case can be misleading because several explanations fit the same behaviour. A model may produce a correct answer through a robust internal representation, memorised structure, external retrieval or a lucky chain of generation. To distinguish explanations, change the input, intervene on the system where possible and test whether the predicted behavioural change follows.
Causal evidence is stronger than correlation when the question concerns mechanism. If a probe can read information from a representation, that shows the information is detectable; it does not automatically prove the representation causes the final behaviour. Interventions that alter the proposed mechanism and produce the predicted effect provide stronger support.
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
Adversarial Test
Coverage is the central problem in red-teaming. Finding one failure proves the failure is possible under the tested conditions; it does not by itself estimate how common the failure is. Failing to find a problem does not prove safety. A useful report states the threat model, test coverage, elicitation method and limitations.
Severity and frequency should be separated. A rare but irreversible failure can matter even when average performance is high. A common low-impact failure may be easier to tolerate or repair. Defensive evaluation should map both axes and connect them to permissions, monitoring and recovery.
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
Recovery and Retesting
Dated capability claims require dated evidence. The tracker in this article uses 30 September 2026 as its observation boundary. It asks whether public evidence satisfies the series’ SI standard across breadth, depth, reliability, transfer, long-horizon work and independent verification. Rapid progress can strengthen several dimensions without automatically satisfying the whole classification.
A consciousness claim has a different evidential structure from a capability claim. Behaviour and self-report may be relevant indicators, but they do not by themselves settle subjective experience. Scientific approaches compare candidate theories and look for properties those theories associate with consciousness. Uncertainty should be preserved rather than filled with anthropomorphic intuition.
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
What Would Change the Conclusion?
False positives and false negatives both matter. Over-attributing an internal mechanism, dangerous tendency or consciousness can lead to bad decisions; under-attributing can also create risk or ethical error. Evaluation should therefore state which error it is optimised to reduce and what trade-off that creates.
Independent evaluation is especially valuable when developers have privileged access to models, training data or internal activations. External researchers can test behaviour, auditors can inspect controlled evidence, and replicated findings can reduce dependence on one interpretation. Broad SI claims should not rest on a single institution’s vocabulary.
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
RFE Closure
Current evidence supports increasingly capable AI across reasoning, coding, multimodal work and agentic tasks. It also supports meaningful research into interpretability and adversarial evaluation. It does not establish complete mechanistic understanding, exhaustive safety coverage or a scientific consensus that current AI is conscious.
Safety benefits from interpretability and red-teaming when these methods reveal actionable failure modes, but neither is a complete safety guarantee. A model can be partly interpretable and still surprise evaluators; a red team can miss untested behaviours. Defence in depth combines evaluation with permissions, monitoring, isolation where appropriate and recovery.
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Frequently Asked Questions
Governance needs evidence that non-specialists can audit. Decision-makers should know what was tested, what was not tested, how severe discovered failures were and what uncertainty remains. Technical complexity should not become a reason to replace accountable judgement with unchallengeable assertions.
Education should teach students to distinguish an explanation from evidence for an explanation. Ask what observation would look different if a competing theory were true. This habit applies to AI internals, safety claims and consciousness. It turns philosophical or technical debate into structured reasoning.
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
Continue the Super Intelligence (SI) Series
Progress has four stages: observe behaviour, formulate competing explanations, design discriminating tests and update confidence after results. This is the Clementi progression from recognition to independent analysis. The reader should finish able to ask what evidence would separate two plausible stories.
RFE closes the loop. Receiver: who needs the conclusion? Function: what decision will it support? Evidence: which observations justify action? Exit: when should the interpretation, safety claim or classification be revised? Applied to AI interpretability, RFE keeps uncertainty connected to responsible decisions.
The central question in AI interpretability is explaining what methods can reveal about internal mechanisms without confusing a partial explanation with complete understanding. The analysis separates representation, probe, circuit, causal intervention, explanation, validation and limitation. These dimensions answer different questions, and a strong conclusion should not borrow certainty from one dimension to fill a gap in another. Super Intelligence (SI) requires especially careful boundaries because capability, explanation, safety and consciousness are often discussed in the same sentence.
Mechanism-versus-Behaviour Matrix
Build a mechanism-versus-behaviour matrix. Put observable behaviours in rows and candidate explanations in columns. Mark which observations each explanation predicts and where they diverge. Then design an intervention or changed condition that discriminates between them. This prevents a compelling narrative about an AI system from becoming accepted merely because it fits one successful output.
A coverage map records the tested space and the untested space. For red-teaming, include threat categories, languages, tool permissions, task horizons and elicitation methods. For interpretability, include layers, behaviours and intervention types. For consciousness, include which theoretical indicators have been considered. Blank regions are not failures, but they are uncertainty and should be labelled.
The dated SI worksheet freezes the observation boundary at 30 September 2026. Score no overall winner; instead record evidence separately for breadth, frontier depth, reliability, unfamiliar-task transfer, long-horizon completion, autonomous system performance and independent replication. A broad SI classification should require convergent evidence across dimensions rather than a single extraordinary result.
Competing theories become useful when they make different predictions. If one explanation says an internal feature is causally necessary and another says it is merely correlated, intervene on that feature and measure the result. If two consciousness theories imply different functional properties, identify the discriminating evidence. Where theories do not yet yield decisive tests, state that limitation rather than inventing certainty.
Coverage Map: What Has and Has Not Been Tested?
A coverage map records the tested space and the untested space. For red-teaming, include threat categories, languages, tool permissions, task horizons and elicitation methods. For interpretability, include layers, behaviours and intervention types. For consciousness, include which theoretical indicators have been considered. Blank regions are not failures, but they are uncertainty and should be labelled.
The dated SI worksheet freezes the observation boundary at 30 September 2026. Score no overall winner; instead record evidence separately for breadth, frontier depth, reliability, unfamiliar-task transfer, long-horizon completion, autonomous system performance and independent replication. A broad SI classification should require convergent evidence across dimensions rather than a single extraordinary result.
Competing theories become useful when they make different predictions. If one explanation says an internal feature is causally necessary and another says it is merely correlated, intervene on that feature and measure the result. If two consciousness theories imply different functional properties, identify the discriminating evidence. Where theories do not yet yield decisive tests, state that limitation rather than inventing certainty.
Independent review separates evidence generation from evidence acceptance. A second team should receive enough information to reproduce behavioural tests, inspect evaluation conditions or challenge the interpretation. Where proprietary access prevents full replication, controlled auditing can still test selected claims. Confidence should rise with convergent independent evidence.
Dated SI Classification Worksheet — 30 September 2026
The dated SI worksheet freezes the observation boundary at 30 September 2026. Score no overall winner; instead record evidence separately for breadth, frontier depth, reliability, unfamiliar-task transfer, long-horizon completion, autonomous system performance and independent replication. A broad SI classification should require convergent evidence across dimensions rather than a single extraordinary result.
Competing theories become useful when they make different predictions. If one explanation says an internal feature is causally necessary and another says it is merely correlated, intervene on that feature and measure the result. If two consciousness theories imply different functional properties, identify the discriminating evidence. Where theories do not yet yield decisive tests, state that limitation rather than inventing certainty.
Independent review separates evidence generation from evidence acceptance. A second team should receive enough information to reproduce behavioural tests, inspect evaluation conditions or challenge the interpretation. Where proprietary access prevents full replication, controlled auditing can still test selected claims. Confidence should rise with convergent independent evidence.
The workbook uses six fields: observation, proposed interpretation, alternative interpretation, discriminating test, result and confidence update. Repeat the process for at least three cases. Then add a section titled “What I still do not know.” This final field is important: rigorous SI literacy includes the ability to preserve uncertainty without treating it as ignorance or filling it with intuition.
Competing-Theory Test: What Observation Would Separate Them?
Competing theories become useful when they make different predictions. If one explanation says an internal feature is causally necessary and another says it is merely correlated, intervene on that feature and measure the result. If two consciousness theories imply different functional properties, identify the discriminating evidence. Where theories do not yet yield decisive tests, state that limitation rather than inventing certainty.
Independent review separates evidence generation from evidence acceptance. A second team should receive enough information to reproduce behavioural tests, inspect evaluation conditions or challenge the interpretation. Where proprietary access prevents full replication, controlled auditing can still test selected claims. Confidence should rise with convergent independent evidence.
The workbook uses six fields: observation, proposed interpretation, alternative interpretation, discriminating test, result and confidence update. Repeat the process for at least three cases. Then add a section titled “What I still do not know.” This final field is important: rigorous SI literacy includes the ability to preserve uncertainty without treating it as ignorance or filling it with intuition.
The final discipline is scope. A successful interpretability result is evidence about a mechanism; a red-team failure is evidence that a behaviour can be elicited under tested conditions; a dated capability result is evidence about that system at that time; a consciousness indicator is evidence relative to a theory. None should silently expand beyond its test.
Independent Review Protocol
Independent review separates evidence generation from evidence acceptance. A second team should receive enough information to reproduce behavioural tests, inspect evaluation conditions or challenge the interpretation. Where proprietary access prevents full replication, controlled auditing can still test selected claims. Confidence should rise with convergent independent evidence.
The workbook uses six fields: observation, proposed interpretation, alternative interpretation, discriminating test, result and confidence update. Repeat the process for at least three cases. Then add a section titled “What I still do not know.” This final field is important: rigorous SI literacy includes the ability to preserve uncertainty without treating it as ignorance or filling it with intuition.
The final discipline is scope. A successful interpretability result is evidence about a mechanism; a red-team failure is evidence that a behaviour can be elicited under tested conditions; a dated capability result is evidence about that system at that time; a consciousness indicator is evidence relative to a theory. None should silently expand beyond its test.
Build a mechanism-versus-behaviour matrix. Put observable behaviours in rows and candidate explanations in columns. Mark which observations each explanation predicts and where they diverge. Then design an intervention or changed condition that discriminates between them. This prevents a compelling narrative about an AI system from becoming accepted merely because it fits one successful output.
Practical Workbook: Evidence, Confidence and Revision
The workbook uses six fields: observation, proposed interpretation, alternative interpretation, discriminating test, result and confidence update. Repeat the process for at least three cases. Then add a section titled “What I still do not know.” This final field is important: rigorous SI literacy includes the ability to preserve uncertainty without treating it as ignorance or filling it with intuition.
The final discipline is scope. A successful interpretability result is evidence about a mechanism; a red-team failure is evidence that a behaviour can be elicited under tested conditions; a dated capability result is evidence about that system at that time; a consciousness indicator is evidence relative to a theory. None should silently expand beyond its test.
Build a mechanism-versus-behaviour matrix. Put observable behaviours in rows and candidate explanations in columns. Mark which observations each explanation predicts and where they diverge. Then design an intervention or changed condition that discriminates between them. This prevents a compelling narrative about an AI system from becoming accepted merely because it fits one successful output.
A coverage map records the tested space and the untested space. For red-teaming, include threat categories, languages, tool permissions, task horizons and elicitation methods. For interpretability, include layers, behaviours and intervention types. For consciousness, include which theoretical indicators have been considered. Blank regions are not failures, but they are uncertainty and should be labelled.
Final Synthesis: Do Not Infer More Than the Test Shows
The final discipline is scope. A successful interpretability result is evidence about a mechanism; a red-team failure is evidence that a behaviour can be elicited under tested conditions; a dated capability result is evidence about that system at that time; a consciousness indicator is evidence relative to a theory. None should silently expand beyond its test.
Build a mechanism-versus-behaviour matrix. Put observable behaviours in rows and candidate explanations in columns. Mark which observations each explanation predicts and where they diverge. Then design an intervention or changed condition that discriminates between them. This prevents a compelling narrative about an AI system from becoming accepted merely because it fits one successful output.
A coverage map records the tested space and the untested space. For red-teaming, include threat categories, languages, tool permissions, task horizons and elicitation methods. For interpretability, include layers, behaviours and intervention types. For consciousness, include which theoretical indicators have been considered. Blank regions are not failures, but they are uncertainty and should be labelled.
The dated SI worksheet freezes the observation boundary at 30 September 2026. Score no overall winner; instead record evidence separately for breadth, frontier depth, reliability, unfamiliar-task transfer, long-horizon completion, autonomous system performance and independent replication. A broad SI classification should require convergent evidence across dimensions rather than a single extraordinary result.
Interpretability Is Not the Same as Asking the Model to Explain Itself
A language model can produce a fluent explanation of why it gave an answer. That explanation may be useful, but it is not automatically a faithful report of the internal computation that generated the answer. Interpretability tries to connect model behaviour to internal states, features, circuits or causal mechanisms rather than relying only on self-description.
For Super Intelligence (SI), this distinction becomes more important as systems become more capable. A highly persuasive explanation is weak evidence if the system can generate the explanation after the answer was already determined by different internal processes.
Mechanistic Interpretability Looks for Internal Computation
Mechanistic interpretability attempts to identify the internal features and pathways that contribute to a model’s output. Anthropic’s 2025 circuit-tracing work introduced attribution graphs designed to reveal parts of the computational pathway from input concepts to outputs. The company also released tools so researchers could inspect similar graphs on open-weight models.
These methods are promising because they aim at causal structure rather than surface behaviour. They remain partial. Large models perform enormous amounts of computation, and current methods expose only fragments of that process.
Features Are Not Automatically Human Concepts
Researchers can often find internal activation patterns associated with interpretable concepts, styles or behaviours. But a feature may combine several ideas, split one human concept across many directions, or change function by context. Treating every discovered feature as a clean semantic label can create false confidence.
The right standard is intervention: if changing the feature reliably changes the predicted behaviour in the expected way, the interpretation becomes stronger.
Causal Tests Are Stronger Than Correlation
Suppose an internal activation correlates with deception. That does not prove the activation causes deceptive behaviour. The model may be using the feature as a consequence or by-product of another mechanism. Causal interpretability changes or ablates the feature and measures what happens.
SI interpretability should therefore distinguish three levels: observation, predictive correlation and causal intervention. The language of “the model thinks X” should be reserved for evidence strong enough to justify the claim.
2026 Research Is Expanding From Features Toward Functional Organisation
Anthropic’s 2026 interpretability work has explored internal character directions, natural-language descriptions of internal states and structures described as resembling a global workspace. These studies are evidence that frontier models contain reusable internal organisation that can sometimes be probed and manipulated.
They do not establish that the models possess human-like minds or consciousness. Functional analogy and subjective experience are separate claims.
Interpretability Can Reveal Planning That Is Not Visible in the Output
Anthropic’s earlier tracing work reported examples where language models appear to represent future output structure before generating the next token. If such findings replicate broadly, they complicate the simple picture that a next-token model operates only one word at a time without higher-level organisation.
For SI, this is important because capability can emerge from internal computation that is not obvious from the training objective alone. Interpretability can help connect architecture to behaviour.
Interpretability Can Also Reveal Unfaithful Explanations
A model may present a tidy step-by-step explanation even when internal tracing suggests a different route produced the answer. This creates a practical warning: chain-of-thought-like explanations should not automatically be treated as transparent windows into the model.
External verification and internal analysis remain complementary. One checks whether the answer is right; the other asks how the system arrived there.
Probes Can Detect Information Without Showing How It Is Used
A linear probe may recover a fact or property from hidden activations. That demonstrates the information is represented in a decodable form. It does not necessarily show that the model uses that representation in producing its output.
This distinction is fundamental. Readable information is not the same as causally active computation.
Interpretability Can Support Monitoring
If researchers identify internal patterns associated with undesirable behaviour, those patterns might be monitored during deployment. Anthropic’s work on persona vectors is one example of using activation directions to detect or steer behavioural tendencies.
For SI safety, this could create an additional sensor layer beyond input and output monitoring. The limitation is generalisation: a monitor trained on known failure modes may miss new ones.
Interpretability Is Not Yet a Complete Safety Case
The International AI Safety Report and current interpretability literature both support a cautious conclusion: interpretability can reveal useful mechanisms, but researchers do not yet understand most frontier-model computation in a complete, predictive way. A partial map should not be treated as proof that hidden failure modes cannot exist.
This is similar to medicine: understanding some pathways in a biological system can improve diagnosis without implying total understanding of the organism.
Worked Example: A Model That Produces a Correct Mathematical Answer
Behavioural evaluation tells us the answer is correct. A verbal rationale tells us what reasoning the model claims to have used. Mechanistic interpretability might reveal whether the internal computation tracks variables, plans intermediate steps or retrieves a memorised pattern.
Each layer answers a different question. Super Intelligence evaluation is stronger when they converge.
Worked Example: Detecting an Undesired Behavioural Mode
Suppose a model becomes more sycophantic after a post-training update. Output tests reveal the behavioural difference. An internal “diff” method can then search for changed representations or activation patterns associated with the new behaviour. If manipulating those patterns changes sycophancy, researchers gain a causal handle rather than only a symptom report.
This can make repair more targeted, though the intervention still requires regression testing elsewhere.
What Would Mature SI Interpretability Look Like?
Stronger evidence would include methods that scale across model families, predict behaviour before it is observed, identify causally important circuits, support reliable intervention, expose hidden goals or planning when present, and maintain low false-positive rates under distribution shift.
Interpretability becomes operationally valuable when it changes what developers can detect, predict and repair.
RFE Closure: Interpretation Must Earn Predictive or Causal Power
The problem is confusing understandable stories with internal understanding. The function of interpretability is to reveal model mechanisms well enough to predict, diagnose or alter behaviour. The receiver is the researcher or operator deciding whether a system can be trusted or repaired. Evidence includes causal interventions, replication and predictive validity.
The exit condition is to reject an interpretation when it fails intervention tests, does not generalise or merely redescribes behaviour after the fact. SI interpretability should reduce uncertainty, not decorate it.
