Verification is the difference between a plausible result and a checked result. A Super Intelligence system can generate an answer, propose a calculation, write code, retrieve a document or request an external action. Verification asks a separate question: what evidence shows that the result is correct, supported, complete and actually happened?
This distinction is foundational because generation and verification fail differently. A model can produce fluent prose that misstates a source. A calculator can return a mathematically correct result for the wrong expression. A tool can execute successfully against the wrong file. A saved document can exist but contain the wrong draft.
This article explains how verification works in Super Intelligence: source checks, recomputation, executable tests, type and schema validation, consistency checks, independent critics, formal methods, outcome verification, regression tests, confidence calibration and human review. The central rule is simple: match the verification method to the kind of claim being made.
Anthropic’s guide to agent evaluations distinguishes an agent’s trajectory from the state it leaves in the environment. That is a useful starting point: saying “I saved the file” is a trajectory claim; opening the actual saved file and checking its contents is outcome evidence.
Previous: 037 — Search. This article begins with the output and works backwards toward evidence.
The Hidden Transition: Candidate Output to Verified Result
Suppose an SI assistant calculates that 90 exercise books remain in a stock room. The number can be generated directly by the model, calculated by a tool, copied from a source or inferred from several records. The visible answer does not reveal which route produced it.
Verification makes the route explicit. First check that the correct records were selected. Then check the arithmetic. Then check that the answer describes the scope honestly: 90 according to the supplied log, not 90 confirmed by a physical count.
A result is therefore verified only relative to a specific proposition. “The arithmetic is correct” is narrower than “the inventory record is correct”, which is narrower than “90 books physically exist”.
Verification Starts by Naming the Claim
You cannot verify “the answer” as one undifferentiated object. Break it into claims. Claim A: movement M05 is pending. Claim B: pending movements should be excluded. Claim C: 80 + 35 − 12 − 8 − 5 = 90. Claim D: the source record was not changed.
Each claim requires different evidence. Claim A needs the source row. Claim B needs the task rule. Claim C needs recomputation. Claim D needs the tool boundary or resulting system state.
This decomposition prevents a common failure: one successful check is incorrectly treated as proof of everything.
Source Verification
Source verification asks whether the cited or retrieved material actually supports the claim. It is not enough for a page to discuss the same topic.
If the answer says “the current policy allows three school days”, inspect the current authoritative policy and locate the exact rule. Check document identity, version, date and status.
A citation that points to an archived five-day rule is evidence against the answer, even if the link is real.
Quotation Verification
A quotation should match the source text accurately and preserve enough context to avoid reversing meaning. Ellipses, omitted conditions and selective fragments can distort a source.
A verification routine can search the source for the quoted span, compare wording and inspect nearby context. For consequential work, exact quote matching is stronger than asking the model whether the quote “looks right”.
Numerical Verification
Arithmetic is often best verified by recomputation with deterministic software. If the model says 19 × 27 = 513, a calculator can confirm it.
The more important check is whether 19 and 27 were the correct quantities. A calculator verifies the operation, not the semantic selection.
Therefore, numerical verification usually has two layers: expression validation and computation validation.
Worked Numerical Example
A class has 24 students. Three are absent. The model says attendance is 87.5%. First verify the expression: present = 24 − 3 = 21. Attendance rate = 21 / 24 × 100.
Then recompute: 21 / 24 = 0.875, so 87.5% is correct. If the model had divided 24 / 21, a calculator would faithfully return the wrong interpretation’s result.
Code Verification
Generated code can be verified by syntax checks, type checks, unit tests, integration tests, static analysis, security scanners and execution in a bounded environment.
A model saying “this code works” is not equivalent to the test suite passing. Execution creates external evidence.
The right tests come from the specification. A function can pass five tests and still be wrong on an untested edge case.
Worked Code Example
Requirement: function mean_non_missing(values) returns the arithmetic mean while ignoring None values, and returns None when no numeric values remain.
Tests: [10,20,30] → 20; [10,None,30] → 20; [None,None] → None; [] → None. A generated implementation that divides by the original list length fails the second test.
The test does not prove universal correctness, but it falsifies one bad implementation and protects the repaired behaviour in future changes.
Schema Verification
Structured outputs can be checked against a schema. If the system promises fields date, time, room and status, a validator can ensure all required fields exist and have permitted types.
Schema validity does not guarantee semantic validity. The model can put “Room 4” into the date field while still producing syntactically valid JSON.
Structural validation and factual validation therefore complement each other.
Constraint Verification
Many tasks contain hard constraints: no more than three students per class, dates must be in the future, total percentage must equal 100, booking room must exist, user must own the resource.
Traditional software can enforce these invariants after model generation. A prompt can encourage compliance; a validator can reject violations.
Consistency Verification
A long answer can contradict itself. One paragraph says the deadline is Friday; another says Monday. Consistency checks compare repeated facts and dependencies.
Consistency is weaker than truth. Two paragraphs can agree on the same wrong date. External evidence is still needed.
Cross-Source Verification
When several independent sources support the same claim, confidence can increase. But ten pages copying one original report are not ten independent confirmations.
Verification should track source lineage and independence, not only count links.
Primary Versus Secondary Sources
A primary source directly establishes a fact within its authority: an official policy, original paper, recorded measurement or database state. Secondary sources explain or report primary material.
Secondary sources can provide context and critique. Primary sources are often preferable for exact rules, dates and specifications.
Freshness Verification
A source can be accurate historically and wrong for the present. Verify publication date, effective date and whether a newer version exists.
This is especially important for software documentation, policies, prices, schedules and laws.
Identity Verification
Before checking content, verify that you have the right object. A file named “Final Policy” can be an old copy. A person’s display name can match another account.
Stable IDs, version hashes and canonical URLs reduce ambiguity.
Tool-Call Verification
A model can propose a tool call. The application validates the operation, arguments, target and permission before execution.
After execution, inspect the returned result. A proposed call is not evidence of execution. An executed request is not always evidence of successful outcome.
Outcome Verification
The strongest check for an external action is often reading the resulting state. After saving a document, open it. After moving a booking, query the schedule. After publishing a page, retrieve the public resource.
This closes the loop between intention and reality.
The Unknown-Outcome Problem
A request can time out after the server applied it. The application sees no response and cannot tell whether the action succeeded.
The correct state is unknown, not failed. Use status checks, idempotency support or resource lookup before retrying a state-changing operation.
Verification Before Action and After Action
Precondition verification checks that the action should occur: correct draft, valid approval, target exists, user has permission.
Postcondition verification checks that the intended state now exists: correct page published, old booking replaced, file contents preserved.
Consequential workflows need both.
Self-Verification by the Same Model
A model can review its own output, spot contradictions and revise mistakes. This can improve results, especially when the task is checkable.
But self-review is not independent evidence. The same blind spot can affect generation and review.
Use self-verification as one layer, not a universal replacement for tools, sources or human review.
Independent Model Critics
A second model or separate prompt can critique the first output. Diversity can reveal errors one pass missed.
If both critics share the same underlying model and training data, correlated failure remains possible. Agreement is useful evidence about consistency, not proof of truth.
Majority Vote
Generate several candidate answers and choose the majority result. This can improve performance on some reasoning tasks.
Majority vote fails when all candidates share the same systematic misconception. It also increases compute.
Best-of-N Selection
Generate N candidates, score them with a verifier and select the best. The verifier may be a rule, reward model, test suite or another model.
The method only works if the scoring signal correlates with actual task quality. A weak verifier can select confidently wrong outputs.
Search as Verification
Search can test a factual claim by retrieving current sources. It is especially valuable when the answer depends on public information outside model parameters.
A search result alone is not verification. Open the relevant source and compare the claim.
Database Verification
For operational state, query the authoritative database. If the user asks whether a payment is recorded, the billing system is stronger evidence than conversation memory.
Database state can still be wrong relative to reality, so reconciliation paths remain necessary.
Formal Verification
Formal methods use mathematical specifications and proofs to establish properties of systems under defined assumptions. They are stronger than ordinary testing for the properties proved.
Formal verification is expensive and task-specific. It is practical for some critical algorithms, protocols and hardware, not every generated paragraph.
Proof Assistants
Interactive theorem provers can mechanically check formal proofs. An SI model can propose proof steps while the proof assistant validates them.
This is a powerful hybrid because creative search comes from the model while correctness of the formal derivation comes from the checker.
Type Systems as Verification
A type checker proves certain classes of program properties under its type rules. It can reject incompatible operations before execution.
Passing a type checker does not prove business correctness, but it removes one class of error.
Property-Based Testing
Instead of writing only specific test cases, property-based testing generates many inputs and checks invariants.
For a sorting function, a property can be: output length equals input length; output is ordered; output contains the same multiset of elements.
This explores a larger input space than a handful of examples.
Metamorphic Testing
Some systems lack an easy exact answer but have known relationships. A translation followed by back-translation should preserve meaning approximately. A ranking should not change when irrelevant whitespace changes.
Metamorphic tests verify relationships between runs rather than one fixed expected output.
Differential Testing
Run the same input through two independent implementations and compare outputs. Disagreement identifies cases needing investigation.
Agreement does not prove correctness if both implementations share the same bug, but it is useful for finding regressions.
Regression Testing
Every discovered failure can become a future test. If a model once counted an archived policy as current, preserve that case after the repair.
Regression sets convert incidents into organisational memory.
Evaluation Versus Verification
Verification checks a particular claim or outcome. Evaluation measures performance across a collection of tasks.
One answer can be verified without establishing system reliability. A system can have strong average evaluation while failing one specific verified case.
Verification Coverage
Not every sentence needs the same level of checking. A casual creative paragraph and a bank transfer deserve different verification budgets.
Coverage should follow consequence, uncertainty, reversibility and user expectation.
Risk-Based Verification
Low-impact formatting changes may be checked automatically. Medium-impact factual reports may require sources. High-impact external actions may require approval plus post-action verification.
The purpose is proportional control, not maximal bureaucracy.
Human Review
Humans are valuable for context, values, domain judgment and ambiguous cases. Human review is not infallible.
A reviewer needs a clear object: source, proposed change, evidence and consequence. Asking “approve?” without showing what changes creates ceremonial oversight rather than informed control.
Human Automation Bias
People can over-trust polished machine output, especially when it arrives quickly and repeatedly appears correct.
Verification interfaces should surface evidence and uncertainty rather than merely presenting a green “AI verified” badge.
The Reviewer’s Workload Matters
If the model produces 1,000 claims and asks a person to check every one, the workflow may fail operationally even if human review is theoretically present.
Good SI reduces verification burden by structuring claims, prioritising high-risk items and automating checks where evidence is machine-readable.
Worked Example: Verify a Research Summary
Claim: “Paper A found a 12% improvement.” Check the original paper’s metric, baseline, population and uncertainty. A 12% relative improvement is different from 12 percentage points.
Claim: “Paper B confirms the finding.” Check whether Paper B studies the same intervention and population. Similar topic is not replication.
Claim: “The evidence proves the intervention works everywhere.” This exceeds the sources unless external validity is established.
Worked Example: Verify a Saved Document
Task: create “Revision Plan October” in a study folder. Model drafts content. File tool returns object ID 872.
Read object 872. Confirm title, destination, body and visibility. If the body is truncated, the create call succeeded but the task did not.
Worked Example: Verify a Published Page
Task: publish reviewed draft V3 to the events section. Before action, confirm V3 is still current. After action, fetch the public page.
Check title, canonical URL, status, body and featured image. A successful API status code is evidence of request handling, not necessarily proof of correct content.
Worked Example: Verify a Calculation With Source Selection
Source rows: opening 100; +20 completed; −5 completed; +10 pending. Rule: completed only. Correct expression is 100 + 20 − 5 = 115.
If the calculator returns 125 from 100 + 20 − 5 + 10, arithmetic is correct for the wrong selection. Verify semantics before computation.
Worked Example: Verify a Classification
Message: “The screen flashes every few seconds.” Model label: DISPLAY. Verification can compare with a labelled test set or domain rule.
If no ground-truth label exists yet, human review may create one. That label later becomes evaluation data.
LLM-as-a-Judge
A language model can score another model’s answer against criteria. This scales evaluation of open-ended outputs.
Judge models can have biases toward verbosity, position, style or familiar phrasing. Calibration against human judgments and adversarial tests is important.
Use judge models where they are validated, not because grading by AI is automatically objective.
Pairwise Evaluation
Instead of assigning an absolute score, a judge compares two outputs and chooses which better satisfies the rubric.
Pairwise comparisons can be easier and more consistent, but order effects and ties need handling.
Rubric-Based Verification
A strong rubric defines dimensions separately: factual support, completeness, constraint compliance, clarity, tool outcome and citation quality.
This prevents polished style from compensating for unsupported facts in one opaque overall score.
Ground-Truth Construction
Verification depends on trustworthy reference answers. Build them from authoritative sources, deterministic calculations or expert consensus.
If the reference is wrong, the evaluation punishes correct outputs and rewards incorrect ones.
Gold Sets and Their Limits
A gold set is a carefully curated set of inputs and expected outputs. It is valuable for regression and benchmarking.
Gold sets age. Policies change, software versions update and language evolves. Version and review them.
Adversarial Verification
Test cases can deliberately target known weaknesses: prompt injection, ambiguous dates, duplicate records, conflicting sources, malformed files.
The goal is not to make the system fail theatrically. It is to find boundaries before real users do.
Red Teaming Versus Ordinary QA
Ordinary QA checks expected workflows. Red teaming explores how the system behaves under adversarial or unusual pressure.
Both create useful verification evidence and should feed regression suites.
Verification Latency
Some checks are instant; others require external queries or human review. Latency affects usability.
Applications can perform fast checks first, return a provisional result when appropriate and continue deeper verification only when the task allows it. Consequential claims should not be presented as final before required checks finish.
Verification Cost
Generating one answer can be cheap while verifying it is expensive. Research may require reading several sources; code may require full integration tests.
System economics should measure cost to a trustworthy result, not only cost per generated token.
Verification and Uncertainty
Verification can reduce uncertainty but rarely removes all uncertainty. A source can itself be wrong. A test suite can be incomplete. A human can overlook a defect.
The final report should match the scope: “passed the defined tests”, not “guaranteed correct in every circumstance”.
A Verification Failure Map
Level 1: claim not defined. Level 2: wrong evidence source. Level 3: evidence stale or incomplete. Level 4: checker validates structure but not meaning. Level 5: self-check shares the same blind spot. Level 6: external action not read back. Level 7: regression set misses the failure. Level 8: human review lacks context. Level 9: final language overstates the evidence.
This eduKateSG map turns “verify it” into a sequence of inspectable responsibilities.
Worked Diagnosis: Citation Exists, Claim Still Wrong
The article cites a real policy page, but the page never states the claimed reason for a rule change. Verification failure is claim-source mismatch.
Repair by finding a source that states the reason or rewriting the claim as unknown.
Worked Diagnosis: Tests Pass, Program Still Wrong
The function passes all existing tests but fails on empty input. The test suite did not cover the edge case.
Add the failure as a regression test and repair the implementation.
Worked Diagnosis: Publish Tool Says Success, Wrong Page Changed
The API call succeeded, but target ID pointed to another page. Tool execution succeeded while task outcome failed.
Verify resource identity before action and read back the specific object after action.
A Practical Verification Worksheet
Write the claim. Name the authoritative evidence. Name the checker. Define what success proves. Define what success does not prove. Record the result. If the task changes external state, add a postcondition check.
This six-line worksheet is enough to improve many SI workflows because it stops vague “double-checking” and replaces it with a specific test.
Independent Exercise 1: Calculation
Model says 17 × 23 = 391. Calculator confirms 391. Is the entire word problem verified?
Answer
No. Only the multiplication is verified. Check that 17 and 23 were the correct quantities and that multiplication was the correct operation.
Independent Exercise 2: Citation
An answer cites an official page about school admissions but claims a current tuition fee. Is the citation sufficient?
Answer
No. The page must support the tuition-fee claim. Topic relevance is not enough.
Independent Exercise 3: Tool Action
A model generates a “send email” tool call but the application never executes it. Can the assistant say the email was sent?
Answer
No. A proposed call is not an executed action.
Independent Exercise 4: Successful API
The save API returns success, but the resulting document is empty. Was the task successful?
Answer
No. The operation succeeded at a narrow level, but the postcondition failed. Read-back exposed the mismatch.
Independent Exercise 5: Self-Check
The same model generates and reviews an answer twice. Both agree. Is the fact proven?
Answer
No. Agreement improves consistency evidence but does not replace independent sources or deterministic checks where available.
Verification by Claim Type: A Practical Matrix
A strong SI workflow classifies the claim before choosing the check. Factual claim about a document: compare against the source. Numerical claim: recompute. Code claim: execute tests. Structural claim: validate schema. Operational claim: inspect external state. Permission claim: check the authorisation record. Historical claim: inspect dated sources. Current claim: use a live authoritative source.
This matrix prevents a generic verifier from being applied everywhere. Asking a second language model whether a bank balance “looks right” is weaker than querying the bank record. Running a calculator on an unsupported statistic does not establish where the number came from.
Verification gets stronger when the checker is naturally matched to the proposition.
Claim Decomposition Before Verification
Long answers often combine dozens of propositions. A useful verifier first identifies atomic claims. “The school changed the policy in September because demand increased, reducing the normal loan from five days to three” contains at least three claims: timing, reason and numerical change.
Each claim can have a different evidence status. The current policy may prove three days. An archived version may prove five. A principal’s announcement may explain the reason. One citation at the end of the paragraph should not be assumed to support all three.
Claim decomposition makes citations, uncertainty and correction more precise.
Entailment: Does the Evidence Actually Support the Claim?
A source can mention the same topic without entailing the answer. Verification asks whether the evidence logically supports the proposition under ordinary reading.
Source: “The programme serves Primary 5 and Primary 6 students.” Claim: “The programme is available to all primary levels.” The source does not entail the claim. The generated statement broadened the scope.
Entailment checking can be performed by humans, specialised models or rules, but the criterion must remain visible.
Contradiction Checking
Verification also looks for evidence that contradicts the output. A research agent should not search only for confirmation.
If one official page says the policy starts 1 October and another current official page says 15 October, the answer should surface the conflict and identify which source has higher authority or newer status.
Contradictory evidence is useful information, not an inconvenience to hide.
Negative Evidence and Absence Claims
Claims such as “there is no policy”, “no study found an effect” or “the database contains no record” are harder to verify because absence depends on search scope.
A failed search does not prove nonexistence. A bounded database query over a complete table can establish absence more strongly than a web search.
Verification should state the scope: “No matching record was found in the current scheduling database” is stronger and more accurate than “No booking exists anywhere”.
Temporal Verification
Time can invalidate otherwise correct evidence. Verify not only what a source says but when it applied.
A 2025 software manual can accurately document a feature that was removed in 2026. A law can be enacted on one date and take effect later. A schedule can be current when retrieved and stale minutes later.
Time-sensitive verification should capture retrieval time, effective date and version where relevant.
Unit Verification Versus System Verification
A unit check isolates one component. Example: does the calculator return 115 for 100 + 20 − 5? A system check starts from the user request, selects rows, builds the expression, executes the calculator and reports the result.
Unit verification helps locate defects. System verification establishes whether the entire task works across handoffs.
Both are needed. A perfect calculator inside a broken workflow still produces a broken service.
Trace Verification
For tool-using systems, a trace records actions, arguments and observations. Verification can inspect whether the sequence followed allowed operations and used the expected resources.
Trace review is useful for debugging, but it should not be confused with verifying final state. A trace can say “publish called” even if publication later failed.
Outcome verification closes the loop.
State-Difference Verification
For updates, compare the before and after state. What fields were supposed to change? Which fields actually changed?
If a user asks to update one date, a diff should show one date change. An unexpected title or visibility change becomes immediately visible.
State diffs are powerful because they verify both the intended positive change and unintended side effects.
Hash Verification
Cryptographic hashes can verify that a file or content object has not changed between two moments. If the hash differs, the bytes differ.
A matching hash does not prove the file is correct, only unchanged relative to the reference. Hashes are identity and integrity tools, not semantic truth checks.
Version Verification
A user may approve draft V2 while V3 exists at execution time. Verification compares the approved version identifier with the current object.
If they differ, the system should not silently transfer approval. This principle applies to documents, code commits, database records and configuration.
Checksum Versus Semantic Equivalence
Two files can have different hashes while meaning the same thing because whitespace changed. Two files can look similar while one contains a crucial semantic difference.
Choose the check that matches the need: byte identity, structured diff or semantic comparison.
Data-Lineage Verification
A report can contain a number derived from several transformations. Verification should be able to trace the number back through source rows, cleaning rules and calculations.
This is especially important in analytics and machine-learning pipelines where one field may be aggregated, filtered and transformed before reaching the final chart.
Reproducibility
If another analyst reruns the same data pipeline with the same inputs and versioned code, they should obtain the same deterministic result where the process is intended to be deterministic.
Reproducibility is a powerful verification property for data analysis, code and scientific computation.
Randomised Systems Need Seed and Distribution Checks
Some SI processes involve sampling. Exact byte-for-byte reproduction may not be appropriate. Instead, record model version, seed when available, generation parameters and acceptance criteria.
Verification can focus on stable facts and distributional behaviour rather than identical wording.
Statistical Verification
A model evaluation result such as 82% accuracy is itself a claim. Verify sample size, population, metric definition and uncertainty.
An 82% score on 50 examples is different evidence from 82% on 50,000 representative examples. Confidence intervals and repeated runs can quantify measurement uncertainty.
A/B Test Verification
When an SI feature claims to improve user outcomes, compare treatment and control groups under a defined experiment.
Check randomisation, sample size, pre-specified metrics, dropouts and whether observed differences are practically meaningful.
Generated anecdotes are not a substitute for experimental evidence.
Causal Claims Need Stronger Evidence
“Users who used SI completed work faster” is an association if faster users chose SI voluntarily. “SI caused faster completion” requires a causal design or stronger assumptions.
Verification method should match claim strength. The more causal the language, the stronger the evidence needed.
Verification of Summaries
A summary should preserve central claims, numbers, conditions and uncertainty from the source. Verification can map each summary sentence to source spans.
Also check omission. A summary that leaves out the one exception that changes the conclusion can be misleading even if every included sentence is accurate.
Verification of Translations
A translation can be checked through bilingual review, back-translation, terminology glossaries and consistency across repeated technical terms.
Back-translation is useful but imperfect because two wrong translations can round-trip into plausible text. Domain review remains important for consequential material.
Verification of Extracted Structured Data
Document extraction can be verified field by field against the source. Dates, names, quantities and IDs deserve exact comparison.
Confidence scores can prioritise low-certainty fields for human review, but the score itself needs calibration.
Verification of OCR and Visual Reading
A model reading an image may misrecognise small text, superscripts or table boundaries. Verification can compare visual crops, use redundant OCR or require manual review for low-resolution regions.
A clean natural-language explanation does not prove the visual input was read correctly.
Verification of Long-Document Analysis
Long context creates omission risk. A model can answer from one section while missing a later exception.
Verification should include targeted retrieval of relevant sections, coverage checks and source references. “The whole document was in context” is not proof that every needed part influenced the answer.
Verification of Plans
A plan can be logically coherent but impossible in the current environment. Verify prerequisites, resource availability, dependencies and permissions.
For example, “send invitations, collect RSVPs, confirm catering” assumes an invitation list and sending authority exist. Planning verification turns prose into executable requirements.
Verification of Agent Progress
An agent can take many steps without getting closer to the goal. Progress verification asks whether each action reduces uncertainty, creates required state or satisfies a subgoal.
Repeated search queries returning the same sources are activity, not progress.
Stop Conditions as Verification
A workflow should know what counts as done. Completion criteria can include all required fields present, source checks passed, external state verified and unresolved exceptions surfaced.
Without a stop condition, agents can continue indefinitely or declare completion too early.
Verifier Independence
Independence is a spectrum. A calculator is highly independent from a language model’s arithmetic generation. A second prompt to the same model is less independent. A different model trained on similar data sits somewhere in between.
When consequences rise, increase independence where practical.
Multiple Independent Checks
Critical results can combine checks: database query + business rule validation + human approval + post-action read-back.
Layering reduces the chance that one failure mode passes through unchecked.
Verification Cascades
Cheap checks can run first: schema, required fields, simple arithmetic. Expensive checks run only when needed: web research, code execution, expert review.
This creates a verification cascade that balances cost and reliability.
Escalation Rules
Define when automated verification is insufficient. Examples: sources conflict, confidence below threshold, high-impact action, new domain, or repeated verifier disagreement.
Escalation makes uncertainty actionable instead of silently lowering quality.
Verification Records
For important tasks, preserve enough evidence to audit later: input version, source IDs, tool results, checker outcomes and final state.
Do not store unnecessary sensitive data. Verification records should be proportional and privacy-conscious.
Verification and Privacy
Checking a result can create new data exposure. A verifier should not send confidential text to an unrelated external service unless authorised.
Sometimes local deterministic checks are preferable precisely because they avoid additional data sharing.
Verification and Security
Adversarial inputs can target verifiers. A malicious document can contain text attempting to persuade an LLM judge that false content is correct.
Security-sensitive verification should isolate untrusted content and use software-enforced boundaries.
Verifier Gaming
If a generator knows exactly how a weak verifier scores answers, it may optimise for the score rather than the real objective.
This problem appears in reward models, benchmarks and automated graders. Rotate tests, use hidden evaluation and inspect real outcomes.
Goodhart’s Law in Verification
When a measure becomes a target, it can stop being a good measure. If “number of citations” becomes the target, systems may add irrelevant citations.
Verification metrics should remain connected to the underlying outcome, not become decorative compliance.
A Verification Budget
Assign verification effort based on risk. A harmless brainstorming list may receive basic sanity checks. A financial transfer may require identity, amount, account, permission, transaction status and reconciliation.
Budgeting makes reliability economically sustainable.
The Clementi Pattern Applied to Verification
Diagnose the exact error rather than saying “careless AI”. Was the source wrong? Was the calculation wrong? Did the tool target the wrong resource? Was the result correct but overclaimed?
Then repair from the first unstable point, re-run the same case and add it to regression testing.
Independent Exercise 6: Version Mismatch
A user approves document V4. Before execution, V5 is created. What should verification do?
Answer
Detect the version mismatch and stop or request review under the workflow’s rules. Approval for V4 does not automatically transfer to V5.
Independent Exercise 7: Search Absence
Web search finds no official notice. Can the assistant verify that no notice exists?
Answer
No. It can report that no notice was found through the searched sources. Nonexistence requires a stronger bounded source or authoritative confirmation.
Independent Exercise 8: Statistical Claim
A model says a new workflow is “20% faster” from five observations. What should be checked?
Answer
Metric definition, baseline, sample size, variation, selection method and whether the difference is reliable. Five observations are weak evidence for a broad performance claim.
Independent Exercise 9: Judge Model
An LLM judge prefers longer answers even when shorter answers are equally correct. What failure appears?
Answer
Judge bias. Calibrate the judge against human rubrics, control length effects and use multiple evaluation methods.
A Final Verification Standard for the Series
Before calling an SI result verified, answer seven questions. What exact claim or action is being checked? What evidence source has authority over that claim? Is the evidence current enough? Is the checker independent enough for the consequence? What does a successful check establish? What does it leave unresolved? What final state can the user inspect?
For a factual answer, this may mean a current primary source and claim-level citation. For a calculation, it means correct source selection plus recomputation. For code, it means an executable test suite under the relevant environment. For a database action, it means the authorised transaction and the resulting state. For a publication, it means the intended version at the intended destination and a live read-back.
The standard should also contain a failure route. When a check fails, the system should identify whether the problem is missing evidence, contradictory evidence, wrong resource identity, failed computation, incomplete action or an overstrong conclusion. Verification is useful only when failure produces a repairable diagnosis.
What Mastery of Verification Looks Like
A reader has mastered this article when they no longer ask only, “Is the AI answer correct?” They ask, “Which part can I verify, with what evidence, and what would that check actually prove?” That shift turns SI from a persuasive interface into an inspectable system.
They can distinguish a valid citation from a relevant-looking citation, a successful tool call from a successful outcome, a passing unit test from universal correctness, a model critic from independent evidence, and a database record from physical reality.
That is the Clementi floor applied to SI: definition, mechanism, worked example, diagnostic failure map, independent practice, answer key and observable completion criteria. Verification is not an extra paragraph after generation. It is a first-class layer of the architecture.
One final habit makes verification durable: preserve the failing case. When a source mismatch, stale version, malformed tool call or unsupported claim is discovered, turn that exact pattern into a regression test. The repair then becomes measurable. Future model, prompt, retrieval or software changes must continue passing the case that once broke the system.
This converts verification from a one-time inspection into a learning loop for the application itself. The system does not merely correct today’s answer; it improves the evidence needed to stop the same class of error from returning unnoticed.
Frequently Asked Questions About Verification in SI
Is verification the same as fact checking?
Fact checking is one form of verification. Verification also covers calculations, code, schemas, permissions, actions and resulting system state.
Can AI verify AI?
Yes, models can critique or score other outputs, but correlated errors remain possible. Independent evidence is stronger when available.
What is the best verifier?
It depends on the claim. Use source comparison for quotations, calculators for arithmetic, test suites for code, databases for current state and humans for context-sensitive judgment.
Does passing tests prove code correct?
Only for the properties and cases covered by the tests. Formal verification can prove stronger properties under explicit assumptions.
Why read back an external action?
Because a successful request can target the wrong resource, produce incomplete content or leave an unexpected state. Postcondition checks verify the actual outcome.
Should every answer be verified deeply?
No. Verification effort should match consequence, uncertainty and reversibility. Casual creative tasks need less than financial, medical or operational decisions.
What is regression testing?
Re-running representative successes and known failures after system changes to ensure repairs remain fixed and old capabilities do not break.
What is an LLM judge?
A model used to score or compare another model’s output against a rubric. It needs its own validation because judge models can be biased or inconsistent.
Verification Makes SI Inspectable
Generation creates possibilities. Verification turns some possibilities into justified results. The process is not one universal “check” but a toolbox matched to claim type: sources, calculation, execution, tests, constraints, formal proofs and human judgment.
The strongest SI systems make verification visible. They show what was checked, what remains uncertain and what evidence supports external actions. This is how capability becomes dependable work.
Continue through the How Super Intelligence Works hub. Next: 039 — Mathematics and Code.
