VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Superintelligence Works | Prediction, Probability and Uncertainty — Why a Likely Output Is Not the Same as a Certain Answer

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

Prediction, probability and uncertainty are at the centre of how many Super Intelligence systems work. A model rarely begins with certainty. It receives information, computes scores or probabilities, and produces a prediction, ranking, token, class, estimate or action candidate. The crucial question is what those numbers mean—and what they do not mean.

This matters because a confident-looking output can be generated from a probabilistic process, while a probability can itself be well measured, poorly calibrated, based on incomplete evidence or attached to the wrong question. Understanding SI therefore requires a clean separation between prediction, probability, uncertainty, confidence and verification.

This article develops those distinctions from first principles. It explains binary probability, multi-class probability, next-token probability, decoding, calibration, thresholds, missing information, distribution shift and decision consequences. It also shows why the probability of a token is not the same thing as the probability that an entire answer is true.

In this eduKateSG series, Super Intelligence, or SI, is our umbrella term for AI technologies. The article uses standard terms such as machine learning, classification, large language model, logits, softmax, calibration and sampling so readers can connect the explanation to current technical documentation.

For the larger architecture, begin with The Whole Stack, then read From Input to Output. Here we focus on one question: how does SI move from incomplete information to a prediction, and how should a human interpret the uncertainty that remains?


The Hidden Transition: A Score Is Not Yet a Decision

Suppose a school uses a simple model to estimate whether a classroom device is likely to fail during the next month. The model outputs 0.82. That number does not by itself say what action should follow. It first needs an interpretation: 0.82 of what, under which model, for which time window and against which definition of failure?

If 0.82 means an estimated probability of failure within 30 days, a maintenance team still needs a decision rule. Does it inspect every device above 0.50? Only devices above 0.80? Does the threshold change when the device is safety-critical or when inspection is expensive? Prediction and decision are different stages.

Google’s Machine Learning Crash Course explains this distinction clearly in logistic regression: the model can output a probability between 0 and 1, while a later classification decision converts that probability into a category using a threshold. See Calculating a probability with the sigmoid function and the associated classification module.

The same principle reappears in language models. The model produces scores for possible next tokens. A decoding strategy determines how one of those tokens is selected. The selected token then becomes part of the growing output. At no stage does a high token probability automatically certify the truth of the final sentence.

Five Meanings That Are Often Mixed Together

1. Prediction

A prediction is the model’s output about an unknown or future quantity. It might be a class, numerical value, next token, ranking score or estimated event probability. A prediction can be useful even when uncertainty remains.

2. Probability

A probability is a numerical representation of uncertainty under a defined model or process. In a binary problem, probabilities for the two mutually exclusive outcomes sum to one. In a multi-class problem, a softmax-style distribution can assign probability mass across several alternatives.

3. Confidence

Confidence is an overloaded word. Sometimes people use it informally to describe how certain a person or system sounds. Sometimes a model exposes a score that is called confidence. Those are not automatically equivalent. A forceful sentence can be wrong, and a numerical score can be poorly calibrated.

4. Uncertainty

Uncertainty is the part of the situation that remains unresolved. It can arise because the world itself is variable, because evidence is missing or noisy, because the model is imperfect, because the input is outside the training distribution, or because the task is ambiguous.

5. Verification

Verification is a separate process that checks a claim, calculation, source or resulting state. Verification can reduce uncertainty about a particular proposition. It does not erase every uncertainty in the entire system.

A Simple Binary Probability From First Principles

Consider a fictional model that estimates whether an email is spam. It computes a real-valued score z and transforms that score with the sigmoid function. The standard logistic function maps any real number into a value between 0 and 1. Google’s current documentation uses exactly this structure when introducing logistic regression.

If z = 0, the sigmoid output is 0.5. If z is strongly positive, the output approaches 1. If z is strongly negative, the output approaches 0. The model’s learned weights determine z from the input features. The sigmoid converts that score into a probability-like output suitable for binary classification.

Now imagine the model outputs 0.73 for one message. That can be read as an estimated 73% probability of the positive class under the model, assuming the model’s output is being interpreted probabilistically. It still does not prove the message is spam. One specific email will ultimately be spam or not spam; the probability describes uncertainty before the outcome is known.

A threshold converts the probability into an action. At threshold 0.5, 0.73 becomes “spam”. At threshold 0.9, the same message would not be automatically classified as spam. The right threshold depends on the costs of false positives and false negatives.

Decision Thresholds Encode Consequences

Suppose falsely moving an important parent email into spam is costly, while allowing one promotional message into the inbox is merely annoying. The system may use a high threshold before auto-blocking. A lower threshold might instead mark a message for review.

Reverse the consequences and the threshold can change. In a screening system where missing a rare dangerous event is far more costly than sending extra cases for inspection, a lower alert threshold may be justified. The model’s probability does not determine the organisation’s values or risk tolerance.

This distinction matters for SI because people often ask a model to “decide” when the hidden requirement is actually a policy choice. The model may estimate probabilities or rank options. The institution must still define what consequences justify which actions.

A useful operational question is therefore: what does this probability control? Does it trigger a warning, reorder a list, request more information, produce a draft, or directly change an external state? The more consequential the action, the more important the threshold and verification design become.

From Logits to Probabilities in Language Models

Large language models commonly produce a score for every token in the vocabulary at each generation step. These raw scores are often called logits. A softmax transformation converts the collection of logits into a probability distribution whose values sum to one.

For a teaching example, imagine the partial sentence “The capital of France is”. Suppose the model’s illustrative next-token probabilities after softmax are: “Paris” 0.93, “Lyon” 0.02, “the” 0.01 and all remaining tokens together 0.04. These values are invented for explanation, not measured from a real model.

The high probability on “Paris” means the model strongly prefers that token as the next continuation under the current input and parameters. It does not mean the system has calculated a 93% probability that the proposition “Paris is the capital of France” is true in an epistemological sense.

That distinction becomes clearer with a sentence whose next word is stylistic rather than factual. After “The sunset was”, a model may distribute probability across “beautiful”, “orange”, “bright”, “stunning” and many other continuations. Token probability describes continuation likelihood, not a single external fact.

Temperature Changes the Distribution, Not the Evidence

Generation systems can transform logits before sampling. The temperature parameter changes how sharp or flat the distribution becomes. Lower temperature generally makes high-scoring tokens relatively more dominant; higher temperature spreads more probability mass toward alternatives.

Hugging Face’s current generation documentation describes temperature as a value used to modulate next-token probabilities, and explains top-k and top-p sampling as methods for restricting the candidate set. See Generation strategies.

Suppose our invented distribution is Paris 0.60, Lyon 0.15, Marseille 0.10 and others 0.15 before a temperature adjustment. A lower-temperature transformation could make Paris more dominant. That does not add a new source, check a database or improve factual grounding. It changes the selection distribution.

This is why “turn the temperature down to make it accurate” is incomplete advice. Lower randomness can make output more repeatable, but repeatable errors remain errors. Evidence quality, task interpretation and verification operate at different layers.

Top-k and Top-p Sampling: Restricting Candidate Tokens

Top-k sampling keeps only the k highest-scoring candidates before sampling. Top-p, also called nucleus sampling, keeps the smallest set of high-probability tokens whose cumulative probability reaches a chosen threshold. Hugging Face documents both as generation strategies.

These techniques are useful for controlling diversity. They should not be confused with fact checking. If the correct token received very low probability because the model misunderstood the input, restricting the candidate set can make the correct token even less likely to appear.

For creative writing, diversity may be desirable. For a structured answer where multiple phrasings are acceptable but the factual content must remain fixed, the application may use lower randomness and external checks. The task determines the useful generation regime.

The Probability of One Token Is Not the Probability of a Whole Answer

A long answer contains many token decisions. Even if each chosen token is locally likely, the conjunction of all factual claims can still contain errors. The sequence probability also reflects wording choices, grammar and style—not only factual correctness.

Imagine an answer of 200 tokens containing three factual claims. The model may generate high-probability connective language around one unsupported claim. Summing or multiplying token probabilities does not magically produce a reliable probability that all three claims are true.

This is why applications that expose token log probabilities should be interpreted carefully. They can help with tasks such as ranking candidate continuations or detecting unusual generations. They are not a universal truth meter.

For factual work, the better route is claim-level evidence. Identify the claim, find the supporting source or calculation, and check whether the evidence actually entails the statement. Token probability and evidence support answer different questions.

Two Broad Sources of Uncertainty

Uncertainty in the world or data-generating process

Some outcomes are inherently variable. A weather forecast can remain uncertain even with excellent sensors because the future depends on a complex evolving system. A student may perform differently across days because attention, task difficulty and conditions vary.

In statistics and machine learning, this kind of irreducible or data-related uncertainty is often discussed using terms such as aleatoric uncertainty. The important practical idea is that better modeling can improve estimates without making a genuinely variable world deterministic.

Uncertainty from limited knowledge or model limitations

Other uncertainty arises because the system lacks information, the model has not learned the relevant pattern, the input is unfamiliar or the source is incomplete. This is often discussed under terms such as epistemic uncertainty.

For example, if a school timetable assistant has never received the updated timetable, uncertainty comes from missing knowledge. A live timetable lookup may reduce it. If the timetable itself contains a room marked “TBC”, the uncertainty is in the source state and cannot be eliminated by retrieving the same record more often.

The two categories are useful teaching tools, but real systems can contain both at once. The practical question is what new evidence or model improvement could reduce the uncertainty, and what uncertainty would remain even after that improvement.

Calibration: Does 80% Really Behave Like 80%?

A model can rank cases well and still be poorly calibrated. Calibration asks whether predicted probabilities correspond to observed frequencies over many comparable cases. If a well-calibrated system assigns probability 0.8 to many events, roughly 80% of those events should occur in the relevant evaluation setting.

Consider 100 comparable predictions, each given probability 0.8. If only 50 positive outcomes occur, the system is overconfident in that region. If 95 occur, it was underconfident. One individual prediction cannot establish calibration; calibration is evaluated across a collection.

The phrase “relevant evaluation setting” matters. Calibration measured on one population may not transfer unchanged to another. A model calibrated on one school, device type or time period can become miscalibrated after conditions shift.

Therefore, a numeric confidence score deserves questions about how it was validated. Was it calibrated? On which data? During which period? Does the present task resemble that evaluation?

A Worked Calibration Example

Imagine a fictional maintenance model produces ten predictions near 0.7 for ten classroom projectors. Over the evaluation window, seven projectors actually fail. That small set is consistent with a 0.7 probability but far too small to prove strong calibration by itself.

Now imagine 1,000 predictions grouped around 0.7, with about 700 positive outcomes under stable conditions. That is stronger evidence of calibration around that range. We would still inspect uncertainty intervals, subgroup behaviour and whether the same relationship holds after deployment conditions change.

The lesson is that calibration is empirical. A model cannot prove its own probability quality merely by returning a number with decimal places.

Distribution Shift: When Yesterday’s Probabilities Stop Meaning the Same Thing

Machine-learning systems are evaluated on data drawn from particular conditions. Deployment can change those conditions. New user behaviour, revised policies, sensor changes, economic shocks, new vocabulary or adversarial adaptation can alter the relationship between inputs and outcomes.

Suppose a classifier estimates whether a school device will fail based partly on its age. The school replaces one manufacturer with a new model whose failure pattern differs. Predictions learned from the old devices may no longer be calibrated for the new devices.

This is distribution shift. It does not automatically mean the model becomes useless, but it weakens the assumption that historical evaluation directly represents current performance. Monitoring and re-evaluation are system responsibilities, not properties that a fixed model can guarantee forever.

Google’s Production ML Systems module emphasises monitoring the broader deployed system because real-world conditions and inputs can change after launch.

Out-of-Distribution Inputs and the Problem of Familiar-Looking Confidence

A model may produce a normal-looking prediction even when the input is far outside the situations on which it was evaluated. The interface can therefore give a false sense that the system “recognised” the case simply because it returned a number.

For example, a classifier trained on ordinary classroom projector telemetry might receive data from a completely different device whose fields have been mapped incorrectly. The model can still output 0.62. The existence of the probability does not establish that the input is meaningful.

Applications can add checks for schema, range, missing fields, source identity and known operating conditions. These checks belong before or around model inference. They are not a replacement for better models, but they can prevent meaningless inputs from being treated as normal cases.

Uncertainty in Retrieval-Augmented Systems

A retrieval-augmented system adds another uncertainty layer. The model can be uncertain about the answer, while the retriever can be uncertain about which document is relevant. The source itself can also be outdated or ambiguous.

Suppose the query asks for a current policy and retrieval returns three passages: an approved current policy, an old archived version and a discussion note. Semantic similarity may be high for all three. The system needs metadata and source governance, not only a model that can write a synthesis.

A useful result can preserve this uncertainty explicitly: “The current approved policy says X; an archived version says Y; the discussion note is not authoritative.” The system has not eliminated uncertainty by hiding disagreement. It has structured the evidence so the user can see which source controls the decision.

Uncertainty in Tool-Using Systems

Tools create uncertainty about external state. A model may intend to save a document, but the operation can fail. A request can time out after the server already applied it. A file can be saved under the wrong identity. These are not next-token uncertainties; they are system-state uncertainties.

The correct response is to inspect tool outcomes and resulting state. If the save outcome is unknown, say so. If the document exists, return its actual identity. If the tool failed, preserve the draft in the conversation when possible.

This is why the article Capability versus Autonomy treats verification and permission as separate from intelligence. A confident model prediction about external state is not a substitute for observing the external state.

Worked Example: Choosing a Review Threshold

Consider a fictional school IT model that predicts whether a laptop needs manual inspection. During evaluation, each laptop receives a probability between 0 and 1. The school has capacity to inspect only a subset each week. The model does not decide the capacity constraint; the organisation does.

Assume five laptops receive probabilities 0.92, 0.81, 0.67, 0.44 and 0.20. If the review threshold is 0.80, two laptops are inspected. If the threshold is 0.60, three are inspected. The threshold changes workload and error trade-offs without changing the model.

Now add asymmetric costs. Missing a failing laptop used for an examination is more costly than inspecting one extra healthy laptop. The school might choose a lower threshold for examination devices and a higher threshold for low-priority spare devices. That is a policy layer using model output, not a change in the probability calculation itself.

Finally, suppose the device manufacturer changes. Before trusting the same threshold, the school should examine whether the probability estimates remain calibrated for the new hardware. A threshold that was sensible under the old distribution may no longer produce the same practical trade-off.

Expected Value: Probability Meets Consequence

A probability becomes decision-relevant when combined with consequences. A simple expected-value calculation multiplies outcomes by their probabilities and sums them. This can help structure decisions, although real institutional choices may include constraints, fairness, legal duties and qualitative factors that do not fit into one number.

Suppose an inspection costs one unit of staff time, while an uninspected failure during an important event costs ten units. If a device has a 0.2 failure probability, the simple expected failure cost is 0.2 × 10 = 2 units, which exceeds the one-unit inspection cost. Under this simplified fictional model, inspection has lower expected cost.

This is only a teaching example. Real decisions require better estimates and fuller consequence models. The point is that model probability and organisational utility are distinct inputs. SI can help calculate expected values, but it should not silently choose which human outcomes count as valuable.

The Difference Between Uncertainty and Ambiguity

Uncertainty means the outcome or state is not fully known. Ambiguity means the task or meaning is not fully specified. These often look similar at the interface but require different repairs.

“Will this device fail next month?” can be a well-defined uncertain question. “Is this device safe?” may be ambiguous if “safe” has no defined criterion. More probabilistic modeling does not solve an undefined target.

For a language model, an ambiguous user request can generate several plausible interpretations. Asking a targeted clarifying question, using context or presenting the interpretations may be better than assigning probabilities to undefined alternatives.

The Difference Between Uncertainty and Error

A correct probability can lead to a prediction that turns out wrong. If a well-calibrated model gives an event probability of 0.2, the event will still occur in some cases. That outcome does not by itself prove the model was wrong.

Conversely, a model can make a badly justified prediction that happens to be correct. Accidental correctness should not be confused with reliability. Evaluation needs many cases and clearly defined metrics.

This distinction is essential when judging SI. One surprising failure can reveal an important defect, but it does not automatically quantify overall performance. One impressive success can demonstrate capability, but it does not establish dependable behaviour across the domain.

What “I’m 90% Sure” Should Mean in an SI Product

An interface that displays “90% confidence” should define the quantity. Is it a calibrated probability of a class? A normalized similarity score? A token probability? A heuristic from several signals? A self-reported model estimate in natural language? These are not interchangeable.

Natural-language self-assessment can still be useful as a behavioural signal, especially when it triggers verification or escalation. It should not be presented as a calibrated statistical probability unless there is evidence supporting that interpretation.

A better interface can explain the evidence directly: “The current policy document supports this answer” or “The available sources disagree on the date.” Evidence descriptions often help users more than an unexplained percentage.

How SI Should Communicate Uncertainty

Useful uncertainty communication is specific. “I don’t know” can be appropriate, but it can often be improved: “The source states the normal rule but does not define weekends”, “The calculation is verified, but the save outcome is unknown”, or “The current document conflicts with an archived version.”

These statements identify what is uncertain and what is known. They also point toward the next useful action: retrieve another source, ask the responsible owner, run a status check or collect more data.

The goal is not to fill every answer with disclaimers. It is to ensure that consequential uncertainty remains visible at the point where it matters.

Repair, Stabilise and Extend: Three Probability Pathways

Repair

If a system outputs probabilities that do not correspond to the intended event, first repair the target definition, labels or data pipeline. A perfectly calibrated model for the wrong target remains useless for the actual decision.

Stabilise

If the target is correct but probabilities drift over time, monitor calibration and input distributions. Recalibration, retraining or revised thresholds may be appropriate depending on the cause.

Extend

If the system is stable on known cases, extension might include subgroup evaluation, abstention mechanisms, uncertainty-aware routing or additional evidence collection for difficult cases. Expansion should preserve the original calibration tests so new complexity does not hide regression.

Independent Exercise 1: Token Probability Versus Truth

A model generates the sentence “The meeting begins at 3 PM.” The token corresponding to “3” had the highest next-token probability. Does that establish the meeting time? Explain the difference between generation probability and factual support.

Answer

No. The token probability says the model preferred that continuation under the supplied context. To establish the meeting time, inspect the calendar, invitation or other authoritative source. A high-probability token can be unsupported if the relevant evidence was missing or misunderstood.

Independent Exercise 2: Thresholds

A classifier outputs 0.76. System A uses a threshold of 0.5 and labels the case positive. System B uses a threshold of 0.9 and labels the same case negative. Has the model contradicted itself?

Answer

No. The model output is the same. The applications apply different decision thresholds. The contradiction is only apparent if probability estimation and decision policy are treated as the same layer.

Independent Exercise 3: Calibration

A model assigns approximately 0.9 probability to 200 comparable events, but only 120 occur. What concern does this raise?

Answer

It raises a calibration concern in that range because the observed frequency, 60%, is far below the predicted 90%. The next step is to inspect the evaluation design, sample comparability, distribution shift and whether the probability output is intended to be interpreted as calibrated.

Independent Exercise 4: Missing Information

A timetable assistant says there is a 70% chance that tomorrow’s room is B12, but the source document marks the room “TBC”. What should the system do?

Answer

It should preserve the source uncertainty. If no authoritative update exists, inventing a probability can mislead the user. A better response is that the room remains to be confirmed, possibly followed by a live lookup or instruction to check the responsible timetable source.

Confusion Matrices: Probability Becomes Observable Error

Once a probability is thresholded into a class, the resulting decisions can be organised into a confusion matrix. There are four possibilities in binary classification: true positive, false positive, true negative and false negative. These outcomes let us measure the consequences of a threshold rather than discussing “accuracy” as one undifferentiated number.

Suppose a fictional laptop-inspection model evaluates 100 devices. Twenty devices really do fail during the evaluation period. At one threshold, the system correctly flags 16 of them and misses 4. It also flags 12 healthy devices and correctly leaves 68 healthy devices unflagged. The confusion matrix is therefore TP = 16, FN = 4, FP = 12 and TN = 68.

Accuracy is (16 + 68) / 100 = 84%. Recall is 16 / 20 = 80%. Precision is 16 / 28, or about 57%. Those numbers describe different properties. Google’s current classification metrics guide emphasises that the useful metric depends on the task and the relative cost of different errors.

Now raise the threshold. The model may flag fewer devices, reducing false positives but also missing more true failures. Precision can rise while recall falls. Lower the threshold and the opposite can happen. The threshold changes the decision regime even though the underlying model scores stay the same.

Why Accuracy Alone Can Be Misleading

Consider a rare event. Suppose only one laptop in 100 is likely to suffer a serious battery fault during the period being studied. A useless classifier that predicts “no fault” for every laptop achieves 99% accuracy while detecting none of the actual faults.

This is why class imbalance matters. For rare positives, recall, precision, false-positive rate and precision-recall curves can be more informative than raw accuracy. The right choice follows the operational consequence, not a universal preference for one metric.

The lesson generalises beyond classification. A single aggregate score can hide the type of failure that matters most to the user. SI evaluation should preserve the structure of the task instead of compressing every outcome into one headline number.

ROC, AUC and the Difference Between Ranking and Thresholding

Receiver operating characteristic curves examine model behaviour across many thresholds by plotting true-positive rate against false-positive rate. Area under the ROC curve, or AUC, summarises how well a binary classifier ranks positive examples above negative ones across thresholds.

Google’s current ROC and AUC guide explains that AUC is useful for comparing ranking quality, while the actual classification still depends on the chosen threshold. A strong ranking model does not decide the operating point for you.

Imagine two devices with scores 0.81 and 0.63. If the higher-scored device truly fails and the lower-scored device does not, the model ranked that pair correctly. AUC aggregates this ranking behaviour over many positive-negative pairs. It does not tell the school how many technicians are available for inspection next Tuesday.

That operational capacity enters later. A system might choose the top ten devices for inspection, or use a fixed threshold, or combine probability with device importance. Ranking quality and action policy remain separate layers.

Precision–Recall Trade-offs in a Worked Example

Return to our 100-laptop example. At threshold A, suppose TP = 18, FN = 2, FP = 30 and TN = 50. Recall is 90%, but precision is only 37.5%. The system catches most failures but creates many false alarms.

At threshold B, suppose TP = 12, FN = 8, FP = 4 and TN = 76. Precision becomes 75%, but recall falls to 60%. The system’s positive alerts are more often correct, but it misses more actual failures.

Neither threshold is automatically “better”. If a missed failure can disrupt an examination while a false alarm only costs a five-minute inspection, threshold A may be more appropriate. If inspections are extremely expensive and failures are low-impact, threshold B may be preferable.

The model is not choosing the school’s values. It supplies a predictive signal. Human policy converts that signal into a decision using consequences, capacity and institutional priorities.

Brier Score: One Way to Evaluate Probabilistic Predictions

Accuracy evaluates thresholded classes. A probability model can also be evaluated directly. One simple metric is the Brier score for binary events: take the squared difference between each predicted probability and the observed outcome, then average those squared errors.

Suppose three events receive probabilities 0.9, 0.7 and 0.2, and the outcomes are 1, 1 and 0. The squared errors are (0.9−1)² = 0.01, (0.7−1)² = 0.09 and (0.2−0)² = 0.04. The average is 0.14 / 3, or about 0.047.

Now compare a more cautious set of predictions 0.6, 0.6 and 0.4 for the same outcomes. The squared errors are 0.16, 0.16 and 0.16, averaging 0.16. Under the Brier score, the first set is better because its probabilities are closer to what happened.

A Brier score combines calibration and discrimination effects and should be interpreted in context. It is not a universal substitute for all other metrics. The value of introducing it here is conceptual: probabilities deserve evaluation as probabilities, not only after thresholding them into categories.

Selective Prediction: Sometimes the Best Output Is “Route This Case”

An SI system does not always need to answer every case. A selective system can abstain or route a case for additional review when the evidence is weak, the input is unfamiliar or the consequence is high.

For example, a document classifier might automatically route ordinary invoices but send unusual layouts or low-confidence extractions to a person. The abstention mechanism can improve reliability on the cases the system does handle automatically, although it also changes workload.

The important point is that abstention should be evaluated. If the system abstains on every difficult case, it may appear accurate while shifting nearly all useful work to humans. Coverage—the fraction of cases handled automatically—becomes part of the evaluation.

A practical evaluation can therefore report both error rate and coverage. “98% accuracy at 20% coverage” describes a different operational system from “94% accuracy at 90% coverage”. The user needs to know which regime applies.

Uncertainty-Aware Routing

Routing is one of the most useful consequences of uncertainty. A simple question with a current source can go directly to generation. A source conflict can be routed to a comparison step. A missing critical field can be routed back to the user. A high-impact external action can be routed to approval.

This is more constructive than treating uncertainty as a reason for paralysis. The system asks: what additional evidence or authority would make the next step justified? Uncertainty becomes an input to workflow design.

For the laptop example, a medium probability might route to a quick diagnostic test rather than immediate replacement. A very high probability on a critical exam device might trigger manual inspection. A very low probability could leave the device in normal service while monitoring continues.

Uncertainty in Natural-Language Questions

Language itself can make a question uncertain. “Which laptop is best?” has no defined criterion. Best for battery life, cost, graphics, repairability or school exams? A model can generate an answer, but the target is ambiguous.

The correct repair is not necessarily a probability distribution over brands. It may be a clarification: “Best for which use?” Once the criterion is defined, the system can retrieve specifications, compare options and present uncertainty about missing or variable measurements.

This illustrates a broader pattern: some uncertainty comes from the world, some from the model and some from the task specification. A mature SI system tries to locate the uncertainty before choosing the repair.

Probability in Ranking Systems

Search engines, recommender systems and retrieval components often use scores to rank candidates. These scores may or may not be calibrated probabilities. A relevance score of 0.82 does not automatically mean an 82% probability that the document is correct.

The same caution applies to embedding similarity. A cosine similarity score can indicate closeness in a representation space, but it is not itself a probability of factual truth or legal authority. An old policy can be highly similar to the current policy because most of the text is identical.

This matters for RAG. The retriever can rank passages by similarity, then the application can filter by metadata such as date, owner or approval status. The model can synthesise the selected evidence. Each score must be interpreted within the component that produced it.

Probability in Multi-Class Problems

Not every classification is binary. Suppose a document classifier chooses among four classes: invoice, receipt, timetable and letter. A softmax layer can transform model scores into a distribution across the four classes.

Imagine an illustrative output: invoice 0.52, receipt 0.40, timetable 0.05, letter 0.03. The top class is invoice, but the margin over receipt is small. A system could accept the top class, ask for more information or route the case for review depending on the task.

Compare another document: invoice 0.98, receipt 0.01, timetable 0.005, letter 0.005. The distribution is much more concentrated. Concentration can be useful, but it still does not prove the input belongs to invoice if the model is seeing an unfamiliar document type not represented among the four classes.

This is an example of closed-set assumptions. If the true class could be “medical form” but the model only knows four labels, softmax still distributes probability across those four. The application needs a way to recognise or manage out-of-scope inputs rather than treating the largest probability as proof that one listed class must be correct.

Entropy as a Measure of Distribution Spread

Entropy is one mathematical way to describe how spread out a probability distribution is. A distribution concentrated on one outcome has lower entropy than a nearly uniform distribution across many outcomes. In some systems, this can help identify uncertain predictions.

For an illustrative four-class distribution [0.97, 0.01, 0.01, 0.01], uncertainty is concentrated. For [0.28, 0.26, 0.24, 0.22], the model has no strong preference. Entropy captures that difference in spread.

But low entropy is not the same as correctness. A model can be confidently wrong. Entropy describes the shape of the model’s distribution, not whether the underlying world agrees with it.

Uncertainty and Ensembles

One way to estimate model uncertainty is to compare predictions from multiple models or multiple runs. If several independently trained models strongly agree, that can be informative. If they disagree widely, the case may deserve review.

However, ensembles are not magic. Models can share the same training data, architecture assumptions or blind spots and therefore agree on the same mistake. Agreement is evidence about model consistency, not proof of truth.

The cost is also higher because multiple models or passes require more computation. Whether the extra uncertainty signal is worthwhile depends on task importance and operational constraints.

A Full Evaluation Sheet for a Probabilistic SI Component

A strong evaluation sheet begins with the target: exactly what event or class is being predicted, over what time window, and for which population? Then record the probability output definition and the data used for evaluation.

Next evaluate ranking and classification behaviour. Include confusion matrices at relevant thresholds, precision, recall and any domain-specific cost measures. If probability values will be shown or used directly, evaluate calibration and a proper scoring rule such as Brier score or log loss.

Then test shift and edge cases. Include new device types, missing fields, corrupted inputs and cases outside the known classes. Measure how the system behaves when evidence is incomplete. If it can abstain, report both error and coverage.

Finally test the system context: what action does each probability trigger, who sets the threshold, what evidence is shown to users, and how will the model be monitored after deployment? A probability is never operationally isolated from the workflow that consumes it.

Worked Mini-Audit: Is This “90% Confidence” Meaningful?

Suppose a dashboard displays “90% confidence: policy is current.” Ask four questions. First, what generated the 90%? A classifier probability, similarity score, language-model self-report or heuristic? Second, was that value calibrated against real current-versus-outdated policy labels?

Third, what population was used? Policies from the same organisation, arbitrary web pages or a synthetic benchmark? Fourth, what action follows? Does 90% merely highlight the document for review, or does it automatically publish an answer to users?

If the product cannot answer these questions, the percentage may still be internally useful, but it should not be interpreted as a validated probability of truth. Precision in display does not create precision in meaning.

Clementi-Style Diagnostic Pattern for Uncertainty

When an SI probability looks wrong, diagnose from the first unstable point. Step 1: target definition—are we predicting the right event? Step 2: input validity—does the data represent the intended case? Step 3: model output—are scores computed correctly? Step 4: calibration—do probabilities correspond to observed frequencies?

Step 5: decision threshold—does the chosen threshold match consequences and capacity? Step 6: action routing—does the workflow respond proportionately? Step 7: reporting—does the interface explain what the probability means without overstating certainty?

This sequence mirrors the Clementi article’s emphasis on diagnosing the exact mathematical weakness instead of labelling every mistake “careless”. In SI, “the AI is uncertain” is too broad. We want to know which uncertainty, produced where, and what evidence could reduce it.

Frequently Asked Questions About Probability and Uncertainty in SI

Is a language model always probabilistic?

Many language models produce probability distributions over next tokens. The application can then use deterministic or stochastic decoding strategies. Other AI systems may use different mechanisms. Do not generalise one generation architecture to every form of AI.

Does temperature measure how uncertain the model is?

No. Temperature is a generation control that reshapes the token distribution used for selection. It is not a calibrated measure of factual uncertainty.

Does the highest-probability answer have to be correct?

No. The model can strongly prefer a wrong continuation when evidence is missing, misleading or outside its learned capability. Probability is conditional on the model and input, not an external truth certificate.

Can a deterministic model still be uncertain?

Yes. A deterministic decision rule can output the same answer every time even when the underlying evidence is weak. Repeatability and justified certainty are different properties.

Can retrieval eliminate uncertainty?

Retrieval can reduce uncertainty by supplying relevant evidence. It can also introduce source-selection errors, conflicting documents and freshness problems. Retrieval changes the evidence environment; it does not guarantee certainty.

What is calibration?

Calibration compares predicted probabilities with observed frequencies over comparable cases. A calibrated 0.8 region should produce positive outcomes at roughly the corresponding frequency under the evaluated conditions.

Why not always choose a 0.5 threshold?

Because thresholds depend on decision consequences, class prevalence, operational capacity and other requirements. The threshold is a policy choice layered on top of model estimates.

Should SI show percentages to users?

Only when the percentage has a defined meaning and appropriate validation. In many cases, direct evidence statements or categorical uncertainty labels can be more useful than a precise-looking but poorly understood number.

What is the difference between uncertainty and ambiguity?

Uncertainty concerns an unknown outcome or state. Ambiguity concerns an unclear task or meaning. Uncertainty may call for better evidence; ambiguity may call for a clearer question.

What should happen when the system is outside its evaluated range?

A well-designed application can flag the unfamiliar condition, request more information, route to a different process or require human review. The appropriate response depends on the task and consequences.

Prediction Becomes Useful When Its Meaning Is Visible

The central lesson is not that probabilistic systems are unreliable. Probability is one of the most powerful tools for reasoning under uncertainty. The lesson is that a number becomes useful only when we know what event it refers to, how the model was evaluated and how the decision layer will use it.

In language models, token probabilities guide generation. In classifiers, probabilities can support thresholds and prioritisation. In forecasting, probabilities can represent possible futures. Across all of these, uncertainty must remain connected to evidence, calibration and the consequences of action.

Continue through the How Super Intelligence Works hub. Previous: 004 — Capability versus Autonomy. Next: 006 — Why Language Models Can Do More Than Language.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading