VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Superintelligence and Space: Exploration, Planetary Protection and Long-Horizon Decisions

eduKate Secondary students reviewing open books for How Super Intelligence Works: the SI Failure Map.

A reader’s guide to exploration, planetary protection and decisions that outlast a mission.

Distinguish demonstrated AI from hypothetical superintelligence. Protect the evidence, clarify authority and keep future choices open.

Three students studying together with open books at a classroom table.
Learning to connect ambitious questions with careful evidence and responsible choices.

Space exploration makes the distinction between intelligence and responsibility unusually clear. A system might become very good at selecting observations, comparing mission concepts or finding patterns in images. None of those achievements, on its own, gives it the authority to disturb an environment, expose people to a risk, spend public resources or decide what future generations should inherit. Better reasoning can improve a decision without becoming the owner of that decision.

This guide explains how to think about powerful AI in scientific exploration beyond Earth. Its central question is: how could increased intelligence expand what we learn while preserving the evidence, environments and human choices that make exploration worthwhile? It focuses on scientific aims, planetary protection, bounded autonomy and decisions whose consequences last beyond a single mission. Transport capacity, settlement supply chains and orbital computing have separate problems; useful connections to those topics do not make them substitutes for this one.

Artificial superintelligence is hypothetical in this article: an imagined system whose intellectual capabilities broadly exceed human capabilities across many important domains. Current AI, mission automation and demonstrated autonomy are described separately. A successful specialised system is evidence about its tested task and conditions. It is not proof that artificial superintelligence exists, that a spacecraft can govern itself, or that a mission can safely dispense with expert judgement.

The worked programme below is entirely fictional. Its money-like budget tokens, outcomes, probabilities, time periods and scores are teaching inputs, not mission estimates. Its arithmetic is deliberately reproducible with a calculator. It includes no spacecraft command sequences, contamination-control recipes, sample-handling instructions or biological experimentation procedures. Real missions require qualified teams and the applicable approval processes.

1. Start with the question the mission should answer

The first decision in exploration is not which impressive machine to send. It is what uncertainty matters and which observations could reduce it. A beautiful image, a large dataset and a confident model response can all leave the original question unresolved. A mission intended to distinguish two explanations needs observations that those explanations predict differently. If both predict the same measurement, collecting that measurement more precisely may add little to the particular decision.

Consider a fictional icy world named Pelagos. Researchers want to know whether a surface feature varies with a repeating environmental cycle or merely appeared different because the images used different lighting. One proposal collects many more images at similar lighting. Another collects fewer images under deliberately varied, documented conditions. The second could be more informative about the competing explanations even if its image count is lower. This is an original reasoning example, not a description of an existing mission.

AI could help by listing alternative hypotheses, organising previous observations and suggesting measurements that would discriminate between them. The important review question is whether it has represented the alternatives fairly. A system can make its preferred explanation look strong by comparing it only with weak alternatives. Reviewers should ask whether the list includes mundane instrument effects, selection effects and environmental explanations, rather than only exciting scientific possibilities.

Separate discovery from confirmation. A broad search may appropriately tolerate many false alarms because its aim is to find candidates worth examining. Confirmation demands stronger controls and independent evidence. Giving both stages the same success metric encourages one of two mistakes: missing unusual phenomena by filtering too aggressively, or announcing discoveries from a preliminary screen. The word “discovery” should not silently change meaning halfway through a proposal.

Scientific return is also more than the number of positive findings. A well-designed null result can eliminate explanations or bound an effect. An unsuccessful search may still be useful if the search region, instrument sensitivity and selection process are documented. A null result with unknown sensitivity tells us much less. A reader evaluating an AI-assisted claim should therefore ask what the mission could have detected, what it could not have detected, and which conclusions survive either result.

A mission question should end in an evidence contract written in ordinary language: we will measure this quantity, under these conditions, to distinguish these explanations; these outcomes support this inference; these other outcomes leave the question open. That contract does not freeze scientific curiosity. It prevents an attractive result from retrospectively becoming the question that the mission supposedly answered all along.

Contents · Previous chapter · Continue to chapter 2

2. Distinguish demonstrated AI from imagined superintelligence

There is already a concrete reason to take AI-assisted exploration seriously. JPL reported on 30 January 2026 that Perseverance completed two drives, on 8 and 10 December 2025, using waypoints produced by generative AI. The engineering team checked the commands using a digital twin before transmission. The two reported distances were 210 metres and 246 metres. This is evidence of a bounded, reviewed application in rover route planning, not evidence of unconstrained mission leadership. JPL’s account of the AI-planned drives

The useful distinction is between a component doing a task and an institution entrusting that component with authority. An image classifier can label candidate features. A planning aid can suggest an observation sequence. A flight system can execute approved activities within its engineered limits. A programme authority decides the mission purpose and acceptable commitments. These roles can interact without becoming interchangeable. Increased capability in the first role does not automatically transfer the last role to the model.

A hypothetical superintelligence might generate more inventive hypotheses, compare a larger range of designs or detect subtle inconsistencies across scientific fields. These are conditional possibilities. Their value would depend on whether the ideas remain testable, whether experiments can be performed, whether the system understands its limits and whether reviewers can inspect the evidence. Faster speculation without better validation could increase the burden on science rather than increase reliable knowledge.

Claims about general capability need evidence across different settings. Success at a well-instrumented navigation task does not establish competence in interpreting an unfamiliar chemical signal. Success with a simulation does not establish performance under unmodelled physical conditions. Success in analysis does not establish that an agent will respect boundaries when an attractive action is outside its authorisation. Transfer must be measured, not inferred from the word “intelligent”.

Readers can use three labels when examining a proposal. “Demonstrated” means a named system performed a specified task under reported conditions. “Proposed” means there is a design or argument that still needs testing. “Hypothetical” means the discussion explores a capability or arrangement that has not been established. A single article or presentation may contain all three. Honest communication labels the transition instead of allowing evidence for one modest achievement to support every later possibility.

The opposite mistake is dismissing every current result because it is not superintelligence. A bounded tool that saves expert time while maintaining evidence quality can be valuable. Its usefulness does not require a claim of universal intelligence. Careful optimism recognises real gains, identifies their conditions and leaves room for new evidence. It neither treats a demonstration as destiny nor demands an impossible level of generality before acknowledging progress.

Contents · Previous chapter · Continue to chapter 3

3. Let physical limits shape the role of autonomy

Space changes the practical meaning of supervision. Communication takes time, contact can be intermittent, equipment has limited resources, and many failures cannot be repaired by a person walking over to the machine. A decision procedure must therefore explain what happens while advice from Earth is unavailable. Merely promising that a human remains “in the loop” is incomplete if the action unfolds faster than a message can return.

An original classroom example makes the timing issue concrete. Suppose a fictional probe has a twelve-minute one-way signal delay. A request takes twelve minutes to reach Earth and a reply takes at least twelve more minutes to return. Even with immediate review and continuous contact, the round trip is twenty-four minutes. If the decision must be made in five minutes, live approval cannot be the sole protective mechanism. This says nothing about a real mission’s delay; the numbers are stipulated for learning.

That limitation creates a need for pre-agreed authority, rather than a licence for unrestricted discretion. The system may be allowed to choose among approved observations while preserving mandatory reserves. It may be required to suspend an optional activity when information is missing. It may be prohibited from crossing a scientifically sensitive boundary regardless of the predicted gain. These are conceptual examples of bounded delegation, not flight procedures.

A strong proposal distinguishes an action envelope from an objective. “Collect useful science” is an objective. “Choose only among these approved observation classes, subject to independently enforced limits” describes an envelope. If the objective alone is supplied, the system could discover an efficient but unacceptable route to it. The envelope records what the institution has actually authorised and what it has deliberately reserved for later judgement.

More autonomy can reduce one risk while increasing another. Waiting for distant advice may waste a brief observation opportunity. Acting locally may create a mistake that reviewers cannot undo. The comparison should include both, instead of treating autonomy as automatically safe or automatically reckless. Different actions deserve different arrangements according to urgency, reversibility, consequence and the quality of available local evidence.

A useful communication failure test asks the team to describe the expected behaviour without using the phrase “ask a human”. What information will the system possess? Which commitments remain valid? Which optional activities can stop? How will it record uncertainty for later review? A proposal that answers those questions has turned supervision from a slogan into an inspectable design obligation. No level of hypothetical intelligence removes the need to answer them.

Contents · Previous chapter · Continue to chapter 4

4. Preserve observations before optimising the story

Scientific data passes through several transformations: an instrument produces a signal, processing makes the signal usable, analysis infers a property, and interpretation connects the property to a broader explanation. An AI-generated account often compresses these stages into one smooth narrative. Readers need to be able to separate them again. Otherwise an uncertain interpretation can acquire the apparent authority of a direct measurement.

Imagine a fictional instrument records a change in brightness. The raw signal is one object. A calibrated brightness estimate is another. A conclusion that the surface composition changed is a third. A claim that a biological process caused the change is a fourth. Each transition adds assumptions. A fluent paragraph that jumps from the first to the fourth hides the scientific work precisely where scrutiny matters most.

AI may help identify unusual patterns, but unusual relative to what? A model trained mainly on common terrain may flag a rare lighting condition. A model trained on processed images may respond to a processing artefact. A system tuned after seeing the target examples may perform much worse on new observations. Evaluation should record data provenance, training boundaries, comparison samples and the conditions under which predictions are meaningful.

Independence also needs substance. Two models trained on similar data and using the same preprocessing pipeline can agree because they share the same error. An expert who sees the model’s preferred answer before reviewing the evidence may become anchored to it. Independent review can require different instruments, different assumptions, blinded analysis or a genuinely separate measurement route. Counting reviewers is easier than establishing independent information.

Preservation matters because future scientists may ask questions the original team did not foresee. NASA’s Planetary Data System describes itself as a long-term archive of digital data products from planetary missions and other acquisitions. That establishes an important practical precedent: scientific value can continue after collection, through reusable, documented records. NASA Planetary Data System

The governance implication here is an original inference: a powerful exploration assistant should be evaluated partly on whether it leaves future investigators a usable trail. Keeping only a polished summary can destroy the ability to revisit a claim. Keeping every file without descriptions can also fail, because nobody can determine what the files mean. Durable evidence combines observations, context, transformations, uncertainty and enough documentation to reconstruct the reasoning without trusting the original narrator.

Contents · Previous chapter · Continue to chapter 5

5. Understand what planetary protection protects

NASA describes planetary protection as protecting Solar System bodies from contamination by Earth life and protecting Earth from possible life returned from elsewhere. Its stated objectives include controlling forward contamination and preventing harmful backward contamination. The scientific purpose includes preserving the integrity of the search for life. NASA planetary protection overview

Forward and backward name directions, not rankings of importance. Forward concerns material carried from Earth towards another world. Backward concerns possible hazards brought towards Earth. These are distinct questions with different evidence needs. A proposal can address one while leaving the other unresolved. A reader should not accept a generic statement that “contamination was considered” without understanding which pathway and consequence were assessed.

The argument is broader than keeping an instrument clean. Suppose an exploration programme introduces material that resembles the signal a future mission hopes to detect. Even if that introduction creates no demonstrated ecological harm, it can make an important scientific question harder to answer. The observation site is part of the evidence. Changing it can change what later measurements mean. Preserving the capacity to learn is therefore a substantive benefit.

Planetary protection should also not be confused with planetary defence. Protecting against hazardous asteroid impacts is a different problem from preventing interplanetary contamination. Both concern safety, but similar names do not imply identical institutions, methods or permissions. Nor does planetary protection exhaust every environmental concern in space. Debris, effects on astronomical observation, resource use and cultural questions can require other frameworks.

COSPAR’s policy page states that its new version was published in January 2026 and introduced a unified policy for icy worlds. Readers should consult the current policy rather than assume an older summary remains sufficient. This article does not assign a category to an actual mission or translate the policy into an operational checklist. Those judgements depend on the mission and the applicable authorities. COSPAR policy page

The key lesson for AI is that a protective condition is not merely an inconvenience that a clever optimiser should route around. Some conditions exist because the consequence is irreversible or because the relevant uncertainty cannot be priced confidently. A system can help organise evidence about compliance, identify missing documentation and compare authorised alternatives. It cannot make an absent approval exist by describing the projected science as exceptionally valuable.

Contents · Previous chapter · Continue to chapter 6

6. Separate feasibility gates from preference scores

Mission proposals often combine unlike considerations in a single ranking. Scientific value, cost, schedule and institutional readiness are converted into scores and added together. Such scoring can make assumptions visible, but only if its limits are respected. Some requirements determine whether an option is admissible at all. Others help compare options already considered admissible. Mixing these layers can allow a high benefit score to conceal a missing essential condition.

In the fictional Pelagos programme, imagine three concepts. A remote observation concept stays within a previously approved study envelope. A second concept requires a new interaction with the environment. A third would eventually return material. For teaching purposes, only the first has complete evidence for the programme’s initial review gate. A score of ninety for the third does not place it above a score of sixty for the first, because the third has not entered the same decision set. These three physical mission concepts illustrate a separate feasibility gate; they are not packages A and B in the analysis-only worked programme in chapters 9–11.

This is not a claim that the other concepts must be rejected forever. It means that the current decision is about what can be responsibly selected now. The board can separately fund evidence needed for later consideration, if that research itself is authorised. “Not currently admissible”, “scientifically unpromising” and “permanently prohibited” are different conclusions. Accurate labels preserve future options while preventing premature commitment.

A gate should have an owner, a basis and evidence of satisfaction. An empty box marked “safe” is not sufficient. A model-generated assurance is not independent evidence simply because it is written in formal language. If an underlying record is missing, the correct state is unresolved. If the requirement is disputed, the dispute should be escalated through the relevant process rather than silently converted into a favourable assumption.

Preferences also need accountable ownership. Who chose to value early results twice as much as broad coverage? Who decided that a particular cost increase was worthwhile? A ranking can conceal these choices beneath arithmetic. Reviewers should be able to vary weights and see whether the preferred option changes. When small reasonable changes reverse the ranking, the decision is value-sensitive; reporting that sensitivity is more honest than presenting one winner as mathematically inevitable.

The simplest safe reading order is therefore: identify the decision, check the gates, state the permissible options, compare their consequences, test the assumptions and record who may choose. A hypothetical superintelligence could improve each analytical stage. It would still need to preserve the distinction between answering “which option scores highest?” and answering “which option is authorised, justified and chosen by the responsible people?”

Contents · Previous chapter · Continue to chapter 7

7. Treat long horizons as several different problems

“Long horizon” can mean a long calculation, a long journey, a long operating life or a long-lasting consequence. These are not the same. A system might execute a complicated plan over several days while making a decision that affects an environment for centuries. Another programme might last decades but preserve the ability to revise its aims every year. Duration alone does not reveal the kind of governance required.

One horizon is technical continuity: can the evidence, software, instruments and expertise remain usable? A second is institutional continuity: who inherits responsibility when a team changes or an organisation closes? A third is consequence duration: how long can an action affect people, environments or future scientific opportunity? A fourth is moral representation: whose interests are considered when some affected people are not yet born or cannot participate?

Hypothetical superintelligence could improve modelling without resolving disagreement about values. A model may estimate an outcome distribution under stated assumptions. It cannot discover a universally accepted moral exchange rate between a current scientific gain and an irreversible future loss merely by calculating longer. The choice of what counts as a benefit or a protected boundary remains part of public and institutional judgement.

Discounting illustrates the distinction. In a simple teaching calculation, a benefit of 100 score units received ten years later has a present-weighted value of about 82.03 at a two per cent annual discount rate, because 100 divided by 1.02 to the tenth power is about 82.03. At five per cent it is about 61.39. These are mathematical consequences of chosen rates, not proof that future people matter by either amount.

A responsible long-horizon comparison should show how conclusions change under several defensible assumptions. It should also present important consequences that are poorly captured by a score. An irreversible alteration, a scientific reference site and a distribution of risk across communities may need separate discussion. A spreadsheet is useful for exposing choices; it becomes misleading when its categories are mistaken for the complete moral world.

The practical outcome is a decision record with two layers. The first states what follows from the model. The second states why the responsible authority accepts, rejects or postpones the option despite the model’s limitations. Keeping the layers separate permits later learning. Future reviewers can see whether an error came from bad measurements, an implausible forecast, an explicit value choice or a process that excluded affected voices.

Contents · Previous chapter · Continue to chapter 8

8. Use reversible stages to buy better information

Staging is valuable when an early action produces information that can change a later decision. It is not valuable merely because a plan contains many milestones. If every stage automatically commits the programme to the next, the apparent sequence may offer little real flexibility. A genuine stage gate preserves the option to change, postpone or stop after the new evidence arrives.

Suppose a fictional programme can perform a limited remote survey before selecting a more expensive research concept. The survey is useful if it improves the decision enough to justify its own cost and delay. It is less useful if the result cannot alter the choice, if its uncertainty is too large, or if the institution has already become politically committed to proceeding. Information has value through the decisions it can change.

Reversibility is often partial. A design study can be revised, but its expenditure cannot be recovered completely. A remote observation may preserve the environment while consuming scarce observing time. A physical intervention may create a record that cannot be restored exactly. Instead of calling an entire programme reversible, ask which commitments can be reversed, which can merely be mitigated and which leave a lasting change.

AI can make the option map more detailed. It might reveal a cheaper preliminary test, identify a hidden dependency or show that two apparently competing concepts share an early research need. Yet it can also generate a false sense of flexibility by proposing later repairs that have never been demonstrated. “We can fix it later” is a claim about capability, resources and authority. Each part requires evidence.

There is an important distinction between learning more and delaying indefinitely. A demand for perfect certainty can prevent useful exploration because perfect certainty is usually unattainable. The better question is what remaining uncertainty is decision-relevant, which feasible study could reduce it, and what stopping rule determines when evidence is sufficient. If no feasible result would change the choice, another study may not be the best use of resources.

A staged programme should therefore publish an honest menu of future branches. A favourable result leads to one authorised review; an unfavourable result leads to another; an ambiguous result has its own path. Each branch needs a budget boundary and a responsible decision-maker. This structure turns curiosity into disciplined adaptation. It allows ambition to continue without pretending that the initial plan already knows everything the programme will later learn.

Contents · Previous chapter · Continue to chapter 9

9. A complete fictional decision: the Pelagos research programme

Pelagos is an invented icy world used only in this lesson. The Pelagos Research Board has thirty budget tokens for a bounded preparatory programme. It is choosing between two admissible research packages, A and B. Both consist of analysis and simulated observation planning using fictional data. Neither authorises a launch, landing, environmental intervention, sample return or operational command. The board’s decision is which research package to commission, not which physical mission to fly.

Package A is a broad comparison study. Package B is a focused study whose value depends more strongly on an uncertain feature of the fictional environment. A separate proposal, C, would require a physical interaction outside the study envelope. C is excluded from the present comparison because the necessary evidence and authorisation are absent. The board records that exclusion before looking at any scores. It does not assign C a penalty and allow a sufficiently exciting projected benefit to overcome the missing gate.

For the arithmetic, the world has two simplified states: H, meaning high relevance for B’s focused question, and L, meaning low relevance. These labels do not mean high or low probability of life, contamination or danger. The board stipulates a prior probability of 0.40 for H and 0.60 for L. These are invented probabilities for a classroom model. They are not estimates of anything on Mars, Europa, Enceladus or another real body.

The board assigns research-value scores on a teaching scale. A yields 64 points in H and 48 in L. B yields 96 points in H and 24 in L. A costs eight tokens and B costs twelve. For this toy model only, the board subtracts one point for each token of package cost. This exchange rate is explicitly a choice, not a natural law. It allows arithmetic practice while leaving the board responsible for whether the scoring rule expresses its aims.

The resulting net scores are A: 56 in H and 40 in L; B: 84 in H and 12 in L. All four are derived by subtracting the package cost from the corresponding research-value score. Archive work and a protected contingency reserve are held separately in the budget and are not subtracted again from these net scores. Because they are the same for either package, they do not change the comparison. Confusing comparison scores with total budget expenditure would double-count costs.

Before receiving additional information, A’s expected net score is 0.40 × 56 + 0.60 × 40 = 22.4 + 24 = 46.4. B’s is 0.40 × 84 + 0.60 × 12 = 33.6 + 7.2 = 40.8. Under the stated expected-score rule, A is preferred by 5.6 points. This is a conditional result: choose A if these probabilities, values, costs and decision rules are accepted. It is not a universal claim that broad studies beat focused studies.

The board also writes down what it does not know. The two-state model may omit an important intermediate state. The scores may not capture distributional fairness. The prior may be poorly supported. A and B might not be the best conceivable packages. These limitations do not make arithmetic useless. They specify the boundary within which the arithmetic answers the question and identify evidence that could change the board’s judgement.

Contents · Previous chapter · Continue to chapter 10

10. Calculate when a preliminary survey is worth buying

The board can purchase a fictional preliminary survey for three tokens. Within this preparatory programme, the survey is a simulated information-gathering analysis using fictional data, not a flight or a new observation of a real world. It produces either a positive or a negative signal about state H. Assume the survey returns positive in 80 per cent of H cases and positive in 25 per cent of L cases. Equivalently, its sensitivity is 0.80 and its specificity is 0.75 within this invented two-state model. A positive signal is not certainty; a negative signal is not exclusion.

The four joint probabilities are straightforward. H and positive: 0.40 × 0.80 = 0.32. H and negative: 0.40 × 0.20 = 0.08. L and positive: 0.60 × 0.25 = 0.15. L and negative: 0.60 × 0.75 = 0.45. They sum to one. The positive-signal probability is 0.32 + 0.15 = 0.47; the negative-signal probability is 0.08 + 0.45 = 0.53.

After a positive result, the probability of H is 0.32 divided by 0.47, or approximately 0.680851. The probability of L is 0.15 divided by 0.47, approximately 0.319149. After a negative result, the probability of H is 0.08 divided by 0.53, approximately 0.150943; the remaining 0.849057 is L. These are Bayesian updates using stipulated likelihoods. They are not a demonstration that a model’s verbal confidence is calibrated.

With a positive signal, A’s expected net score is (0.32 × 56 + 0.15 × 40) divided by 0.47 = 23.92 divided by 0.47, approximately 50.8936. B’s is (0.32 × 84 + 0.15 × 12) divided by 0.47 = 28.68 divided by 0.47, approximately 61.0213. The board would choose B. With a negative signal, A scores 22.48 divided by 0.53, approximately 42.4151, while B scores 12.12 divided by 0.53, approximately 22.8679. The board would choose A.

The value of using the survey is easiest to reproduce without rounding any posterior probabilities. On positive branches the board selects B, giving the probability-weighted contribution 0.32 × 84 + 0.15 × 12 = 28.68. On negative branches it selects A, giving 0.08 × 56 + 0.45 × 40 = 22.48. Together these yield 51.16 before paying for the survey. Subtract the survey cost of three points to obtain 48.16.

Compare 48.16 with the best no-survey value, 46.4. The expected improvement is 1.76 points. The maximum survey cost that would leave the board indifferent under this model is 51.16 − 46.4 = 4.76 tokens, using the stipulated one-point-per-token conversion. At five tokens, the survey strategy would score 46.16, which is 0.24 below choosing A immediately. A useful survey can therefore become a poor purchase if its cost changes.

This result assumes the survey arrives before the package decision and that the board genuinely follows the positive-to-B, negative-to-A rule. If the board purchases the survey but always chooses A, the information adds no decision value here and only consumes resources. If results arrive after package commitment, they cannot improve that earlier selection. Timing, willingness to revise and the credibility of the signal are all part of the value of information.

Contents · Previous chapter · Continue to chapter 11

11. Test the fictional choice before calling it robust

Expected value is one lens. Sensitivity analysis asks how much the answer relies on uncertain inputs. Let p be the probability of H before any survey. A’s expected net score becomes 40 + 16p. B’s becomes 12 + 72p. Setting them equal gives 28 = 56p, so the crossover is p = 0.50. Below one half, A wins; above one half, B wins. At one half they tie at 48 points.

This is a useful decision boundary because it tells researchers which uncertainty matters. If the plausible prior range were 0.20 to 0.40, A would remain preferred throughout that range under this scoring rule. If the plausible range were 0.40 to 0.60, the winner would change. The second case calls for greater attention to evidence about p or for an alternative decision criterion. Reporting only the result at 0.40 would hide that sensitivity.

The survey model is also uncertain. Suppose the signal were unrelated to the state: it returned positive with the same probability under H and L. Bayes’ rule would leave the posterior equal to the prior. There would be no expected selection gain, however sophisticated the report looked. A classifier that cannot discriminate the decision-relevant states is not valuable for this choice merely because it produces confident labels.

Another check is regret. In H, B’s net score exceeds A’s by 28 points. In L, A exceeds B by 28 points. Thus each fixed package has a maximum regret of 28 across the two states. A minimax-regret criterion does not uniquely select A or B in this toy case. That difference from the expected-value ranking is not an arithmetic error. The criteria express different ways of handling uncertainty and deserve an explicit institutional choice.

The resource check uses actual token costs, not expected scores. If the board buys the survey and later chooses B, spending is three for the survey, twelve for B and four for archive work, totalling nineteen. Keeping six tokens protected for contingency brings the total committed-or-reserved amount to twenty-five, leaving five unallocated from thirty. With A, the analogous amount is three + eight + four + six = twenty-one, leaving nine unallocated. Both branches fit the budget.

The six-token reserve is not a discount applied to risk, and unused reserve is not additional scientific return. It remains unavailable for optional expansion under the board’s rule. An AI that “improves” the plan by silently spending it has changed the authorisation, even if the expected score rises. This is why budget feasibility and objective optimisation must be tested separately. A valid numerical optimum can still be an invalid programme choice.

The board’s conditional decision is now complete: buy the three-token survey, choose B after positive and A after negative, fund the four-token archive task, preserve the six-token reserve, and keep C outside scope. Before commissioning, the board must accept the scoring assumptions and verify that the survey’s stated performance is an appropriate input. If that evidence fails, the decision returns to review rather than continuing under a fabricated certainty.

Contents · Previous chapter · Continue to chapter 12

12. Make the evidence package outlive the model

A programme decision is only as inspectable as the records behind it. The Pelagos board therefore requires a small evidence package. It includes the question, the admissible options, all numerical inputs, the source or stipulated status of each input, the equations, the decision rule and the approval record. A future reader should be able to reproduce 46.4, 40.8, 51.16 and 48.16 without access to the original AI conversation.

Versioning matters because a correct calculation can be applied to the wrong inputs. Suppose the survey cost changes from three to five after a supplier clarification. An assistant that continues to display the old expected improvement of 1.76 has created a stale conclusion. The corrected result is negative 0.24 relative to immediate A. The record should show both the changed input and the reason the recommendation changed, rather than overwrite history as though the new recommendation had always been obvious.

The package should preserve disagreement as well as consensus. A reviewer may accept the arithmetic but reject the one-point-per-token exchange rate. Another may question the prior. A third may identify a missing package. These are different objections with different remedies. Collapsing them into “review completed” loses information. A useful decision log states which objections were resolved, which remain and who accepted the consequences of proceeding.

For AI-assisted work, reproducibility includes the distinction between generated suggestions and verified facts. The model might propose that a particular archive format is suitable. A responsible person must check that claim against actual requirements. The model might calculate a posterior. A simple independent calculation can verify it. The model might describe a mission approval as complete. That requires an authoritative record, not another paraphrase of the same unsupported assertion.

A durable archive also needs an audience. What can a scientist outside the original team understand? Can a student identify which numbers are fictional? Can an administrator see what was authorised? Can a later reviewer reconstruct the selection without inferring intentions from file names? Documentation quality is partly the distance between what the authors remember and what an unfamiliar reader can establish independently.

The lesson for hypothetical superintelligence is demanding but straightforward. Its outputs should increase the quality of the evidence left behind. A system that generates brilliant recommendations but makes the institution dependent on its undocumented judgement can reduce long-term agency. A system that exposes assumptions, supports independent reconstruction and preserves revision history may be less theatrical while making the programme much more trustworthy.

Plan the end of responsibility, not only the end of funding

A long-lived programme can lose its institutional memory before it loses its physical effects. People retire, contracts end and software suppliers change. A decision made carefully at the beginning can become difficult to interpret later if nobody inherits the rationale. The answer is not to assume that an AI will remember everything indefinitely. The programme should define the human and organisational responsibility for preserving, reviewing and transferring its records.

In the fictional Pelagos study, closure requires more than receiving a final report. The board checks that the numerical inputs are archived, the chosen package has delivered its specified analysis, unresolved questions are labelled and the unused budget is accounted for. Because this is a study rather than a flight, there is no spacecraft disposal decision. That boundary should remain explicit; a classroom programme should not acquire imaginary operational responsibilities merely to make its story more dramatic.

For a real physical programme, the same reasoning prompts different, specialist questions. Which obligations survive completion of the primary science? Who maintains the evidence needed to assess later effects? What happens if the original organisation ceases to exist? What resources have been committed to continuing responsibilities? This article does not answer those mission-specific questions, but a reader can recognise when a proposal omits them.

Ending one activity can also create a new scientific opportunity. A preserved dataset may support a later reinterpretation. A documented null result may prevent another team from repeating an uninformative search. A record of why an option was rejected may become useful when technology or policy changes. Closure should therefore preserve the distinction between a completed obligation and an unanswered question. The first can end; the second should remain intelligible.

A model update should not silently rewrite an institution’s past. If a new assistant gives a different interpretation of an old decision, that interpretation belongs beside the original evidence and rationale. It does not replace them as historical fact. This is especially important when future systems become more persuasive: improved analytical capability should make the record easier to interrogate, rather than make earlier uncertainty disappear from view.

Contents · Previous chapter · Continue to chapter 13

13. Decide who has authority, including when people disagree

Space programmes involve scientific teams, engineers, funders, regulators, partner organisations and affected publics. These groups do not necessarily ask the same question. A researcher may seek a discriminating observation; an engineer may worry about verified performance; a funder may prioritise affordability; a community may question how benefits and risks are distributed. A capable assistant should make these differences clearer, rather than present one technical objective as if it represented everyone.

The Outer Space Treaty is part of the legal background. Articles VI and IX of the treaty address state responsibility, harmful contamination and international consultation. The particular obligations for a real activity depend on the applicable legal and institutional setting. This educational article is not legal advice or a finding that any proposed activity is permitted. Mission teams should use current authoritative documents and qualified legal review. United Nations Treaty Series: Outer Space Treaty

The broader governance argument here is normative: a technical recommendation should not erase the people responsible for deciding. A model cannot acquire public legitimacy by being unusually persuasive. Nor should a board outsource its explanation to the model’s asserted superior understanding. If people cannot evaluate every internal calculation, they still need evidence standards, independent checks, bounded commitments and accountable reasons for proceeding.

Disagreement should be structured, not treated as a defect to be optimised away. One group may support a study only if its results remain openly accessible. Another may require a stronger case for protecting future research opportunities. A third may prefer investment in a different scientific question. The purpose of deliberation is not always to discover that one group was mathematically wrong. It can be to negotiate priorities within legitimate constraints.

A useful review separates factual disputes, forecast disputes and value disputes. A factual dispute might concern whether a record exists. A forecast dispute might concern how likely an outcome is. A value dispute might concern whether a particular irreversible consequence is acceptable. The first calls for evidence retrieval, the second for modelling and uncertainty analysis, and the third for authorised judgement and participation. Using the wrong remedy can waste time or conceal a political choice.

Institutions also need a way to hear bad news. If the assistant’s success is judged only by whether the programme advances, it may be rewarded for removing obstacles in the narrative. A better evaluation rewards accurate identification of missing evidence and legitimate reasons to pause. The desired outcome is not maximum forward motion. It is a decision that remains defensible when its assumptions, trade-offs and dissenting views are visible.

Contents · Previous chapter · Continue to chapter 14

14. Examine powerful future scenarios without turning them into promises

A future system might analyse scientific literature far faster than a research team, design unfamiliar instruments or propose observation strategies nobody had considered. Those possibilities can motivate preparation. They should not be attached to a prediction calendar without evidence. The relevant question is which capability would have to be demonstrated, under what conditions, before a particular responsibility could reasonably be delegated.

Consider three deliberately hypothetical futures. In the first, AI becomes a much better research assistant but physical experimentation remains slow and expensive. The main gain is better selection among limited observations. In the second, analysis and robotic experimentation both improve, but international coordination remains difficult. The limiting factor may become agreement on acceptable action. In the third, capable systems become cheap and widespread, creating many competing proposals and a greater burden on shared environments and institutions.

These scenarios suggest different preparations. The first rewards strong experimental design and data preservation. The second rewards transparent decision procedures and cooperation. The third rewards enforceable boundaries, credible identity and accountability, and the ability to reject activities that impose costs on others. None follows automatically from the phrase superintelligence. The point is to identify dependencies that remain important across several plausible futures.

Self-expanding infrastructure deserves particular caution in discussion. Being able to design a machine is not the same as demonstrating that an entire system can source materials, manufacture replacements, maintain itself and expand reliably. Even a demonstrated ability to expand would not settle whether expansion should be authorised. This article does not provide designs or instructions for autonomous replication, environmental release or bypassing supervision. It treats those possibilities as governance questions requiring separate evidence and authority.

A headline such as “AI will make humanity multiplanetary” bundles many claims. It assumes technical capability, material systems, sustained resources, reliable operation, human agreement and an acceptable relationship with environments beyond Earth. Breaking the headline into dependencies does not eliminate ambition. It turns an impressive sentence into a research agenda whose steps can be examined and revised.

The most useful preparation is often robust across futures: maintain scientific literacy, preserve evidence, distinguish permission from capability and insist on clear accounts of uncertainty. These are valuable whether progress is fast, uneven or disappointing. Preparing for a hypothetical superintelligence should not require pretending that its arrival date, motives, reliability or consequences are already known.

Contents · Previous chapter · Continue to chapter 15

15. Exercises: use the distinctions independently

Exercise one concerns evidence labels. A report says that an AI produced rover waypoints in a reviewed demonstration, then concludes that future spacecraft will choose their own scientific and environmental priorities safely. Identify the demonstrated claim, the extrapolation and the missing evidence. Rewrite the report’s conclusion so that it recognises the achievement without presenting the extrapolation as established. Your answer should mention the different scope of route planning and mission governance.

Exercise two concerns timing. In a fictional scenario with a twelve-minute one-way delay, a local decision has a five-minute deadline. Calculate the fastest possible message round trip, ignoring processing time. Explain why “we will ask Earth” does not fully define the response. Give one high-level design requirement that belongs in the pre-agreed authority envelope. Do not invent an operational command or a real mission rule.

Exercise three concerns the Pelagos baseline. Use A’s net scores of 56 and 40, B’s of 84 and 12, and probabilities 0.40 and 0.60. Recalculate both expected scores, select a package under the expected-score criterion and state the size of the advantage. Then change the probability of H to 0.60. Does the same package remain preferred? Show the arithmetic rather than relying on the earlier recommendation.

Exercise four concerns the survey. Starting from the four joint probabilities 0.32, 0.08, 0.15 and 0.45, calculate the positive probability and the posterior probability of H given a positive signal. Explain why 0.80 is not the posterior probability. Then calculate the posterior after a negative signal and state which package is selected in each branch under the stipulated model.

Exercise five concerns information value. Calculate the expected score of the branch-dependent strategy before the survey cost, then after a cost of three. Compare it with the best no-survey option. Repeat using a survey cost of five. Explain what changes and what does not: the survey’s information quality remains the same, but the purchase decision may differ. Include the maximum break-even survey price in your answer.

Exercise six concerns authority. A system proposes spending the protected six-token reserve to add an attractive optional study. Its calculation is correct and the total still fits thirty tokens. Has it remained within the board’s stated decision? Explain the difference between budget feasibility and permission. Identify the next step without pretending that a higher score provides the missing authorisation.

Exercise seven concerns planetary protection. A fictional press release says a mission has addressed contamination because its instrument produces clean images. Explain why this is insufficient. Distinguish instrumental image quality, forward contamination and backward contamination. Identify what kind of evidence the reader should seek while remaining at the policy and assurance level. Do not propose laboratory procedures.

Exercise eight concerns future generations. A proposal shows that a distant benefit has a low present-weighted score under one discount rate and concludes that irreversible impacts need no further discussion. Explain the unsupported step. Suggest two additional views of the decision that could inform deliberation without claiming that either automatically supplies the morally correct answer.

Contents · Previous chapter · Continue to chapter 16

16. Full answers and what each answer teaches

Answer one: the demonstrated claim is that a named AI-assisted process produced waypoints used in a bounded, reviewed rover demonstration. The extrapolation is that future spacecraft can safely set scientific and environmental priorities independently. Missing evidence includes transfer across tasks, reliable adherence to constraints, performance under unfamiliar conditions and legitimate authority to make those choices. A careful conclusion is: the result supports further evaluation of AI-assisted route planning within tested safeguards; it does not establish autonomous mission governance. The lesson is to preserve the size of the evidence rather than stretch it to match the excitement of the headline.

Answer two: twelve minutes outward plus twelve minutes back gives a minimum twenty-four-minute round trip. Five minutes is nineteen minutes shorter than that minimum, so a new response from Earth cannot arrive in time under the assumptions. The programme needs an approved local response envelope, including what decisions may occur without new approval and what optional activity must pause when conditions are unresolved. This is a governance and engineering requirement, not a recommendation for a particular spacecraft action. Real systems must define their own validated behaviour.

Answer three: at p = 0.40, A scores 0.40 × 56 + 0.60 × 40 = 46.4; B scores 0.40 × 84 + 0.60 × 12 = 40.8. A leads by 5.6. At p = 0.60, A scores 33.6 + 16 = 49.6; B scores 50.4 + 4.8 = 55.2. B now leads by 5.6. The reversal shows that the recommendation depends on the probability input. A correct calculation at one prior is not evidence that the same recommendation is robust across all plausible priors.

Answer four: the positive probability is 0.32 + 0.15 = 0.47. The posterior H probability is 0.32/0.47, approximately 0.680851. The value 0.80 is the probability of a positive signal given H; it is not the probability of H given positive. These conditional probabilities have different denominators. After a negative signal, the H probability is 0.08/(0.08 + 0.45), approximately 0.150943. The stipulated expected-score rule selects B after positive and A after negative. This is a compact demonstration of why a test’s advertised sensitivity does not directly tell a reader what a particular result means.

Answer five: using B on positive branches contributes 0.32 × 84 + 0.15 × 12 = 28.68. Using A on negative branches contributes 0.08 × 56 + 0.45 × 40 = 22.48. The total is 51.16 before survey cost. At cost three, the net is 48.16, exceeding 46.4 by 1.76. At cost five, the net is 46.16, falling below 46.4 by 0.24. The break-even cost is 4.76. The information is unchanged, but its net decision value changes with price. This teaches why “more evidence is useful” and “this evidence is worth buying” are separate propositions.

Answer six: no. The reserve was explicitly protected from optional expansion. A proposal can fit the total budget while violating how that budget was authorised. The system should flag the proposed change and return it to the board for a new decision, preserving the original reserve until approval is actually recorded. If the board later changes the rule through its proper process, the new authorised plan can be evaluated. The initial model should not pretend that arithmetic already supplied that institutional action.

Answer seven: clean images concern the quality of an instrument’s output, not whether Earth-derived material reaches another world or possible hazards return to Earth. Forward contamination concerns the first direction and backward contamination the second. The reader should seek the mission-specific protection assessment, responsible authority, applicable requirements and evidence that relevant conditions have been addressed. The answer is not to infer compliance from a polished image or to improvise technical procedures. This teaches the importance of matching assurance evidence to the actual claim.

Answer eight: discounting applies a selected weighting convention to a modelled benefit. It does not prove that all distant interests are negligible or that irreversible impacts can be ignored. Additional views could include a sensitivity analysis across discount rates, a separate account of irreversible changes, a distributional analysis of who receives benefits and risks, or a staged alternative that preserves future choice. These views inform deliberation rather than replace it. A responsible decision records why its chosen trade-offs remain acceptable despite the limitations of any single numerical representation.

Contents · Previous chapter · Continue to chapter 17

17. A practical review before believing the next headline

Begin with the verb. Did the system propose, simulate, classify, plan, execute or independently verify? These describe different achievements. Next identify the object: an image, a route, an instrument design, a scientific inference or an entire mission. Then identify the conditions: available data, human preparation, expert intervention, permitted actions and the setting of the test. A headline often becomes much clearer once those three pieces are written down.

Ask what the comparison excludes. If the report says a model produced plans faster, does it include the time experts spent checking them? If it says science return increased, does it define science return before measuring the result? If it says the system was autonomous, which choices remained fixed by people? These questions are not hostile. They are how a reader gives a real achievement its proper meaning.

For environmental claims, ask whether the proposal distinguishes a modelled low risk from demonstrated satisfaction of the applicable requirements. Ask which uncertainty remains unresolved and who has authority to accept it. For long-horizon claims, ask who preserves the evidence, who inherits the obligation and what future decision remains open. A plan with a clear end date can still leave responsibilities after that date.

Finally, look for a counterfactual. What would happen without the AI, without the proposed mission or with a smaller staged alternative? If every comparison is between a spectacular success and doing nothing, the option set may be artificially narrow. Good reasoning includes credible alternatives and explains why the chosen one is preferable under explicit assumptions.

A student can practise this on a short article without becoming an aerospace specialist. Write four sentences: what was shown, what was not shown, what would justify the next claim and who decides whether the next action is acceptable. That exercise links scientific literacy with civic judgement. It rewards accurate curiosity rather than reflexive enthusiasm or reflexive dismissal.

Contents · Previous chapter · Continue to chapter 18

18. The durable conclusion: explore without erasing the next question

The promise of more capable AI in space is not merely that it might make a machine move faster or generate more observations. It could help people ask better questions, notice overlooked alternatives and make limited resources more informative. Those gains matter when they survive independent checking and remain connected to the purposes of exploration.

The boundary is equally important. Intelligence does not automatically confer authority, and a high expected score does not dissolve a protective requirement. A fictional model can teach reasoning without becoming a mission plan. A current demonstration can support confidence in a particular application without proving a hypothetical future capability. A long time horizon can widen responsibility rather than excuse vague promises.

The Pelagos example ended with a bounded decision, a reproducible calculation, a protected reserve and an excluded option whose prerequisites were incomplete. It did not end with certainty about the world. That is a realistic form of intellectual progress: make a choice within a stated model, preserve the reasons, identify what would change it and keep the next decision genuinely open.

For readers, the central habit is to protect the question as carefully as the answer. Preserve the observations, the alternative explanations, the environments that future researchers may need and the people’s ability to revise their aims. If hypothetical superintelligence eventually helps exploration, one measure of its value should be whether it expands that capacity for responsible learning rather than using present capability to close the future prematurely.

Sources and further reading

The concrete current-AI example comes from JPL’s report, dated 30 January 2026, of Perseverance’s AI-planned drives. It supports the bounded historical demonstration described here, not claims about general superintelligence.

The definitions of forward and backward protection are grounded in NASA’s planetary protection overview. For current international scientific policy context, consult COSPAR’s planetary protection policy page, which identifies the January 2026 revision. These sources do not supply the fictional Pelagos numbers.

For preserved scientific data, see NASA’s Planetary Data System. For treaty context, consult the United Nations Treaty Series text of the Outer Space Treaty. The legal discussion is a limited educational orientation; current applicability and implementation require authoritative mission-specific review.

The Pelagos programme, utility scale, priors, survey characteristics, budgets, review questions and exercises are original teaching constructions. The governance recommendations are the article’s reasoned proposals. They should not be read as NASA or COSPAR endorsements, official mission requirements or predictions of future capabilities.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading