VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How to Run a Super Intelligence Readiness Audit for Your Workplace | SI Readiness Checklist

eduKate Secondary students reviewing open books for How Super Intelligence Works: Attention.

A Super Intelligence readiness audit asks whether a workplace has enough process clarity, reliable information, human ownership, technical access, verification and governance to run a useful SI workflow without pretending uncertainty has disappeared. The audit is not a score of how “advanced” a company looks. It is a practical test of whether one proposed workflow is ready for the next step.

A team can own the newest models and still be unready if nobody knows which policy is current, who owns the process or how the output will be checked. Another team with simple tools may be highly ready because the workflow is stable, data is accessible, responsibility is clear and the outcome is measurable.

This article gives a complete Super Intelligence workplace readiness audit for organisations, managers, operations teams and technical builders. It covers outcome readiness, workflow readiness, task fit, knowledge, data, permissions, security, privacy, human skills, review capacity, governance, measurement, recovery and change management.

It is Article 18 in the Workplace Super Intelligence series, following time-to-value.


The Short Answer

A workplace is ready to pilot Super Intelligence when it can answer five fundamental questions: What outcome are we improving? What information does the task require? What may the system do? How will we know when it is wrong? Who owns the result?

A deeper audit expands those questions into twelve domains: Outcome, Workflow, Task, Knowledge, Data, Tools, Permissions, Human Capability, Verification, Governance, Measurement and Recovery.

Do not use the audit to produce false precision. Use it to find the first readiness gap. Repair that gap, retest the workflow and then decide whether the project should remain assistive, enter a pilot, deepen integration or stop.

Why Readiness Is Workflow-Specific

A company is not simply “AI-ready” or “not AI-ready.” Readiness depends on the use case. The same organisation may be ready to use SI for internal meeting summaries but unready to let an agent modify customer records.

The difference may come from consequence, data sensitivity, source quality, technical integration or review capacity. This is why a readiness audit should begin with a specific workflow rather than a generic maturity questionnaire.

Organisation-wide readiness still matters because common capabilities—identity, knowledge, security, governance and training—can support many workflows. But the final go/no-go decision belongs at the workflow level.

The Four Readiness Outcomes

Ready to test

The workflow is bounded, the required context is available, the output can be verified and the pilot can run with acceptable consequence. Begin a controlled pilot and collect evidence.

Ready with conditions

The central mechanism can be tested, but one or more boundaries must remain manual or tightly controlled. For example, SI may prepare an action while a human continues to execute it.

Repair before pilot

A foundational gap prevents a meaningful test. The current source may be unclear, the workflow may be unstable or the acceptance standard may not exist. Repair the underlying process first.

Not appropriate for this design

The proposed SI role creates little value, cannot be verified, depends on unavailable data or transfers inappropriate authority. Choose a different role, narrower task or conventional software solution.

Domain 1 — Outcome Readiness

The first audit asks whether the organisation knows what useful change it wants. “Use AI in HR” is not an outcome. “Reduce the time required to prepare an onboarding answer while keeping policy accuracy and source traceability” is closer to an auditable objective.

  • Is the receiver of value known?
  • Can success be observed?
  • Is the outcome more meaningful than usage volume?
  • Does the outcome belong to an existing process owner?
  • Can the team explain why this outcome matters now?

If the outcome is unclear, stop. A powerful system cannot rescue an undefined objective.

Domain 2 — Workflow Readiness

The organisation should understand how the work currently moves from trigger to completion. This does not require perfect process documentation, but the main states, handoffs and decisions should be visible.

  • What starts the work?
  • Which tasks occur?
  • Where does information enter?
  • Who makes consequential decisions?
  • Where does the work wait?
  • What marks completion?
  • Which exceptions leave the normal path?

If different employees describe completely different processes, the SI project may be exposing process instability rather than a machine-intelligence opportunity. Stabilise enough of the workflow to run a meaningful test.

Domain 3 — Task Readiness

The task should be bounded at the level of the cognitive operation. Summarise, extract, compare, classify, draft, retrieve, plan, monitor and act are different operations with different verification methods.

Task readiness is stronger when frequency is meaningful, context is available, output standards are clear, mistakes are detectable and the action is reversible.

Use What Should Super Intelligence Do at Work? for the deeper task-fit method.

Domain 4 — Knowledge Readiness

Knowledge readiness asks whether the organisation knows which information is authoritative. Intelligent retrieval cannot compensate for an organisation that maintains several conflicting versions of the same policy without ownership.

  • Are canonical sources identified?
  • Does each important source have an owner?
  • Can obsolete versions be distinguished from current ones?
  • Are important decisions recorded?
  • Is internal terminology defined where necessary?
  • Can the system identify when no source supports an answer?

Low knowledge readiness is one of the most important reasons to keep early workflows assistive. The pilot may help reveal documentation gaps, but the system should not silently manufacture institutional knowledge.

Domain 5 — Data Readiness

Data readiness asks whether required information is available in a usable, permitted and sufficiently current form. The data does not need to be perfectly structured. Modern systems can interpret unstructured material, but missing identity, date, version or ownership can still make the result unreliable.

  • Is the required data accessible?
  • Is currentness known?
  • Are key identifiers stable?
  • Are definitions consistent across systems?
  • Can missing fields be detected?
  • Can sensitive data be limited to what the task needs?

A large dataset is not automatically a ready dataset. Quality, relevance, permission and provenance matter.

Domain 6 — Tool and Integration Readiness

Some workflows need only an assistant. Others need email, calendar, document repositories, databases, CRM, project tools, code repositories or business systems. Every new connection adds capability and operational dependency.

  • Is the integration available and supported?
  • Can it authenticate the correct user or service identity?
  • Can read and write permissions be separated?
  • Does the tool return a reliable success or failure state?
  • Can failed actions be retried safely?
  • Can actions be reversed when necessary?

If the central value can be tested manually, do that before building expensive integration. Integration readiness matters most when connected action is the thing being tested.

Domain 7 — Permission Readiness

Permission readiness separates what the system may see from what it may change. An SI workflow should not inherit broad access simply because it runs under an employee’s account.

  • What may the system read?
  • What may it transform?
  • What may it recommend?
  • What may it prepare?
  • What may it execute?
  • What requires human approval?
  • Which actions are prohibited?
  • What happens when the required permission is unavailable?

Agentic AI makes these questions especially important because a system can take actions across tools. IMDA’s updated 2026 Model AI Governance Framework for Agentic AI recommends assessing and bounding risks, placing limits on powers, defining meaningful human checkpoints and implementing technical controls. See IMDA’s updated framework.

Domain 8 — Security and Privacy Readiness

A workflow that handles public information and a workflow that handles customer records should not share the same readiness bar. Security and privacy readiness depends on the data, integrations, threat model and legal obligations.

  • Is the tool approved for the data class involved?
  • Is access limited to authorised users?
  • Are secrets, credentials and privileged data protected?
  • Can untrusted content influence the system?
  • Are external actions constrained?
  • Are logs available for important actions?
  • Are retention and deletion requirements understood?
  • Can the organisation respond to an incident?

Where personal data is used, organisations should follow applicable privacy law and policy. In Singapore, the PDPC has advisory guidelines on the use of personal data in AI recommendation and decision systems. See PDPC guidance.

Domain 9 — Human Capability Readiness

Employees need enough domain understanding to direct and verify the system. A novice may obtain a fluent answer but lack the knowledge to recognise a subtle failure. An expert may use SI effectively with less explanation but can also become overconfident in a system that often performs well.

  • Do users understand the task without SI?
  • Do they know which sources are authoritative?
  • Can they recognise common failure modes?
  • Do they know when to escalate?
  • Can they explain what remains their responsibility?
  • Can critical skills be preserved for verification and recovery?

Training should focus on task understanding, context, evidence and review—not only prompt techniques.

Domain 10 — Verification Readiness

Verification readiness asks whether the workflow can detect important errors at a reasonable cost. Different claims deserve different checks.

  • Can source-backed claims be traced?
  • Can calculations be checked deterministically?
  • Can structured outputs be validated against a schema?
  • Can code be tested?
  • Can classifications be compared with labelled examples?
  • Can human reviewers focus on material judgment rather than rereading everything?

If verification requires repeating the entire task, the project may have weak economics. Narrow the SI role or improve the checking mechanism.

Domain 11 — Governance and Accountability Readiness

Someone must own the workflow. Responsibility should not be distributed so widely that nobody can decide whether the system should continue operating.

  • Who owns the business outcome?
  • Who owns the authoritative knowledge?
  • Who owns the technical integration?
  • Who reviews material output?
  • Who approves consequential actions?
  • Who decides when to pause or change the system?
  • Who receives incident reports?

NIST’s AI Risk Management Framework and Generative AI Profile emphasise lifecycle risk management and organisational responsibility across design, development, use and evaluation. The Generative AI Profile is voluntary, but its lifecycle framing is useful for readiness audits. See NIST AI 600-1.

Domain 12 — Measurement and Recovery Readiness

A workplace should know what it will measure and what it will do when the workflow fails. These questions are linked: measurement reveals degradation; recovery limits consequence.

  • Is there a baseline?
  • Is the first-value milestone defined?
  • Can cycle time, human effort, quality and rework be measured?
  • Can exception and incident rates be tracked?
  • Can the workflow be paused?
  • Can external actions be reconstructed?
  • Can the organisation return to a known manual or previous process?
  • Is there a trigger for reassessment?

A workflow without measurement cannot demonstrate value. A workflow without recovery can become brittle as autonomy increases.

The Readiness Audit Sequence

  1. Outcome
  2. Workflow
  3. Task
  4. Knowledge and data
  5. Verification
  6. Human capability
  7. Permissions and security
  8. Tools and integration
  9. Governance
  10. Measurement
  11. Recovery
  12. Adoption

Run the audit in sequence rather than treating all domains as equal. The first failed foundation may make later questions premature. Do not spend weeks auditing agent permissions for a task whose acceptance standard is still undefined.

The First Weak Link Rule

Readiness is constrained by the earliest missing condition that prevents a meaningful test. A workflow with perfect security but no authoritative source is not ready. A workflow with excellent data but no accountable owner is not ready for consequential action.

Find the first weak link, repair it and rerun the audit. This creates progress without requiring the entire organisation to become “AI mature” before any useful work begins.

Readiness by Autonomy Level

Assist

The lowest readiness threshold. The human initiates the task, supplies context and reviews the output. Strong domain knowledge and safe tool use may be sufficient for a bounded low-risk pilot.

Collaborate

Requires reusable context, clearer output standards and a meaningful review process. The human and SI divide substantial parts of the work.

Automate

Requires stable triggers, exception handling, monitoring, tool reliability and clear ownership. The process can progress without constant manual initiation.

Operate or agentic

Requires stronger permissions, logging, tool controls, checkpoints, recovery and governance because the system can select or execute actions across multiple steps.

A workplace can be ready at one autonomy level and not ready at the next. This is normal and useful.

Worked Example: Policy Assistant

An internal policy assistant seems straightforward, but the audit finds that multiple document versions are stored in shared drives and no one owns retirement of obsolete copies. Knowledge readiness is weak.

The correct first project is not a more advanced retrieval model. It is establishing canonical policy sources and ownership. Once repaired, the organisation becomes more ready for both human search and SI retrieval.

Worked Example: Customer Email Drafting

The team has approved knowledge and clear tone guidelines. The task is reversible because humans review before sending. Outcome, task and verification readiness are strong.

However, customer-specific facts live in a restricted system the model cannot access. The pilot can still proceed with manually supplied context. Deeper integration waits for approved access.

Worked Example: Automated Refunds

An organisation wants an agent to issue refunds. The audit must examine policy rules, transaction limits, fraud risk, identity, account data, approval thresholds, logging and rollback. A system may be ready to recommend refunds long before it is ready to execute them.

The readiness audit therefore lowers the autonomy level rather than rejecting all SI involvement.

Worked Example: Software Coding Agent

A coding agent can create value in a sandbox quickly, but production readiness depends on repository permissions, testing, secrets management, code review, deployment controls and rollback.

The audit should separate permission to read code, create a branch, open a pull request and deploy. These are distinct authority levels.

Worked Example: Recruitment Workflow

The organisation may be ready to use SI for scheduling, job-description drafting and interview-note organisation while being unready to automate candidate ranking. Different tasks inside the same HR workflow receive different readiness outcomes.

This is why readiness should remain task-specific even within one department.


What a Super Intelligence Readiness Audit Is Actually For

A readiness audit is not a certification exercise and it is not a way to assign a fashionable maturity score. Its purpose is to find the first organisational condition that would make a proposed SI workflow unreliable, slow, unsafe or impossible to measure.

The audit therefore looks at the whole operating environment: workflow clarity, source quality, data, human capability, permissions, verification, exceptions, measurement, governance, recovery and ownership. A single severe gap can matter more than strong performance elsewhere.

Audit Scope: Organisation, Department or Workflow?

The audit can be run at three levels. An organisation-wide audit asks whether common infrastructure and policy exist. A department audit asks whether a function can support several related use cases. A workflow audit asks whether one specific process can safely move to a higher level of SI participation.

Start with the smallest scope that matches the decision. If the immediate question is whether to automate one reporting workflow, a company-wide transformation audit may create more work than insight.

The Ten Evidence Categories

  1. Workflow evidence: current-state maps, SOPs, examples and actual case paths.
  2. Knowledge evidence: canonical documents, source owners, version history and retirement rules.
  3. Data evidence: systems of record, field quality, currentness and reconciliation.
  4. People evidence: user skills, reviewer capability, tacit knowledge and workload.
  5. Permission evidence: current access, write authority, approval thresholds and identities.
  6. Verification evidence: tests, rules, review checklists, source references and audit trails.
  7. Exception evidence: escalation queues, edge cases, ownership and case volume.
  8. Measurement evidence: baselines, cycle times, errors, rework, cost and outcome metrics.
  9. Governance evidence: approved tools, data rules, incident procedures and accountability.
  10. Recovery evidence: rollback, fallback, outage procedures and continuity.

An audit should prefer observable evidence over confidence. “We have good documentation” is weaker than identifying the current policy, its owner and the last update date.

Audit Dimension 1 — Outcome Clarity

Ask whether the team can state the real-world result of the workflow. Vague goals such as “use AI for customer service” are not enough. A better statement identifies the accepted result, receiver and timing.

Evidence can include service definitions, accepted deliverables and business metrics. A weakness here means the project risks optimising activity rather than outcome.

Audit Dimension 2 — Workflow Visibility

Can the team map trigger, intake, context, tasks, decisions, handoffs, verification, action, closure and learning? If not, the workflow may contain hidden states that later become automation failures.

Use the workflow-mapping article as the evidence standard.

Audit Dimension 3 — Source Authority

Which documents and systems are authoritative? Can obsolete versions be identified? Who owns updates? Can the SI system retrieve the right source without relying on personal memory?

A workplace with strong models and weak source authority is not ready for source-grounded automation.

Audit Dimension 4 — Data Reliability

Which structured fields matter for the workflow and how reliable are they? Are customer records current? Do project systems reflect real milestones? Do financial systems reconcile?

The audit should focus on the data required for the use case rather than attempting to judge every dataset in the enterprise.

Audit Dimension 5 — Task Fit

Is the proposed SI role well matched to the work? The system may be strong at extraction, comparison and drafting while the real task depends on negotiation or institutional judgment.

Use task cards and the repetitive-versus-judgment framework to show why SI belongs where proposed.

Audit Dimension 6 — Human Capability

Can users frame the task, supply context, detect unsupported output and escalate? Can reviewers perform the checks assigned to them? Do employees know which critical skills must remain alive for recovery?

Training certificates are weaker evidence than real exercises with flawed outputs and exception cases.

Audit Dimension 7 — Permission Control

Can the organisation separate read, interpret, recommend, prepare and execute permissions? Is least privilege practical? Are service accounts and human identities distinguishable?

Broad all-or-nothing access is a readiness gap, especially for tool-using agents.

Audit Dimension 8 — Verification

How are important claims and actions checked? Source evidence, deterministic calculation, schemas, tests, specialist review and world-return confirmation may all play roles.

The audit should ask whether verification cost is sustainable at projected volume.

Audit Dimension 9 — Exceptions

Which cases leave the normal path, how often and who owns them? Does the exception packet preserve verified context, or must specialists rebuild the case from scratch?

High exception volume can make a nominally automated process operationally unready.

Audit Dimension 10 — Measurement

Is there a baseline for cycle time, touch time, error, rework, queue length, receiver effort or outcome quality? If not, the organisation may be unable to prove value after deployment.

A readiness audit should identify the smallest set of metrics required for the proposed project.

Audit Dimension 11 — Governance

Are approved tools, data boundaries, action rules, human checkpoints and incident procedures explicit? Do written policies reflect actual behaviour?

Governance should become more specific as autonomy and failure radius increase.

Audit Dimension 12 — Recovery

Can the workflow be stopped, downgraded or rolled back? Is there a fallback during outage? Can uncertain transactions be reconciled?

Readiness for action requires more recovery capability than readiness for drafting.

Audit Dimension 13 — Ownership

Who owns the business outcome, technical implementation and knowledge source? One person may hold multiple roles, but ownership must be visible.

A workflow with no owner tends to accumulate local fixes without systematic improvement.

Audit Dimension 14 — Change Control

What happens when the model, source set, permissions, tools or policy changes? Important workflows should know what counts as a material change and when representative tests must be rerun.

Audit Dimension 15 — Continuity

What happens if the SI service or connector is unavailable? Critical workflows need a fallback proportionate to their importance.

Continuity is also a test of human capability. If nobody can operate the fallback, the organisation may have created brittle dependence.

The Four Audit Ratings

Use qualitative states rather than false precision: Unknown, Fragile, Stable and Proven.

Unknown

The condition has not been established or evidence is missing.

Fragile

The condition exists but depends on one person, informal practice or inconsistent execution.

Stable

The condition works repeatedly across normal cases and users.

Proven

The condition has survived representative edge cases, material changes, recovery or sustained operating volume.

Do not average these states into one readiness percentage. The weakest required dimension is often the one that limits the next autonomy level.

The Red-Flag Rule

Some findings should stop a move to higher autonomy regardless of strengths elsewhere: unclear authority over consequential action, no way to verify important output, sensitive data in an unapproved environment, no exception owner, or inability to determine whether external actions succeeded.

A red flag does not require abandoning SI. It usually means lowering the role to assistance or preparation while the control is repaired.

The Green-Light Rule

A workflow is a stronger candidate when the current process is visible, sources are current, task fit is strong, review is affordable, permissions can be scoped, exceptions are owned and the outcome can be measured.

Green light means ready to pilot or advance one level, not permission for unlimited autonomy.

Audit Interviews: Process Owner

  • What outcome does this workflow exist to produce?
  • Which step limits performance today?
  • What errors matter most?
  • Which decisions require your authority?
  • What would make you stop the SI project?
  • Which metric would prove useful improvement?

Audit Interviews: Frontline Users

  • Where do you search for information?
  • What do you copy between systems?
  • What cases make you ask for help?
  • What corrections happen repeatedly?
  • What part of the process is missing from the SOP?
  • What would you never want the system to do automatically?

Audit Interviews: Downstream Receivers

  • What do you need from the upstream output?
  • What is usually missing?
  • What causes clarification or rework?
  • What evidence do you need before acting?
  • Would faster upstream output help or overwhelm you?

Audit Interviews: IT and Security

  • Which systems and data would the workflow access?
  • Can permissions be scoped by role and action?
  • What logs are available?
  • What changes or outages would affect the workflow?
  • How would credentials or tool access be revoked during an incident?

Audit Interviews: Knowledge Owners

  • Which sources are authoritative?
  • How are versions managed?
  • What content should not be retrieved?
  • How are gaps reported?
  • Who decides when a source becomes obsolete?

Audit Interviews: Governance or Risk

  • Which laws, policies or professional rules apply?
  • Which decisions require human authority?
  • What incidents require escalation?
  • What evidence would be needed after a failure?
  • How often should the workflow be reviewed?

The One-Hour Workflow Audit

  1. Name the workflow and outcome.
  2. Draw the normal path.
  3. Identify the top three sources.
  4. Mark the consequential decisions.
  5. Mark permissions and sensitive data.
  6. Identify verification and exceptions.
  7. Write the baseline metrics available today.
  8. Name the three owners: process, technical and knowledge.
  9. Identify one recovery method.
  10. Rate each required dimension Unknown, Fragile, Stable or Proven.

This is enough for an initial go/no-go decision on a bounded pilot.


The One-Day Department Audit

Select three to five important workflows rather than auditing every task. For each, run the one-hour method, then compare common gaps. If several workflows suffer from the same knowledge, identity, review or data problem, that gap may deserve shared infrastructure.

The department audit should finish with a readiness map and a short list of repair priorities, not a catalogue of every possible AI use case.

The 30-Day Organisational Audit

Week 1 — Discover

Inventory existing SI use, important workflows and approved tools. Identify where employees are already creating private workflows, prompt libraries or personal knowledge stores. The goal is to understand the real operating surface, not only the systems formally approved.

Week 2 — Evidence

Collect source, data, permission, review and incident evidence for the highest-value workflows. Verify what is actually current and what merely appears in policy documents.

Week 3 — Test

Run representative cases, including should-stop examples, on selected workflows. Evaluate both the machine output and the human review process. A reviewer who cannot detect a seeded failure is itself an audit finding.

Week 4 — Prioritise

Separate readiness repairs from SI projects. Some of the highest-value findings may be knowledge cleanup, source ownership, access controls or workflow simplification. Not every audit recommendation should involve another model or agent.

Audit Output 1 — Readiness Gap Map

List the material gaps by workflow and dimension. Highlight the first condition that blocks the intended autonomy level.

Do not hide weak areas inside an average score. A single inability to verify a consequential action can outweigh several strong dimensions.

Audit Output 2 — Repair Backlog

Create a backlog of concrete repairs: designate source owner, remove obsolete documents, create task rubric, add review checklist, scope permissions, define exception queue, create rollback or establish baseline.

Every repair should have an owner and observable completion evidence. “Improve documentation” is not enough; “publish one canonical travel policy, archive two obsolete versions and assign Finance Operations as owner” is actionable.

Audit Output 3 — Use-Case Portfolio

Classify proposed workflows into four broad states: ready now, ready after repair, assist-only for now and not worth pursuing. This prevents the organisation from treating every idea as equally mature.

The portfolio should also record time-to-value and transfer learning. A modest use case may deserve priority if it builds a retrieval or evaluation capability that later supports several workflows.

Audit Output 4 — Governance Actions

Record policy changes needed: approved tools, data restrictions, action limits, incident triggers, user responsibilities and role ownership.

Governance actions should map directly to observed workflow risk rather than repeat generic principles.

Audit Output 5 — Measurement Plan

For each workflow, name a baseline, primary outcome metric, quality metric, exception metric and review-burden metric.

This creates the minimum evidence needed to decide later whether the system should be retained, promoted, repaired or retired.

Audit Output 6 — Review Date and Triggers

Readiness changes. Set a reassessment trigger or date for important workflows, especially after model, data, tool, permission, regulation or policy changes.

A periodic review can be light when the system is low risk. High-impact systems need more explicit lifecycle control.

Audit Output 7 — Ownership Map

Name the process owner, technical owner, knowledge or policy owner, review owner and incident contact. This map makes responsibility visible when the system crosses departments.

Audit Output 8 — Dependency Map

List external providers, model services, retrieval systems, identity platforms, business applications and human roles the workflow depends on. Mark which dependencies can stop the process or change its behaviour materially.

This helps distinguish a local workflow issue from a systemic risk shared across many SI systems.

Audit Output 9 — Permission Register

Record what the system can read, recommend, prepare and execute. Include service accounts, delegated user identities and high-impact write permissions.

The register should make it easy to answer whether a new use case actually needs the authority it is requesting.

Audit Output 10 — Evidence Pack

  • Current-state workflow map
  • Representative inputs
  • Accepted outputs
  • Failure cases
  • Current instructions
  • Source list and owners
  • Permission inventory
  • Review checklist
  • Exception examples
  • Baseline metrics
  • Evaluation results
  • Incident or near-miss notes
  • Recovery procedure
  • Change ledger

The evidence pack becomes the durable basis for promotion, transfer or reassessment.

Audit Readiness for Assist

For Assist, focus on approved tools, data rules, user skill and verification. Tool connection and shared infrastructure may be unnecessary.

A workflow can be ready for Assist even when broader organisational readiness is low because the human remains close to every interaction.

Audit Readiness for Collaborate

For Collaborate, add shared context, reusable instructions, review consistency and task ownership. The method should work for more than one user and the handoff between human and SI should be explicit.

Audit Readiness for Automate

For Automate, add triggers, eligibility, exceptions, scoped tools, observability, world return and recovery. The workflow should already be stable manually or collaboratively.

A trigger should not start a process whose normal path the organisation cannot explain.

Audit Readiness for Operate

For Operate, add shared identity, portfolio governance, dependency mapping, incident response, change control and cross-workflow monitoring.

Several proven workflows should already exist. Operating infrastructure should emerge from repeated need rather than be built speculatively.

Audit Readiness for Agents

For agents, examine objective boundaries, tool inventory, action permissions, memory, time horizon, stop conditions, external content, world return and human checkpoints.

The audit should ask whether the task genuinely requires dynamic planning. If a fixed workflow can perform the job, simpler architecture is easier to govern.

Audit Readiness for Multi-Agent Systems

When several agents interact, audit shared state, handoff evidence, conflicting instructions, permission inheritance and who resolves disagreement.

Agent-to-agent delegation can create hidden authority chains. The organisation should be able to reconstruct why a consequential action occurred.

Audit Readiness for Computer Use

Computer-use systems can operate graphical interfaces. Audit account scope, interface brittleness, confirmation for high-impact actions, untrusted content and whether visual success matches the underlying system-of-record state.

Audit Readiness for Long-Running Workflows

Long-running tasks need checkpoints because facts and permissions can change while the system works. Audit refresh intervals, deadline handling, cancellation and whether the objective remains valid over time.

Audit Readiness for Sensitive Data

Where personal, confidential, privileged or security-sensitive data appears, audit purpose, minimisation, access, approved environment, retention and downstream use.

A system should not receive broad sensitive context merely because “more context helps”.

Audit Failure Mode: Scoring Without Evidence

Teams fill in a maturity spreadsheet based on perception. The result looks precise but does not expose operational reality.

Require at least one observable artefact or real example for every material rating.

Audit Failure Mode: Company-Wide Averaging

A strong engineering function and weak knowledge process average into a reassuring middle score. The real blocker disappears.

Keep readiness profiles by dimension and workflow.

Audit Failure Mode: Auditing Tools Instead of Work

The organisation assesses which model features it has rather than whether workflows, knowledge and controls are ready.

Capabilities matter only in relation to a use case and a desired autonomy level.

Audit Failure Mode: No Frontline Input

Managers describe the official process while hidden workarounds and exceptions remain invisible. Include the people who actually perform the work.

Audit Failure Mode: No Receiver Input

The upstream team optimises production without understanding downstream needs. Include the receiver and measure clarification, rework and usefulness.

Audit Failure Mode: Treating Policy as Control

A written rule is not evidence that systems technically enforce the boundary or that employees understand it. Check actual permissions, interfaces and behaviour.

Audit Failure Mode: Ignoring Review Capacity

A workflow appears ready because it includes human review, but projected volume overwhelms reviewers. Model queue capacity explicitly.

A reviewer who rubber-stamps because of workload is not a meaningful control.

Audit Failure Mode: Ignoring Exceptions

The normal path is excellent while edge cases have no owner. An automation is only as operable as its exception path.

Audit Failure Mode: No Recovery Test

The system can act but nobody has rehearsed rollback or outage fallback. Recovery readiness should be tested before high-impact scale.

Audit Failure Mode: Permanent Green Status

A workflow passes once and is considered ready forever. Model changes, data drift, new permissions and incidents can move it back to Fragile.

Audit Failure Mode: Confusing Adoption With Readiness

High usage does not prove readiness. Employees may use tools frequently while source quality, permissions and measurement remain weak.

Audit Failure Mode: Treating a Vendor’s Controls as the Whole Audit

Enterprise products may provide strong security and administration, but workplace readiness also depends on local task definition, source authority, review, exceptions and ownership.

The deploying organisation still owns the operating design.

Audit Failure Mode: Auditing Once Before Launch

Readiness changes after launch. Real use reveals new edge cases, workarounds and dependencies. Treat the first audit as the beginning of lifecycle evidence, not the final certificate.

Audit Failure Mode: No Demotion Path

Teams design how autonomy will increase but not how it will decrease. A mature audit defines the conditions that move the workflow back toward more human control.

Audit Failure Mode: No Retirement Path

A workflow becomes obsolete but permissions, triggers and stored data remain. Audit lifecycle should include retirement and cleanup.

The Readiness Audit and Singapore’s Current Adoption Context

Singapore’s Ministry of Manpower reported on 30 April 2026 that 28.5% of firms had begun adopting AI while 3.8% were integrating AI into core processes. The figures are a dated national snapshot, but they illustrate why readiness matters: access to tools and integration into core work are different states. See MOM’s report.

The same report identified implementation cost, expertise, strategy, trust, integration complexity and data security among barriers. A readiness audit turns those broad barriers into workflow-specific evidence.

The Readiness Audit and NIST

NIST’s Generative AI Profile is a voluntary companion to the AI Risk Management Framework for incorporating trustworthiness considerations into generative-AI design, development, use and evaluation. Its lifecycle orientation supports the audit principle that readiness should examine governance, context, measurement and management around the whole system. See NIST AI 600-1.

The Readiness Audit and Agentic AI

When workflows can pursue goals and use tools, the audit should examine objective scope, third-party agents, tool access, significant human checkpoints, automation bias, world return and recovery.

IMDA’s Model AI Governance Framework for Agentic AI and its May 2026 update provide current Singapore guidance on bounding agent powers and maintaining meaningful human accountability. See IMDA’s framework factsheet.

The Audit and Personal Data

If the workflow uses personal data, include purpose, access, data minimisation, retention and decision consequence in the audit. Singapore organisations should consider applicable PDPA obligations and PDPC’s advisory guidelines on AI recommendation and decision systems.

See PDPC’s advisory guidelines.

The Audit Promotion Gate

A workflow should move to higher autonomy only when the dimensions required for that level are at least Stable and the organisation has evidence from representative cases.

Promotion should be tied to conditions, not a project calendar or vendor roadmap.

The Audit Demotion Gate

Move the workflow back toward human control when material errors rise, sources become stale, review capacity collapses, permissions expand unexpectedly, a serious incident occurs or the operating environment changes.

The Audit Retire Gate

Retire the use case when value remains small, review cost stays high, task fit is poor or conventional software solves the problem better.

Readiness work should eliminate weak projects as well as enable strong ones.

The Audit Transfer Test

A workflow judged ready in one department should not automatically be copied elsewhere. The receiving team should compare object, data, authority, consequence, verification and receiver needs.

Reuse common infrastructure where appropriate; reassess local task readiness.

The Audit Owner Triangle

Every audited workflow should identify the process owner, technical owner and knowledge or policy owner. Governance and security may add oversight, but these three roles make operating responsibility concrete.

The Audit Decision Tree

  1. Is the outcome clear? If no, clarify before SI.
  2. Is the workflow visible? If no, map it.
  3. Are sources authoritative? If no, repair knowledge.
  4. Is task fit strong? If no, choose another use case.
  5. Can output be verified? If no, lower autonomy.
  6. Can permissions be scoped? If no, avoid broad tool access.
  7. Are exceptions owned? If no, design the queue.
  8. Is baseline measurable? If no, establish one.
  9. Can the system recover? If no, protect external action.
  10. Does the next level remove meaningful friction? If no, stay where you are.

What This Article Owns

This page owns the readiness audit method: the evidence, interviews, dimensions, ratings, outputs and decision gates used to assess whether a workplace or workflow can support Super Intelligence responsibly.

The earlier readiness article defines what a ready workplace looks like. This article turns that definition into an auditable process. The next article will use audit and value evidence to prioritise SI projects.

Frequently Asked Questions

What is a Super Intelligence readiness audit?

It is a structured review of workflow clarity, sources, data, people, permissions, verification, exceptions, measurement, governance, ownership and recovery for a proposed or existing SI workflow.

How long should an audit take?

A bounded workflow can receive an initial audit in an hour. A department may take a day. Organisation-wide readiness may require several weeks because evidence must be gathered across systems and teams.

Do we need external consultants?

Not necessarily. Process owners, frontline users, technical teams, knowledge owners and governance functions can run a strong audit if they use evidence rather than opinion.

Should we produce a numeric readiness score?

A qualitative profile is usually more useful. A single number can hide a severe gap in one dimension.

What is the minimum readiness for a pilot?

An approved environment, a bounded familiar task, a named owner, current context, meaningful review and a measurable baseline are often enough for a low-risk pilot.

What is the minimum readiness for automation?

A stable normal path, known exceptions, scoped permissions, sustainable verification, observability, world return and recovery.

What is the minimum readiness for agents?

In addition to automation readiness, the task should genuinely need dynamic multi-step action, with bounded objectives, tool limits, stopping and stronger monitoring.

How often should audits be repeated?

After material changes to models, data, sources, tools, permissions, policy or incident history, and periodically for important workflows.

What if the audit finds many gaps?

Narrow the use case and repair the first limiting condition. Do not attempt to fix every dimension at once.

What comes next?

Continue to How to Prioritise Super Intelligence Projects, which uses value, readiness, time-to-value, bottleneck impact and risk to order the SI portfolio.

The Core Audit Rule

Audit the operating system around the intelligence, not just the intelligence itself. A strong model inside a weak workflow is not readiness. A modest model inside a clear, measurable and governable workflow can create much more reliable value.

The audit is successful when it tells the organisation what to repair, what to pilot, what to delay and what not to build.


Audit Dimension 4 — Data Reliability

Which structured fields matter for the workflow and how reliable are they? Are customer records current? Do project systems reflect real milestones? Do financial systems reconcile? Are required fields routinely missing?

The audit should focus on data required for the selected workflow rather than attempting to judge the entire enterprise. A company can be ready for one bounded workflow while other datasets remain immature.

Audit Dimension 5 — Task Fit

Does the proposed SI role match the actual work? Retrieval, extraction, comparison, drafting and summarisation are different from negotiation, professional judgment or high-consequence authority. Use the job-to-task decomposition rather than the job title.

Strong task fit means the model’s capability directly addresses the repeated cognitive burden and the resulting output can be checked.

Audit Dimension 6 — Human Capability

Can users define outcomes, provide relevant context, verify important claims and recognise when the system should stop? Can reviewers perform the specific checks assigned to them?

Training completion is weak evidence by itself. A stronger audit uses realistic cases, including a deliberately flawed output, and observes whether employees detect unsupported claims, missing context and incorrect action.

Audit Dimension 7 — Permission Control

Can the organisation separate read, interpret, recommend, prepare and execute permissions? Is least privilege practical? Can an agent access one repository without receiving broad access to unrelated systems?

A workflow that needs only context retrieval should not inherit write authority merely because the platform makes it easy.

Audit Dimension 8 — Verification

How is correctness checked? Factual claims may need source evidence. Structured extraction can use schemas. Arithmetic can be recomputed. Code can be tested. Consequential interpretation may need qualified professional review.

The audit should also examine cost. If every output needs complete human reconstruction, the workflow may be technically possible but economically unready.

Audit Dimension 9 — Exception Handling

Which cases leave the normal path? Missing information, source conflict, high value, sensitive categories, unfamiliar situations and failed tool calls should have explicit routes.

An exception queue needs an owner, expected response time and enough context that the human does not rebuild the case from zero.

Audit Dimension 10 — Measurement

Is there a baseline for cycle time, touch time, rework, queue age, errors, exception rate, receiver effort or outcome quality? Without a baseline, teams may mistake activity for improvement.

The audit should identify one primary outcome metric and a small number of quality or risk metrics for each proposed workflow.

Audit Dimension 11 — Governance

Are approved tools, data rules, action limits, human checkpoints and incident procedures explicit? Does written policy reflect the way employees and agents actually work?

Governance readiness becomes more important as autonomy and failure radius increase. A drafting assistant and a tool-using financial agent should not share one generic rule.

Audit Dimension 12 — Recovery

Can the workflow be stopped, downgraded or rolled back? Can uncertain transactions be reconciled? Does a manual fallback exist for important work?

Recovery requirements should scale with reversibility. An unsent draft needs little recovery; a production deployment or financial transaction may need strong rollback or reconciliation.

Audit Dimension 13 — Ownership

Who owns the business outcome, technical implementation and authoritative knowledge? These may be different people. Every important workflow should have enough ownership that someone can decide to change, pause or retire it.

A workflow without an owner tends to accumulate local workarounds while the root problem remains.

Audit Dimension 14 — Change Control

What happens when the model, source set, tool access, prompt, policy or orchestration changes? Important workflows should define what counts as a material change and when representative tests must be rerun.

A model upgrade should not automatically inherit the same action authority if its behaviour on the workflow has not been checked.

Audit Dimension 15 — Continuity

What happens if the SI service, connector or knowledge system is unavailable? Critical workflows need a fallback proportionate to their importance.

Continuity also tests capability atrophy. If the manual fallback exists on paper but nobody remembers how to operate it, the organisation is not truly ready.

The Four Audit Ratings

Unknown

The condition has not been established or the team has no evidence.

Fragile

The condition exists but depends on one person, informal workarounds or inconsistent execution.

Stable

The condition works repeatedly across normal users and cases.

Proven

The condition has survived representative exceptions, material changes, recovery or sustained operating volume.

Do not average these ratings into one reassuring percentage. A single Unknown or Fragile dimension can block higher autonomy when that dimension is critical to safety or outcome.

The Red-Flag Conditions

  • Consequential authority is unclear.
  • Important output cannot be verified independently.
  • Sensitive information would enter an unapproved environment.
  • External action can occur without reliable world-return evidence.
  • No one owns exceptions.
  • Review volume is expected to exceed human capacity.
  • Source authority is disputed.
  • The workflow cannot be stopped or recovered after a material failure.

A red flag does not necessarily mean abandoning SI. It usually means reducing the role to preparation, research or drafting while the control is repaired.

The Green-Light Conditions

  • Workflow and outcome are visible.
  • Required sources are current and owned.
  • Task fit is strong.
  • Users understand the system’s role.
  • Important output can be verified efficiently.
  • Permissions can be scoped.
  • Exceptions are recognisable and owned.
  • Baseline metrics exist.
  • Recovery is proportionate to consequence.

Green light means ready to pilot or advance one level—not permission for unlimited autonomy.

Audit Interviews: Process Owner

  • What outcome does this workflow exist to produce?
  • Which state constrains performance today?
  • Which errors matter most?
  • Which decisions require your authority?
  • What would make you reduce SI autonomy?
  • Which metric would prove useful improvement?

Audit Interviews: Frontline Users

  • Where do you search for information?
  • What do you copy between systems?
  • What cases make you stop and ask someone?
  • What corrections happen repeatedly?
  • What part of the real process is missing from the SOP?
  • What should the system never do automatically?

Audit Interviews: Downstream Receivers

  • What do you need from the upstream output?
  • What is usually missing?
  • What causes clarification or rework?
  • What evidence do you need before acting?
  • Would faster upstream production help or overwhelm you?

Audit Interviews: IT and Security

  • Which systems and data does the workflow access?
  • Can permissions be scoped by role and action?
  • What logs are available?
  • What changes or outages would affect the workflow?
  • How would credentials or agent tool access be revoked during an incident?

Audit Interviews: Knowledge Owners

  • Which sources are canonical?
  • How are versions managed?
  • What content should not be retrieved?
  • How are gaps and contradictions reported?
  • Who decides when a source becomes obsolete?

Audit Interviews: Governance and Risk

  • Which laws, policies or professional rules apply?
  • Which decisions require human accountability?
  • What events constitute an SI incident?
  • What evidence would be required after a failure?
  • How often should the workflow be reviewed?

The One-Hour Workflow Audit

  1. Name the workflow and real-world outcome.
  2. Draw the normal path from trigger to closure.
  3. Identify the top three authoritative sources.
  4. Mark consequential decisions and external actions.
  5. Mark sensitive data and permission boundaries.
  6. Identify verification and exceptions.
  7. Write the available baseline metrics.
  8. Name process, technical and knowledge owners.
  9. Identify at least one recovery or fallback method.
  10. Rate every required dimension Unknown, Fragile, Stable or Proven.

This one-hour format is enough for an initial pilot-readiness decision on many bounded workflows.

The One-Day Department Audit

Select three to five important workflows rather than auditing every employee task. Run the one-hour method for each, then compare common gaps.

If several workflows suffer from the same knowledge, identity, review, connector or permission problem, that gap may deserve shared infrastructure.

The 30-Day Organisational Audit

Week 1 — Discover

Inventory important workflows, approved tools and existing SI use. Include shadow practices where employees have created personal workflows outside shared standards.

Week 2 — Gather evidence

Collect source, data, permission, review, incident and measurement evidence for the highest-value workflows.

Week 3 — Test

Run representative cases, including missing-data, conflicting-source and should-stop examples. Evaluate both the system and the human review process.

Week 4 — Prioritise

Separate readiness repairs from SI projects. Some of the most valuable outputs may be source ownership, workflow simplification, access control or better evaluation rather than a new agent.

Audit Output 1 — Readiness Gap Map

List each material gap by workflow and dimension. Highlight the first limiting condition that blocks the intended autonomy level.

Do not hide severe weaknesses inside an average score.

Audit Output 2 — Repair Backlog

Convert gaps into actions: designate source owner, remove obsolete documents, create a rubric, add a review checklist, scope permissions, define the exception queue, build a rollback path or establish the baseline.

Every repair should have an owner and observable completion criterion.

Audit Output 3 — Use-Case Portfolio

Classify use cases into ready now, ready after repair, assist-only for now and not worth pursuing. This prevents the organisation from treating every AI idea as equally mature.

Audit Output 4 — Governance Actions

Record required policy and control changes: approved environments, data restrictions, action thresholds, human checkpoints, incident routes and ownership.

Audit Output 5 — Measurement Plan

For each workflow, name the baseline, primary outcome metric, quality measure, exception measure and human-review burden.

Audit Output 6 — Reassessment Trigger

Define when the audit must be reopened: model change, source change, new data, expanded permissions, incident, regulatory change, major volume increase or ownership change.

Audit for Assist

For Assist, focus on approved environment, data rules, user skill and verification. Shared infrastructure may be unnecessary.

Audit for Collaborate

For Collaborate, add shared context, reusable instructions, review consistency and ownership. The method should work for more than one expert user.

Audit for Automate

For Automate, add trigger control, eligibility, exceptions, scoped tools, observability, world return and recovery. The workflow should already be stable manually.

Audit for Operate

For Operate, add portfolio governance, shared identity, dependency mapping, incident response, change control and cross-workflow monitoring.

See The Four Levels of Workplace Super Intelligence for the autonomy model.

Audit Failure Mode: Scoring Without Evidence

Teams fill in a readiness spreadsheet based on confidence and opinion. The result looks quantitative while hiding uncertainty.

Require at least one real artefact, example or observed behaviour for every material rating.

Audit Failure Mode: Company-Wide Averaging

Strong engineering practice and weak knowledge governance average into a comfortable middle score. The actual blocker disappears.

Keep readiness profiles by workflow and dimension.

Audit Failure Mode: Auditing Tools Instead of Work

The company inventories model features rather than asking whether workflows, sources and human controls can support them.

Capabilities matter only in relation to the work.

Audit Failure Mode: No Frontline Input

Managers describe the official process while hidden workarounds, missing fields and exceptions remain invisible. Include the people who perform the work.

Audit Failure Mode: No Receiver Input

Upstream teams optimise production without understanding downstream needs. Include the receiver and measure clarification or rework.

Audit Failure Mode: Treating Policy as Enforcement

A written rule says the agent cannot take a certain action, but the technical permissions still allow it. Audit actual behaviour and access, not only policy text.

Audit Failure Mode: Ignoring Review Capacity

A workflow appears ready because every output has human review, but projected volume would overwhelm reviewers. Model review capacity explicitly.

Audit Failure Mode: Ignoring Exceptions

The normal path is polished while edge cases have no owner. Readiness includes the exception path.

Audit Failure Mode: No Recovery Test

The system can act but nobody has rehearsed rollback, credential revocation or manual fallback. Recovery should be demonstrated before high-impact scale.

Audit Failure Mode: Permanent Green Status

A workflow passes once and is considered ready forever. Models, data, policy and user behaviour change. Readiness must be maintained.

Audit Failure Mode: Confusing Adoption With Readiness

High usage does not prove readiness. Employees can use SI daily while source authority, permissions and outcome measurement remain weak.

Audit and Singapore’s Current Adoption Context

Singapore’s Ministry of Manpower reported on 30 April 2026 that 28.5% of firms had begun adopting AI while 3.8% were integrating AI into core processes. The figures are a dated national snapshot, not a readiness score for any single organisation, but they illustrate the difference between tool access and operational integration. See MOM’s report.

The same report identified implementation cost, expertise, strategy, trust, integration complexity and data security among barriers. A readiness audit turns those broad barriers into workflow-specific evidence.

Audit and NIST

NIST’s Generative AI Profile is a voluntary companion to the AI Risk Management Framework for incorporating trustworthiness considerations into generative-AI design, development, use and evaluation. Its lifecycle framing supports the audit principle that readiness must include governance, context, measurement and ongoing management. See NIST AI 600-1.

Audit and Agentic AI

When workflows can pursue goals and use tools, the audit should examine objective scope, tool access, third-party agents, meaningful human checkpoints, automation bias, world-return evidence and recovery.

IMDA’s Model AI Governance Framework for Agentic AI and its May 2026 update provide current guidance on bounding agent powers and maintaining human accountability. See IMDA’s framework factsheet.

Audit and Personal Data

If the workflow uses personal data, include purpose, access, data minimisation, retention and decision consequence in the audit. Singapore organisations should consider applicable PDPA obligations and the PDPC’s advisory guidelines on AI recommendation and decision systems.

See PDPC’s advisory guidelines.

The Promotion Gate

A workflow should move to higher autonomy only when the dimensions required for that level are Stable or Proven and representative cases show the workflow can operate within the intended envelope.

Promotion should follow evidence, not a project calendar.

The Demotion Gate

Move the workflow back toward human control when material errors rise, sources become stale, review capacity collapses, permissions change, a serious incident occurs or the operating environment shifts.

The Retire Gate

Retire the use case when value remains small, review cost remains high, task fit is poor or ordinary software solves the problem more reliably.

Readiness work should eliminate weak projects as well as enable strong ones.

The Audit Evidence Pack

  • Current-state workflow map
  • Representative inputs
  • Accepted outputs
  • Failure and exception cases
  • Current instructions
  • Source list and owners
  • Permission inventory
  • Review checklist
  • Baseline metrics
  • Evaluation results
  • Incident or near-miss notes
  • Recovery procedure
  • Change ledger

This pack becomes durable evidence for promotion, transfer and reassessment.

The Audit Transfer Test

A workflow judged ready in one department should not automatically be copied elsewhere. The receiving team should compare object, data, authority, consequence, verification and receiver needs.

Reuse shared infrastructure where appropriate; reassess local workflow readiness.

The Audit Owner Triangle

Every audited workflow should identify the process owner, technical owner and knowledge or policy owner. Governance and security can add oversight, but the triangle makes operating responsibility concrete.

The Audit Decision Tree

  1. Is the outcome clear? If no, define it.
  2. Is the workflow visible? If no, map it.
  3. Are authoritative sources known? If no, repair knowledge ownership.
  4. Is task fit strong? If no, choose another use case.
  5. Can important output be verified? If no, lower autonomy.
  6. Can permissions be scoped? If no, avoid broad tool access.
  7. Are exceptions owned? If no, design the queue.
  8. Is there a baseline? If no, establish one.
  9. Can the workflow recover? If no, protect external action.
  10. Does the next level remove meaningful friction? If no, remain at the current level.

What This Article Owns

This page owns the readiness audit method: evidence, interviews, dimensions, ratings, outputs and decision gates used to assess whether a workplace or workflow can support Super Intelligence responsibly.

The earlier readiness article defines what an SI-ready workplace looks like. This article turns that definition into an auditable process. The next article uses readiness and value evidence to prioritise SI projects.

Frequently Asked Questions

What is a Super Intelligence readiness audit?

It is a structured review of workflow clarity, sources, data, people, permissions, verification, exceptions, measurement, governance, ownership and recovery for a proposed or existing SI workflow.

How long should an audit take?

A bounded workflow can receive an initial audit in an hour. A department may take a day. Organisation-wide readiness can take several weeks because evidence must be gathered across systems and teams.

Do we need external consultants?

Not necessarily. Process owners, frontline users, technical teams, knowledge owners and governance functions can run a strong audit if they use evidence rather than opinion.

Should we produce a numeric readiness score?

A qualitative profile is usually more useful. A single number can hide a severe gap in one dimension.

What is the minimum readiness for a pilot?

An approved environment, bounded familiar task, named owner, current context, meaningful review and measurable baseline are often enough for a low-risk pilot.

What is the minimum readiness for automation?

Stable normal path, known exceptions, scoped permissions, sustainable verification, observability, world return and recovery.

What is the minimum readiness for agents?

In addition to automation readiness, the task should genuinely need dynamic multi-step action, with bounded objectives, tool limits, stopping and stronger monitoring.

How often should audits be repeated?

After material changes to models, data, sources, tools, permissions, policy or incident history, and periodically for important workflows.

What if the audit finds many gaps?

Narrow the use case and repair the first limiting condition. Do not attempt to fix every dimension at once.

What comes next?

Continue to How to Prioritise Super Intelligence Projects, which uses value, readiness, time-to-value, bottleneck impact and risk to order the SI portfolio.

The Core Audit Rule

Audit the operating system around the intelligence, not just the intelligence itself. A strong model inside a weak workflow is not readiness. A modest model inside a clear, measurable and governable workflow can create much more reliable value.

The audit is successful when it tells the organisation what to repair, what to pilot, what to delay and what not to build.


Worked Audit: Weekly Management Reporting

A management-reporting workflow wants SI to assemble the weekly pack. The audit finds a stable deadline and clear receiver, but input arrives through several channels. Project milestones live in the project system, while narrative updates arrive through email and chat.

Outcome clarity is Stable. Source authority is mixed: structured milestones are Stable, narrative state is Fragile. Verification is Stable because figures can be reconciled. Exception handling is Fragile because missing updates are chased informally. The recommendation is not “automate reporting” yet. It is to standardise intake and define the missing-update route, then pilot collaborative drafting.

Worked Audit: Customer Support

A support team wants automatic replies for routine cases. The audit finds current policy, reliable account data, clear categories and a specialist escalation queue. The main weakness is permissions: one service account can both read and edit too many customer fields.

The workflow is ready for classification, retrieval and draft preparation. Automation of external action should wait until read, recommend and write permissions are separated. This illustrates how one Fragile dimension can determine the appropriate autonomy level.

Worked Audit: Hiring Support

HR wants SI to screen applications. The audit finds strong administrative workflows but inconsistent selection criteria across hiring managers, unclear exception handling and no agreed evaluation of fairness.

The workflow may still be ready for scheduling, application organisation, job-description drafting and interview-note preparation. The audit narrows the use case rather than rejecting Super Intelligence entirely.

Worked Audit: Contract Review

A legal team wants SI to compare incoming contracts with standard terms. The clause library is current and owned, reviewers can inspect source passages, and the system remains read-only.

The workflow is ready for extraction, comparison and issue-list preparation. Final legal interpretation and commitment remain with authorised professionals. Readiness is strong because the system role is bounded and evidence is visible.

Worked Audit: Coding Agent

An engineering team wants an agent to implement bounded issues and open pull requests. The audit finds version control, branch protection, tests, CI, code review and rollback already in place.

The main gap is requirement ambiguity. The audit recommends a stop condition when acceptance criteria conflict or tests cannot establish expected behaviour. Existing engineering controls make the workflow more agent-ready than a less disciplined process.

Worked Audit: Finance Commentary

A finance team wants SI to draft management commentary. The audit finds reliable source-of-record figures and deterministic reconciliation, but contextual explanations are stored informally in personal notes.

The workflow is ready for a pilot if the system uses authoritative numbers and finance professionals remain responsible for interpretation. The audit also recommends capturing recurring business explanations in maintained context.

Worked Audit: Knowledge Assistant

A company wants a natural-language policy assistant. The audit finds many documents but no retirement process. Several obsolete policy versions remain searchable.

This is a source-authority red flag. The correct readiness repair is to establish canonical current documents, owners and retirement rules before claiming reliable enterprise search.

Worked Audit: Procurement Comparison

A procurement team wants SI to rank suppliers. The audit finds proposals with inconsistent pricing bases, different scopes and incomplete service definitions.

The workflow is ready for extraction and comparison support but not automated ranking. The audit recommends normalising terms and preserving the authorised procurement decision.

Worked Audit: Education Feedback

A training team wants automated feedback. The audit finds a clear rubric, validated answers and strong instructor review, but no measurement of whether learners improve after receiving generated feedback.

The workflow is operationally ready for a pilot but measurement is Fragile. The audit adds a delayed learning-outcome check before scale.

Worked Audit: Incident Response

Operations wants SI to summarise incidents and recommend next steps. The audit finds strong logs and runbooks but high-impact production permissions attached to the same agent identity.

The recommendation is to split read-and-analyse permissions from production action. The system can become a strong incident copilot before it earns autonomous remediation rights.

The Audit Should Test Should-Stop Cases

A readiness audit is incomplete if it tests only successful completion. Include cases where the correct behaviour is to stop: missing data, conflicting sources, unavailable tools, exceeded thresholds, sensitive categories or uncertain action state.

A system that knows how to abstain can be more operationally ready than one that completes every test case confidently.

The Audit Should Test Human Review

Give reviewers a mix of correct and flawed outputs. Ask them to identify unsupported claims, missing context, incorrect calculations or inappropriate actions. This tests the control itself rather than assuming a human checkpoint is automatically effective.

If reviewers miss predictable errors, the workflow may need better evidence presentation, training or a different review design.

The Audit Should Test Queue Capacity

Estimate how many outputs or exceptions the system can create and how many the human team can process. A workflow can be technically correct and operationally unready because its review or exception queue will grow without bound.

Capacity testing belongs in readiness because automation changes volume even when error rate stays constant.

The Audit Should Test World Return

For actions, verify that the workflow can distinguish prepared, attempted, succeeded, failed and uncertain states. A tool timeout should not be treated as proof of success or failure without another check.

This is particularly important for messages, payments, tickets, deployments and record updates.

The Audit Should Test Recovery

Simulate or tabletop a failure. How would the workflow be stopped? Which credentials are revoked? How is manual processing restored? How are uncertain actions reconciled?

A recovery plan that has never been examined may be only a document rather than an operating capability.

The Audit Should Test Currentness

Check whether time-sensitive data is refreshed before action. A policy may remain valid for months, but customer entitlement, inventory, price or access state can change quickly.

Currentness rules should be explicit for every material source.

The Audit Should Test Transfer

If the project will expand to another team, compare the receiving workflow before transfer. A similar title does not prove similar data, consequence or authority.

A readiness audit should make clear which parts are reusable infrastructure and which are local assumptions.

The Audit Should Test Skill Preservation

Ask whether automation removes practice from a skill humans still need for verification or recovery. If yes, the workflow may require deliberate training, explanation or manual drills.

Readiness includes the human system’s ability to operate when SI fails or when a difficult exception falls outside the automated path.

The Audit Should Test Incentives

Employees may be rewarded for speed, volume or use of the new system. Check whether those incentives encourage rubber-stamping, hiding exceptions or producing more low-value output.

Governance can fail through incentives even when technical controls are strong.

The Audit Should Test Documentation Drift

Compare the documented workflow with actual behaviour. If users have workarounds, private prompts or unofficial data sources, include them in the readiness finding.

Shadow practices are evidence about missing capability or poor usability, not merely non-compliance.

The Audit Should Test Dependency Concentration

Identify shared components whose failure would affect many workflows: model gateway, retrieval service, identity provider, document store or connector platform.

A centralised control can improve readiness and increase systemic failure radius simultaneously. Continuity planning should reflect that trade-off.

The Audit Should Test Vendor Change

Ask how the workflow is notified when a provider changes model behaviour, tool availability, data handling or pricing. Important workflows should have a method for retesting material changes.

The Audit Should Test Cost at Scale

Pilot cost can be small while scaled cost is large. Estimate model, connector, storage, review and exception cost at realistic volume.

Readiness includes knowing whether the workflow remains economically useful after scale.

The Audit Should Test Receiver Value

Ask downstream users whether the SI-enabled output reduces or increases their effort. A faster upstream team can still make the system worse if the receiver must correct more.

Receiver effort is a core audit metric for handoff-heavy workflows.

The Audit Should Test Organisational Learning

Review repeated corrections. Have they changed the source, instruction, taxonomy or validation? If not, the organisation may be using humans as a permanent patch.

A ready system learns at the workflow level even when the model itself does not retain every correction.

Readiness Audit Checklist by Risk Level

Low-risk internal work

  • Approved tool
  • Allowed data
  • Named task owner
  • Basic verification
  • Baseline measure

Medium-risk connected workflow

  • Current sources
  • Scoped permissions
  • Structured review
  • Exception owner
  • Logs
  • Recovery path
  • Representative evaluation

High-impact or agentic workflow

  • Explicit objective boundary
  • Least privilege
  • Pre-action verification
  • Meaningful human checkpoints
  • World-return evidence
  • Incident response
  • Dependency map
  • Change control
  • Periodic reassessment
  • Domain-specific legal or professional controls

The Audit Report Structure

  1. Executive summary: what decision the audit supports.
  2. Workflow scope and outcome.
  3. Current autonomy level.
  4. Evidence reviewed.
  5. Readiness profile by dimension.
  6. Red flags.
  7. Repair backlog.
  8. Pilot or promotion recommendation.
  9. Measurement plan.
  10. Owners and review date.

Keep the report focused on operating decisions. The purpose is not to produce a long document; it is to create an actionable readiness state.

The Audit Review Meeting

The final audit should be reviewed by the process owner and the roles responsible for technical and knowledge control. High-impact workflows may also require security, legal, compliance or risk input.

The meeting should end with one of four decisions: proceed, proceed after specific repair, remain at lower autonomy or stop.

The Audit-to-Portfolio Handoff

Once several workflows have been audited, the organisation can compare projects using value, readiness, time-to-value, bottleneck impact and risk.

This prevents readiness from being treated as an isolated compliance exercise. It becomes an input to investment prioritisation.

The Audit-to-Training Handoff

Audit findings should inform employee development. If reviewers struggle with source verification, train that capability. If process owners cannot define exceptions, strengthen workflow literacy.

Training should repair observed readiness gaps rather than follow generic AI curricula.

The Audit-to-Architecture Handoff

Repeated readiness gaps can justify shared infrastructure. If many workflows lack source-grounded retrieval, build a governed knowledge layer. If permission management is repeatedly weak, improve identity and tool access controls.

Architecture should emerge from repeated workflow evidence.

The Audit-to-Governance Handoff

Where several workflows encounter the same policy ambiguity, governance should standardise the rule. This can include approved data classes, human-approval thresholds, incident categories or vendor review requirements.

The Audit-to-Retirement Handoff

A mature programme also retires systems. When a workflow is no longer useful, clean up permissions, triggers, stored data, credentials and documentation.

Readiness includes end-of-life discipline.

The Final Readiness Audit Standard

A good readiness audit answers four questions. Can the workflow use SI effectively? Can it use SI safely? Can it prove value? Can it recover when the system fails?

If any answer is unclear, the audit should specify the repair rather than hide uncertainty inside a score.

The Audit Rule in One Sentence

Do not ask whether the company is ready for AI in the abstract; audit whether this workflow has the evidence, controls, people and recovery required for this level of Super Intelligence.

That framing keeps readiness practical. It converts a broad technology debate into a series of operating decisions that can be improved one workflow at a time.


The Audit Should End With a Decision, Not a Report

A readiness audit is useful only if it changes what the organisation does next. The final meeting should convert evidence into one of four decisions: proceed, repair, constrain or stop. A long document with no operating decision becomes another layer of governance overhead.

Proceed

The workflow has enough clarity, source authority, verification, ownership and recovery for the proposed pilot or autonomy level. Proceed with a bounded scope and measurable success criteria.

Repair

The use case is attractive, but one or more readiness conditions are Fragile. Repair the first limiting condition—such as source ownership, review capacity, exception routing or permission design—then rerun the relevant evidence checks.

Constrain

The task remains useful for SI, but the proposed authority is too broad. Reduce the role from execute to prepare, from automate to collaborate, or from agentic tool use to read-only assistance. This preserves learning while protecting the consequential boundary.

Stop

The use case creates too little value, depends on an unstable process or is better solved by ordinary software, process simplification or human work. Stopping protects attention and prevents sunk-cost thinking.

The Audit Should Distinguish Local and Shared Gaps

Some readiness problems belong to one workflow. Others are shared across the organisation. A missing customer field may be local. Weak identity controls, fragmented knowledge or inconsistent model evaluation may affect many workflows.

Separate local repairs from shared infrastructure investments. This allows central teams to solve common problems once while process owners remain responsible for local outcomes.

The Audit Should Record Uncertainty

A mature audit is allowed to say that evidence is incomplete. Unknown should remain Unknown rather than being upgraded to Stable because the project team wants to proceed. Uncertainty is an operating fact, not a weakness in the document.

Where evidence is missing, design the pilot to collect it. For example, if review burden is unknown, run a representative sample and measure it before automating at scale.

The Audit Should Be Comparable Over Time

Use the same core dimensions at the next review so the organisation can see what changed. A workflow may move from Fragile to Stable on knowledge while moving from Stable to Fragile on review capacity after volume increases.

This longitudinal view is more useful than a one-time readiness badge because SI systems, data, people and tools continue changing after launch.

Audit Evidence by Autonomy Level

Assist evidence

Show that users understand the task, know which data may be shared and can verify the output. Examples of accepted work and known failure cases are usually sufficient.

Collaborate evidence

Show that the method transfers to multiple users, context is repeatable, review is meaningful and the workflow consistently reaches an accepted result.

Automate evidence

Show stable normal-path performance, explicit eligibility, sustainable exception handling, scoped permissions, logs, world-return confirmation and recovery.

Operate evidence

Show that several workflows can share infrastructure without losing local ownership, that dependencies are monitored and that incidents, changes and portfolio-level drift can be managed.

Audit Evidence by Risk Level

Low consequence

A small representative sample, user review and reversible outputs may be enough. The audit can focus on value and usability.

Moderate consequence

Require stronger source grounding, explicit review criteria, exception handling and measurement of correction burden.

High consequence

Require qualified review, strong evidence, explicit authority, tighter permissions, recovery and documented justification for any autonomous action.

The Audit Should Test the Human System

Readiness can fail because humans cannot or will not perform the roles the design assumes. Reviewers may be overloaded. Managers may discourage override. Employees may not know which source is current. Process owners may lack authority to change broken rules.

Test the operating behaviour, not only the technical design. A formally perfect workflow can still fail if the human system around it is unrealistic.

The Audit Should Test the Receiver

Ask the next person in the workflow whether SI-enabled output reduces or increases their work. Faster upstream generation can create more clarification, more review or more noise.

Receiver effort belongs in readiness because a workflow is a chain, not a local productivity metric.

The Audit Should Test the World Return

For every external action, confirm what evidence returns from the real system. A message ID, transaction state, updated record or deployment status is stronger than a conversational statement that the action succeeded.

If completion cannot be observed reliably, autonomous action is not ready.

The Audit Should Test Failure Recovery

Run at least one tabletop scenario: the model gives a bad recommendation, a connector times out, a source becomes unavailable or the system performs an incorrect action. Ask who notices, who stops the workflow, what evidence remains and how the organisation restores a safe state.

A recovery plan that has never been mentally or operationally tested should remain Fragile.

The Audit Should Test Change

Ask what happens when the model, source set, policy, permissions or business objective changes. If the workflow cannot identify which tests should be rerun, it may not be ready for long-lived operation.

The Audit Should Test Transfer

Before calling a use case scalable, move it to a second user, team or case class. Observe which assumptions break. Transfer tests reveal whether the method is organisational or still dependent on one expert.

The Audit Should Test Scale

Projected volume changes queue behaviour, cost and exception load. A system that is ready for fifty cases may not be ready for five thousand. Estimate review and exception capacity before scale.

The Final Audit Checklist

  1. Outcome and workflow are explicit.
  2. Authoritative sources are known and owned.
  3. Required data is reliable enough for the task.
  4. Task fit is demonstrated rather than assumed.
  5. Users and reviewers can perform their roles.
  6. Permissions match the minimum necessary authority.
  7. Important output has a workable verification method.
  8. Exceptions have a named owner and route.
  9. Baseline and outcome measures exist.
  10. Governance reflects actual system behaviour.
  11. Recovery and continuity are proportionate to failure radius.
  12. Process, technical and knowledge ownership are named.
  13. Material changes trigger retesting.
  14. Audit evidence is stored for future comparison.
  15. The audit ends with proceed, repair, constrain or stop.

If several of these conditions remain unclear, the organisation has found the next work to do. Readiness auditing is valuable precisely because it exposes those conditions before autonomy expands.

Institutional Readiness Lab: Schools, Companies and Public Services

Reader task. Decide what a particular institution can responsibly try next. Use the evidence below to distinguish a useful preparation tool from an unsupported expansion of authority. Finish with a decision, the evidence that limits it, a named local owner and a condition for reconsideration. You can complete this exercise on paper; no AI account, deployment or connection to an institutional system is needed.

Everything in the three packets, including institutions, policies, test logs, outputs and timings, is fictional and illustrative. None describes an actual school, company, public authority, child or applicant, and none reports a product test. The task is to audit these supplied records, not to make decisions about real people. Here, SI retains this series’ practical-AI meaning. The exercise supplies no evidence that hypothetical artificial superintelligence exists or that greater capability creates decision-making authority.

Read the decision rules, then the school, company and public-service packets. Make your own findings before checking the worked audit. Then attempt the changed cases before reading their separate answers.

1. Decision rules: evidence first, authority locally

Use the owner article’s qualitative vocabulary without converting it into a percentage. Unknown means the required evidence is absent. Fragile means a condition is partly present but fails, is inconsistent or depends on an untested assumption. Stable means it works repeatedly within the declared scope. Proven requires broader evidence of relevant exceptions, change, recovery or sustained operation. Stable within a paper rehearsal does not mean proven in service. These are editorial working labels, not externally certified maturity levels.

For this lab, write a profile in five fields: purpose and authority; sources and data; checking and capacity; access and challenge; value and recovery. Each field needs a card reference and a limitation. “Access and challenge” means people can use the intended route and ask a responsible human to correct or explain an answer. In a teacher-only preparation exercise it concerns the participating teachers; a learner-facing or public-facing service needs evidence about its actual intended users. A corporate success cannot stand in for school learning evidence or public-service access.

Proceed: the required conditions are evidenced for the exact bounded activity already permitted by the fictional local owner. State the cap and the next check. Repair: a necessary condition is missing or fails; complete a specific repair and retest before the proposed activity. Constrain: a separately evidenced, permitted narrower role remains useful; state what that role excludes. Stop: reject the proposed design when its purpose or authority is unacceptable under the packet’s rules, or its benefit does not justify it. Stop refers to that design, not every possible use of AI.

Check prohibited purposes and missing authority first. Then ask whether an allowed narrower activity has its own supporting evidence. If not, repair before proceeding. “Human review” counts only when the reviewer can check the relevant errors, has time, can withhold an output and has a fallback. A promise of later review cannot repair a capacity deficit today. Reasonable auditors may choose different labels for a mixed proposal, but they must identify the same blocked and permitted actions.

The external frameworks support context-sensitive risk management, not a universal pass mark. NIST AI RMF 1.0 (January 2023) is voluntary; its Generative AI Profile, NIST AI 600-1 (July 2024), addresses generative-AI risks across the lifecycle. The decisions and numerical caps below are supplied local exercise rules, not NIST thresholds. Where agents are actually proposed, IMDA’s May 2026 framework update reinforces bounded powers and human accountability. This lab does not require an agent.

2. School packet: prepare feedback without claiming learning gains

S1 · Purpose and authority. Cedar Practice School wants teachers to prepare clearer fraction-feedback cards. Its teaching lead permits one further teacher-only rehearsal using eight invented work samples. No child uses a chatbot; no real student work, names, grades, learning needs or account data enter a tool. Teachers retain lesson design, feedback selection and all assessment decisions. The school has not authorised classroom use or a claim of improved attainment. Its paper fallback is the existing teacher-written feedback bank.

S2 · Complete sample and key. The invented task is “Calculate 3/4 − 1/6 and explain why your denominator is valid.” The invented learner script says “(3 − 1)/(4 − 6) = 2/−2 = −1.” The teacher key is “3/4 = 9/12 and 1/6 = 2/12, so 9/12 − 2/12 = 7/12. Twelfths are equal-sized units; subtracting denominators changes the units incorrectly.” The first error is the subtraction rule, before the correct arithmetic 2/−2 = −1.

S3 · Two candidate feedback cards. Card A says: “Your calculation shows you cannot do fractions. Copy 7/12 and memorise the answer.” Card B says: “Keep the pieces the same size. Rewrite both fractions in twelfths, then subtract the numerators. What does the denominator tell you?” The local rubric requires a correct diagnosis, mathematically sound help, respectful language and an opportunity for the learner to explain. A fails the rubric despite its correct final number. B is an appropriate first hint; the teacher key in S2 provides the full explanation when needed. Neither card is evidence that a learner has understood.

S4 · Supplied rehearsal log. Two teachers each reviewed four synthetic cards. Teacher A’s four reviews took 4, 4, 4 and 4 staff-minutes; teacher B’s took 4, 4, 4 and 4 staff-minutes. Across the eight cards, three deliberately flawed cards contained, respectively, denominator subtraction, an unsupported statement about a learner’s ability, and an answer with no reasoning opportunity. All three were rejected and repaired using the key and rubric. The other five were accepted. Both teachers completed a paper-fallback task correctly. The eight-card review used 32 combined staff-minutes from a shared budget of 40 combined staff-minutes across both teachers. This is a labour budget, not a claim about elapsed session duration. The sample is tiny and deliberately constructed.

S5 · Evidence still missing. There are no learner responses after feedback, no delayed independent work, no comparison with teacher-only instruction and no classroom accessibility evidence. A slide claims, “Eight reviewed cards prove the tutor improves learning.” That claim is unsupported. The teaching lead’s actual success criterion for the next rehearsal is narrower: both teachers must again identify planted errors, produce explainable feedback and stay within the shared 40-staff-minute budget. Any later learning study would need separate local approval, appropriate safeguards and a suitable learning-outcome design.

School readiness concerns educational purpose and agency, not only fluent output. UNESCO’s 2023 guidance on generative AI in education and research addresses human-centred, age-appropriate use, privacy and ethical and pedagogical validation. It supports asking these questions; it does not validate this fictional tutor or prescribe these eight cards. No product-access age or universal legal-consent threshold is inferred here.

3. Company packet: reporting accuracy is not review capacity

C1 · Purpose and authority. Harbour Components wants internal draft summaries of aggregate packing work. Its operations lead permits staff to prepare and check drafts from the supplied synthetic records. Only the lead can approve a reporting trial and the final internal pack. The reviewer may withhold an unsupported draft and ask the operations lead to resolve it. There are no employee-level performance records, customer records or external recipients. A model may not edit the source ledger or send a report. The intended benefit is less total staff effort at unchanged factual quality, compared with a checked spreadsheet summary.

C2 · Complete input. The source owner’s signed exercise sheet defines completion rate as completed packs divided by planned packs for the same week. Week A: 120 planned, 72 completed. Week B: 150 planned, 90 completed. Both rows have a matching period label and owner confirmation. A separate narrative card N1 reads “The missing parts arrived; all delayed packs are complete,” but its period and approving owner are blank. The source rule is explicit: omit N1 from the substantive report until the knowledge owner verifies its period and content; show the omission as an unresolved input.

C3 · Draft to check. “Completion rate rose 25 percentage points in Week B, from 72 to 90, because missing parts arrived. All delayed packs are complete.” The first sentence confuses counts with rates. Both completion rates are 60%: 72/120 and 90/150. Completed volume rose by 18 packs, which is 25% relative to 72. The completion-rate change is zero percentage points. The causal explanation and all-complete statement depend on unsupported N1.

C4 · Proposed volume and observed effort. The proposal would generate 45 draft summaries each week. In a supplied six-summary rehearsal, each summary took eight staff-minutes in total, including setup, drafting, source and arithmetic checks, correction, acceptance and routine rework. All eight minutes are charged to the single reviewer’s protected budget of 180 staff-minutes per week; eight is not a per-stage allowance. The six-card trial included two planted rate/count errors and one undated explanation; the reviewer found all three. No reviewer cover or overflow arrangement exists. The 180 minutes must also cover unusual exceptions; they are not free extra capacity.

C5 · Value, recovery and repair evidence. Manual checked summaries take ten staff-minutes each on the same six exercise cases. The AI-assisted method’s eight minutes include setup, drafting and review, but software costs and wider-case performance are unknown. The manual spreadsheet remains usable and the lead can withhold the draft pack. A proposed repair would cap generation at 18 summaries a week, reserve the remaining review time for exceptions and send all other summaries through the existing manual route with its existing staffing. This is a proposal, not an approved or tested change. Missing-input N1 remains withheld until the source owner resolves it.

4. Public-service packet: information help, accessible routes and contestability

P1 · Purpose and local authority. Meadow Civic Information Desk wants to explain public service-desk opening hours and direct questions to staff. The service owner permits staff-only preparation from public information and requires a human to approve any public answer. The proposed chatbot would instead answer the public directly. Local policy keeps the staffed counter and telephone route available, permits people to request correction or an explanation from the duty supervisor, and prohibits the assistant from deciding entitlement to benefits. No resident, application, health, financial or identity records are supplied or needed.

P2 · Complete authoritative source card. In this fictional packet, desk manager guide M3 is current: “Monday, Wednesday, Thursday and Friday: 10:00–18:00. Tuesday, Saturday and Sunday: closed. A step-free entrance is on the west side. The duty supervisor handles corrections and questions outside this guide. Benefit eligibility is decided by the authorised benefits team, not by this information desk.” The desk manager owns M3 and withdraws superseded copies. The weekly check log records that staff used M3 on three consecutive rehearsal days. This establishes a source in the exercise, not actual opening hours.

P3 · Complete request and flawed draft. The invented request is “Can I come on Tuesday, and can the assistant tell me if I qualify for support?” The draft says “Yes, come any weekday between 10:00 and 18:00. Our assistant can confirm your entitlement.” Both parts exceed the evidence. The corrected draft is “The desk is closed on Tuesday. M3 lists opening hours of 10:00–18:00 on Monday, Wednesday, Thursday and Friday. The information assistant cannot determine benefit eligibility; that question belongs to the authorised benefits team. You can ask the duty supervisor to explain or correct this information.” No personal details are needed to prepare this answer.

P4 · Checking and access record. Two trained staff members each checked six fictional draft answers in a rehearsal; all twelve checks took three minutes each. Together they found all three seeded errors: an incorrect opening day, invented eligibility authority and an invented service promise. The proposed staff-only scope is 15 drafts a week, at that observed three-minute rate, within 60 protected staff-minutes including an exception reserve. Staff rehearsed withdrawing a wrong draft and using M3 directly. Public-facing chatbot keyboard access, screen-reader operation, language coverage and the supervisor-request route have not been tested. An interface screenshot alone does not establish any of these.

P5 · Separate expansion request. A project note asks to “let the same assistant approve or reject benefit applications.” That violates P1 and supplies no approved criteria, decision authority or safeguards. It is a different use case, not a minor feature of an opening-hours assistant. Do not invent applicant records or make any eligibility judgments in this exercise.

The source of the access and challenge questions matters. The OECD’s 2022 public-service design principles are advisory and support accessible, equitable services and accountability. The OECD AI Principles, updated in 2024, support meaningful information that helps people understand and challenge AI outputs. These frameworks inform this exercise’s chosen safeguards. They are not proof of a universal legal duty, and they do not replace the applicable local law, service standard or authorised decision process.

5. Worked audit: four outcomes from specific evidence

School: proceed with the permitted teacher-only rehearsal. Purpose/authority: Stable for preparation under S1, absent for classroom expansion. Sources/data: Stable for the supplied key and synthetic inputs in S2. Checking/capacity: Stable within S4’s rehearsal, with 8 × 4 = 32 combined staff-minutes used and eight combined staff-minutes remaining from the shared budget of 40. Access/challenge: Stable for the two participating teachers, who can reject a card and consult the teaching lead; learner-facing access is Unknown. Value/recovery: teacher-preparation quality can be checked and the paper fallback was rehearsed, but learning improvement is Unknown under S5. Nothing is Proven at classroom scale. Record the scope beside every rating.

The school decision does not require pretending S5 is complete. Learning effectiveness is not the claim being tested in this teacher-only rehearsal. The teaching lead should reject the attainment slide, retain both the correct key and the inappropriate-feedback example, and review the next eight cards for errors, workload and explanation quality. Stop the rehearsal if either teacher cannot verify a card; use the paper bank. Direct learner use needs a fresh decision with educational, privacy and access evidence. A cautious auditor may call the overall proposal Constrain if its scope includes that unapproved expansion; the permitted and blocked activities remain the same.

Company: repair before the 45-summary trial. Purpose/authority: Stable for drafting, with trial approval still belonging to the operations lead. Sources/data: Stable for the two numerical rows, Unknown for N1. Checking/capacity: Fragile at the proposed volume despite successful error detection in C4. Access/challenge: Stable in the small staff exercise, where the reviewer can withhold and query the lead. Value/recovery: a usable manual fallback exists and the six-case comparison is promising, but broader net value remains Unknown. The failure is not repaired by the presence of a reviewer’s name.

Demand is 45 × 8 = 360 staff-minutes per week against 180 available: a 180-minute weekly shortfall before unusual exceptions. Theoretical capacity is 180/8 = 22.5 summaries, so at most 22 whole summaries fit, leaving only four minutes. That arithmetic is not a sensible operating cap. C5’s proposed cap of 18 uses 144 minutes and leaves 36 for exceptions; the remaining 27 summaries must keep their existing manual staffing rather than vanish from the workload. The lead must confirm that staffing, authorise the cap, and rerun an 18-summary rehearsal with missing-source cases. No automatic trial approval follows from this calculation.

The complete repaired report excerpt is: “Week A completed 72 of 120 planned packs (60%). Week B completed 90 of 150 (60%). Completed volume increased by 18 packs (25%), while completion rate was unchanged. N1 is undated and unapproved, so the reason for the change and the status of delayed packs are not established.” The knowledge owner owns the N1 repair. The reviewer checks every number, denominator, period and explanatory claim. On the six supplied cases the time saving is 6 × (10 − 8) = 12 staff-minutes; it is not an established organisation-wide saving.

Public service: constrain to the permitted staff-only drafting role. Purpose/authority: Stable for preparation, absent for direct automatic replies. Sources/data: Stable for M3’s public facts. Checking/capacity: Stable for the rehearsed staff path, with 15 × 3 = 45 minutes and 15 minutes reserved. Access/challenge: Unknown for the proposed public interface under P4; the existing staffed routes remain available under P1. Value/recovery: draft checking and withdrawal are demonstrated only in the exercise; actual public benefit is Unknown. The service owner can retain the narrower preparation task and the corrected P3 response. Staff continue to approve public wording and route unresolved questions.

Constrain does not certify the untested chatbot interface. Before considering direct public interaction, the service owner would need evidence from the intended access routes, a working correction request, current source handling and appropriate local review. Test with nonsensitive fictional questions first; do not ask residents to submit case records merely to show a prototype. If a current answer cannot be established, use the existing staffed information route instead of guessing.

Public-service expansion: stop automated eligibility decisions. P5 contradicts the local authority boundary in P1. Better accuracy on opening-hours questions does not remove that prohibition. The useful alternate role is approved public-information preparation, with eligibility questions directed to the authorised team. A real high-impact use case would require its own legitimate authority and domain-specific assessment; this lab neither performs nor approves one. All four decisions concern a workflow and scope, never an institution’s overall worth or “AI readiness”.

6. Error clinic: the reassuring average and the missing result

Misleading average. A sponsor assigns four company dimensions 4/4 and gives review capacity 0/4, then announces “16/20 = 80%, so we can proceed.” The arithmetic is correct but the conclusion is not. The scales have no validated equal intervals or compensatory meaning, and a spare source-quality point cannot supply a reviewer-minute. The actual constraint remains 360 minutes of demand against 180 of capacity. Replace the average with the five-field profile, flag the failed capacity gate, and record the capped-volume repair and approval still required.

Missing-evidence error. Another sponsor relabels the school’s missing learning outcomes “Stable” because teachers accepted five cards without repair. The five cards concern feedback acceptability; there is no observed learner result. Repair the claim now to “Teacher review worked on this small synthetic set.” Keep learning effectiveness Unknown. A proposed future delayed no-help task is not a completed result. A stronger model, a larger training budget or an enthusiastic user survey does not supply the absent evidence.

When stopping is a useful result. Suppose a new conventional template generates all required company facts correctly and takes six total staff-minutes per checked summary on the same cases. If no other benefit is shown, the eight-minute AI approach has no demonstrated time advantage for this task. Stop that design or investigate a different, bounded benefit. Readiness can favour a spreadsheet, a source correction, a teacher conversation or an existing staffed service. It does not require moving up an autonomy ladder.

7. Changed cases: make an independent decision first

Cover the answer section and write one short decision for each case. Supply the allowed scope, two evidence references, the limiting condition, any arithmetic, the accountable role, the next observable check and one claim you cannot make. Use only the changed facts below and the original cards that have not been replaced. If you already read the key, describe your attempt as guided practice; a fresh independently answered variant is needed to demonstrate transfer.

T1 · A different school. A receiving school retains the exact S2 key, S3 rubric and S1 teacher-only permission. It has one teacher and a 40-minute session. Its own four-card rehearsal took 6, 6, 6 and 6 minutes, with every seeded error detected and the paper fallback completed. It proposes eight cards next. It has no learner-outcome evidence. What is the decision, a defensible revised cap and the claim boundary? Do not import the original school’s four-minute rate.

T2 · A different company team. The knowledge owner now supplies a dated and approved N1 correction: “Parts arrival is confirmed for Week B; no causal test was performed. Twenty delayed packs remain incomplete at the end of Week B.” The numerical rows are unchanged. The receiving team’s lead has authorised up to 18 internal draft summaries weekly; a separate existing team retains the other 27 manual summaries. Two trained reviewers have 90 protected minutes each. A local 18-summary rehearsal took 144 minutes including routine corrections, caught all seeded errors and successfully used the manual fallback. Exceptional-work reserve is the remainder; software costs and long-term benefit remain unmeasured. Decide whether the bounded trial can proceed and write the corrected interpretation.

T3 · A changed public-service context. Staff have verified keyboard access, screen-reader reading order and the supervisor-request path for a prototype using fictional questions. But an approved replacement guide M4 now states “Wednesday closed for maintenance; other M3 days unchanged.” A cache still gives M3’s Wednesday opening answer. The service owner has not authorised direct automatic replies. The proposed scope is a staff-only information draft for “Can I visit on Wednesday?” Decide what can be used now and what must be repaired. Do the accessibility tests authorise benefit decisions?

8. Separate answers and checking standard

T1 answer: repair the eight-card plan, or constrain to an authorised smaller rehearsal. Local demand is 8 × 6 = 48 minutes, eight minutes beyond the session. Six cards would require 36 minutes and leave only four; a defensible proposed cap is five cards, using 30 minutes and reserving ten. The receiving teaching lead must authorise that revised scope and check it in the next rehearsal. Source/key and demonstrated review skill transfer as useful evidence; capacity does not transfer unchanged. Learning effectiveness remains Unknown. A smaller cap is defensible if its reserve and local purpose are explained. “Proceed with eight because Cedar passed” fails the transfer test.

T2 answer: proceed with the authorised 18-summary trial. Review capacity is 90 + 90 = 180 minutes. Observed demand is 144 minutes, leaving 36 for unusual exceptions; the remaining 27 summaries have explicit separate staffing. The source gap has been repaired without proving causation. Report 60% completion in each week, volume up 25%, and 20 delayed packs still incomplete. Parts arrival is confirmed for Week B but is not established as the cause of the volume increase. The operations lead owns the trial; reviewers withhold any unsupported summary and fall back to the manual process when demand exceeds capacity. Retest on source, staffing or volume changes. Do not claim scale readiness, net monetary savings or automatic sending authority.

T3 answer: repair the stale-source path and retain staff approval. M4 takes precedence: the corrected draft says “The desk is closed on Wednesday for maintenance under M4. Other opening days remain as listed in M3. Ask the duty supervisor for clarification or a correction.” A staff member can prepare that answer directly from M4 now. The stale cached draft cannot be reused. The knowledge owner must withdraw M3 as a standalone current guide, connect M4 to the remaining M3 details and test a Wednesday question plus an unchanged-day question. The service owner then reassesses the narrow tool path. Access evidence has improved, but authority for direct replies is still absent; P5’s benefit-decision proposal remains stopped. Passing one gate does not repair a different gate.

Check your reasoning, not your optimism. A complete answer (1) names the exact scope and local decision-maker; (2) distinguishes a supplied observation, an inference and missing evidence; (3) checks workload using consistent units and reserves time; (4) makes no unsupported learning, causal, access or authority claim; and (5) gives a repair or continuation condition that someone can actually verify. All five are needed for this exercise’s completion criterion; it is not a validated psychometric score. A different decision label is acceptable when it protects the same boundaries and gives a defensible next step.

9. Questions to settle before carrying the method elsewhere

Does a low-risk preparation task need all the machinery of an autonomous service? No. The evidence should match the proposed scope and consequence. These packets can be audited without integration, dashboards or live users. Do not demand a deployment merely to prove readiness for a paper exercise.

Does synthetic practice show readiness for real service? Only narrowly. It can expose arithmetic, authority and checking errors under stated conditions. It cannot establish real-world learning gains, population-level fairness, every accessibility need or reliability under changing demand. Record the transfer boundary before increasing the claim.

Who decides when a public-service answer is wrong? The responsible local source or service owner resolves it through the applicable process; an AI-generated explanation is not the final authority. The exercise explicitly retains a correction and explanation route. Actual legal remedies, service requirements and accessibility obligations depend on the jurisdiction and service.

What should the final audit record say? “For this workflow and this permitted scope, we choose this outcome because these evidence cards establish these conditions; these facts remain unknown; this person owns the next check; this event reopens the decision.” Attach the corrected output and the failed example. This is more useful than awarding an institution a flattering readiness percentage.

For the definition behind the audit, use the existing SI-ready workplace companion. For action, carry the bounded decision into the conclusion immediately below. The local institution retains authority over its own people, sources, services and next steps.

The Readiness Audit in One Sentence

A Super Intelligence readiness audit asks whether the workplace can absorb more machine capability without losing control of context, evidence, authority, outcome or recovery.

That is the standard to carry into project prioritisation: the best SI project is not merely valuable. It is valuable enough, ready enough and governable enough to deserve the organisation’s next unit of attention.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading