Super Intelligence master guide › Energy, work, learning and safety series › Article 0012
AI progress becomes practically useful when a system can complete more of a clearly defined task to an acceptable standard with less supervision. Delegation is the act of handing over some work while retaining an appropriate way to check the result and control important decisions. The meaningful measure is useful completion, rather than how long the assistant runs or how confident its final message sounds.
For adults, students and parents, this changes the question from “How intelligent is this model?” to “What work can I entrust to it, under which conditions?” A tool may answer a short question well and struggle with a longer process containing conflicting information, several dependent steps or a decision requiring local knowledge.
This guide explains task horizons, reliability, review effort and practical delegation briefs. All worked examples are hypothetical. They show how to organise a task and assess the outcome, not a guarantee that any particular product can perform the task reliably.
The series title refers to “Super Intelligence”, but artificial superintelligence remains hypothetical. Current AI capabilities are uneven. A useful increase in delegated work is evidence about a task and its operating conditions; it does not by itself establish general superiority across intellectual activities.
What makes a task genuinely delegated?
A task is more than a prompt. It has a goal, inputs, constraints and a completion condition. Asking an assistant to “research options” leaves much undefined. Asking it to compare supplied options against stated criteria and identify missing evidence creates a task that can be examined.
Delegation also includes a boundary around action. The assistant may be allowed to organise information without sending messages, change a draft without publishing it or propose a schedule without confirming bookings. These boundaries let the user grant useful independence while keeping consequential decisions in the right place.
A task is genuinely delegated when the assistant takes responsibility for an agreed portion of the process and returns a result that can move the work forward. A response that merely explains how the user could do the work may be helpful advice, but it is a different outcome from completion.
Completion needs an observable meaning
“Make it good” is difficult to evaluate. “Include every supplied proposal, use the same comparison criteria and mark missing information” provides observable conditions. The user can then check whether the result meets the task, rather than rely on an overall impression.
Acceptance conditions should be proportionate. A first draft may need completeness and source fidelity while allowing later stylistic editing. A calculation may require exact inputs and repeatability. A task that changes a record may need a check that the correct record was changed. Different work needs different evidence.
Scope is part of capability
An assistant may perform a task well when the inputs are bounded and fail when the scope is open. That difference is useful information. It identifies the conditions under which the capability works, rather than forcing a single label of capable or incapable.
For beginners, scope control can create a better experience. Start with the material already available and a specific result. Expand only when the assistant demonstrates useful completion and the review process remains manageable. This is a practical method for learning what to delegate.
Task horizons: duration and difficulty are different
A task horizon can describe the difficulty of work a system can complete, using the time a human would ordinarily require as a reference. It should not automatically be read as the number of hours an assistant can run unattended. The clock time of a run and the difficulty of the task are separate quantities.
METR’s task-horizon methodology uses human task duration and a defined success threshold. Its commonly discussed fifty-percent horizon describes a fitted level of task difficulty at that threshold, rather than universal reliability for all work of that duration. The evaluated task distribution also matters. METR: Task-completion time horizons.
This gives readers a valuable interpretation rule. A research result about structured software tasks should not silently become a claim that the same system can perform any job for the stated number of hours. Ordinary work may include tacit context, negotiation and success conditions that differ from a benchmark.
A long task is not a pile of short tasks
Completing a hundred independent exercises is different from completing one process with a hundred dependent stages. The independent work can often be divided and checked separately. In the dependent process, an early mistake may influence later steps.
Imagine a comparison task that requires reading records, resolving a contradiction, choosing an interpretation and then calculating a result. The stages connect. An incorrect interpretation can produce a neat calculation based on the wrong premise. Evaluating only the final arithmetic misses the more important failure.
Duration alone does not show value
A task may take a human several hours because of slow access to records rather than intellectual difficulty. Another may take minutes but involve a delicate judgment. Human duration is useful context, but it should be combined with the task’s structure and consequences.
For practical delegation, ask what makes the task difficult. Is it length, ambiguity, exactness, missing information or a need for contextual judgment? The answer helps determine the brief and review process. It also explains why a model can perform differently on tasks that look equally long.
Reliability changes what you can entrust
A system that sometimes completes a task can be useful for exploration, but a recurring operational process may need much higher confidence. The acceptable level depends on what happens when the result is wrong and how easily the mistake can be detected and repaired.
A draft recommendation that is clearly marked for review allows a different tolerance from a decision that creates a commitment. A user should therefore consider the consequences alongside the observed performance. The same output quality can be acceptable in one role and insufficient in another.
Several steps create several opportunities for failure
Consider an idealised process with five required steps. If each step had a ninety-percent chance of being correct and the events were independent, multiplying gives about fifty-nine percent for all five being correct. Real task errors need not be independent, so this is a teaching calculation rather than a model-performance estimate.
The example shows why strong local performance can coexist with weaker end-to-end completion. It also shows the value of checking critical intermediate results. A checkpoint may prevent an early error from influencing the rest of the task, provided the check actually detects the relevant problem.
Repeated success should cover meaningful variation
A task performed well once may have been unusually straightforward. Evaluation should include variations that matter: incomplete records, different formats, an ambiguous instruction or a conflict between sources. These cases reveal whether the assistant understands the task’s boundary or simply handles a familiar pattern.
This does not require inventing endless tests. Choose cases that reflect actual work and its important difficulties. Once the evidence is sufficient for the intended role, use the system within that role and revisit the assessment when the task or tool changes materially.
Supervision is part of the cost
An assistant can generate an output quickly while requiring substantial briefing and review. The net benefit depends on the complete process. Count the time spent preparing inputs, checking the result, correcting mistakes and integrating the output into the next stage.
Suppose a hypothetical task previously required ninety minutes. With assistance, preparation takes fifteen minutes, generation five, review twenty and correction ten. The assisted process takes fifty minutes and saves forty under those assumptions. Describing it as a five-minute task would misrepresent the user’s contribution.
The result might still be valuable if it saves time, improves quality or makes previously impractical work possible. Honest accounting makes the benefit clearer. It also helps identify where a better brief or more suitable task could reduce supervision further.
Distinguish review from redoing
A useful review checks a result against defined conditions. Redoing reconstructs the work independently because the output cannot be trusted or understood. If the user routinely has to redo everything, the apparent delegation may not be delivering much benefit.
The distinction can be practical. A report with sources tied to individual claims may be easier to check than a polished report with no traceable evidence. An organised output can reduce review effort even when the amount of generated text is smaller. Design the result for examination, not merely presentation.
Attention matters as well as minutes
A user may value fewer interruptions even when total task time changes little. Conversely, constant checking can fragment attention and make the process harder to manage. The review arrangement should fit how the user works.
For a hypothetical longer task, a checkpoint after the comparison stage may be more useful than repeated minor questions about formatting. The assistant can make routine choices within the brief while pausing at a decision that materially changes the result. This concentrates supervision where it has value.
The economic consequence is net value after supervision and correction; a longer autonomous attempt is useful only when the accepted result justifies the total effort and resources.
Write a delegation brief that does real work
A strong brief begins with the intended outcome. State who will use the result and what they need to do with it. This context helps determine the appropriate format and depth. A report for deciding between options has a different purpose from notes for learning about a topic.
Then define the materials and constraints. Specify which records should be used, which information should be treated as authoritative and what the assistant should do when something is missing. An instruction to mark uncertainty is more useful than encouraging a complete-looking answer at any cost.
Finally, define the action boundary and completion conditions. May the assistant edit a local draft? May it contact anyone? Does the task end with a proposal or an implemented change? Clear answers prevent the user and system from silently working toward different outcomes.
A practical brief in connected prose
For a hypothetical comparison, the user might ask the assistant to examine three supplied proposals for a school activity, compare cost, timing and staffing, identify missing information and return a short recommendation with reasons. The user can reserve the final choice and all external messages.
That brief has a purpose, bounded input, shared criteria and a reviewable result. It does not need complex terminology. The value comes from making the task’s logic explicit and giving the assistant a clear route when the records do not support a conclusion.
Explain what should trigger a question
A brief can distinguish routine decisions from material uncertainty. The assistant may choose a clear paragraph structure without asking, while a missing cost that changes the recommendation should be flagged. This reduces unnecessary interruption without treating silence as permission for every action.
The trigger should follow the task’s purpose. Ask for clarification when the choice changes scope, meaning or a consequential decision. For a low-impact formatting detail, a reasonable choice may be sufficient. The user gets a more useful process when attention is directed toward the important uncertainties.
Effective delegation requires a clear brief and an acceptance standard; the user must distinguish missing information, an unclear goal and a failure of execution.
Worked example: comparing activities for a family
Imagine a family choosing among three weekend activities. The assistant receives the activity descriptions, stated prices, travel requirements and the family’s available period. The task is to prepare a comparison and identify unanswered questions.
The first output should preserve the difference between known and missing information. If one activity does not state whether equipment is included, the assistant should not assume that it is. It can explain how the missing information affects cost comparison and propose a question to resolve it.
Suppose the assistant recommends the lowest listed price. The user checks whether travel cost, duration or an unstated extra has been ignored. This is not merely proofreading. It examines whether the recommendation follows the task’s criteria and whether the data supports the comparison.
The family may then refine the task. One child values time outdoors, while another needs an activity with less physical effort. These preferences were not in the original material. The assistant can update the comparison once they are provided, but should not be credited with knowing them beforehand.
No booking is made as part of this task. The action boundary is comparison. That keeps the exercise reversible and lets the user learn how the assistant handles constraints before considering a wider role. A useful result saves preparation while leaving the family’s preferences and final decision visible.
The completed task is not simply a table. It is a comparison that the family can use, with uncertainty marked and no invented commitment. This is the level at which delegation becomes meaningful.
Worked example: an administrative research report
Consider a hypothetical team asked to compare three approaches to organising an internal training programme. It has approved descriptions, costs and participation requirements. The assistant’s task is to produce a report using those materials and a shared set of evaluation questions.
The team asks for an initial outline before the full report. This checkpoint tests whether the assistant has understood the decision and the information available. If the outline assumes a fourth approach that was never supplied, the team can correct scope before substantial work is done.
During preparation, the assistant identifies conflicting participation figures. It should preserve the conflict and ask which source controls, or describe both if that is appropriate to the brief. Quietly choosing one figure would conceal a judgment that may materially affect the result.
The final report ties each comparison to the source material. The reviewer checks the high-impact claims first: total cost, time commitments and requirements that exclude an option. This is more efficient than giving equal attention to every stylistic phrase.
Now count the supervision. If the reviewer spends substantial time locating sources because the report does not identify them, the output format needs improvement. If the sources are clear and the report accurately organises the material, review can focus on decisions rather than reconstruction.
The team assesses completion by asking whether the report makes its decision easier while preserving the evidence. It does not judge success by length or by whether the assistant produced a confident recommendation. The useful work is a faithful comparison that supports a real next step.
Worked example: delegating a software change
A hypothetical club maintains a simple registration form. It wants the form to recognise that a required field is missing and display a clear message. The assistant is asked to prepare the change in a reviewable copy, with the original available for comparison.
The brief names the intended behaviour and gives examples: a completed field should proceed, an empty field should show a message and unrelated form behaviour should remain suitable. These conditions describe a result rather than instruct every implementation detail.
The assistant may write a change, but the reviewer examines whether it meets the examples and whether it affects other paths. A form can behave correctly in the demonstration while failing with a different input. The evaluation should therefore cover the behaviour that matters to users.
The action boundary also matters. Editing a reviewable copy is different from deploying it to a live service. The assistant’s ability to prepare the change does not establish that publication should happen without an authorised review. Scope and authority remain separate dimensions of delegation.
Now consider a failure. The assistant may report completion even though it changed the wrong file or could not save the edit. The reviewer needs evidence of the resulting state, not only a narrative about what was attempted. This principle applies to any task that changes something beyond producing text.
For delegation, the core question remains whether the intended change was completed and verified within the agreed boundary.
The broader application is delegating a software change across its lifecycle; implementation, integration, verification and maintenance remain parts of the completed job.
Longer tasks need structure and state
A longer task may require the assistant to retain decisions, intermediate findings and unresolved questions. Without a clear record, it may repeat work, contradict an earlier choice or lose an important constraint. The task should produce a state that can be understood as work progresses.
For a research report, that state might be a list of sources examined, questions answered and gaps remaining. For a software change, it might include the intended behaviour, edits made and checks completed. The record should serve the task rather than become a detailed diary of every minor action.
A handover should preserve what matters
If a user or another process takes over, they need to know the current result, the basis for important decisions and what remains unresolved. A useful handover makes continuation possible without reconstructing the whole history.
This also helps the original user review completion. A final report that explains a limitation clearly can be more useful than one that presents unfinished work as finished. The assistant should identify the actual boundary reached and the remaining requirement in plain language.
Recovery is a capability
A system may encounter an unavailable record, a failed operation or an unexpected format. Its response matters. Can it explain what failed, preserve completed work and continue the parts that remain possible? Or does it abandon the task or conceal the problem?
Recovery should be evaluated as part of longer delegation. A task rarely becomes more useful merely because the assistant kept running. The value lies in whether it moves toward the intended outcome while maintaining an accurate account of its state.
The operating mechanism is permissions, monitoring and recoverable actions; a task’s scope should be enforced by what the agent can reach, change and restore.
Worked example: a task that should return an unresolved question
Imagine a hypothetical comparison of three membership plans. The user asks for the lowest total cost over a year. Two plans supply all the required prices. The third has a lower monthly fee but does not state an initial charge. The assistant cannot establish the requested ranking completely from the supplied material.
A useful result calculates the totals supported by the records and identifies the missing charge. It can also explain the condition under which the third plan would be cheaper. For example, if its known yearly amount is lower by twenty hypothetical units, an initial charge above that difference would change the ranking.
This conditional result is more useful than inventing a charge or selecting the plan with the lowest visible monthly price. The assistant has performed the available work and shown exactly what is needed to finish. The task’s uncertainty becomes smaller and more actionable.
The user can then obtain the missing information or choose a narrower question. If the purpose is a first screening, the conditional comparison may be sufficient. If the purpose is a final purchase decision, the missing information needs resolution. Completion depends on the intended use, not on whether the output can be made to look final.
This example also clarifies what persistence should mean. The assistant should continue useful work that does not depend on the missing charge, but should not pretend that repeated reasoning creates the absent fact. Productive persistence combines progress with an honest boundary.
For a student, the exercise is similar to solving a word problem with an unspecified value. A good answer can state a formula or condition and explain why a single numerical result is unavailable. Understanding that boundary is evidence of sound reasoning, not a failure to be helpful.
Decide how much independence a task needs
Independence can be granted in stages. An assistant might first prepare material, then propose a change, then implement an approved reversible change. These are different roles with different completion evidence. The user should know which stage the brief authorises.
A practical comparison is between choosing a report layout and changing the criteria used to rank options. Layout can often be handled within routine discretion. Changing criteria alters the meaning of the task and may require user direction. Treating both choices as equally important creates either unnecessary interruption or inappropriate independence.
The same principle applies to communication. An assistant preparing a message can choose clear wording while preserving the intended meaning. Sending that message to someone else is an additional action. The user’s authority boundary should be explicit enough that successful drafting does not quietly become permission to communicate.
For longer tasks, the best checkpoint is often located where a consequential choice occurs. The assistant can complete the surrounding preparation and present a concrete result for review. This gives the user something meaningful to assess and keeps attention focused on the decision that needs it.
The arrangement should reflect the actual task, rather than a blanket rule that every step requires approval or that no step does. Appropriate independence is the ability to make routine progress while preserving meaningful control. That is a skill both users and systems need as delegation expands.
Evaluate the complete result
A delegation assessment should compare the intended task with the delivered state. Did the assistant use the right material? Did it respect the constraints? Does the result meet the acceptance conditions? Were important uncertainties and failures visible?
NIST’s generative AI risk profile identifies confident erroneous output as a risk. This reinforces the need to inspect evidence and resulting behaviour rather than use fluency as a reliability measure. NIST: Generative AI profile.
Match checks to failure modes
If the main risk is omitted information, check completeness against the source set. If the main risk is arithmetic error, verify inputs and calculations. If the task changes a record, inspect the record. Generic proofreading may miss the failure that actually matters.
The check should be independent enough to detect the problem. Asking the same assistant to repeat that it is correct does not establish correctness. A source comparison, a reproducible calculation or observation of the resulting state can provide a more meaningful basis for confidence.
Do not confuse a benchmark with your workflow
A benchmark can inform expectations, but a local task has its own inputs, conventions and consequences. Try representative work and examine the complete process. The goal is not to recreate a research study; it is to establish whether the assistant has a suitable role in the user’s actual activity.
This principle supports responsible expansion. When a task works with reasonable review effort, a user can consider a larger scope. When it fails in a meaningful way, the response may be a narrower task, better inputs or a different tool. Capability should be learned through evidence rather than assumed from a general score.
Common mistakes in AI delegation
The first mistake is treating an open-ended instruction as an agreed task. The assistant may produce something plausible while answering a different question from the one the user intended. A clear goal and completion condition make this mismatch easier to prevent.
The second is omitting review from the time-saving claim. Generation may be fast while verification remains demanding. Include the whole process so that the claimed benefit corresponds to the user’s experience.
The third is using output length as evidence of thoroughness. A long report can repeat unsupported claims, while a shorter structured result can preserve all the important information. Measure coverage against the task rather than word count alone.
The fourth is granting broader authority because a smaller task worked. An assistant that drafts a good message has not thereby demonstrated suitability to send it, negotiate a commitment or change a record. Each extension needs a clear boundary and appropriate evidence.
The fifth is accepting a declaration of completion without examining the state. For a practical task, the result may be a saved file, changed schedule or reviewed comparison. The assistant’s statement should correspond to that result, and uncertainty should remain visible.
Building delegation skills through learning
Students can begin with a task they understand well enough to check. Ask an assistant to organise supplied notes, then compare the result with the original. Identify any missing point, changed meaning or invented detail. This teaches source fidelity alongside task design.
Adults can use a similar exercise at work with suitable material. Define one repeated activity and record the complete time and correction involved before and after assistance. The exercise can reveal both useful capability and hidden review effort.
Parents can encourage children to explain why they accepted or rejected an output. “The assistant said so” is not a reason. A reason may be that the result matches the supplied facts, the calculation can be reproduced or the explanation satisfies the task. This preserves the learner’s judgment.
The practical aim is increasing the amount of useful work that can be entrusted while maintaining understanding of the result. Delegation is strongest when the user knows the task’s purpose, the assistant’s limits and the evidence that supports completion.
Frequently asked questions about delegating work to AI
Does a longer task horizon mean an assistant can work unattended for that many hours?
Not necessarily. A horizon may refer to task difficulty measured through human duration at a stated success threshold. It should be read with the benchmark’s task distribution and method, rather than as a universal unattended operating period.
For everyday use, examine the actual task and supervision required. The result you need may differ substantially from the work used in a benchmark.
How do I know whether a task is suitable for delegation?
Look for a clear goal, available inputs and a result you can check. Consider what happens if the assistant makes an error and whether the action can be reversed. A bounded, reviewable task is usually easier to assess than an undefined responsibility.
Suitability is specific. A task can become more appropriate when scope is narrowed or the input process improves, even if the model itself has not changed.
Should I write every step in the brief?
You need to define the outcome, constraints and important boundaries. Specifying every minor step can limit useful independence and create extra work. The assistant may be able to choose routine methods within the task.
Give more guidance where a choice affects meaning, evidence or authority. A clear acceptance condition often does more work than a long set of procedural instructions.
Is human review evidence that delegation has failed?
Review can be part of a successful arrangement. The question is whether the complete process produces useful value with an appropriate amount of attention. A task that saves substantial preparation while remaining easy to check can be worthwhile.
If the user must routinely redo the work, the benefit may be weak. Distinguish checking from reconstruction when measuring the result.
What is the best measure of AI progress for ordinary users?
A practical measure is the useful work completed to the required standard, including the supervision and repair needed. Capability, quality and user effort belong together. A system that takes on more work while making review harder may not improve the user’s outcome.
Continue with AI adoption across industries or return to the Super Intelligence guide to connect delegation with infrastructure, software and learning.
Previous: 0011 — Space Transport and the Logistics of a Larger Civilization · Next: 0013 — The Economics of Software Abundance
