Teacher planning has a peculiar cost structure. The work is intellectually important, but many parts are repetitive: generating practice questions, adapting a text, drafting a quiz, finding another example, creating a first version of a worksheet, or producing a model answer that will later be checked and revised. These are exactly the kinds of tasks generative AI can perform quickly.
The obvious question is whether speed comes at the price of quality.
A Teacher Choices trial from the Education Endowment Foundation, completed in September 2026, gives a useful answer for one bounded use case. In 68 English secondary schools, 259 mainly Science teachers were asked to use ChatGPT with a practical guide for lesson and resource preparation. After a five-week familiarisation period, teachers using ChatGPT spent an average of 56.2 minutes per week preparing lessons and resources for the relevant classes, compared with 81.5 minutes for the comparison group: a saving of 25.3 minutes, or 31 percent. An expert panel, blind to which resources were AI-assisted, did not detect a reduction in resource quality. The time-saving result received a high security rating.
That is promising. It is not a licence to outsource planning.
The real mechanism is narrower: AI can compress low-level drafting and variation work when a knowledgeable teacher remains responsible for the learning goal, curriculum fit, factual accuracy, difficulty, safeguarding and final judgement.
The 50-second answer
AI-assisted lesson preparation works best as a teacher-controlled drafting loop. The teacher defines the instructional job, asks the model for a bounded output, checks it against subject knowledge and learner needs, revises it, and keeps only what improves the lesson.
The strongest current evidence is about teacher time, not direct student attainment. In the EEF trial, ChatGPT-assisted teachers saved about 25 minutes per week on relevant lesson and resource preparation, a 31 percent reduction, while blinded expert reviewers found no apparent quality loss in the resources sampled. Teachers most commonly used the tool for creating questions or quizzes and generating activity ideas rather than handing over complete lessons.
The caveat is crucial. Fluent output can be wrong, mispitched or pedagogically weak. Quality control is not an optional final polish; it is the teacher’s core role. AI is useful when it removes drafting friction and leaves professional judgement more available—not when it replaces the reasoning that makes a lesson appropriate for a real class.
1. Start with the instructional job, not the tool
A planning session should begin with what students need to learn, what evidence will show it and where the likely bottleneck lies. Opening a chatbot before naming the job encourages generic activity generation.
A concrete example: the teacher writes, ‘Students can expand brackets but confuse equivalent forms; I need three examples that isolate that distinction’ before prompting the model.
The caveat: A technically impressive output can still be irrelevant if the learning problem was never specified.
2. Use AI for variation after the core example is secure
Models are good at generating surface variation quickly. That is valuable once the teacher knows what structure must remain invariant.
A concrete example: after designing one valid ratio problem, the teacher asks for six variants that change context while preserving the same multiplicative relationship.
The caveat: Generated variants can accidentally change difficulty or introduce a second concept. Every item needs review.
3. Question generation is a high-value use case
Teachers repeatedly need examples, distractors, quizzes and retrieval prompts. AI can accelerate first drafts, especially when the teacher supplies the target misconception.
A concrete example: generate four options for a hinge question where each wrong option maps to a known error about gradient.
The caveat: Distractors invented without misconception knowledge may be implausible or diagnostically useless.
4. AI can help produce first-pass explanations
A model can offer several ways to explain an idea, giving the teacher alternatives to compare.
A concrete example: ask for three explanations of electric potential difference: one verbal, one analogy-based and one diagram-oriented.
The caveat: Analogies can import misconceptions. The teacher must decide which explanatory trade-offs are acceptable.
5. Adaptation is useful when the target is preserved
Teachers often need to simplify language, change context or reduce reading load without lowering the intellectual goal.
A concrete example: rewrite a dense Science prompt using shorter sentences while keeping the causal reasoning requirement unchanged.
The caveat: Simplification can quietly remove disciplinary vocabulary or clues students need to learn. Compare before and after.
6. AI can help create worked examples
A model can draft step-by-step examples that the teacher then verifies and edits for instructional sequencing.
A concrete example: generate a quadratic example that makes a sign error tempting, then ask the teacher to review each step.
The caveat: Mathematical and scientific errors are especially dangerous because polished formatting can make them look authoritative.
7. AI can help create model answers, but provenance matters
A draft model answer can save time when the teacher already understands the success criteria and will revise the response.
A concrete example: generate a first version of a short evidence-based paragraph, then annotate where it meets or misses the rubric.
The caveat: Students should not be told a machine-generated answer is an authoritative exemplar merely because it sounds polished.
8. Prompt quality improves when constraints are explicit
Models respond better when teachers name year level, prerequisite knowledge, learning goal, forbidden shortcuts, output form and checks.
A concrete example: ‘Create five short-answer questions for students who know linear equations but confuse gradient and intercept; include answers and likely errors.’
The caveat: Long prompts do not guarantee good pedagogy. Constraints should clarify the job, not bury it.
9. The teacher should ask for alternatives, not one answer
A single generated plan creates anchoring. Several options allow comparison and make professional judgement visible.
A concrete example: request three activity structures with different trade-offs in time, participation and cognitive demand.
The caveat: Too many options create choice overload. Generate enough to compare, then stop.
10. Verification is a planning phase
Checking factual accuracy, calculations, quotations, links, curriculum claims and answer keys must be treated as part of the workflow.
A concrete example: recalculate every generated Mathematics answer and verify every Science claim against trusted materials.
The caveat: Verification time can erase savings on high-risk tasks. Use AI where checking is cheaper than creating from scratch.
11. Curriculum alignment cannot be inferred from fluency
Models know broad educational language but may not reliably match the exact sequence, specification or local assessment conventions.
A concrete example: the teacher cross-checks a generated question against the actual syllabus statement and previously taught content.
The caveat: Do not let a model invent curriculum authority or claim that a topic is required without verification.
12. Difficulty calibration needs human judgement
AI can make text shorter or problems larger, but difficulty depends on prerequisite knowledge, representation, language and decision load.
A concrete example: a teacher rejects a generated ‘challenge’ question because it merely adds arithmetic rather than deeper reasoning.
The caveat: ‘Harder’ prompts often produce bigger numbers instead of qualitatively different thinking.
13. The model does not know the class
It does not see last week’s misconceptions, the student who needs a visual scaffold, the class’s pacing history or the emotional residue of a failed assessment unless the teacher supplies relevant non-sensitive context.
A concrete example: the teacher uses anonymous misconception patterns, not personal student records, to request targeted examples.
The caveat: Privacy matters. Do not paste identifiable learner data into systems unless policy and permissions clearly allow it.
14. AI should not decide what evidence counts as mastery
Assessment design requires construct judgement: what capability is being measured and what would contaminate the result.
A concrete example: a teacher decides that spelling should not dominate a Science explanation rubric, even if the model suggests it.
The caveat: Generated rubrics can reward easy-to-count features rather than the intended construct.
15. Time saving should be measured end-to-end
Fast generation is not the same as reduced workload if teachers spend equal time repairing weak outputs.
A concrete example: track prompt, review, correction, formatting and finalisation time across several weeks.
The caveat: The EEF trial measured planning time after a familiarisation period, which matters because beginners may initially work more slowly.
16. Familiarisation has a setup cost
Teachers in the EEF trial had five weeks to learn the guide and practise before recorded planning weeks. Tool fluency changes the economics.
A concrete example: a new user first learns a few stable prompt patterns for quizzes, examples and adaptations.
The caveat: Schools should not expect instant efficiency from a tool teachers have never used.
17. Light use can be enough
Teachers in the trial often used ChatGPT for one or two activities rather than handing over complete lessons. This is instructive.
A concrete example: use AI to generate practice questions while retaining teacher-designed explanation and sequence.
The caveat: A whole-lesson generation workflow can create more checking work and reduce coherence.
18. A stable prompt library can reduce repeated work
Teachers can keep proven prompt patterns for recurring jobs rather than rebuilding instructions every time.
A concrete example: store a prompt for ‘generate near-miss examples with misconception labels’ and update only the topic.
The caveat: Prompt libraries should be reviewed as tools and curricula change. Old prompts can preserve outdated assumptions.
19. Subject expertise makes AI more useful
Experts can detect implausible answers, weak examples and hidden misconceptions more quickly. AI often amplifies the value of existing knowledge rather than replacing it.
A concrete example: an experienced Chemistry teacher spots that a generated analogy confuses bond breaking with energy release.
The caveat: Novices may be most vulnerable to accepting fluent errors because they lack the knowledge needed to verify them.
20. AI can support novice teachers only inside stronger systems
Beginning teachers may benefit from examples and planning scaffolds, but they also need mentoring, curriculum materials and opportunities to understand why a lesson is sequenced as it is.
A concrete example: a mentor reviews AI-assisted resources and asks the novice to explain each design decision.
The caveat: A chatbot should not become a substitute mentor that never observes classroom consequences.
21. Resource quality is not lesson quality
The EEF panel reviewed lesson resources, not every dimension of live instruction. A good worksheet can sit inside a weak explanation or poor pacing.
A concrete example: a teacher uses saved planning time to anticipate misconceptions and prepare checks for understanding.
The caveat: Do not overstate the trial as proof that AI-generated lessons improve student learning.
22. Saved time has to go somewhere
Workload benefits become educationally meaningful when reclaimed time is used for feedback, diagnosis, rest, collaboration or other high-value work.
A concrete example: a teacher spends twenty saved minutes reviewing student errors before the next lesson.
The caveat: Savings can simply be absorbed by additional tasks. Organisational expectations influence whether workload actually improves.
23. AI can support accessibility drafting
Models can create alternate wording, examples or formats that teachers review for accessibility and curriculum fit.
A concrete example: produce a plain-language first draft of instructions and a vocabulary-supported version for review.
The caveat: Accessibility needs are individual. Generated adaptations should not be assumed appropriate without learner-specific judgement.
24. Bias can enter through examples and contexts
Generated names, occupations, family structures and cultural examples may reproduce stereotypes or narrow assumptions.
A concrete example: the teacher audits a set of word problems for whose lives and contexts are represented.
The caveat: A diverse surface does not guarantee unbiased content. The underlying assumptions also matter.
25. Copyright and source boundaries still matter
A model may produce text resembling common source material or fail to provide trustworthy provenance.
A concrete example: teachers use generated original practice items and verify any quoted or attributed material separately.
The caveat: Do not ask the model to reproduce copyrighted textbook passages or answer keys.
26. The best prompt may ask the model to critique its own draft
Second-pass prompts can surface assumptions, missing prerequisite knowledge or possible errors, giving the teacher a review checklist.
A concrete example: ‘List three ways this worksheet could mislead a novice and identify every item requiring factual verification.’
The caveat: Self-critique is not independent validation. The same model can confidently repeat its original error.
27. Human review should be risk-weighted
Not every output needs the same scrutiny. A brainstormed activity idea is lower risk than a medical example, safeguarding scenario or high-stakes assessment item.
A concrete example: teachers use quick review for low-stakes warm-up ideas and full verification for assessment resources.
The caveat: Risk categories should be defined by school policy rather than individual intuition alone.
28. AI can reduce blank-page planning cost
Sometimes the main benefit is initiation. A rough first draft gives the teacher something concrete to improve.
A concrete example: the model produces a skeletal sequence that the teacher reorganises around actual class needs.
The caveat: Starting from a machine draft can anchor thinking. Teachers should sometimes plan key decisions before seeing suggestions.
29. Planning with AI can expose the teacher’s own criteria
Rejecting a generated output forces the teacher to articulate what good looks like.
A concrete example: a teacher says, ‘This question is too easy because the method is named in the wording.’
The caveat: This benefit appears only if the teacher actively evaluates rather than passively accepts.
30. Shared team use can improve consistency
Departments can agree on safe, useful prompt patterns and verification routines for common resource jobs.
A concrete example: a Science team shares a checked prompt for generating retrieval questions aligned to its curriculum sequence.
The caveat: Standardisation should not eliminate teacher adaptation to real classes.
31. School policy should distinguish permitted jobs
Clear boundaries help teachers know what can be generated, what data may be entered and what must remain human-authored or independently verified.
A concrete example: policy allows quiz drafting and adaptation but forbids identifiable student data and unsupervised high-stakes grading.
The caveat: Overly vague policy creates hidden use; overly rigid policy may block low-risk efficiency gains.
32. The quality-control loop is the real mechanism
The strongest workflow is teacher intention → bounded generation → verification → adaptation → classroom evidence → revision.
A concrete example: after using generated questions, the teacher notes which distractor failed diagnostically and improves the prompt next time.
The caveat: Without the loop, AI use becomes content vending rather than professional planning.
33. Usage frequency should follow value, not novelty
The EEF trial observed declining frequency of use over time. Teachers may learn where the tool genuinely saves time and stop using it elsewhere.
A concrete example: a teacher keeps AI for quiz variation but returns to manual planning for complex conceptual explanations.
The caveat: Lower frequency is not failure if it reflects better task selection.
34. The final authority must remain accountable
A school needs to know who is responsible when a resource is wrong. The answer cannot be ‘the model produced it.’
A concrete example: the teacher signs off the final resource as they would any other material used with students.
The caveat: Accountability without adequate time or training is unfair. Organisations must support verification if they expect responsible use.
Practical route for teachers
Use a five-part loop: define the learning job, generate narrowly, verify, adapt, observe. Keep a small prompt library for recurring low-risk tasks. Decide in advance which outputs require full factual checking and which only need pedagogical review.
Track actual end-to-end planning time for several weeks rather than assuming AI is saving time. Spend reclaimed time deliberately. The strongest implementation is not maximal use; it is selective use where generation is faster than creation and verification is cheaper than repair after the lesson.
Practical route for students
Students are not the direct target of lesson-preparation tools, but they benefit when teachers use saved time well. You can also learn from the same principle: use AI for drafts, alternatives and practice generation, but keep yourself responsible for checking facts and understanding why an answer works.
Fluency is not authority. If a generated explanation seems perfect, test it against your notes, textbook, worked examples or another trusted source.
Practical route for parents and families
Parents should care less about whether a teacher used AI at some stage of planning and more about whether the final materials are accurate, appropriate and coherent. A teacher using AI to draft quiz questions is different from outsourcing professional judgement.
Schools should have clear privacy and quality-control rules. Families can reasonably ask how identifiable student data are protected and who reviews AI-assisted resources before classroom use.
Common failure modes
- Opening the AI tool before defining the learning problem.
- Generating complete lessons when only a small resource task needed help.
- Trusting fluent explanations without factual verification.
- Using identifiable student information without clear policy and permission.
- Treating bigger numbers or longer text as genuine difficulty progression.
- Allowing AI-generated rubrics to redefine what mastery means.
- Claiming the EEF time-saving trial proved improved student attainment.
- Ignoring the five-week familiarisation cost in implementation planning.
- Saving time but filling the saved minutes with more low-value work.
- Treating self-critique by the same model as independent validation.
Frequently asked questions
What did the 2026 EEF trial actually find?
Teachers using ChatGPT with a guide spent 56.2 minutes per week on relevant lesson/resource preparation versus 81.5 minutes in the non-GenAI group, an average saving of 25.3 minutes or 31 percent. Expert blind review did not detect a quality reduction in the sampled resources.
Did the trial show students learned more?
No. The primary outcome was teacher planning time, not student attainment.
What were teachers mostly using ChatGPT for?
Often one or two bounded tasks such as generating questions, quizzes or activity ideas rather than whole-lesson automation.
Should schools require every teacher to use AI?
No. The evidence supports a useful option for some planning tasks, not mandatory maximal use. Task fit, subject expertise, policy and verification cost matter.
What is the safest default rule?
Use AI to draft; use professional judgement to decide. The accountable teacher remains responsible for the final resource.
The final idea
The most useful way to think about AI in lesson preparation is not as a replacement planner. It is a drafting machine sitting inside a professional planning system.
The evidence now gives us a concrete reason to take that seriously: in one well-defined trial, teachers saved meaningful time without an observed drop in resource quality. That is valuable. But the result becomes educationally useful only when we preserve the part the tool cannot own—responsibility for what students should learn, whether the material is correct, whether the difficulty is appropriate, and what happened when the resource met a real class.
The rule is: automate the draft, not the judgement. Save time where checking is cheap, and spend the recovered attention on the parts of teaching only a teacher can see.
Sources and further reading
- Education Endowment Foundation — ChatGPT in lesson preparation Teacher Choices trial
- UNESCO — Guidance for generative AI in education and research
- Education Endowment Foundation — Using Digital Technology to Improve Learning
Continue exploring on eduKateSG
- How X Works Hub
- When Does Classroom Technology Cost More Attention Than It Saves? | How EdTech Tool Friction Works
- How to Think Properly | Use AI Without Handing Over the Final Judgement
- How Teacher Workload Works | Protect Teaching Quality by Designing the Work, Not Just Asking Teachers to Cope
A risk-tier for AI planning tasks
Low-risk drafting jobs
Brainstorming contexts, generating additional practice variations, proposing warm-up questions, producing formatting ideas and offering alternative examples are relatively low risk when the teacher can quickly recognise weak output. These are good candidates for automation because verification is fast.
The teacher should still check for stereotypes, curriculum mismatch and accidental repetition. “Low risk” means the cost of a mistake is limited and easy to catch, not that review is unnecessary.
Medium-risk instructional jobs
Worked examples, model answers, differentiated text, assessment distractors and explanatory analogies require stronger review. A subtle mathematical error, misleading analogy or over-simplified adaptation can directly teach the wrong thing.
For these jobs, require an explicit verification pass. Recalculate, compare with trusted curriculum materials, inspect every distractor and ask whether the generated version changed the construct. Medium-risk work is often where AI saves useful time—but only because teacher expertise makes checking efficient.
High-risk jobs
High-stakes assessment, safeguarding content, personalised decisions about identifiable students, clinical or mental-health claims, legal advice and any output that could materially affect student access or progression should have much stricter controls. In many schools, some of these jobs should not be delegated to public generative systems at all.
The planning rule is simple: the higher the consequence, the less acceptable it is to rely on fluent plausibility.
A seven-minute quality-control pass
A short review routine can preserve much of the time saving.
First, check the learning target: does the resource actually require the intended knowledge? Second, check facts and answers. Third, inspect difficulty: did the model accidentally add an extra prerequisite? Fourth, inspect language and inclusion. Fifth, check instructions for ambiguity. Sixth, ask whether one generated item contains a hidden shortcut. Seventh, remove anything the teacher would not be comfortable defending to a colleague or parent.
This routine should become faster with practice. The EEF trial’s familiarisation period matters because efficient AI use is itself a learned professional routine.
Three examples of useful rejection
A generated quiz can be grammatically perfect and still be rejected because every wrong option is obviously silly. A rewritten reading passage can be rejected because it removed the technical vocabulary students actually need to learn. A “harder” Mathematics question can be rejected because it only increased calculation length.
These rejections are signs of professional value, not tool failure. The model creates options cheaply; the teacher protects the instructional standard. The better teachers become at naming why an output fails, the more precisely they can prompt and the more intelligently they can decide not to use AI at all.
When not using AI is the faster choice
If the teacher already has a trusted resource, if the curriculum sequence is highly specialised, if verification will take longer than creation, or if the concept requires a carefully designed representation, manual planning may be quicker and safer.
Selective non-use is part of AI competence. The goal is not to maximise model interactions. It is to reduce total planning cost while preserving or improving quality.
A practical prompt architecture for teachers
Use four fields: context, learning job, constraints, verification.
Context tells the model the subject, age range and prerequisite knowledge. Learning job states what students must understand or practise. Constraints specify what the output must and must not do: number of items, reading level, required vocabulary, prohibited shortcuts, output format. Verification asks the model to flag every factual claim, calculation or answer key that the teacher should independently check.
This structure is intentionally boring. Good planning prompts should reduce ambiguity, not display prompt-writing cleverness. Teachers can save the pattern and reuse it.
What schools should measure after adoption
Measure more than login counts. Useful indicators include end-to-end planning time, teacher-reported usefulness by task type, correction rate of generated materials, frequency of serious errors, proportion of staff who feel pressured to use the tool, and how saved time is reallocated.
If usage rises but planning time does not fall, the implementation may be adding work. If planning time falls but correction incidents rise, quality control needs strengthening. If both improve, the school can identify which task categories are producing value.
The strongest implementation question is not “How much AI are we using?” It is “Which professional jobs became cheaper without becoming worse?”
A school implementation sequence
Schools can introduce AI planning more safely in stages. Start with a small set of low-risk jobs and volunteer teachers. Provide one concise guide covering privacy, verification, copyright and prohibited uses. Let teachers practise long enough to develop fluency before measuring time or judging value.
Next, collect examples of both good and bad outputs. Professional learning should include rejection practice: teachers compare generated resources and explain why one is misleading, mispitched or weak. This is more useful than only demonstrating impressive outputs because real competence includes knowing when not to accept the draft.
Then measure end-to-end planning time and correction burden. Ask which jobs are genuinely cheaper. A school may find that quiz variation and first-pass differentiation save time while conceptual explanations and high-stakes assessment remain faster to design manually.
Finally, keep policy revisable. Models change, features change, privacy terms change and staff expertise changes. A fixed policy written around one product can become obsolete quickly. The stable principles should be accountability, data protection, curricular alignment, verification and risk-weighted use.
A teacher-facing red-team routine
Before using an AI-assisted resource, deliberately try to break it. Ask: What assumption does this question make about prior knowledge? Can a student get the correct answer for the wrong reason? Is there an unintended cue? Does the model answer contain a claim that sounds plausible but is not necessary? Does the adaptation preserve the vocabulary students are supposed to learn?
For a worksheet, solve the problems yourself in an order different from the answer key. For a reading resource, check every factual claim and cited detail. For a rubric, test it against two deliberately contrasting student responses to see whether it rewards the intended construct.
Red-teaming turns quality control from passive proofreading into adversarial inspection. That mindset is especially useful because generative AI’s greatest strength—fluent plausibility—is also what can make weak outputs easy to trust.
The economics of verification
The value of AI planning depends on a simple inequality: the time needed to prompt, inspect and repair the output must remain lower than the time needed to create an equally good resource from scratch. That inequality changes by task and by teacher expertise.
A five-question retrieval quiz may take seconds to verify. A complex worked solution with several equations may take long enough that manual creation is faster. A culturally sensitive text adaptation may require such careful reading that the time saving disappears. Teachers should learn these task economics from experience rather than assuming every generated draft is efficient.
This is why the EEF finding is useful but bounded. It demonstrates that, across the tested planning work, a meaningful average time saving was possible. It does not say every planning task should be generated.
Why saved time can improve teaching only indirectly
The trial showed reduced preparation time, not what every teacher did with the reclaimed minutes. The educational value of workload reduction may appear indirectly: more feedback, better diagnosis, stronger collaboration, less evening work or simply better-rested teachers.
Those outcomes are not trivial. Sustainable attention is part of teaching quality. But they should not be converted into unsupported claims that AI automatically raises attainment.
A mature school implementation can value workload savings on their own terms while continuing to study whether classroom learning changes.
A final planning rule
If you cannot explain why the AI-assisted resource is better than the resource you would otherwise have used, do not use it. Novelty is not a planning objective. The useful output is the one that saves real time, preserves the intended learning and survives professional review.