eduKateSG Learning Node Series · 0004
Sometimes a lesson is not too difficult because the idea is beyond the learner. It is too difficult because the learner is trying to identify the parts and understand their relationships at the same time.
Imagine being shown an animation of a car’s braking system while a narrator explains pressure, pistons, cylinders, brake fluid and force transmission. If half the component names are unfamiliar, the learner is solving several problems at once: What is that part called? Where is it? What state can it be in? What does it connect to? What changes when the pedal moves?
Pretraining changes the order.
Before the full system begins to move, the learner first meets the important components. Names become known. Locations become familiar. Basic characteristics become available. Then, when the complex explanation starts, working memory can spend more of its limited capacity on relationships and causality instead of basic identification.
The principle is simple: learn enough of the parts that the whole becomes thinkable.
Quick Read: The Pretraining Principle
Richard Mayer’s multimedia-learning work describes the pretraining principle as the finding that people can learn more deeply from a complex multimedia message when they already know the names and characteristics of the main concepts or components. In Mayer’s classic example, learners understand an animation of a car braking system better when they first receive training on the important parts and their basic states.
A summary in Multimedia Learning explains the theoretical rationale: a learner watching a fast causal system must build both component models and a causal model. Pretraining moves some component-building work earlier so the main lesson has more capacity available for understanding how the system operates.
Pretraining does not make the final idea simpler. It makes the learner more ready to process its complexity.
The Airport Problem
Imagine arriving at a huge airport for the first time with six minutes to make a transfer.
You are trying to read gate numbers, understand terminal letters, recognise security zones, find the train, interpret arrows, remember the destination and decide whether you are moving in the right direction. Each symbol is simple in isolation. The combined system is not.
Now imagine you saw a map beforehand. You already know that Terminal 1 and Terminal 3 are connected by a train, that gate letters correspond to concourses, and that transfer security sits between two zones.
The airport did not become smaller.
You became better prepared to parse it.
Pretraining works in a similar way. It gives the learner a few stable landmarks before the full flow begins.
Why Complex Systems Overload Novices
Complex material often contains many interacting elements. To understand a process, the learner must represent not only each element but also what it does to other elements.
Consider cellular respiration. A novice may be meeting mitochondria, glucose, oxygen, ATP, carbon dioxide, water, enzymes and energy transfer almost simultaneously. If every term is unstable, the learner has little capacity left for the central causal question: how does the process transform stored chemical energy into usable cellular energy?
An expert experiences a different task. Many components are already chunked. The words are familiar. Several relationships have become one organised structure. What looks like ten separate elements to the novice may behave like two or three meaningful units for the expert.
Pretraining does not give the novice expertise. It begins reducing the element-identification tax.
Names Are Not Trivial
Teachers sometimes avoid vocabulary at the beginning of a lesson because they do not want learning to become memorisation. The concern is reasonable. But unnamed components are difficult to think about together.
A name creates a handle.
Once a learner knows “denominator,” a teacher can say, “Notice that the denominator is changing while the numerator is fixed.” Without the label, the same instruction becomes “notice that the number underneath the line in the fraction is changing.” More words are required to point at the same object, and the learner must repeatedly reconstruct which part is being referenced.
Names can therefore reduce communication cost—provided the learner connects the name to the correct concept rather than memorising an empty label.
Characteristics Matter as Much as Names
Pretraining is not a vocabulary list.
If a learner knows that a “piston” is a word but does not know what it is, where it sits or how it can move, the label provides little help. The component needs a minimal model.
A useful pretraining packet answers questions such as:
- What is this part called?
- What does it look like or where is it located?
- What is its basic role?
- What states can it be in?
- What does it connect to?
- What distinction must not be confused with a nearby concept?
The goal is not complete understanding. It is enough stable structure that later relationships can be processed without constantly pausing to decode the parts.
Pretraining Is a Sequencing Decision
Many instructional debates are really sequencing debates disguised as philosophical arguments.
Should students learn vocabulary or concepts? Should they see worked examples or solve problems? Should they receive explicit instruction or explore?
Often the answer is: both, but not simultaneously and not in the same proportion.
Pretraining says that certain component knowledge belongs earlier because it changes the cognitive cost of what comes next.
This is not front-loading everything. It is moving a small amount of strategically chosen prerequisite processing before the complex event.
The Two Models the Learner Must Build
In a dynamic explanation, learners often need at least two kinds of model.
Component Model
What are the parts? What are their properties? What can each part do?
System Model
How do the parts interact through time? Which event causes the next? What changes, moves, transfers, accumulates or constrains?
Trying to build both models from zero during a fast explanation can overwhelm a novice. Pretraining starts the component model early so the system model gets more room.
Working Memory Is Not a Warehouse
Students are sometimes told to “focus harder” when a lesson contains too many interacting unfamiliar elements.
Focus matters. But attention cannot create unlimited working-memory capacity.
If six unfamiliar terms, three transformations, two diagrams and a causal chain all have to be coordinated at once, the problem may be instructional load rather than motivation.
Pretraining reduces one source of load by making selected elements familiar enough to stop behaving as completely new objects.
For the wider load mechanism, see How Cognitive Load Works During Revision.
The Best Pretraining Is Small
There is a temptation to respond to pretraining research by creating a forty-page prerequisite packet.
That defeats the point.
Pretraining should usually be the minimum useful preparation that makes the next layer more learnable. If the main lesson is about how a cyclone forms, students may need to know pressure, warm moist air, condensation and rotating airflow. They do not need an entire meteorology course before the cyclone can be introduced.
The design question is not “What could students know first?” It is “Which missing components would otherwise consume so much attention that the system relationship becomes hard to understand?”
Pretraining Versus Prerequisite Teaching
The terms overlap, but they are not identical.
Prerequisite teaching can involve substantial earlier knowledge that is logically required. To solve simultaneous equations, a learner may need algebraic manipulation skills developed over months or years.
Pretraining is often narrower. It introduces names, characteristics, states or basic relations immediately before a complex lesson so that the learner can process the lesson more effectively.
A prerequisite may be a foundation. Pretraining is often a ramp.
Pretraining Versus Previewing
Previewing a chapter can orient the learner: headings, diagrams, questions, summary.
Pretraining goes further. It aims to make selected components usable before the main explanation begins.
A preview says, “You are going to learn about the heart.” Pretraining says, “Before we trace circulation, locate the atria, ventricles, valves and major vessels and understand the direction each structure controls.”
The first creates expectation. The second reduces component uncertainty.
Pretraining Versus Advance Organisers
An advance organiser gives the learner a high-level conceptual structure before detailed learning. Pretraining may focus more on the parts that will later participate in that structure.
The distinction is useful because complex lessons can need both.
Before teaching plate tectonics, an organiser might say that Earth’s surface is divided into moving plates whose interactions explain several geological phenomena. Pretraining might identify crust, mantle, plate boundaries, convection, subduction and the basic motion types.
One provides the skyline. The other labels the buildings.
Pretraining and Vocabulary
Vocabulary can function as infrastructure for a subject.
In science, words such as evaporation, condensation, diffusion, density and pressure are not decorative terminology. They are compressed handles for recurring processes and relationships.
In mathematics, coefficient, gradient, factor, vector and reciprocal let teachers refer precisely to structures. In English, inference, connotation, register, counterargument and tone allow students to classify language decisions.
Pretraining vocabulary works when the word is attached to a meaningful minimum model. Memorising definitions without examples or distinctions can create labels without structure.
Use the Vocabulary Learning Hub for the subject-specific vocabulary system.
Pretraining for Mathematics
Mathematics lessons often fail because students meet a new procedure while basic representations remain unstable.
Before teaching gradient from coordinate geometry, a learner may need stable ideas of coordinates, horizontal and vertical change, ratio and signed difference. Before completing the square, the learner may need coefficient language, expansion structure and square identities. Before trigonometric modelling, the learner may need the geometry of right triangles and the meaning of ratios.
Good pretraining does not solve the target problem in advance. It stabilises the objects the target problem will manipulate.
That distinction protects the learner from being “prepared” so completely that no new learning remains.
Continue through the Mathematics Learning Hub.
Pretraining for Science
Science is full of systems whose components interact dynamically.
Before students understand an electric circuit, identify current, potential difference, resistance, source, conductor and load. Before a lesson on the digestive system, stabilise the main organs and their basic roles. Before a lesson on inheritance, establish gene, chromosome, DNA, allele and trait.
The target lesson can then ask how the parts interact, what changes, what is conserved and what evidence supports the model.
Pretraining turns the later explanation from a vocabulary storm into a causal problem.
Continue through the Science Learning Hub.
Pretraining for English
English can also overload learners when metalanguage and task demands arrive together.
Before analysing rhetorical effect, students may need clear distinctions among audience, purpose, tone and technique. Before a lesson on argument structure, they may need claim, evidence, reasoning, counterargument and rebuttal. Before a comprehension task on writer’s craft, they may need to distinguish literal meaning, inference and effect.
The lesson then becomes about relationships and judgement rather than decoding the terminology in the question.
Continue through the English Learning Hub.
Pretraining for Writing
A writing task can overwhelm a novice because planning, content generation, vocabulary, sentence construction, organisation, audience control and transcription all compete for attention.
Pretraining can isolate one part before the full composition begins. Teach the structure of a counterargument before asking for a complete argumentative essay. Practise dialogue punctuation before a narrative where dialogue matters. Stabilise a small vocabulary field before a descriptive piece. Rehearse paragraph architecture before expecting whole-essay coherence.
The point is not to reduce writing to parts. It is to make selected parts less expensive so the learner can coordinate the whole.
Pretraining for Examinations
Exam performance contains its own component knowledge: paper structure, command words, time constraints, answer forms, mark allocation and common task types.
A student who meets all of these for the first time during a mock paper is learning the interface while trying to demonstrate subject knowledge.
Pretraining the examination interface can reduce that unnecessary load. Students should know where sections begin, what “state,” “explain,” “justify” or “compare” require, how many marks imply how much response, and what resources are allowed.
Then the mock paper can test subject performance rather than interface discovery.
Use the Examinations & Assessment Hub for the performance layer.
Pretraining for Digital Tools
Software can create the same overload problem as academic content.
Imagine teaching spreadsheet modelling while the learner does not know cell references, formulas, worksheets, ranges or basic navigation. The modelling lesson becomes a software-navigation lesson.
A short pretraining sequence on the interface can release capacity for the actual intellectual task.
This principle generalises to laboratory equipment, graphing calculators, learning platforms, coding environments and professional software. Tool fluency can be a prerequisite for domain fluency when the tool mediates the task.
Pretraining for Procedures
Procedures often contain components, states and transitions.
Before learners execute the full procedure, pretraining can identify instruments, controls, indicators, safety states and sequence markers.
However, where safety is involved, pretraining is not optional enrichment. Essential safety information must be explicitly taught and verified before performance.
The “let them discover it” philosophy ends where preventable harm begins.
A Five-Minute Pretraining Design
Pretraining can be very short.
- Minute 1: show the whole system once without explaining every relationship.
- Minute 2: identify four to six key components.
- Minute 3: give one defining characteristic or state for each component.
- Minute 4: ask learners to identify or retrieve the components from a clean diagram or example.
- Minute 5: state the driving question the full lesson will answer.
Then begin the main instruction.
This small investment can change the processing demands of the next twenty minutes.
A Twenty-Minute Pretraining Design
For a denser topic, use a longer ramp.
- Introduce the component names.
- Show each component in at least two representations.
- Give one defining property and one common confusion.
- Use a quick retrieval check.
- Ask learners to classify examples and non-examples.
- Connect only the simplest relationships.
- Stop before the central system explanation is fully taught.
The final instruction still needs somewhere to go. Pretraining should prepare the bridge, not secretly complete the journey.
Pretraining Should Be Retrieved, Not Merely Seen
If the purpose of pretraining is to make component knowledge available during the main lesson, simply displaying the labels may not be enough.
After introducing the components, remove the labels and ask learners to identify them. Give the function and ask for the name. Show the name and ask for the defining characteristic. Use a quick matching or oral recall check.
This connects pretraining with retrieval practice. The learner should enter the complex lesson carrying the components, not merely having encountered them.
For durable return, Series 0001 explains How Successive Relearning Works.
Pretraining Can Include Generative Work
Pretraining is often presented as teacher-supplied preparation, but learners can also generate part of the component model once enough information is available.
Label a blank diagram. Sort components by function. Draw the part from memory. Explain the difference between two terms. Construct a simple concept map that the main lesson will later expand.
The key is timing. Generation should not recreate the overload that pretraining is trying to prevent.
Series 0002 develops the broader mechanism in How Generative Learning Works.
Pretraining Can Prepare Productive Failure
Productive Failure requires enough prior knowledge for a learner to generate meaningful attempts. Pretraining can create that minimum floor.
Before asking students to explore a complex comparison problem, teach the names and basic properties of the representations they are likely to use. Before asking learners to build a causal model, establish the components whose relationships they must investigate.
This is an important connection: guidance and exploration are not enemies. Good pretraining can make later exploration more productive by shrinking irrelevant search.
Series 0003 develops that sequence in How Productive Failure Works.
The Expertise Reversal Problem
Instructional support that helps novices can become redundant for experts.
A learner who already knows the components may gain little from being forced through a basic pretraining sequence. Worse, redundant instruction can consume attention and create boredom.
Research on prior knowledge in multimedia learning has long emphasised that instructional principles can interact with expertise. A Cambridge Handbook chapter on the prior knowledge principle notes that designs helpful for low-knowledge learners may be less helpful or even interfere for high-knowledge learners.
The practical rule is simple: pretrain what is unstable, not what is already automatic.
How to Diagnose Whether Pretraining Is Needed
Before building a pretraining module, inspect learner state.
- Can the learner name the major components?
- Can the learner identify them in a diagram, example or interface?
- Can the learner state their basic characteristics?
- Can the learner distinguish commonly confused components?
- Can the learner retrieve these facts without continuous prompting?
- Does the learner still lose the main explanation because basic labels consume attention?
If these answers are strong, skip or compress pretraining. If they are weak and the main lesson depends on them, build the ramp.
For larger learning diagnosis, use the Diagnostics & Recovery Hub.
The Front-Loading Failure
Pretraining can be ruined by excess.
A teacher thinks, “Students need background knowledge,” and assigns a massive chapter before the lesson. The chapter contains the very complexity pretraining was meant to manage. The learner now experiences overload earlier instead of less overload overall.
Good pretraining is selective. It should identify the few components that unlock the next model.
The best ramp is not the longest ramp. It is the shortest ramp that safely reaches the platform.
The Vocabulary-Only Failure
Another bad design is a list of terms detached from the system they will enter.
Students memorise “stomata,” “guard cells,” “transpiration” and “diffusion” as definitions. Then the lesson begins and the terms still feel disconnected.
Pretraining should preserve enough functional information that later relationships have hooks.
A guard cell is not just a definition. It changes shape and controls the opening of a stoma. That minimal functional characteristic makes later explanation possible.
The Over-Explaining Failure
At the opposite extreme, pretraining can accidentally become the whole lesson.
The teacher explains every relationship before the main animation, worked example or investigation. When the “real lesson” begins, learners are merely seeing the same explanation again.
Sometimes repetition is useful. But the special value of pretraining comes from distributing processing: components first, system relationships later.
Keep that division visible.
The Static-Only Failure
Learners may recognise a component on one labelled diagram but fail to identify it when the representation changes.
Pretraining should therefore use at least modest variation where the later task requires it. Show the component in a clean diagram and a realistic image. Use the formal term and a plain-language description. Rotate a geometry figure. Change the surface form of a grammar example.
The aim is stable identity across changing appearance.
The “Easy Means Learned” Failure
Pretraining often feels easy because the learner handles isolated components rather than the full system.
Do not confuse that local success with mastery of the whole topic.
A student who can label every organ of the digestive system may still fail to explain digestion. A learner who knows every algebra term may still fail to solve an equation. A student who identifies claim and evidence may still write a weak argument.
Pretraining prepares the learner for complexity. It does not replace complexity.
The Main Lesson Must Use the Pretraining
A common curriculum problem is teaching prerequisite vocabulary on Monday and never explicitly connecting it to Tuesday’s mechanism.
The main lesson should call back to the pretraining.
“Remember the valve can open or close. Watch what happens to pressure when it changes state.” “Yesterday we distinguished numerator and denominator. Today notice which one is fixed.” “You already know what connotation means. Now we are going to see how connotation shifts tone.”
The ramp needs to meet the road.
Pretraining and Segmenting
Pretraining is one way to manage essential processing. Segmenting is another.
Pretraining moves selected work earlier. Segmenting slows the flow by breaking a complex explanation into learner-paced parts.
The two can be combined. First identify the important components. Then present the system in stages so the learner has time to assemble causal relationships.
Research summarised in the Cambridge Handbook of Multimedia Learning has reported support for both principles across multiple experimental tests, while later meta-analytic work also reminds us that effect sizes and boundary conditions vary across media, domains and learner groups.
Pretraining and Worked Examples
Worked examples show how an expert solution unfolds. Pretraining can make the example easier to parse.
Before a worked trigonometry solution, identify opposite, adjacent and hypotenuse relative to the chosen angle. Before a worked argument analysis, identify premise, claim, evidence and qualifier. Before a worked coding example, establish variable, function, parameter and return value.
Then the learner can attend to why the expert chooses each step rather than repeatedly asking what the symbols mean.
See How Worked Examples Work for Performance.
Pretraining and Fading
Once components are familiar, pretraining should shrink.
The first lesson may label every part. The second may show only initials. The third may present an unlabeled diagram. The fourth may ask learners to generate the diagram themselves.
Support that never disappears can become dependency.
This is where pretraining connects to How Fading Works. The aim is not permanent simplification but temporary preparation for independent complexity.
Pretraining and Transfer
A learner may perform beautifully when every component appears exactly as pretraining presented it.
Transfer asks whether the model survives a changed surface.
After learning components, vary the representation. Rotate the diagram. Change the context. Use a new example. Remove labels. Ask which component plays the same functional role in a different system.
If pretraining creates only picture recognition, it has not yet built robust conceptual access.
See Why Transfer Is the Real Proof of Learning.
A Student Protocol: Build the Component Floor
Students can pretrain themselves before difficult chapters.
- Scan the chapter for recurring technical terms.
- Identify the five to ten components that appear most often.
- Learn one accurate sentence about each.
- Find or draw a simple representation.
- Distinguish commonly confused pairs.
- Close the source and retrieve the components.
- Then begin the full explanation.
This is especially useful when a learner repeatedly stops in the middle of explanations to look up basic terms.
A Parent Protocol: Check the Vocabulary of the Task
When a child says, “I do not understand this chapter,” the problem may not be the whole chapter.
Ask the child to point to and explain the key components. Can the learner tell the difference between evaporation and boiling? Between area and perimeter? Between evidence and explanation? Between numerator and denominator?
If the component map is unstable, more full-length practice may be premature.
Repair the floor first, then return to the system.
A Teacher Protocol: Pretrain Only the Bottleneck
Before each complex lesson, ask one question: Which unfamiliar component is most likely to steal attention from the central relationship?
That component is a candidate for pretraining.
Do not teach ten items if two are the real bottleneck. Do not review an entire previous chapter when one representation is missing. Good pretraining has surgical precision.
This keeps the learner moving without turning every lesson into remediation.
A Tutor Protocol: Read the Pause
Experienced tutors learn to notice where a learner’s attention is being consumed.
A student may pause not because the target reasoning is hard but because one term is unstable. The tutor explains the term; the student resumes immediately. That is evidence that the bottleneck was component knowledge.
Another student knows every term but cannot connect them. More pretraining will not help. The weak link is relational.
This is why pretraining should be diagnostic, not automatic.
Pretraining and the First Weak Link
The first weak link in a complex task often appears before the task itself.
A learner cannot interpret a graph because axis language is unstable. Cannot understand a chemistry mechanism because particle terms are confused. Cannot answer a literature question because “effect” and “purpose” are treated as synonyms. Cannot follow a proof because the definition being used is not retrievable.
These failures look large at the final task but may originate in a small missing component.
Pretraining is one way to prevent that small missing component from consuming the entire lesson.
How Much Pretraining Is Enough?
Enough means the learner can recognise, name and state the defining characteristics required to follow the main explanation without continuous repair.
Not necessarily fluency. Not necessarily long-term retention. Not necessarily full application.
The pretraining criterion is readiness for the next learning event.
Later retrieval and successive relearning can convert temporary readiness into durable knowledge. Later practice and transfer can convert component knowledge into flexible performance.
Pretraining is one stage in the route, not the whole route.
What the Research Does and Does Not Say
Research on multimedia learning has repeatedly supported pretraining in contexts where learners face complex, often fast-paced explanations. A 2014 Cambridge Handbook summary reported support in 13 of 16 experimental tests with a median effect size of 0.75 in the reviewed studies. Earlier Mayer work reported strong transfer advantages in several tests.
But these results should not be converted into a universal claim that every lesson should begin with a glossary.
A 2025 meta-analysis of Mayer’s multimedia-learning research found that effects vary across principles, media, domains, age groups and outcome types. Research is strongest when we use it to understand mechanisms and boundary conditions, not when we turn one finding into a classroom ritual.
In 2026, Mayer’s new chapter on the pretraining principle in Teaching with Instructional Video continues to frame pretraining around presenting names and characteristics of key terms before the video, while explicitly considering empirical rationale, boundary conditions and applications.
A Decision Tree for Pretraining
Use this simple route before adding pretraining:
- Is the upcoming material complex and highly interactive? If no, ordinary instruction may be enough.
- Are several key components unfamiliar? If no, skip redundant pretraining.
- Will unfamiliar components compete with understanding relationships? If yes, pretrain.
- Can the component floor be taught briefly? If yes, keep it brief.
- Does the learner need to retrieve those components during the lesson? Include a quick retrieval check.
- Will the main lesson explicitly use the pretraining? If no, redesign the handoff.
- Are students already experts in the components? Fade or remove the pretraining.
Pretraining Across the Micro–Meso–Macro Scale
The Micro–Meso–Macro Learning Control Tower provides another useful view.
Micro: identify the components, terms, symbols or elementary operations.
Meso: connect components into a routine, mechanism, paragraph, procedure or causal chain.
Macro: use the connected model in a larger unfamiliar performance.
Pretraining is strongest at the micro-to-meso gate. It stabilises enough micro structure that meso relationships can be built without overload.
It should not freeze learning at the micro level. The whole point is to cross the gate.
Pretraining as Interface Design
There is a broader principle here.
Every difficult domain has an interface: symbols, labels, controls, conventions and recurring representations. Experts stop noticing the interface because it has become transparent. Novices experience it as part of the problem.
Pretraining teaches enough of the interface that the learner can look through it at the underlying idea.
In mathematics, notation becomes transparent. In science, technical vocabulary becomes transparent. In English, task language becomes transparent. In software, controls become transparent.
When the interface disappears from conscious struggle, attention can move to the model.
The Deep Principle: Complexity Should Arrive in Layers
Human learning often fails not because the final idea is impossible, but because too many layers are introduced at once.
We ask a learner to decode the words, identify the parts, understand the relationships, remember the sequence, infer the mechanism and solve the transfer problem in a single pass.
Experts can sometimes do this because earlier layers have been compressed through years of learning.
Novices need architecture.
Pretraining is one architectural move: separate identification from integration. Make the parts stable enough that the whole can be assembled.
Then, once the whole is understood, return to complexity. Remove labels. Vary the representation. Ask the learner to generate. Retrieve later. Transfer to a new case.
The ramp eventually disappears. The learner keeps the building.
Use This Tomorrow
Before your next difficult chapter or video, identify the five to ten components that the explanation assumes you already recognise. Learn their names, basic characteristics and common confusions. Retrieve them once without looking. Then start the full lesson and notice whether the relationships become easier to follow.
If they do, you have not made the subject easier. You have changed the order in which the difficulty arrives.
Research and Further Reading
- Mayer — Pre-training Principle, Multimedia Learning
- Mayer & Pilegard — Segmenting, Pre-training and Modality Principles
- Mayer — Pretraining Principle, Teaching with Instructional Video (2026)
- Cromley & Chen — Meta-analysis of Mayer’s multimedia-learning research
- Study & Learning Methods Hub
eduKateSG Learning Node Series · 0004 of the continuing series. Previous: 0003 — How Productive Failure Works. Continue through the Study & Learning Methods Hub and the wider eduKateSG Learning Hubs.