VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Multimedia Learning Works | Build One Mental Model From Words and Pictures

eduKateSG Learning Node Series · 0014

Sometimes a picture does not make a lesson prettier. It makes the idea possible to see.

Imagine learning how a bicycle derailleur moves the chain across gears using only a paragraph. The paragraph may be accurate. But the learner must convert nouns, positions and movements into a spatial mechanism internally.

Now add a useful diagram or video showing the chain, sprockets, derailleur cage and direction of movement while the explanation names the relationships.

The value of the visual is not decoration. It carries structure that language would otherwise need many words to reconstruct.

The multimedia principle asks how words and pictures can work together to help a learner build a coherent mental model. The deeper idea is not “more media is better.” It is: use complementary representations when each carries something the learner needs to understand the same system.

Quick Read: The Multimedia Principle

Richard Mayer’s cognitive theory of multimedia learning has long argued that people can learn more deeply from well-designed combinations of words and pictures than from words alone in many instructional contexts. His 2026 book Teaching with Instructional Video includes a dedicated Multimedia Principle chapter focused on the learning value of video plus narration relative to narration alone, while the 2021 Cambridge Handbook of Multimedia Learning places the principle inside a broader evidence-based architecture of learning from multiple representations.

The principle is often misremembered as “add pictures.” That is too shallow. A visual helps when it contributes task-relevant structure and when the learner can integrate it with the verbal explanation.

Words can name a relationship. Pictures can expose its geometry. Learning happens when the learner binds both into one usable model.

The Map-and-Directions Problem

Suppose someone tells you how to walk from a station to a museum.

“Exit on the north side, cross the river, take the second left after the library, continue past the park and turn right at the triangular plaza.”

The description can work. But it asks language to carry a spatial route.

Now place the same instructions on a map. The map shows relative position, branching, distance and landmarks. The words can focus on decisions.

The map and the directions do not need to duplicate each other. Their value comes from complementarity.

Two Representations, One Meaning

A learner can represent the same idea in different forms.

A function can be an equation, graph, table or verbal rule. A historical event can be represented as narrative, timeline, causal map or geographic map. A biological system can be described in prose or shown as a labelled diagram. A sentence can be read normally or parsed as a structure.

Multiple representations become educationally powerful when the learner understands what stays invariant across them.

The equation y = 2x + 3 and its graph are not two unrelated facts. They are two views of the same relationship. The graph makes slope and intercept visible. The equation makes exact symbolic manipulation easy.

Multimedia learning should help the learner cross that bridge.

Why Pictures Can Reduce Verbal Reconstruction

Some relationships are expensive to describe entirely in words.

Try describing a complex circuit, a geometry figure or the movement of tectonic plates without showing it. The listener must build spatial structure from sequential language.

A well-designed visual externalises part of that structure. The learner can inspect relative position instead of continuously reconstructing it.

This does not eliminate thinking. It relocates thinking toward the relationship that matters.

Why Pictures Can Also Make Learning Worse

A visual can be irrelevant, ambiguous, overloaded or misleading.

A decorative photograph of a scientist may add interest but contribute nothing to the mechanism. A colourful infographic may compress so aggressively that causal relationships disappear. A 3D animation may look realistic while hiding the variable the learner needs to compare.

More external representations also create coordination costs. A 2024 meta-analysis in Educational Psychology Review examined the value of using more than two external representations in STEM education and reinforces the importance of considering how multiple representations are coordinated rather than assuming that more views automatically produce better understanding.

The useful question is not “Can I add a picture?” It is “What important relationship becomes easier to understand because this representation exists?”

Multimedia Learning Is Not Learning Styles

The multimedia principle is sometimes confused with the claim that individual learners have fixed visual, auditory or kinesthetic styles that instruction should match.

These are different ideas.

Multimedia design asks what representations fit the structure of the material and the learner’s processing task. A diagram may help everyone understand a spatial mechanism because the mechanism itself has spatial structure, not because some students belong to a “visual type.”

The existing eduKateSG owner How Dual Coding Works | Words, Pictures and the Limits of Two Channels covers the wider representation question and explicitly preserves the boundary against simplistic learning-styles claims.

The Integration Problem

Words and pictures do not integrate themselves.

A student can look at the diagram and read the paragraph without knowing which sentence belongs to which visual feature. The learner may memorise two representations separately.

This is why the surrounding principles matter.

  • Spatial contiguity keeps corresponding information physically close.
  • Temporal contiguity keeps corresponding information aligned in time.
  • Signaling directs attention to the relationship being explained.
  • Coherence removes irrelevant competition.
  • Redundancy warns that repeating the same message in unnecessary forms can overload rather than help.

The multimedia principle is therefore not a standalone decorating instruction. It lives inside an integration system.

The Representation-Selection Problem

Different ideas deserve different visuals.

Use a timeline for sequence through time. Use a map for spatial relationships. Use a graph for quantitative change. Use a flowchart for process and branching. Use a diagram for parts and relationships. Use a photograph when realistic appearance is important. Use animation when change through time is itself the concept.

The visual format should be selected from the structure of the problem.

A pie chart does not become educational merely because it is colourful. A map does not help explain an argument whose structure is logical rather than geographic.

When Animation Earns Its Cost

Animation is powerful when the movement or transformation is difficult to infer from static states.

Show how an engine cycle changes, how a geometric transformation maps one shape to another, how a wave propagates or how components in a system interact through time.

But animation disappears. Static diagrams remain available for inspection.

That means a strong lesson may combine both: animation for dynamic relationship, still image for later inspection, and words for naming and explanation.

The Pace Problem

Video creates a pace even when the learner can pause it.

A 2025 meta-analysis on accelerated video learning in Frontiers in Psychology examined how playback speed affects learning outcomes and reinforces an obvious but often neglected point: changing speed changes the learner’s processing conditions.

Faster is not simply more efficient. The relevant question is whether the learner can still select, organise and integrate the information at that pace.

This links the multimedia principle to segmenting and learner control.

Multimedia in Mathematics

Mathematics is already a multimedia subject. Symbols, diagrams, graphs, tables and language coexist constantly.

Weak instruction sometimes treats these representations as separate chapters. Strong instruction shows that they are translations.

For a linear function, move between equation, graph, table and verbal interpretation. For geometry, connect the drawing with symbolic relationships and proof language. For statistics, connect a dataset with plots and numerical summaries.

The learner should not only know each form. The learner should know how information travels between them.

Continue through the Mathematics Learning Hub.

Multimedia in Science

Science constantly crosses scales that ordinary perception cannot access.

Words describe molecules, fields, energy transfers and geological time. Visuals make hidden structures inspectable. Graphs reveal patterns that prose would obscure. Models make causal mechanisms manipulable.

But scientific visuals also carry assumptions. A particle diagram is not a photograph. A colour gradient may encode a variable rather than literal colour. A simplified cell diagram may exaggerate proportions.

Good multimedia learning therefore teaches both the phenomenon and how to read the representation.

Continue through the Science Learning Hub.

Multimedia in English

English education also uses multiple representations, even when the page appears mostly verbal.

A paragraph can be represented as prose, a claim-evidence-reasoning map, a sentence-flow diagram or a rhetorical structure. A narrative can become a timeline. A comprehension passage can be annotated to show pronoun reference, lexical chains or shifts in viewpoint.

The visual should expose language structure rather than replace reading.

Continue through the English Learning Hub.

Multimedia in Vocabulary

Concrete vocabulary can benefit from pictures because images make referents visible. Abstract vocabulary needs more care.

A picture can show an apple. It cannot directly show justice, irony, obligation or ambiguity. For abstract words, diagrams, examples, contrasts and contextual scenes may be more useful than literal illustration.

This is why multimedia vocabulary design should begin with meaning structure rather than a rule that every word deserves an image.

Use the Vocabulary Learning Hub for the subject-specific system.

Multimedia and Worked Examples

A worked example can become clearer when the reasoning is attached to the representation where the decision occurs.

In mathematics, show the equation and annotate the transformation. In science, pair the causal statement with the mechanism diagram. In writing, display the paragraph and map its argumentative structure.

Here multimedia is not additional content. It is a way of exposing expert structure.

Multimedia and Generative Learning

Series 0002, How Generative Learning Works, asks the learner to create representations rather than merely receive them.

That is the natural next step.

First, use words and pictures to help build the model. Then remove part of the support and ask the learner to reconstruct the diagram, explain the graph, redraw the process or translate the verbal rule into another representation.

Receiving multiple representations can support understanding. Generating multiple representations tests ownership.

Multimedia and Retrieval

Once the model is understood, a visual can become a retrieval cue.

Show the unlabeled diagram and ask for the mechanism. Show the graph and ask for the equation. Show the equation and sketch the graph from memory. Show one stage of a process and ask what happens next.

The representation now changes role—from explanation support to retrieval challenge.

Multimedia and Accessibility

Instructional media must also remain accessible.

A learner who cannot hear the narration needs captions or transcripts. A learner who cannot see the diagram needs meaningful verbal description. Colour should not be the only carrier of a distinction. Controls must be usable with assistive technologies.

Accessibility can alter the ideal arrangement of channels, but exclusion is not an acceptable “cognitive optimisation.” The design problem includes all learners who need access.

Multimedia and AI-Generated Content

Generative AI can produce diagrams, videos, voices, animations and interactive explanations quickly. That lowers production cost. It does not automatically improve instructional design.

A generated visual can be factually wrong. It can imply a relationship that does not exist. It can include decorative complexity. It can mislabel a component or use inconsistent scale.

Educational multimedia still needs evidence, verification and an explicit learning purpose.

The question is not “Can AI make this visual?” It is “What learning job will this visual perform, and how will we verify that it represents the domain correctly?”

The Search-Cost Test

One practical way to audit multimedia is to measure search.

How often must the learner ask:

  • Which part is the narrator talking about?
  • Where is that label?
  • Which graph line matches this sentence?
  • What did the animation show a moment ago?
  • Why is this picture here?

Every unnecessary search consumes attention. Good multimedia reduces search for correspondences while preserving search for meaning.

The Translation Test

Ask the learner to translate between representations.

Can the student describe the graph in words? Draw the process from the paragraph? Write the equation represented by the table? Explain what each arrow means in the diagram? Turn the timeline into a causal narrative?

If the learner cannot translate, the representations may be stored side by side without being integrated.

The Removal Test

After understanding is established, remove one representation.

Hide the labels. Remove the graph. Close the paragraph. Ask the learner to reconstruct the missing view.

This reveals whether multimedia built a deeper model or merely made the lesson easier to follow while all supports were present.

The First Weak Link Test

When a learner fails with multimedia material, diagnose the integration layer.

Can the student understand the words alone? Can the student read the visual alone? Can the student map one onto the other? Does the learner know what the visual conventions mean? Is the display too crowded? Is timing misaligned?

A weak final answer may come from missing subject knowledge, weak representation literacy or failed integration. Those are different repair jobs.

Use the Diagnostics & Recovery Hub when the source is unclear.

A Teacher Multimedia Protocol

  • Start with the learning relationship, not the medium.
  • Choose a visual representation that exposes something words alone would make expensive to reconstruct.
  • Keep words and visuals spatially and temporally aligned.
  • Signal the correspondence when the display is complex.
  • Remove decorative material that competes with the model.
  • Avoid redundant text when narration and visuals already carry the same message, unless accessibility or task conditions require it.
  • Pause or segment when the learner needs integration time.
  • Ask learners to translate across representations.
  • Later remove supports and require reconstruction.

A Student Multimedia Protocol

  • Do not collect diagrams you cannot explain.
  • Write one sentence beside each important visual relationship.
  • Turn paragraphs into diagrams only when the diagram clarifies structure.
  • Turn diagrams back into words to test integration.
  • Pause videos when the visual changes faster than you can explain it.
  • Redraw important diagrams from memory.
  • Use multiple representations to compare, not to decorate.

What the 2026 Instructional-Video Work Adds

Instructional video is now a normal part of school, tutoring, professional learning and online education. Mayer’s 2026 book is useful because it reframes classic multimedia principles for this environment.

Video brings words, visuals, timing, narration, captions, instructor presence and pacing into one moving system. That means multimedia design cannot be separated from modality, contiguity, synchrony, signaling, voice and embodiment.

The modern learning object is not simply a page with a picture. It is often a timed interface.

The Deep Principle: Representation Should Expose Structure

The best reason to use multimedia is not engagement.

It is representational power.

Words are excellent for naming, sequencing, qualifying and arguing. Pictures are excellent for showing spatial relations, shapes, patterns and states. Graphs compress quantitative change. Animation exposes transformation through time. Audio carries rhythm, pronunciation and temporal information.

Each medium has affordances.

World-class instructional design chooses the representation whose affordance matches the structure of the learning problem, then helps the learner integrate the views into one model.

The learner should leave with more than several media files. The learner should leave with one connected understanding.

Use This Tomorrow

Take one difficult idea you currently study only in words. Ask whether a diagram, graph, timeline, map or animation could expose a relationship that the prose makes hard to hold. Add only that representation. Then explain how the words and visual describe the same model. Finally, remove one and reconstruct it from the other.

Research and Further Reading


eduKateSG Learning Node Series · 0014 of the continuing series. Previous: 0013 — How Temporal Contiguity Works. Continue through the Study & Learning Methods Hub.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading