VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

How Modality Works in Learning | Share the Load Between Eyes and Ears

eduKateSG Learning Node Series · 0009

A learner can run out of attention before the lesson runs out of information.

Picture a science animation showing blood moving through the chambers of the heart. The diagram is changing every second. Arrows appear. Valves open and close. Labels point to structures. At the same time, a paragraph of explanation sits beneath the animation.

The learner has two jobs for the eyes: inspect the changing visual and read the explanation. Both may be important. Both are competing for the same limited visual attention.

Now change only one thing. Keep the meaningful animation, but speak the short explanation aloud instead of placing every sentence on the screen. Under appropriate conditions, the learner can distribute processing across visual and auditory channels rather than forcing two streams through the same visual bottleneck.

That is the intuition behind the modality principle in multimedia learning. But the useful version of the principle is not “audio is better than text.” The useful version asks a harder question: which information should arrive through which channel, for this learner, in this task, at this speed?

Quick Read: What the Modality Principle Actually Says

Research associated with cognitive load theory and multimedia learning has repeatedly examined situations in which learners must integrate graphics with verbal explanation. The 2021 Cambridge Handbook of Multimedia Learning includes a dedicated chapter on the modality principle, while earlier handbook syntheses describe conditions under which spoken words paired with graphics can outperform printed words paired with the same graphics.

The proposed mechanism is straightforward. Working memory is limited. When both the explanatory words and the visual representation require intensive visual processing, they compete. Moving appropriate verbal information into speech can reduce that competition and make more capacity available for integrating the explanation with the picture.

The point is not to maximise audio. The point is to stop one channel from becoming the bottleneck when another usable channel is available.

The Whiteboard Problem

A teacher draws a graph while explaining how gradient changes. If the teacher writes a full paragraph on the board and asks students to read it while simultaneously tracing the changing line, the eyes must keep switching.

Look at the sentence. Find the graph. Read the next phrase. Find the corresponding point. Return to the sentence. The learner may understand each element individually but lose the relation between them because attention is consumed by coordination.

If the teacher instead points to the graph while saying, “Notice that the vertical change is getting larger for the same horizontal change,” the visual channel can stay on the graph while the verbal information arrives through hearing.

The mathematics did not become easier. The interface became less expensive.

Why This Is a Working-Memory Problem

Complex learning requires temporary coordination. The learner may need to hold a label, a relationship, a changing diagram and the current goal long enough to bind them into a coherent model.

Working memory does not behave like an unlimited warehouse. When a task contains many unfamiliar interacting elements, capacity can become the limiting constraint. This is why learners sometimes understand every sentence when read separately yet fail to understand the whole explanation.

Modality is one design lever. Instead of reducing the amount of essential information, it changes how some of that information is delivered.

For the broader mechanism, see How Cognitive Load Works | When Working Memory Becomes the Bottleneck.

Eyes and Ears Are Not Two Independent Hard Drives

The popular simplification says: “Use both channels and you double capacity.” Human cognition is not that tidy.

Visual and auditory processing interact. Learners must still integrate the information. Speech disappears through time. Visuals can often be inspected again. Language proficiency changes how expensive listening is. Environmental noise changes the reliability of audio. Hearing differences matter. Technical vocabulary may be easier to inspect in print. Learner control matters. Prior knowledge changes everything.

So the modality principle is conditional. Its value is strongest when spoken words can reduce visual competition without creating a new auditory or language-processing bottleneck.

When Spoken Explanation Helps Most

Modality is especially useful when the visual itself deserves continuous visual attention.

  • a dynamic scientific process;
  • a labelled anatomical diagram;
  • a geometry construction;
  • a map whose spatial relationships matter;
  • a graph changing over time;
  • a physical demonstration;
  • a worked problem where each spoken sentence points to a specific visual step;
  • or an animation where reading a separate text would force repeated visual switching.

In these cases, concise narration can leave the eyes free to inspect the representation that carries the spatial or dynamic structure.

When Printed Text May Be Better

Speech has one serious limitation: it vanishes.

If the information is dense, symbolic, unfamiliar or exact, a learner may benefit from text that remains available for inspection. A long algebraic condition, a legal definition, a technical formula, a list of constraints or a sentence containing unfamiliar vocabulary can be easier to revisit in print than in fleeting narration.

Printed text can also support self-pacing. The learner decides when to stop, reread, annotate or compare.

The question is therefore not “speech or text?” It is “which representation gives the learner the best control over the information this task requires?”

The Second-Language Boundary

Listening is not equally cheap for every learner.

A student learning through a second or additional language may decode written academic language more reliably than rapid speech. An accent, unfamiliar pronunciation, reduced forms or fast delivery can increase processing cost. In that case, moving text into narration can remove one bottleneck and create another.

Good design therefore considers language proficiency. Slower narration, replay controls, transcripts, key-term labels and pretraining can all matter.

This is one place where How Pretraining Works connects directly with modality: if key technical terms are already known, spoken explanation becomes easier to process.

Accessibility Changes the Design

An instructional principle is not a reason to remove access.

Learners who are deaf or hard of hearing may require captions or transcripts. Learners with auditory-processing difficulties may need persistent text. Learners studying in noisy spaces may not have reliable audio. Others may use screen readers or assistive technologies.

The instructional goal is not to obey a purity rule about channels. It is to make the essential structure available with the lowest unnecessary processing cost for the actual learner.

This means captions can be educationally necessary even where a narrow laboratory version of the redundancy principle might predict a cost from duplicated words. Accessibility requirements are not “extraneous.” They are part of the learner’s route to the content.

Modality and the Redundancy Trap

A common slide contains a diagram, a full paragraph of text, and a narrator reading that paragraph word for word.

It feels generous: the learner can see it and hear it. But the design may force the learner to process the same verbal stream twice while also trying to inspect the diagram.

This is where modality and redundancy interact. Sometimes narration plus a meaningful graphic is cleaner than narration plus the same full printed narration plus the graphic. Series 0011 examines that problem separately in How Redundancy Works in Learning.

The key is not “never duplicate.” Short labels, formulas, unfamiliar names and accessibility text may still be useful. The problem is unnecessary duplication that competes for the same attention needed to understand the model.

Modality and Signaling

Speech becomes more useful when the learner knows where to look.

If narration says, “The pressure rises here,” but the screen contains twelve unlabeled components, the learner may spend the crucial second searching for “here.”

Visual signaling can solve the reference problem: highlight the relevant line, point to the component, change emphasis, or reveal the label at the moment it matters.

The two principles therefore work together. Modality can distribute processing; signaling can direct processing.

See How Signaling Works in Learning.

Modality and Segmenting

Even excellent narration becomes difficult if it moves faster than the learner can integrate it.

A learner-paced video, animation or demonstration allows processing to stop at natural boundaries. The learner can replay the sentence, inspect the diagram and continue when the relationship is understood.

That is why modality should not be separated from pacing. A badly paced spoken explanation can disappear before it has been integrated.

See How Segmenting Works.

Modality in Mathematics

Mathematics alternates between visual-symbolic information and verbal explanation.

A teacher can point to an equation while saying why a transformation preserves equality. The learner sees the symbolic change while hearing the reason. This can be cleaner than forcing the learner to read a paragraph beside the equation while the relevant line changes.

But mathematics also contains information that should remain visible: formulas, assumptions, definitions, diagrams and multi-step expressions. The strongest design often uses speech for transient explanation and persistent text for the symbolic objects that must be inspected repeatedly.

The medium follows the job.

Modality in Science

Science often asks learners to understand dynamic causal systems. These are natural candidates for well-timed narration because the visual representation carries motion, position, direction or transformation.

During an animation of convection, spoken explanation can direct attention to rising warm material and sinking cooler material without making the learner repeatedly leave the animation to read a separate paragraph.

But technical vocabulary should not disappear. Key terms can remain as concise labels while the mechanism is narrated. The goal is not an empty screen with a voice. It is a coherent division of labour between channels.

Modality in English and Language Learning

Language learning complicates modality because the sound of language may itself be the object of learning.

If students are learning pronunciation, rhythm, stress or listening comprehension, audio is not merely a delivery channel. It is part of the content.

At other times, written form is essential. Spelling, punctuation, syntax and visual word recognition require persistent text.

A strong language lesson therefore switches modality deliberately. Hear the sentence. Read it. Compare the forms. Hide one representation and retrieve the other. Use sound when sound carries the learning target and print when print carries the target.

Modality in Note-Taking

Students often copy what a teacher says while also trying to understand the explanation. This creates a different channel conflict.

Writing can support encoding, but verbatim transcription may consume the attention needed to follow the causal structure. If every word feels important, the learner becomes a stenographer.

One solution is to provide stable reference material for exact details while the live explanation focuses on relationships. Students can then take selective notes rather than trying to preserve the entire speech stream.

The best note is not the one that captures the most words. It is the one that supports later reconstruction.

The Replay Button Changes the Principle

Many classic multimedia experiments were conducted with tightly controlled presentations. Modern learners often have pause, rewind, playback speed, captions and transcripts.

Learner control can reduce the cost of transient speech. A difficult sentence can be replayed. A fast animation can be paused. A caption can be turned on only when needed.

This means digital design should not ask only which modality to use. It should ask what control the learner has over that modality.

Control converts a disappearing stream into something closer to an inspectable resource.

The Expertise Boundary

Novices and experts do not experience the same display.

An expert sees a familiar graph and immediately chunks several features into one structure. A novice may process axes, labels, scale, units and trend separately. A spoken explanation that helps the novice coordinate these elements may feel redundant to the expert.

As expertise grows, support should change. Narration may become shorter. Labels can disappear. Learners can be asked to generate the explanation themselves.

This is why instructional design should be state-sensitive rather than permanently attached to one format.

The Attention Test

A useful diagnostic question is: where do the learner’s eyes need to be while this sentence matters?

If the answer is “on the moving diagram,” consider speaking the explanation. If the answer is “on the exact wording,” keep the text. If the answer is “switching between two distant visual sources,” repair the layout. If the answer is “nowhere in particular,” modality may not be the main issue.

This small question turns a broad learning principle into an observable design decision.

A Teacher Protocol

  • Identify the essential visual representation.
  • Mark which information must remain visible for inspection.
  • Move short explanatory relations into speech when reading would compete with looking.
  • Keep technical labels, equations or unfamiliar names visible when persistence matters.
  • Signal the visual element being discussed.
  • Pause at conceptual boundaries.
  • Provide captions, transcripts or alternate access where required.
  • Check whether learners can reconstruct the model after the explanation ends.

A Student Protocol

Students can manage modality even when the lesson design is fixed.

  • If a video is visually dense, stop reading unrelated notes while the animation is moving.
  • Pause after an important explanation and restate it.
  • Use captions when audio is unclear, but do not feel obliged to stare at them if they pull attention away from the visual structure.
  • Write down technical terms that need a persistent form.
  • After the explanation, draw or explain the model without the video.
  • Replay only the segment containing the broken link instead of restarting the entire lesson.

The Common Failure: Turning Every Slide Into a Script

Presentation software makes it easy to place every sentence on the screen. Teachers then read the sentences aloud.

The slide becomes a teleprompter for the teacher and a competing text source for the learner.

A better slide often contains the representation, a few essential labels and enough structure to preserve orientation. The spoken explanation provides the temporary verbal layer. Detailed reference material can exist elsewhere for later study.

The live presentation and the revision document do not need to be the same artefact.

What Recent Synthesis Adds

Multimedia-learning principles have accumulated a large research base, but modern synthesis also emphasises variation across contexts. A 2025 meta-analysis of Richard Mayer’s multimedia-learning research examined effects across principles, media, domains and learner characteristics, reinforcing an important lesson: design principles are useful as evidence-informed defaults, not context-free laws.

That matters for modality. The classic effect is strongest under specific processing conditions. It should be tested against accessibility, language proficiency, pacing, prior knowledge and the actual informational role of the visual.

The Deeper Principle: Give Each Channel a Job

The most useful way to think about modality is not “use multimedia.” It is “assign information to channels deliberately.”

The eyes may inspect structure. The ears may carry transient explanation. Text may preserve exact definitions. Labels may anchor reference. Gesture may signal location. The learner’s own speech may become retrieval.

When every channel carries the same thing, the system can become noisy. When every channel carries unrelated things, the system can become distracting. Good design gives each representation a reason to exist.

Learning improves not because more media were added, but because less attention was wasted on the interface.

Use This Tomorrow

Take one explanation you are learning from. Ask where your eyes need to be while the important relationship is being explained. If reading the explanation pulls your eyes away from the thing you need to inspect, try a short spoken explanation while keeping the visual in view. Then pause and reconstruct the relationship in your own words.

Do not add audio because audio is fashionable. Use it when it removes a genuine processing conflict.

Research and Further Reading


eduKateSG Learning Node Series · 0009 of the continuing series. Earlier nodes include 0005 — How Segmenting Works and 0006 — How Signaling Works in Learning.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading