eduKateSG Learning Node Series · 0065
How Implementation Fidelity Works | Preserve the Active Ingredients Without Turning Teaching Into a Script
A school finds an intervention with promising evidence. The staff attend training. Slides are shared. Materials are printed. A timetable is adjusted. Everyone agrees that the new approach will begin on Monday.
Three months later, the programme has the same name but no longer has the same mechanism.
One teacher shortened the sessions because the class was busy. Another removed the difficult questioning sequence because students disliked it. A third kept the worksheet but stopped the feedback loop. A fourth improved the examples but changed the frequency. A fifth followed every line exactly even when the learner clearly needed a different surface example.
Everyone says the intervention was implemented.
But was it?
Implementation fidelity is the discipline of keeping the parts that make an intervention work while allowing enough intelligent adaptation for the intervention to survive contact with real classrooms.
The 50-Second Read
- Implementation fidelity asks whether an intervention was delivered in a way that preserves its intended mechanism, not merely whether its name appeared on a timetable.
- Fidelity is multidimensional. Content, frequency, duration, coverage, delivery quality and participant responsiveness can all matter.
- A programme can look faithful on paper while losing the active ingredient that produced the effect.
- Literal replication is not always the goal. Good implementation distinguishes core components from surface features that can be adapted intelligently.
- Too little fidelity can turn a tested intervention into a different intervention. Too much rigidity can make implementation brittle and context-blind.
- Monitoring should be light enough to sustain, precise enough to detect drift and connected to learner outcomes rather than becoming compliance theatre.
- When outcomes disappoint, fidelity data helps distinguish at least two possibilities: the intervention was delivered poorly, or it was delivered well and still did not work sufficiently in this context.
- The practical question is not “Did we do the programme?” It is “Did we preserve the mechanism, deliver enough of it, and verify that learners actually received what the design intended?”
Canonical Owner Boundary
This page owns the question of implementation fidelity: whether a teaching approach, intervention or training design is delivered closely enough to its intended mechanism that the original theory of change still has a fair chance to operate. How the Enacted Curriculum Works owns the gap between intended and actually taught curriculum. How Instructional Routines Work owns repeatable classroom routines. How Instructional Dosage Works owns the amount and distribution of exposure. How Formative Assessment Works owns the use of evidence while learning can still change. Implementation fidelity asks a different systems question: did the intervention that reached the learner still contain the parts that were supposed to make it work?
1. A Programme Name Is Not a Mechanism
Schools often speak in labels: retrieval practice, explicit instruction, tutoring, formative assessment, mastery learning, peer tutoring, reading intervention.
Labels are useful for communication. They are dangerous for evaluation.
Two classrooms can both say they use retrieval practice while one asks every learner to retrieve from memory before feedback and the other lets students copy answers while a teacher calls the activity a quiz. Two tutoring programmes can both run three sessions a week while one preserves small-group diagnosis and rapid correction and the other becomes worksheet supervision.
If the mechanism changes, the evidence attached to the original label cannot simply be carried over unchanged.
2. Carroll’s Framework: Fidelity Is More Than Adherence
Christopher Carroll and colleagues proposed an influential implementation-fidelity framework in 2007. Their model treats fidelity as more than simple adherence to instructions. It includes dimensions such as content, frequency, duration and coverage, while also recognising moderators including intervention complexity, facilitation strategies, quality of delivery and participant responsiveness.
That distinction matters in education because the learner does not receive a checklist. The learner receives an enacted experience.
Source: Carroll et al., A Conceptual Framework for Implementation Fidelity.
3. Fidelity Has Several Dimensions
- Content: were the intended instructional components actually used?
- Frequency: did the intervention occur often enough?
- Duration: did each encounter last long enough for the mechanism to operate?
- Coverage: did learners receive the relevant parts, or only a convenient subset?
- Quality: was the intended practice delivered competently?
- Responsiveness: did learners participate in the way the intervention depends on?
A programme can be strong on one dimension and weak on another. Perfect attendance does not rescue poor delivery. Excellent delivery cannot compensate indefinitely for half the intended sessions. Complete coverage can still fail if learners never actually do the thinking the design requires.
4. The Active Ingredient Is the Part You Cannot Casually Remove
An active ingredient is the feature that is tightly linked to the theory of how the intervention changes learning.
For retrieval practice, an active ingredient is the attempt to retrieve before simply restudying. For worked examples, the relevant mechanism includes studying a sufficiently explicit solution before responsibility is transferred. For formative assessment, evidence must actually change teaching or learning action; collecting an exit ticket and filing it away preserves the object but loses the mechanism.
This is why fidelity cannot be defined only by visible materials.
5. EEF: Be Tight on Core Components and Intelligent About Adaptation
The Education Endowment Foundation’s third edition of A School’s Guide to Implementation, published in April 2024, places strong emphasis on specifying core components before implementation. The guidance argues that schools need to know which elements must remain consistent and where intelligent adaptation is acceptable.
This creates a useful distinction: fidelity does not require every teacher to sound identical. It requires the intervention to remain faithful to what matters.
Source: Education Endowment Foundation, A School’s Guide to Implementation.
6. Surface Adaptation Can Preserve Deep Fidelity
A Primary Mathematics teacher and a Secondary English teacher should not be expected to use the same examples, vocabulary or classroom timing simply because both are implementing a common principle.
They may preserve the same deep structure while adapting the surface.
For example, a common instructional routine might require: activate prerequisite knowledge, model one difficult decision, obtain a response from every learner, diagnose the misconception, provide immediate correction and retest independently. The Mathematics teacher can use fractions. The English teacher can use inference. The core sequence remains recognisable while the content changes completely.
7. Literal Fidelity Can Become Bad Teaching
Suppose a programme says to spend twelve minutes on guided practice. After seven minutes, every learner demonstrates secure independent performance. A teacher who continues mechanically for five more minutes may preserve the schedule while wasting learning time.
Or suppose a scripted example assumes vocabulary the class does not understand. Repeating the script more faithfully can reduce comprehension.
Fidelity is not obedience to arbitrary surface details. It is disciplined preservation of the causal logic of the intervention.
8. Adaptation Can Also Quietly Destroy the Programme
The opposite failure is common because teachers are practical problem-solvers. They remove awkward components first.
Unfortunately, the awkward component may be the learning mechanism.
- The difficult independent attempt is removed because students complain.
- Feedback is delayed because marking takes time.
- Spacing becomes massed because timetabling is easier.
- Small groups become large groups because staffing is expensive.
- Diagnostic questioning becomes volunteer questioning because it feels faster.
- Retesting disappears because the correction looked successful.
The intervention becomes easier to deliver and less able to do its intended work.
9. Fidelity Drift Is Usually Gradual
Programmes rarely collapse in one dramatic moment. They drift.
A five-minute shortcut becomes normal. A skipped check becomes routine. A new teacher inherits the materials without the rationale. A term later, the visible shell remains but the underlying sequence has changed.
That is why implementation needs memory. The system must preserve not only what to do but why the component exists.
10. Training Without Follow-Through Is a Weak Fidelity System
A launch workshop can create shared vocabulary. It cannot guarantee months of competent enactment.
Teachers need opportunities to rehearse, compare interpretations, see examples, receive feedback and resolve edge cases after implementation begins. Otherwise the intervention slowly becomes whatever each person remembers from the launch.
Implementation support is therefore not separate from fidelity. It is one of the mechanisms that protects fidelity.
11. Dosage Fidelity Is Necessary but Not Sufficient
It is tempting to monitor what is easiest to count: number of sessions, minutes delivered, worksheets completed.
Those measures matter. They answer whether enough opportunity existed.
But a programme can hit its dosage target while missing its instructional target. Fifteen sessions of weak diagnosis are not equivalent to fifteen sessions of the intended feedback loop.
Quantity must therefore be paired with a measure of mechanism.
12. Quality of Delivery Changes the Meaning of Adherence
Two teachers can perform the same sequence with radically different educational effects.
Both ask the diagnostic question. One waits, samples every learner, interprets the responses and changes the next example. The other asks, accepts one volunteer answer and moves on.
The observed checklist may show the same box ticked. The quality of enactment is different.
13. Participant Responsiveness Is Part of the System
An intervention may depend on learners explaining, attempting, retrieving, discussing or practising independently.
If learners are physically present but do not engage in the intended cognitive activity, the intervention has not fully reached them.
This does not mean blaming learners for implementation failure. It means designing implementation so that the required participation is feasible, understandable and supported.
14. Measuring Fidelity Can Become a Distortion
The moment a system measures fidelity, staff may optimise for the measure.
If the measure is “three quizzes per week,” teachers can deliver three low-quality quizzes. If the measure is “students discussed the answer,” a brief superficial exchange can satisfy the record.
Fidelity measures should therefore sample the mechanism, not merely count visible artefacts.
15. Reading-Intervention Research Shows How Messy Fidelity Measurement Can Be
A 2023 systematic review by van Dijk, Lane and Gage examined reading-intervention studies in pre-K–12 settings that included implementation-fidelity measures in their analyses. Across fifty studies, dosage, adherence and quality were common, but the review found substantial variation in how fidelity was conceptualised and measured, and varied estimates of its relationship with student outcomes.
The lesson is important: fidelity is not one universally standardised number. The measure must fit the mechanism.
16. Fidelity Data Protects Evaluation From the Wrong Conclusion
An intervention produces no meaningful improvement.
Without fidelity evidence, two very different explanations remain entangled:
- The intervention was implemented substantially as designed and was not sufficiently effective here.
- The intervention was not implemented in a way that gave its mechanism a fair test.
Those lead to different decisions. The first may justify replacement. The second may justify implementation repair.
17. High Fidelity Does Not Prove the Intervention Is Good
A system can faithfully implement a weak idea.
Fidelity answers whether the intended intervention was delivered. It does not prove that the intervention’s theory is correct, that the evidence generalises to this population, or that the opportunity cost is justified.
This is why fidelity and outcome evidence must travel together.
18. Mathematics Example: The Missing Independent Step
A department adopts an example–problem routine: teacher models one problem, students immediately solve a structurally similar problem independently, teacher diagnoses the result.
Over time, teachers begin solving the second problem together with the class because it feels supportive.
The routine still contains two problems. The active ingredient—the handover into independent execution—has disappeared.
19. English Example: The Discussion That Stops Sampling Thinking
An English department introduces structured discussion to improve inference. The design requires all students to commit to an interpretation before public discussion so the teacher can see the distribution of thinking.
Later, classes revert to the same confident volunteers answering first. Discussion remains. Diagnostic coverage disappears.
Again, the visible form survives while the mechanism thins.
20. Science Example: Inquiry Without the Decision
A practical investigation is designed so students choose a variable, justify a control and interpret whether evidence supports a claim.
To save time, the teacher gives the method, names every control and tells students what the graph should show.
The experiment still happens. The reasoning task does not.
21. CivDJ Cross-Domain Comparison: Medicine, Aviation and Software
Medicine distinguishes between a treatment protocol and the actual treatment received. Dose, timing, adherence and clinical judgement all matter. Education should not copy medical dosing literally, but the comparison shows why “prescribed” and “received” are different states.
Aviation uses standard operating procedures because some actions are too safety-critical to reinvent every flight. Yet pilots still exercise judgement when conditions differ. The system is tight around critical controls and flexible around context.
Software deployment offers another analogy. A tested feature can behave differently if configuration, dependencies or permissions change in production. The code name may be identical while the operational system is not.
Implementation fidelity is education’s version of configuration control around the learning mechanism.
22. Rainbolt Missing-Node Scan: Where Fidelity Commonly Leaks
- The programme name is preserved but the active ingredient is not specified.
- Staff receive launch training but no follow-up coaching.
- Dosage is counted while delivery quality is invisible.
- Materials are monitored while learner responsiveness is ignored.
- Teachers adapt the hardest component away.
- Leaders demand literal conformity to surface details that do not matter.
- New staff inherit resources without the rationale.
- Implementation data is collected but never used.
- Outcome failure is blamed on the intervention without checking whether the intervention was actually delivered.
- Outcome success is credited to the programme without checking which components were present.
23. Build a Core-Component Map Before Launch
Before implementation, write down three categories:
- Must preserve: components tightly linked to the mechanism.
- Can adapt: surface features that can change without damaging the mechanism.
- Unknown: features where the evidence is not strong enough to know how much adaptation is safe.
This map is more useful than simply telling staff to “follow the programme.”
24. Use Minimum Viable Monitoring
Monitoring that is too heavy becomes its own implementation burden.
A useful fidelity system samples just enough evidence to answer four questions:
- Was the core component present?
- Was enough of it delivered?
- Was it delivered competently enough to function?
- Did learners engage in the intended activity?
Those can be sampled through short observations, work samples, teacher logs, learner responses, digital traces or targeted coaching conversations. The method depends on the intervention.
25. Separate Support From Surveillance
If fidelity monitoring is experienced only as inspection, teachers may hide problems instead of surfacing them.
Early implementation should make drift discussable: Which component is hard to deliver? Which assumption does not fit this class? Which adaptation was made, and why? Did the change preserve the mechanism?
A learning implementation system treats variance as information before it treats variance as misconduct.
26. A Practical Fidelity Audit
- Name the intended learner outcome.
- State the theory of change in plain language.
- Identify the two or three core components most necessary for that mechanism.
- Specify minimum dosage where evidence justifies it.
- Define what competent delivery looks like.
- Define what learners must actually do.
- Name surface features teachers may adapt.
- Collect a small amount of implementation evidence.
- Compare implementation evidence with learner-state evidence.
- Repair drift before scaling.
- Revisit the core-component map when evidence shows an assumption was wrong.
- Do not call the programme ineffective until you know what was actually implemented.
27. Evidence and Limits
Implementation science gives education useful language for adherence, dosage, quality, responsiveness, core components and adaptation. Education-specific guidance, including the EEF’s 2024 implementation guide, increasingly emphasises structured but flexible implementation rather than simplistic replication.
But fidelity–outcome relationships are not mechanically uniform. Different studies define fidelity differently, different interventions depend on different mechanisms, and high-fidelity implementation can still produce modest outcomes if the intervention itself is weak or poorly matched to context. Research on implementation should therefore not be used to justify rigid scripting or to blame practitioners whenever outcomes disappoint.
The strongest practical use of fidelity is diagnostic. It helps a school know what intervention actually reached learners and which part of the implementation chain should be changed next.
28. The Return Path
Return to the school that launched the promising intervention.
The question after three months is not whether teachers still use the programme name.
Ask whether the active ingredients are still visible. Ask whether learners receive enough of them. Ask whether teachers understand which parts can flex. Ask whether the intervention is becoming more competent through use or quietly mutating away from its own mechanism.
That is what implementation fidelity protects.
Good implementation is not copying every surface detail. It is knowing what must survive, what may change and how to tell whether the mechanism still reaches the learner.