HSW-0119 · How Studying Works
A student decides to become more disciplined.
She starts tracking questions completed each night.
The first week, she finishes 120 questions.
The second week, she reaches 180.
The dashboard looks excellent.
Then a teacher gives her an unfamiliar paper.
She struggles badly.
What happened?
Nothing dishonest had to happen. Once “questions completed” became the visible success number, the student naturally began choosing short, familiar questions. She skipped problems that looked slow. She repeated question types she already knew because they increased the count efficiently. Difficult corrections were postponed because one correction could consume the time needed to add ten easy questions to the total.
The metric improved.
The capability the metric was supposed to represent did not improve at the same rate.
This article calls that pattern study metric gaming: behaviour adapts to improve the visible measure in ways that weaken the measure’s relationship to the real learning outcome.
A measure is useful while it helps us see the goal. It becomes dangerous when the goal quietly becomes the measure.
This is deliberately narrower than Measurement Error, which asks why marks can move even when underlying capability has not moved; Benchmarking, which asks how to compare performance without turning comparison into the goal; Learning Evidence Density, which asks how much useful signal a task produces; and How Test Expectancy Works, which owns the legitimate way future assessment changes present study. Study Metric Gaming owns another question: what happens when people optimise the proxy instead of the capability the proxy was meant to reveal?
The systems route: every measurement system creates incentives
Once a measure becomes consequential, people pay attention to it.
That is often the point.
Schools use marks to direct effort. Apps use streaks to encourage consistency. Teachers use completion checks to make practice visible. Organisations use targets because invisible goals are hard to manage.
The difficulty is that a metric is almost always a compressed representation of something larger.
- A test score represents some sample of knowledge and performance.
- A homework completion rate represents some amount of attempted work.
- A reading log represents some amount of reading activity.
- A revision streak represents some pattern of consistency.
- A practice accuracy percentage represents performance on a selected set of items.
Compression creates convenience. It also creates a gap between the number and the thing the number stands for.
Metric gaming lives in that gap.
Campbell’s law is a warning about pressure on indicators
Social scientists have long warned that quantitative indicators can become distorted when they carry heavy decision pressure. Discussions of Campbell’s law make the broad point that the more a quantitative social indicator is used for consequential decision-making, the more pressure exists to corrupt or distort the processes it is intended to monitor.
The National Academies has repeatedly cautioned against treating a single test score as a definitive educational judgment. Its work on high-stakes testing and test-based accountability examines unintended consequences that can arise when incentives become concentrated around measured outcomes.
National Academies: High Stakes — Testing for Tracking, Promotion, and Graduation
National Academies: Incentives and Test-Based Accountability in Education
Brookings has similarly discussed how high-stakes test accountability can produce narrowing, strategic behaviour and other distortions when the measured outcome becomes too dominant.
Brookings: The costs of misusing test-based accountability in schools
This does not mean measurement is bad.
It means measurement changes behaviour, so the measurement system itself must be designed.
Gaming does not always mean cheating
The word “gaming” can sound accusatory. In this article it has a broader systems meaning.
A learner can game a metric without breaking a rule and without consciously trying to deceive anyone.
If an app celebrates the number of flashcards reviewed, the learner may naturally choose very small cards. If a school praises homework completion, a student may prioritise tasks that can be finished visibly and avoid slower diagnostic work. If a parent asks only for the mock-exam score, the learner may repeat familiar papers because scores rise faster.
People adapt to the scoreboard they can see.
The system designer therefore has a responsibility to ask what behaviour the scoreboard makes attractive.
Legitimate test preparation is not automatically metric gaming
Examinations are real performance environments. Students should understand instructions, time constraints, answer formats, mark allocation, question conventions and the kinds of thinking the examination actually requires.
That is not gaming. It is alignment.
The boundary appears when preparation raises the observed measure without building the capability needed for new, legitimate instances of the task.
- Learning how to structure a response under the real time limit is alignment.
- Memorising an exact answer that happens to reappear is not evidence of broad capability.
- Practising representative question types is alignment.
- Repeating the same paper until recall of the paper dominates the score weakens its value as a readiness measure.
- Learning vocabulary that supports reading and writing is capability building.
- Learning to recognise answer keys by surface cues may raise a narrow score without improving transfer.
The question is not, “Did the score rise?”
Ask, “Would the learner still perform if the legitimate surface details changed?”
The learning route: proxy metrics are most dangerous when the learner forgets the target
Learning targets are usually richer than one number.
- Can the learner retrieve important knowledge without prompts?
- Can the learner choose an appropriate method?
- Can the learner explain why it works?
- Can the learner detect and correct errors?
- Can the learner transfer the idea to a new context?
- Can the learner perform under realistic time and support conditions?
- Can the learner still do it later?
A single metric may sample one or two of these dimensions.
That is fine if everyone remembers the sample is not the whole capability.
Metric gaming begins when the sampled dimension becomes the easiest route to looking successful.
Metric 1: study hours
Study hours can be useful for answering one question: how much time was allocated?
They cannot answer every question about learning.
If “two hours studied” becomes the target, the learner can satisfy the number by watching explanations passively, reorganising notes, remaining logged into a platform, or rereading comfortable material.
None of those activities is automatically useless. The problem is that elapsed time is being asked to stand in for independent capability.
Pair time with evidence: what could the learner do at the end that could not be done at the beginning?
Metric 2: questions completed
Question count rewards speed and volume.
That can be valuable during fluency practice.
It can also make one difficult misconception look economically irrational. Why spend fifteen minutes repairing one hard question when fifteen easy questions could be added to the total?
If the purpose of the session is fluency, volume may be appropriate. If the purpose is diagnosis and repair, a lower count can represent more learning.
The metric must match the job.
Metric 3: accuracy percentage
Accuracy seems closer to capability.
But accuracy depends on the questions selected.
A student can maintain 95% accuracy by practising only mastered items. Another student can fall to 62% because she deliberately enters weak topics and unfamiliar variants.
Which learner had the better study session?
The number alone cannot tell us.
Accuracy needs context: item difficulty, novelty, support level, timing and purpose.
Metric 4: streaks
Streaks can help make a habit visible.
They can also cause the learner to protect the streak rather than protect the learning.
A five-minute trivial review keeps the streak alive even when a meaningful forty-minute repair session is repeatedly avoided.
The streak has done its behavioural job—encouraging return—but it should not be mistaken for evidence that important capability is improving.
Metric 5: mock examination score
Mock scores are valuable because they compress many dimensions into a realistic performance sample.
They become less informative if the paper is no longer meaningfully unseen.
A student who repeats the same paper may learn from the corrections, which is useful. But the rising score on the repeated paper should not be interpreted as if each repetition were a fresh estimate of exam readiness.
The paper has changed role. It has moved from measurement instrument to training material.
That is perfectly legitimate as long as the label changes too.
The Mathematics route: an answer rate can hide method fragility
Suppose a student learns to recognise that every worksheet question in a section requires the quadratic formula.
Accuracy rises.
Then an examination mixes factorisation, completing the square, graphs and formula-based solutions.
The learner’s real challenge was method selection, not formula execution.
Blocked practice had made the method choice invisible. The local accuracy metric therefore overestimated independent capability.
A better measure includes mixed and unseen problems in which the learner must first decide what kind of problem is present.
The English route: word count can improve while writing quality falls
Word count is useful when a task has a minimum development requirement.
But if length becomes the dominant metric, a writer can add repetition, inflated phrasing, redundant examples and decorative vocabulary.
The essay becomes longer without becoming clearer, more coherent or more persuasive.
Measure what the writing is meant to do: communicate, develop, organise, support and control language for an audience and purpose.
Length is one constraint, not the owner of quality.
The Science route: lab completion can hide weak causal reasoning
A student can complete every practical step correctly by following a procedure.
If the measured target is only “experiment completed,” the student may never need to explain why a variable is controlled, why repeated readings matter, what uncertainty means, or whether the conclusion follows from the evidence.
The completion metric is not wrong. It is incomplete.
Add questions that expose the scientific reasoning behind the procedure.
The financial route: quarterly targets show why time horizons matter
Finance and corporate management often wrestle with the tension between short-term targets and long-term value.
A manager can sometimes improve a near-term metric by postponing maintenance, reducing investment or pulling activity forward from the future. The short-term number improves while the long-term system becomes more fragile.
Students can do the same thing academically.
- Cramming can improve tomorrow’s quiz while displacing sleep and later retention.
- Repeating familiar papers can improve mock scores while reducing time for unfamiliar transfer.
- Finishing visible homework can protect completion statistics while leaving a foundational weakness unrepaired.
The useful question is: what future cost is being hidden by today’s metric?
The school route: one number cannot safely own a complex educational outcome
Schools need measurement. Without it, improvement becomes impressionistic.
But the more complex the outcome, the more dangerous it is to let one proxy dominate.
A school that judges reading only by pages logged may encourage page accumulation. A school that judges teaching only by short-cycle scores may unintentionally discourage slower foundational work. A school that judges support only by programme attendance may optimise enrolment rather than impact.
A robust measurement system therefore asks both:
- Did the indicator improve?
- Did the underlying educational capability improve in ways the indicator could not easily manufacture?
The teacher route: make the metric harder to optimise without learning
Good assessment design reduces easy proxy exploitation.
- Use fresh examples rather than exact repeats when measuring transfer.
- Mix problem types so method selection remains visible.
- Ask for explanation as well as answer where reasoning matters.
- Include delayed checks so short-term memorisation is not the only route to success.
- Sample across the curriculum rather than announcing exactly which tiny subset will be measured.
- Separate training scores from readiness scores.
The goal is not to trick students.
The goal is to keep the easiest way to raise the measure reasonably aligned with the capability we actually want.
The tutor route: diagnostic honesty matters more than impressive dashboards
A tutor can make a student look better quickly by choosing familiar material, prompting heavily and reporting only the final supported accuracy.
That may feel encouraging.
It is dangerous if the parent or student mistakes supported classroom success for independent performance.
Separate the conditions.
- supported attempt;
- guided correction;
- independent retry;
- delayed independent retest;
- unseen transfer.
One learner can have five different accuracies under those five conditions. That is not a problem. It is a more honest map of where capability currently lives.
The parent route: ask what the number can and cannot prove
A mark can be good news without being complete news.
When a result rises, ask:
- Was the task genuinely new?
- Was support reduced?
- Was timing realistic?
- Was the content representative?
- Did the learner make fewer serious errors or simply meet easier questions?
- Can the learner explain the improvement?
- Does the gain survive a later retest?
These questions do not take joy away from improvement. They protect the meaning of improvement.
The AI route: dashboards can become easier to optimise than the mind
Digital learning systems can measure almost everything: minutes active, questions attempted, hints used, streaks, levels, XP, badges, accuracy, speed and completion.
More data does not automatically mean better evidence.
A learner can keep an app open without thinking. A learner can request hints until an answer becomes obvious. A learner can choose easier quiz settings. An AI system can generate polished output that makes the student’s workflow look productive while independent retrieval remains weak.
Digital systems should therefore distinguish activity telemetry from capability evidence.
The first tells us what the learner did with the system. The second tells us what the learner can now do without the system doing the crucial thinking.
Use a metric portfolio instead of one king metric
One defence against gaming is to avoid giving one proxy total control.
A useful study metric portfolio might include:
- output: what was completed;
- accuracy: how much was correct;
- difficulty: how challenging the items were;
- independence: how much support was required;
- transfer: whether performance survived a changed context;
- retention: whether performance survived time;
- error severity: whether the remaining mistakes are superficial or structural;
- coverage: whether important domains are being omitted.
No learner needs a giant dashboard for every session. The point is conceptual: when one number matters too much, pair it with another number or qualitative check that is hard to improve through the same shortcut.
Use cold samples
A cold sample is a short check that the learner has not just rehearsed in the exact same form.
It can be a fresh question, a new text, a changed number set, a delayed retrieval prompt, an unfamiliar application or an oral explanation without notes.
Cold samples are valuable because they are harder to optimise through memorising the measurement instrument itself.
They do not replace normal practice. They audit whether normal practice is still connected to generalisable capability.
Use delayed evidence
A metric can be gamed by timing as well as task selection.
If the learner is always measured immediately after practice, short-lived accessibility can look like durable mastery.
Return later.
Ask the learner to retrieve again after enough time has passed for the immediate rehearsal context to weaken.
This is why durable learning systems care about retention and relearning, not only end-of-session performance.
The training route: workplace KPIs have the same structural risk
Professional training may track modules completed, certificates earned, videos watched and quiz passes.
Those measures can show participation.
They do not automatically show transfer to real work.
A training system becomes more credible when it asks whether behaviour changed on the job, whether errors fell, whether decisions improved and whether capability survives outside the training platform.
Students who learn to question proxy metrics are therefore learning a workplace skill as well as an academic one.
The world route: civilisation runs on proxies, so proxy literacy matters
GDP, inflation, unemployment, credit ratings, hospital waiting times, crime statistics, school results, customer-satisfaction scores and productivity measures all compress complicated realities.
Societies need such measures because complexity cannot be managed without abstraction.
But mature systems do not worship the abstraction.
They ask what the metric leaves out, what behaviour it encourages, how it can be manipulated, and what complementary evidence is needed.
Study metric gaming is therefore a small classroom version of a much larger civic skill: learning to use numbers without letting numbers replace reality.
A practical metric-audit protocol
- Name the real capability. What do we ultimately want the learner to be able to do?
- Name the proxy. Which number or visible indicator is standing in for that capability?
- Explain the link. Why should improvement in the proxy usually indicate improvement in the capability?
- Find the cheap shortcut. How could someone raise the proxy while learning less than intended?
- Check whether the shortcut is already happening. Look at task selection, support levels, repetition and timing.
- Add a guardrail. Use fresh items, mixed practice, delayed checks, explanation or transfer.
- Use more than one evidence channel. Do not let one metric become sovereign.
- Separate training measures from readiness measures. A repeated practice score and a cold assessment score answer different questions.
- Retire corrupted metrics. If a number can no longer be interpreted honestly, stop treating it as evidence merely because it is easy to collect.
The center-to-edge route
- Learner: Which number am I trying to improve, and could I improve it without becoming more capable?
- Peer: Are comparisons changing what we choose to practise?
- Teacher or tutor: Does my measurement system reward the behaviour I actually want?
- Family: Have we made marks, hours or completion the only visible definition of progress?
- School: Which indicators could be improved through narrowing rather than genuine learning?
- Education system: Are consequential decisions based on a portfolio of evidence or an overloaded proxy?
- Training organisation: Are course-completion metrics connected to real workplace transfer?
- World: Which public metrics remain useful only because institutions continuously audit their relationship to reality?
The improvement route: pair every important metric with an anti-gaming check
Create a simple table with four columns:
- metric;
- what it is supposed to represent;
- how it could improve without the target improving;
- one check that would expose the shortcut.
For example:
- Questions completed → fluency and practice volume → choose only easy questions → include a mixed cold sample.
- Accuracy → correct performance → practise only mastered material → report accuracy by difficulty and novelty.
- Study hours → sustained effort → remain active without retrieval → add an independent exit task.
- Mock score → exam readiness → repeat the same paper → use a fresh paper or changed item set.
The exercise changes the learner’s relationship with data. Metrics stop being trophies and become instruments.
The final rule
Measure learning.
But keep asking whether the measure still deserves to be believed.
If the easiest way to improve the score is no longer the same as the best way to improve the learner, redesign the score.
Previous in the numbered series: HSW-0118 · Study Livelock.