The question
What happens when improving the measurement becomes more important than achieving the underlying goal?
A person begins tracking their daily steps because they want to become healthier. The number helps. It makes an otherwise vague intention visible, provides feedback and encourages them to move more.
Over time, however, the relationship can change. Reaching the target begins to matter in itself. The person may choose activities that produce more recorded steps, or treat a beneficial recovery day as failure because it breaks a streak. The number still measures something real, but it no longer serves only as evidence about the original goal. It has begun to define success.
This does not happen whenever people track their behaviour. Activity trackers can help people become more active, and visible feedback can support genuine change. The problem is not that the number exists or influences behaviour. Influence is often why the metric was introduced.
The problem begins when improving the number becomes more important than improving what the number was meant to represent.
A Metric Is Evidence About a Goal
Many important goals cannot be observed directly.
Learning, health, service quality and public benefit contain several dimensions. To make them visible, people select indicators.
A test score provides evidence about learning. A step count measures a genuine component of physical activity. The number of completed cases represents one aspect of workplace output. Waiting time captures a real dimension of access to a service.
These measures are not necessarily separate from the goal. They may measure important parts of it. But one measured part rarely contains the whole purpose.
A student's score may reflect relevant knowledge while omitting reasoning, transfer or material outside the test. A step count records movement while saying little about strength, sleep, pain or recovery. A case-completion measure records volume while omitting complexity and quality.
The measure can be accurate, informative and incomplete at the same time.
This is why a metric has to be judged in relation to its purpose and use. A measure that helps identify broad trends may be insufficient for evaluating one person. A score that helps a teacher find an area for review may require greater scrutiny when it determines funding, promotion or employment.
The metric provides evidence about the goal. Problems emerge when its role expands beyond what that evidence can justify.
A Measure Changes When Consequences Depend on It
A measure can begin as a way of observing performance. Once rewards, penalties, rankings or access depend on it, it also begins directing performance.
A school's test results may affect its reputation or funding. A hospital's waiting-time figures may determine whether it satisfies a public target. An employee's case count may influence promotion. A platform's engagement statistics may guide investment and product decisions. A person's tracked streak may shape how they judge themselves.
The formal definition of the metric may remain unchanged. Its behavioural role does not.
People now have a reason to organise action around improving it.
A measure used for learning may invite experimentation. The same measure attached to punishment may encourage caution, selective effort or concealment. A threshold can make performance just above the boundary much more valuable than performance just below it, even when the practical difference is small.
Consequential use does not automatically invalidate a measure. It can, however, change the system being measured or make existing limitations in the measure more important.
Once the number affects what happens next, people begin learning what makes it move.
People Adapt to What Is Counted
Imagine a workplace that evaluates employees primarily by the number of cases they complete.
The measure may initially reveal useful differences in workload or efficiency. Once evaluation depends heavily on it, employees have a reason to prioritise cases that can be closed quickly.
That response is not necessarily harmful. Resolving simple cases efficiently may release time for more complex work. It becomes a distortion when important difficult cases are neglected, quality deteriorates or the allocation of effort conflicts with the service's underlying purpose.
Incentives do not only affect how hard people work. They can affect which work they choose, how they divide their time and whether they continue doing valuable work that is not counted.
Several responses are possible.
Sometimes adaptation produces genuine improvement. Better tools, clearer procedures or greater skill increase both completed volume and service quality.
Sometimes effort shifts toward measured work while unmeasured work receives less attention.
Sometimes people comply narrowly, meeting the formal requirement while neglecting the purpose behind it.
Gaming goes further: the recorded measure is strategically improved without equivalent improvement in what it is supposed to represent.
Falsification is different again. It involves deliberately altering, concealing or misreporting information.
These behaviours should not be collapsed into a choice between honest improvement and cheating. Observing a better metric does not establish which process produced it. An unusual performance pattern also does not, by itself, reveal anyone's motive.
The important question is what changed underneath the number.
A Better Score Is Not Always a Better Outcome
Education makes the distinction especially visible.
Tests can reveal gaps in knowledge, create common standards and show whether students have learned important material. Teaching content represented by a well-designed assessment can produce genuine learning.
But "teaching to the test" can describe very different practices.
A teacher may align instruction with a legitimate curriculum. Students practise relevant knowledge and become more capable. In that case, improved scores may provide valid evidence of progress.
Instruction can also narrow toward predictable formats, repeated item types or content most likely to appear on the consequential test. Attention may move away from valuable but untested material. When a proficiency threshold determines consequences, effort may concentrate on students expected to move from just below to just above it.
In some accountability settings, scores on consequential tests have improved more than performance on assessments not tied to the same consequences. Other research has detected deliberate score manipulation near important thresholds.
These findings do not mean that every improvement on a high-stakes test is false. They show that the headline score cannot explain its own cause.
Higher performance may reflect genuine learning, test-specific preparation, selective effort, manipulation or several of these at once. If the goal is broader learning, additional evidence from unfamiliar tasks, later performance or other assessments may be needed.
The same metric used as the target cannot, by itself, establish that the broader goal improved to the same degree.
A Target Can Help and Distort at the Same Time
Targets are often discussed as though they either work or fail. The same target can produce genuine improvement and problematic adaptation.
England's emergency departments were subjected to a target concerning the proportion of patients whose stay should not exceed four hours. Reported waiting-time performance improved substantially. Some staff described the target as helpful in reducing delays and improving patients' experience.
Researchers also observed strong patterns around the four-hour threshold. Healthcare reports described opportunities to reorganise procedures, admission decisions or recording practices in ways that helped avoid a recorded breach.
These are different forms of evidence. Better reported waiting times, staff perceptions, threshold clustering and documented opportunities for gaming do not prove one another.
Clustering near a threshold is not automatically suspicious. A target is intended to concentrate effort, so the pattern may reflect genuine urgency, strategic adaptation or both.
The case therefore cannot be reduced to either "the target worked" or "the target was gamed." It may have helped address a real problem while also narrowing attention and complicating interpretation.
Evidence of gaming does not make every gain false. Improved target performance does not prove equivalent improvement in every dimension of care.
A complete evaluation must ask what happened to the target, the underlying outcome and the parts of the system the target did not measure.
Goodhart's Law Names More Than One Problem
A familiar warning says that when a measure becomes a target, it ceases to be a good measure. This is commonly called Goodhart's law.
It is useful as a warning, but not as an automatic verdict.
When a proxy appears to fail, three questions help identify what happened.
First, are we selecting noise along with real performance? An unusually high score may partly reflect temporary variation that will not persist.
Second, has optimisation changed the relationship between the metric and the goal? A measure may work well under ordinary conditions but behave differently when pushed to an extreme or used to reorganise the system.
Third, can people improve the metric through a route that bypasses the goal? They may exploit what the measure includes, neglect what it omits or satisfy the formal target without creating the intended result.
Some of these failures involve strategy; others do not. Some make the metric slightly less informative, while others fundamentally alter what it represents.
Goodhart's law is therefore better understood as a family of warnings. Optimising a proxy can weaken its relationship with the underlying goal, but the mechanism and result must still be demonstrated.
Not every target becomes useless.
Metrics Can Improve the System
The possibility of distortion is not an argument against measurement.
Without metrics, serious problems can remain hidden. Decision-makers may rely on memory, reputation or confident impressions that are incomplete and biased. Poor performance can become normal because no one can see its pattern.
Measurement can reveal that some patients wait much longer than others, that professional practice differs from an accepted standard or that an intervention is not producing the expected result. It can expose unequal outcomes that isolated observations fail to detect.
Research on audit and feedback in healthcare demonstrates this positive function. When professionals receive information about their practice and how it compares with a meaningful standard, performance can improve. The average effects vary and are often modest, but measured feedback can support real change when its design and context are appropriate.
Personal tracking offers another example. A step count does not measure the whole of health, yet it can help someone become more active.
The fact that a measure influences behaviour is not the problem. Feedback is often intended to do exactly that.
The relevant question is whether the behaviour organised by the metric remains aligned with the underlying goal.
Abandoning an imperfect measure is not automatically safer. It may return decisions to less visible, less comparable or more biased judgment. Often the better response is to improve the measure's use and interpretation.
More Measures Are Not Automatically Better
When one metric becomes distorted, a common response is to add more.
This can help. Tracking quality as well as quantity may reveal a trade-off that one measure conceals. Combining short- and long-term outcomes may reduce fixation on immediate performance. Listening to affected people may expose consequences that administrative data miss.
But several measures can share the same blind spot. A composite score may hide trade-offs inside one new number. Additional reporting consumes time, and every new measure can become another target.
Qualitative judgment can supply context that numbers omit. A professional may recognise that one case was unusually difficult or that formally successful performance produced a poor human result.
Yet qualitative judgment can also be inconsistent, opaque or biased. It should complement quantitative evidence, not be presumed superior to it.
Lower-stakes measures may support learning with less pressure for selective optimisation. Periodic review may reveal neglected groups, threshold effects or a changing relationship between the measure and its purpose. Neither safeguard guarantees correction.
Return to the workplace case. Adding a quality score to the case count may reveal whether faster completion is undermining service. But if both measures reward only easily documented performance, difficult cases may remain neglected. Speaking with employees and service users may expose that omission, but their reports also require interpretation.
No measurement system interprets itself.
Keep the Metric Answerable to the Goal
A metric remains useful when it continues serving the purpose for which it was created.
Five questions can test that relationship.
First, what is the underlying goal?
Second, what does the metric capture well, and what does it omit?
Third, what behaviour do the consequences attached to the metric invite?
Fourth, what other evidence can test whether the underlying goal improved?
Fifth, does the metric still serve the purpose for which it was created?
That additional evidence does not have to be perfectly independent. It must, however, provide information not reducible to the same target metric. It may include longer-term outcomes, unfamiliar tasks, quality measures, experiences of affected people or direct examination of neglected cases.
These sources also have limitations. The purpose of combining them is not to create a flawless measurement system. It is to make divergence visible.
Sometimes the evidence will show genuine alignment: the number improved because the underlying work improved.
Sometimes it will reveal narrowing: the counted dimension improved while another important dimension deteriorated.
Sometimes it will identify strategic manipulation. Sometimes it will show that the metric remains useful but is being interpreted too broadly.
The appropriate response is not always to abandon the measure. It may be to revise its use, reduce the stakes, add another source of evidence, change the target or clarify the original purpose.
Metrics make important goals visible. They become misleading when visibility is mistaken for completeness and recorded success becomes the final standard of success.
A metric should remain evidence about the goal. When improving the number becomes the final definition of success, the tool has begun replacing its purpose.