Essays · Evidence and Explanation

Measurement Changes the System

Measurement can sometimes do more than record human behaviour: observation, assessment and monitoring can become causal conditions influencing what happens next.

By Yona Ole Lobulu ·

Essay13 min readD2.10

Topic
Evidence and Explanation
Read first
2 pieces should be read before this one
Reading time
About 13 minutes of reading
Difficulty
Advanced reading: this page assumes a fair amount of earlier reading

The question

How can observing, measuring or repeatedly assessing human behaviour alter the very system being studied?

Definition

Measurement reactivity is change in responding, experience or behaviour attributable partly to the process of being measured, assessed, monitored, questioned or observed. Reactivity is not assumed to be universal, large, beneficial, harmful, cumulative, or produced through one mechanism: its existence, direction and magnitude are empirical questions.

Imagine someone agrees to track their exercise every evening.

Each day they answer the same questions:

Did you exercise?

For how long?

How difficult was it?

Did you intend to exercise?

Why did you or did you not?

At first, the questions may seem to do nothing more than record what already happened.

But the assessment itself is also an event.

The person may begin noticing exercise more often during the day. They may anticipate having to report their behaviour that evening. They may become more aware of missed intentions. They may change how they interpret what counts as exercise.

Or nothing detectable may change.

That uncertainty is the point.

For the person being measured, an assessment is not only a data-generating procedure. It can also be an event to which the person responds.

The simplest model of measurement is:

behaviour → measurement

D2.10 adds another possibility:

behaviour → measurement → possible system response → later behaviour or reporting

Once that happens, the measure is no longer only a window onto the system.

It has become one event within it.

Measurement Can Have Two Roles

Measurement begins with an evidential purpose.

Researchers want to know:

what someone did;

what they experienced;

how often something occurred;

whether a state changed;

how strongly a construct was expressed.

That is measurement's representational role.

The procedure generates information intended to stand for something about the person or behaviour.

As established in Measurement and Operational Definition, this requires decisions about how a construct is operationalised and what the resulting measure actually represents.

D2.10 adds a different question:

What does implementing the measure do?

A question may redirect attention.

An observer may alter the social context.

Repeated assessment may create anticipation or familiarity.

Self-monitoring may make a behaviour more noticeable.

Feedback generated from measurement may influence what happens next.

None of these effects is guaranteed.

But when one occurs, the measurement procedure is doing two things at once.

A measure can function simultaneously as a source of evidence and as an event within the causal system generating future evidence.

That does not make the evidence meaningless.

It means the measurement procedure itself may need to be included among the conditions shaping what follows.

Measurement Reactivity Is Conditional

Measurement reactivity refers to change in responding, experience or behaviour attributable partly to the process of being measured, assessed, monitored or observed.

Research supports the existence of such effects.

It does not support a universal law that measurement reliably changes behaviour in one direction.

Studies testing whether simply asking people questions changes their later behaviour have found effects in some contexts, but updated evidence suggests that average effects can be small and heterogeneous.

Research using intensive repeated assessment shows a similarly mixed picture. Some studies detect reactivity. Others find little or none. In many cases, reactivity has not been directly tested.

The appropriate conclusion is therefore conditional:

The existence, direction and magnitude of measurement reactivity are empirical questions.

Measurement may:

increase behaviour;

decrease it;

alter reporting;

change one experience without changing another;

change awareness without detectable behavioural change;

produce little measurable effect.

Awareness of being measured does not determine the direction of the response.

This is why slogans such as:

what gets measured improves

do not belong here.

Measurement can become behaviourally active.

Whether it does, and what follows, must be established.

Observation Changes the Measurement Context

One form of possible reactivity occurs when people know that someone is observing them.

Adding an observer changes the social situation.

Behaviour now occurs in a context containing:

possible evaluation;

expectations about the observer;

altered perceived consequences;

opportunities for self-presentation;

increased self-awareness.

Adding an observer changes the social context; whether behaviour responds to that change is an empirical question.

Sometimes the behavioural response can be substantial.

Sometimes it is small or undetectable.

And a behavioural change during observation does not tell us why it occurred.

Possible explanations might involve:

social evaluation;

demand characteristics;

attention;

expectancy;

impression management;

altered incentives.

A change during observation does not identify why the change occurred.

The familiar label Hawthorne effect is often used for behaviour that changes because people know they are being studied or watched.

But the research grouped under that label is too heterogeneous for it to function as one canonical mechanism.

Observation effects need explanation.

Giving them a familiar name does not provide one.

Assessment Can Redirect Attention

An observer is not necessary for measurement to affect the system.

Questions themselves present content for attention.

If someone is repeatedly asked about:

sleep;

exercise;

mood;

studying;

spending;

stress;

those topics repeatedly enter attention during assessment.

Attention is one plausible route through which assessment may become reactive, not the default explanation for all measurement effects.

Assessment can alter what is attended to at the moment of measurement.

Whether that altered attentional state persists, changes interpretation or contributes to later behaviour is a separate empirical question.

This distinction matters.

The chain:

assessment → attention

does not automatically establish:

attention → later behavioural change.
Assessment can alter what is attended to; whether that shift contributes to later behaviour requires separate evidence.

Repeated assessment may also interact with:

salience;

interpretation;

memory;

motivation.

But these should be treated as possible pathways or affected processes, not as universal mechanisms of measurement reactivity.

An observed change after assessment does not tell us which pathway operated.

This is where Attention Selects Information becomes relevant downstream.

D2.10 needs only the measurement implication: asking a person to attend to something can itself become part of the conditions under which later evidence is generated.

Self-Monitoring Can Blur the Line Between Measurement and Intervention

The distinction becomes especially visible when people measure themselves.

Consider:

tracking steps;

logging spending;

recording food;

rating mood;

monitoring sleep;

keeping a study log.

These procedures may begin as attempts to record behaviour or experience.

But repeated self-monitoring can sometimes alter what happens next.

The tracked behaviour may become more noticeable.

A discrepancy between current behaviour and an intention may become more visible.

The person may begin acting partly in response to the fact that the behaviour is being tracked.

At that point, measurement is no longer performing only a representational function.

But this does not justify the categorical statement:

self-monitoring is an intervention.

Sometimes it is behaviourally active.

Sometimes it is not detectably so.

Recording and intervention are distinguishable functions, but they can coexist in the same procedure.

There is another causal complication.

Many systems described casually as "tracking" do much more than record.

They may also provide:

prompts;

goals;

feedback;

comparisons;

reminders;

incentives.

When these are bundled together, a later behavioural effect cannot automatically be attributed to measurement alone.

When self-monitoring is combined with prompts, goals, feedback or incentives, any behavioural effect cannot automatically be attributed to measurement alone.

The causal components need to be distinguished if the explanation depends on them.

Repeated Measurement Creates a Measurement History

Repeated measurement introduces a temporal dimension.

Suppose someone completes the same assessment twenty times.

By the twentieth assessment, they have a history that did not exist at the first.

They may now:

recognise the questions;

anticipate future assessments;

become practiced at responding;

interpret response categories differently;

think about the topic between measurement occasions.

That history exists even if it has no meaningful causal effect.

Repeated measurement creates a history of prior assessments; whether that history changes subsequent responses is an empirical question.

Repeated assessment therefore creates opportunities for reactivity.

It does not establish reactivity.

The temporal pattern can also differ.

Any effect might:

accumulate;

diminish;

appear mainly at the beginning;

fluctuate;

remain undetectable.

There is no general rule that more measurement produces more change.

Still, when earlier assessments have affected the person or response process:

Later evidence may sometimes come from a system partly shaped by earlier evidence collection.

That possibility is important.

But it must immediately be distinguished from a simpler temporal observation:

A time trend under repeated measurement does not, by itself, show that measurement caused the trend.

Scores may change because the target phenomenon changed independently.

They may change because the person became familiar with the task.

Reporting may change.

Measurement error may contribute.

Several processes may occur together.

A measurement history is real.

Its causal importance must still be demonstrated.

Feedback Adds Another Causal Step

Measurement and feedback often occur together.

They should not be treated as the same event.

Imagine a device recording someone's daily steps.

One version stores the data without showing them to the person.

Another displays the count.

A third sends alerts, trends, comparisons or goals based on the count.

All three measure steps.

They do not create the same causal system.

Once the measurement result is returned to the person, the sequence becomes:

behaviour → measurement → information → feedback → later behaviour
When measurement generates feedback, feedback becomes an additional causal event rather than merely part of the record.

This distinction matters because a change attributed loosely to "measurement" may actually depend on something added after measurement:

feedback;

reminders;

comparison;

goals;

incentives.

Recording alone and recording plus feedback are different causal arrangements.

D2.10 establishes that distinction.

It does not yet ask what happens when the metric itself becomes something behaviour is organised to optimise.

That belongs downstream.

Measurement Error and Measurement-Induced Change Are Different Problems

Measurement can create difficulty in at least two different ways.

The first is measurement error.

The recorded value fails to represent the target accurately.

A sensor may misclassify behaviour.

A questionnaire may be unreliable.

A participant may interpret a scale differently than intended.

The second is measurement-induced change.

The procedure contributes causally to changing later behaviour, experience or responding.

Measurement error concerns representation; measurement reactivity concerns causal influence.

These dimensions are independent enough to produce four conceptual possibilities:

An accurate, non-reactive measure represents the target without materially altering later responding. An accurate but reactive measure represents the target accurately while also influencing later responding. An inaccurate, non-reactive measure misrepresents the target without materially altering it. An inaccurate and reactive measure both misrepresents the target and influences later responding.

This makes an important point clear.

A measure can alter what is measured even when the measurement itself is accurate.

The reverse is also true.

A measure can be inaccurate without changing the target at all.

The two problems should not be collapsed.

Self-report adds another layer.

Suppose repeated ratings change.

That does not tell us automatically what changed.

Measurement reactivity can affect:

The target phenomenon

The underlying behaviour, state or experience changes.

The response process

The person changes how they:

interpret the question;

use the response scale;

recall the target;

categorise the experience;

report it.

Or both may change.

Measurement reactivity can affect the target phenomenon, the response process used to report it, or both.

Therefore:

A change in the measurement response is not automatically a change in the underlying phenomenon.

This distinction is crucial for What Counts as Evidence of Change.

It is possible to observe a changed measure without yet knowing whether the target itself changed in the same way.

Assessment burden is another separate problem.

Repeated questionnaires may become tiring or reduce compliance.

People may skip assessments or leave a study.

That matters scientifically.

But declining participation does not itself establish reactivity in the behaviour or experience being measured.

Burden and target reactivity can interact.

They are not synonyms.

Reactivity Does Not Make Measurement Scientifically Useless

At this point, an overcorrection becomes possible.

If measurement can sometimes alter human behaviour, perhaps measurement itself is hopelessly contaminated.

That conclusion does not follow.

Reactivity can itself be studied.

Researchers can, where appropriate:

compare measurement intensities;

separate recording from feedback;

vary observation conditions;

test assessment effects;

compare different protocols;

interpret results relative to the measurement environment.

Reactivity does not make measurement impossible; it makes the measurement procedure part of the conditions that may need to be explained.

Even more importantly, reactivity is only a threat relative to a particular inference.

Suppose the scientific question is:

How do people behave when they know they are being monitored?

Then responding to observation may be part of the target phenomenon.

Suppose instead the conclusion is:

This is how people behave when they are not being monitored.

Now the same observation-induced change may threaten that inference.

Whether reactivity threatens validity depends on the target inference.

More precisely:

Reactivity threatens an inference when the measurement-induced change is relevant to the target claim and is not appropriately incorporated into the design or interpretation.

This is where D2.10 intersects with Laboratory Effects and Real-World Behaviour.

Measurement procedure is one feature of the conditions under which evidence is generated.

If those conditions differ from the target setting, measurement-induced effects may matter to generalisation.

Reactive measurement is therefore not synonymous with invalid measurement.

Its significance depends on what the evidence is supposed to establish.

Measurement Can Become an Alternative Explanation for Change

Suppose repeated assessments show that someone changed over six weeks.

D2.10 adds one candidate explanation that might otherwise be overlooked:

the measurement procedure itself contributed to what changed.

That does not mean it did.

The evidence may instead reflect:

substantive change independent of measurement;

altered reporting;

practice or familiarity;

ordinary fluctuation;

measurement error;

several processes together.

But once measurement reactivity is plausible:

Measurement-induced change can become an alternative causal explanation for observed change.

This matters because evidence collection is often treated as though it sits outside the causal history of the person being studied.

Sometimes it does for practical purposes.

Sometimes it does not.

When measurement is potentially reactive, the history of measurement becomes part of the evidential context for interpreting change.

Measurement protocols can also differ across studies.

One study may assess participants once.

Another may question them daily.

Another may provide feedback.

Another may involve visible observation.

Those differences could contribute to differences in findings.

They should not be assumed to do so.

Measurement procedure is one possible source of cross-study variation, not a generic explanation for non-replication.

Broader questions about robustness remain with Replication, Robustness and Scientific Confidence.

When Measurement Enters the System

The structure can now be stated simply.

A behaviour or experience occurs.

It is measured.

The measurement may alter:

attention;

the social context;

the response process;

feedback;

later choices.

Then something happens next.

That later state is measured again.

The loop can therefore become:

behaviour or experience → measurement → possible system response → later behaviour, experience or reporting → new measurement

Sometimes the middle link is negligible.

Sometimes it matters.

Measurement can enter causal loops with behaviour.

This is the boundary D2.10 establishes.

It should not be confused with the stronger downstream phenomenon in which a measure becomes something people actively optimise.

D2.10 establishes that measurement can influence behaviour. D2.11 asks what happens when the measure itself becomes an object of optimisation or begins to displace the underlying goal.

That distinction remains strict.

The Measure Is Part of the Conditions

Return to the person answering exercise questions every evening.

Did the questions change their exercise?

The fact that exercise later changed would not tell us.

The assessments may have contributed.

They may have changed only how the person reported their behaviour.

They may have had little effect at all.

That causal question requires evidence.

What D2.10 establishes is the possibility that evidence collection itself belongs inside the system being explained.

In human systems, measurement can become part of the causal conditions affecting what is measured, so observation and assessment cannot always be treated as behaviourally neutral.

The scientific response is not to assume reactivity and not to ignore it.

It is to determine whether it matters for the inference being made.

Sometimes the measure does not merely record the system. It becomes one of the conditions under which the system changes.

Behind this page

The claims this essay makes, the evidence behind them, and the limits it accepts.

Evidence status

High confidence

Strongly supported, though resting on synthesis or principle rather than a single decisive body of evidence.

Claims

  1. Measurement reactivity can occur

    High confidence

    What this does not assert: Being measured, observed or monitored can contribute causally to later responding.

  2. Reactive measurement can remain scientifically informative

    High confidence

    What this does not assert: Reactivity is a threat relative to a particular inference, not to measurement as such.

  3. Self-monitoring can become behaviour-changing

    Strong but context-sensitive

    What this does not assert: Context-sensitive; whether it does so in a particular setting is an empirical question.

  4. Visible observation can produce substantial effects in some contexts

    Strong but context-sensitive

    What this does not assert: Context-sensitive; effects elsewhere may be small or undetectable.

  5. Repeated assessment can influence later responding

    Strong but context-sensitive

    What this does not assert: Context-sensitive; temporal patterns may accumulate, attenuate, appear early, fluctuate or be absent.

  6. Measurement protocols can contribute to differences across studies

    Strong but context-sensitive

    What this does not assert: Context-sensitive; procedure is one possible source of variation, not a generic explanation for non-replication.

  7. Question-behaviour effects are heterogeneous and often small on average

    Strong but context-sensitive

    What this does not assert: Context-sensitive; asking about behaviour does not reliably change it.

  8. Measurement can become part of the causal conditions affecting what is measured

    Canonical inference

    What this does not assert: A Library-level synthesis of the evidence rather than a single reported finding.

  9. A measure can serve simultaneously as evidence and as an event in the causal system

    Canonical inference

    What this does not assert: The representational and causal roles are distinguishable but can coexist.

  10. An accurate measure can still be reactive

    Canonical inference

    What this does not assert: Representation quality and causal reactivity are separate dimensions.

  11. A change in the measurement response is not automatically a change in the underlying phenomenon

    Canonical inference

    What this does not assert: Changes in scores or reports require interpretation.

  12. Measurement reactivity is not universal

    High confidence

    What this does not assert: Many procedures produce no detectable change in the target.

  13. A time trend under repeated measurement does not, by itself, show that measurement caused the trend

    Canonical inference

    What this does not assert: Genuine change, practice, familiarity, altered reporting, error and ordinary variation remain alternatives.

  14. Feedback is an additional causal event rather than part of the record

    Canonical inference

    What this does not assert: Recording alone and recording plus feedback are different causal arrangements.

  15. Bundled tracking effects cannot be attributed to measurement alone

    Canonical inference

    What this does not assert: Prompts, goals, feedback, comparisons and incentives are separable causal components.

  16. Attention is one plausible route to reactivity, not the default explanation

    Canonical inference

    What this does not assert: Assessment can alter attention without that shift contributing to later behaviour.

  17. Naming an observation-related change does not explain its mechanism

    Canonical inference

    What this does not assert: The Hawthorne label covers heterogeneous phenomena and is not one established mechanism.

  18. Assessment burden and target reactivity are different problems

    Canonical inference

    What this does not assert: They can interact; lower compliance is not evidence that the target changed.

  19. Whether reactivity threatens validity depends on the target inference

    Canonical inference

    What this does not assert: Responding to observation may be part of the target phenomenon or a threat to generalisation.

  20. Measurement-induced change can become an alternative causal explanation for observed change

    Canonical inference

    What this does not assert: It becomes a candidate explanation, not the established one.

  21. Measurement can enter causal loops with behaviour

    Canonical inference

    What this does not assert: The existence and magnitude of the middle causal link must still be demonstrated.

  22. The direction and magnitude of reactivity vary

    High confidence

    What this does not assert: Measurement may increase, decrease, alter or leave the target unchanged.

  23. Observation-related behavioural effects occur in some contexts

    High confidence

    What this does not assert: Their presence in one setting does not establish them as a general law.

  24. Repeated assessment is not necessarily neutral

    High confidence

    What this does not assert: It creates opportunities for reactivity without establishing it.

  25. Repeated assessment is not necessarily reactive

    High confidence

    What this does not assert: A measurement history can exist with no meaningful causal effect.

  26. Measurement error and measurement-induced change are distinct problems

    High confidence

    What this does not assert: Error concerns representation; reactivity concerns causal influence.

  27. Response-process change and target change are distinguishable

    High confidence

    What this does not assert: Either, both or neither may occur under a reactive procedure.

  28. Observed reactivity does not identify its mechanism

    High confidence

    What this does not assert: Evaluation, expectancy, attention, demand characteristics and incentives remain competing candidates.

Where to go from here

Next published piece

When Metrics Replace the Goal

Metrics make important goals visible, but they capture only part of what matters. When rewards, penalties or attention become organised around the number, people may improve the measure without producing equivalent improvement in the underlying goal.

Continue through the Library →See where this sits in the graph →

Back to the Library →