Essays · Evidence and Explanation

What Counts as Evidence of Change?

A changed score, behaviour or subjective experience can all be evidence of change. But what each observation allows us to conclude depends on what is claimed to have changed, relative to what, and across what period.

By Yona Ole Lobulu ·

Essay10 min readD2.1

Topic
Evidence and Explanation
Read first
3 pieces should be read before this one
Reading time
About 10 minutes of reading
Difficulty
Advanced reading: this page assumes a fair amount of earlier reading

The question

What evidence justifies the conclusion that a person has changed?

Definition

A claim that a person has changed is justified when the evidence is appropriate to the specified outcome, level and timescale and provides sufficient reason to conclude that the observed difference reflects the change being claimed rather than measurement error, ordinary variability or a plausible alternative interpretation.

A person says they feel different.

Someone close to them notices fewer arguments. A questionnaire score has improved. Their behaviour in a difficult situation looks different from before.

Has the person changed?

Perhaps.

But those observations do not all support the same conclusion, and none can be interpreted properly until we know what change is actually being claimed.

If the claim is that behaviour changed, behavioural evidence matters directly. If the claim is that subjective anxiety decreased, first-person report becomes highly relevant. If the claim is that the person has undergone a lasting psychological transformation, the evidential burden is much greater.

The problem is therefore not simply whether we have evidence.

It is whether the evidence we have is appropriate to the conclusion we want to draw.

Evidence begins with the claim

Different things can change.

Behaviour can change. Experience can change. Cognitive performance can change. Interpersonal patterns can change. Bodily and neural processes can change.

These outcomes may be related, but they are not interchangeable.

Evidence that someone behaved differently tells us something about behaviour. It does not automatically establish that their identity changed.

A person's report that they feel less anxious may be highly informative about their experience. It does not establish that their behaviour, physiology or long-term pattern changed to the same degree.

A neural measure may reveal a neural difference. Its psychological meaning still has to be justified.

The evidential value of an observation therefore depends partly on the claim it is being asked to support.

The first question should be:

What exactly are we claiming changed?

Only then can we ask whether the available evidence fits the claim.

What was actually observed?

Evidence often becomes inflated as we move from observation to interpretation.

Suppose a person's questionnaire score is lower than it was three months ago.

What we know most directly is that the recorded score differs.

We may then infer that the psychological construct represented by the questionnaire has changed. That inference may be well supported, but it still depends on what the measure captures and how confidently the observed difference can be interpreted.

The same logic applies elsewhere.

A person reports feeling calmer. Someone observes fewer arguments. Performance on a cognitive task improves. Physiological activity differs. A neural signal changes.

Each observation can be informative. Broader conclusions require additional justification.

The further a conclusion moves beyond what was observed, the more support is required for that move.

This does not make external observation inherently superior to subjective evidence. Some phenomena are partly defined by experience itself. Pain, meaning or felt anxiety cannot simply be replaced by behavioural or biological proxies without changing the question.

The relevant distinction is between what was observed and what that observation is being used to establish.

Measures do not remove that distinction.

A questionnaire is not identical to the construct it is intended to represent. A cognitive task is not the entirety of a cognitive capacity. A physiological indicator does not carry a fixed psychological meaning. Behavioural observation captures behaviour, not every process that produced it.

A changed measure may therefore support a claim of underlying change without proving it automatically.

The reverse is also possible. A person may change while a particular measure fails to detect that change because it is noisy, insensitive, too narrow or poorly matched to the phenomenon.

Both errors matter: a measure changed, therefore the broader construct definitely changed and the measure did not change, therefore nothing changed go beyond what the observation alone can establish.

Change relative to what?

A claim about within-person change requires comparison.

If someone exercises twice this week, that observation means little by itself.

If they have exercised twice every week for years, there may be no meaningful departure from their prior pattern. If they had not exercised for a year, the same week may tell us something very different.

The reference point matters.

But even "baseline" can sound more secure than it is. A single earlier observation may capture an unusual day rather than a person's ordinary functioning. Mood, sleep, performance, behaviour and physiological measures can all fluctuate.

The better question is:

What prior functioning or trajectory is the new observation being compared with?

Sometimes one earlier observation is informative enough. In other cases, repeated observations are needed to characterise what is normal for that person before the newer pattern can be interpreted.

This is also why between-person difference and within-person change must remain separate.

If one person's score is lower than another person's today, the two people differ.

That does not tell us whether either person changed.

A recorded difference is not automatically meaningful change

Suppose a score moves from 52 to 55.

The recorded values differ.

What remains uncertain is what produced that difference and what conclusion it warrants.

Measurements contain error. People also fluctuate.

Fatigue, stress, recent events, practice effects, context and temporary states can alter an observation without establishing the broader change being claimed.

Methodological approaches to reliable change address one part of this problem by asking whether a score movement is large enough to exceed what would reasonably be expected from measurement unreliability.

But a second question remains.

A statistically reliable difference and a substantively meaningful change are not identical.

A score may move beyond expected measurement error while the difference has little practical or psychological importance. An apparently important difference may also remain difficult to interpret when the measure itself is imprecise.

So we need to distinguish: Is the observed difference sufficiently reliable? from: Does that difference justify the substantive claim being made?

There is no universal threshold that answers both questions across all forms of human change.

The broader standard is more useful: the evidence should give us adequate reason to distinguish the claimed change from ordinary variability, measurement uncertainty and plausible alternative interpretations.

The timescale of the evidence must match the timescale of the claim

Consider one calm response during conflict.

That observation can support: this person responded calmly in this situation.

It cannot by itself establish: this person now generally responds differently during conflict.

Still less: this person's emotional functioning has changed in a lasting way.

Each statement extends the claim further through time.

The evidence has to extend with it.

A temporary state can be supported by evidence appropriate to a temporary state. A recurring pattern requires repeated observations. A claim of stable or lasting change requires evidence capable of supporting stability or durability rather than mere occurrence.

Different outcomes operate on different timescales, so no universal observation period can do this work for every phenomenon.

The principle is simpler:

The timescale of the evidence must match the timescale of the change being claimed.

Without that alignment, the conclusion becomes temporally broader than the evidence.

Different evidence answers different questions

There is a recurring temptation to rank forms of evidence as if one method were intrinsically closer to the truth.

That is usually the wrong comparison.

Self-report can be especially informative when the claim concerns subjective experience, beliefs or perceived states. Its limitations—recall, interpretation, reporting conditions and incomplete introspective access—matter, but they do not make subjective evidence inherently weak. For some questions, subjectivity is part of the phenomenon itself.

Behavioural observation bears closely on behavioural claims. If someone who previously avoided public speaking repeatedly begins giving presentations, that is strong evidence that their behaviour changed. It does not uniquely reveal whether confidence, training, incentives, social pressure or something else produced the change.

Performance measures, physiological measures and neural measures can each provide important evidence at the levels they assess. Their interpretation becomes less direct when they are used to support broader claims about psychological meaning.

A neural difference, for example, can establish a measured neural difference. It does not automatically certify that a psychological change is deeper, more genuine or more important than one supported through another method.

Observer reports can add information about interpersonal or behavioural patterns outside the person's own perspective, while still reflecting the observer's limited contexts and interpretations.

All measurement requires interpretation when it is used to support a broader claim.

That is why there is no universally superior form of evidence independent of the question being asked.

The best evidence is evidence capable of supporting the claim.

Convergence can strengthen evidence

Sometimes several appropriately chosen methods point in the same direction.

A person reports lower anxiety. Their avoidance decreases. Someone close to them notices that they enter situations they previously avoided.

Because these observations arise through different routes, their convergence may strengthen confidence that the broader pattern changed.

This is one reason multimethod evidence can be valuable.

Yet convergence does not mean that every genuine change must appear at every level.

A behavioural change may occur before self-description catches up. A subjective change may be real even when an unrelated biological measure is unchanged. Different aspects of a person can move at different rates or remain partly dissociated.

Additional measurements therefore help only when they actually bear on the claim.

Evidence should converge on the claim, not merely accumulate around the person.

More data is not automatically stronger evidence if the added observations are poorly aligned with the conclusion.

Evidence that change occurred does not tell us why

Suppose someone's anxiety decreases after beginning therapy.

The evidence may strongly support the conclusion that their anxiety changed.

A different question is whether therapy caused the change.

The temporal sequence alone cannot settle that. Other circumstances may have changed during the same period, natural fluctuation may have contributed, and several causes may be involved.

This boundary matters throughout the study of human change.

Evidence that an outcome changed and evidence that identifies its cause are different evidential problems.

This essay is concerned with the first.

The conclusion must not outrun the evidence

Compare these statements:

Her score was lower today.

Her reported anxiety decreased.

Her anxiety has changed.

Her emotional functioning has changed.

She has undergone lasting psychological transformation.

Each conclusion travels farther from the immediate observation.

As the claim becomes broader, more enduring or more abstract, the evidential burden increases.

A score change may require evidence that the measure is reliable and meaningfully related to the construct. A claim of broader psychological change requires evidence extending beyond one indicator. A claim of lasting change requires temporal evidence. A claim of transformation carries further assumptions about the breadth or depth of what changed.

Stronger evidence does not necessarily mean more instruments, more expensive technology or biological measurement.

It means evidence capable of supporting the additional inference.

The strength and breadth of the conclusion should not exceed what the evidence can support.

Human-change language can expand faster than the observations beneath it.

A behavioural difference may be interpreted as identity change. A subjective shift may be treated as evidence of transformation. A short period of improvement may be described as lasting.

Any of those conclusions might eventually be justified.

The evidence has to get there too.

What counts as evidence of change?

The argument can now be stated as a general criterion.

A claim that a person has changed is justified when the evidence is appropriate to the specified outcome, level and timescale and provides sufficient reason to conclude that the observed difference reflects the change being claimed rather than measurement error, ordinary variability or a plausible alternative interpretation.

That criterion does not prescribe one universal method.

It requires alignment.

We need to know what is claimed to have changed, what was actually observed, what comparison makes the difference interpretable and whether the timescale of the evidence matches the timescale of the conclusion.

We also need to recognise how much inference lies between observation and claim.

No single indicator can establish every form of human change because different indicators answer different questions.

Precision rather than certainty

Taking measurement error, ordinary variability and inference seriously can make the evidential standard sound demanding.

It is demanding.

But it does not make human change unknowable.

The purpose of evidential discipline is to make claims about change more precise.

Instead of asking only whether someone has "really changed," we can ask:

What changed?

What evidence supports that conclusion?

Across what time and relative to what prior pattern?

How much does the conclusion extend beyond what was directly observed?

Those questions do not remove uncertainty. They tell us what kind of confidence the evidence justifies.

Once that foundation is established, later questions can be asked more rigorously: how the construct should be measured, what caused the change, what mechanism produced it, and how robust the scientific conclusion is.

Those are downstream problems.

This essay establishes the condition that comes first: before explaining human change, we need to know what our evidence actually allows us to say changed.

Sources and research record6 sources, with findings, strengths and limitations as entered

References

6 sources this piece rests on, as entered in the Library.

  1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

    Theoretical article · Psychological Bulletin, 52(4) · 281–302

    Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

    doi:10.1037/h0040957

  2. Jacobson, N. S., Truax, P. (1991) Clinical significance: A statistical approach to defining meaningful change in psychotherapy research

    Methodological article · Journal of Consulting and Clinical Psychology, 59(1) · 12–19

    Introduces the reliable change index and a criterion for clinical significance, separating the question of whether a score movement exceeds measurement unreliability from the question of whether it is substantively meaningful.

    doi:10.1037/0022-006X.59.1.12

  3. Curran, P. J., Bauer, D. J. (2011) The disaggregation of within-person and between-person effects in longitudinal models of change

    Methodological review · Annual Review of Psychology, 62 · 583–619

    Shows that between-person differences and within-person change are distinct effects that must be separated explicitly, and that conflating them yields conclusions the data do not support.

    doi:10.1146/annurev.psych.093008.100356

  4. Shiffman, S., Stone, A. A., Hufford, M. R. (2008) Ecological momentary assessment

    Review · Annual Review of Clinical Psychology, 4 · 1–32

    Documents how repeated, in-context measurement captures ordinary variability and state fluctuation that single retrospective assessments obscure, bearing directly on what a baseline observation can represent.

    doi:10.1146/annurev.clinpsy.3.022806.091415

  5. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

    Methodological article · Psychological Bulletin, 56(2) · 81–105

    Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

    doi:10.1037/h0046016

  6. Poldrack, R. A. (2006) Can cognitive processes be inferred from neuroimaging data?

    Review · Trends in Cognitive Sciences, 10(2) · 59–63

    Shows that inferring a psychological process from observed brain activity is weak unless the activated region is selective for that process, cautioning against reading psychological meaning directly off a neural measure. Read with the later methodological debate it opened.

    doi:10.1016/j.tics.2005.12.004

Further reading

Behind this page

The claims this essay makes, the evidence behind them, and the limits it accepts.

Evidence status

High confidence

Strongly supported, though resting on synthesis or principle rather than a single decisive body of evidence.

Claims

  1. A measure is an indicator of a construct, not the construct itself

    Established

    What this does not assert: Interpreting a score as evidence about the underlying construct requires justification beyond the score movement.

    1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

      Theoretical article · Psychological Bulletin, 52(4) · 281–302

      Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

      doi:10.1037/h0040957

  2. No measurement modality is universally superior

    High confidence

    What this does not assert: Self-report, behaviour, performance, physiology, neural measures and observer report answer different questions; the appropriate evidence depends on the claim.

    1. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

      Methodological article · Psychological Bulletin, 56(2) · 81–105

      Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

      doi:10.1037/h0046016

    2. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

      Theoretical article · Psychological Bulletin, 52(4) · 281–302

      Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

      doi:10.1037/h0040957

    3. Poldrack, R. A. (2006) Can cognitive processes be inferred from neuroimaging data?

      Review · Trends in Cognitive Sciences, 10(2) · 59–63

      Shows that inferring a psychological process from observed brain activity is weak unless the activated region is selective for that process, cautioning against reading psychological meaning directly off a neural measure. Read with the later methodological debate it opened.

      doi:10.1016/j.tics.2005.12.004

  3. Evidence that change occurred does not identify its cause

    High confidence

    What this does not assert: Temporal sequence alone cannot establish causation; that is a separate evidential problem, treated downstream.

  4. The timescale of the evidence must match the timescale of the claim

    Canonical inference

    What this does not assert: A single occasion supports an occasion-level claim; recurring patterns and durability require evidence extending across the relevant period.

  5. The broader the conclusion, the greater the evidential burden

    Canonical inference

    What this does not assert: Stronger evidence means evidence capable of supporting the additional inference, not more instruments or more expensive technology.

  6. Alignment among claim, observation, level, comparison and timescale is the criterion for evidence of change

    Canonical synthesis

    What this does not assert: The integration of construct validity, reliable change, within-person reasoning, temporal measurement, multimethod convergence and inferential caution into one criterion is The Shifting Point's synthesis rather than a definition belonging to a single research tradition.

  7. Measurement error can produce differences that do not reflect change

    Established

    What this does not assert: Reliable-change methods ask whether a movement exceeds what unreliability alone would produce.

    1. Jacobson, N. S., Truax, P. (1991) Clinical significance: A statistical approach to defining meaningful change in psychotherapy research

      Methodological article · Journal of Consulting and Clinical Psychology, 59(1) · 12–19

      Introduces the reliable change index and a criterion for clinical significance, separating the question of whether a score movement exceeds measurement unreliability from the question of whether it is substantively meaningful.

      doi:10.1037/0022-006X.59.1.12

  8. Reliable change and substantively meaningful change are distinct questions

    Established

    What this does not assert: A movement can exceed measurement error and still carry little practical or psychological importance.

    1. Jacobson, N. S., Truax, P. (1991) Clinical significance: A statistical approach to defining meaningful change in psychotherapy research

      Methodological article · Journal of Consulting and Clinical Psychology, 59(1) · 12–19

      Introduces the reliable change index and a criterion for clinical significance, separating the question of whether a score movement exceeds measurement unreliability from the question of whether it is substantively meaningful.

      doi:10.1037/0022-006X.59.1.12

  9. Between-person difference does not establish within-person change

    Established

    What this does not assert: The two effects must be separated explicitly; conflating them produces conclusions the data do not support.

    1. Curran, P. J., Bauer, D. J. (2011) The disaggregation of within-person and between-person effects in longitudinal models of change

      Methodological review · Annual Review of Psychology, 62 · 583–619

      Shows that between-person differences and within-person change are distinct effects that must be separated explicitly, and that conflating them yields conclusions the data do not support.

      doi:10.1146/annurev.psych.093008.100356

  10. Within-person change requires an appropriate temporal comparison

    Established

    What this does not assert: Longitudinal models require repeated observation of the same person rather than a single cross-sectional contrast.

    1. Curran, P. J., Bauer, D. J. (2011) The disaggregation of within-person and between-person effects in longitudinal models of change

      Methodological review · Annual Review of Psychology, 62 · 583–619

      Shows that between-person differences and within-person change are distinct effects that must be separated explicitly, and that conflating them yields conclusions the data do not support.

      doi:10.1146/annurev.psych.093008.100356

  11. Ordinary variability can make a single baseline unrepresentative

    Established

    What this does not assert: Repeated, in-context measurement reveals state fluctuation that one earlier observation can obscure.

    1. Shiffman, S., Stone, A. A., Hufford, M. R. (2008) Ecological momentary assessment

      Review · Annual Review of Clinical Psychology, 4 · 1–32

      Documents how repeated, in-context measurement captures ordinary variability and state fluctuation that single retrospective assessments obscure, bearing directly on what a baseline observation can represent.

      doi:10.1146/annurev.clinpsy.3.022806.091415

  12. Convergence across differing methods can strengthen an inference

    Established

    What this does not assert: Convergence is informative because the methods carry different sources of error, not because more measurements were taken.

    1. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

      Methodological article · Psychological Bulletin, 56(2) · 81–105

      Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

      doi:10.1037/h0046016

  13. Method-specific variance limits what agreement between measures shows

    Established

    What this does not assert: Measures sharing a method can agree for reasons belonging to the method rather than the construct.

    1. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

      Methodological article · Psychological Bulletin, 56(2) · 81–105

      Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

      doi:10.1037/h0046016

  14. A neural difference does not certify a psychological interpretation

    Established

    What this does not assert: Reverse inference is weak unless the neural measure is selective for the process being claimed; this is a caution about inference, not a dismissal of neural evidence.

    1. Poldrack, R. A. (2006) Can cognitive processes be inferred from neuroimaging data?

      Review · Trends in Cognitive Sciences, 10(2) · 59–63

      Shows that inferring a psychological process from observed brain activity is weak unless the activated region is selective for that process, cautioning against reading psychological meaning directly off a neural measure. Read with the later methodological debate it opened.

      doi:10.1016/j.tics.2005.12.004

Sources

  1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

    Theoretical article · Psychological Bulletin, 52(4) · 281–302

    Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

    doi:10.1037/h0040957

  2. Jacobson, N. S., Truax, P. (1991) Clinical significance: A statistical approach to defining meaningful change in psychotherapy research

    Methodological article · Journal of Consulting and Clinical Psychology, 59(1) · 12–19

    Introduces the reliable change index and a criterion for clinical significance, separating the question of whether a score movement exceeds measurement unreliability from the question of whether it is substantively meaningful.

    doi:10.1037/0022-006X.59.1.12

  3. Curran, P. J., Bauer, D. J. (2011) The disaggregation of within-person and between-person effects in longitudinal models of change

    Methodological review · Annual Review of Psychology, 62 · 583–619

    Shows that between-person differences and within-person change are distinct effects that must be separated explicitly, and that conflating them yields conclusions the data do not support.

    doi:10.1146/annurev.psych.093008.100356

  4. Shiffman, S., Stone, A. A., Hufford, M. R. (2008) Ecological momentary assessment

    Review · Annual Review of Clinical Psychology, 4 · 1–32

    Documents how repeated, in-context measurement captures ordinary variability and state fluctuation that single retrospective assessments obscure, bearing directly on what a baseline observation can represent.

    doi:10.1146/annurev.clinpsy.3.022806.091415

  5. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

    Methodological article · Psychological Bulletin, 56(2) · 81–105

    Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

    doi:10.1037/h0046016

  6. Poldrack, R. A. (2006) Can cognitive processes be inferred from neuroimaging data?

    Review · Trends in Cognitive Sciences, 10(2) · 59–63

    Shows that inferring a psychological process from observed brain activity is weak unless the activated region is selective for that process, cautioning against reading psychological meaning directly off a neural measure. Read with the later methodological debate it opened.

    doi:10.1016/j.tics.2005.12.004

Read this first

This piece assumes them.

What this opens up

What becomes readable once you have this.

Where to go from here

Next published piece

Measurement and Operational Definition

Researchers can measure constructs such as stress, motivation or identity only by deciding how those constructs will become observable. The resulting measure can provide evidence about the phenomenon—but it should never be mistaken for the phenomenon itself.

Continue through the Library →See where this sits in the graph →

Back to the Library →