Concepts · Evidence and Explanation

Measurement and Operational Definition

Researchers can measure constructs such as stress, motivation or identity only by deciding how those constructs will become observable. The resulting measure can provide evidence about the phenomenon—but it should never be mistaken for the phenomenon itself.

By Yona Ole Lobulu ·

Concept6 min readFoundationalD2.2

Topic
Evidence and Explanation
Read first
One piece should be read before this one
Reading time
About 6 minutes of reading
Difficulty
Advanced reading: this page assumes a fair amount of earlier reading

The question

How do measures represent the phenomena being claimed?

Definition

Measurement makes a construct empirically accessible through operational definitions and observable indicators, but conclusions are warranted only to the extent that those indicators reliably and validly represent the aspect of the construct being claimed.

How do researchers measure something like motivation?

They might ask how motivated someone feels, measure how long they persist on a difficult task, or observe which option they choose when one requires more effort.

Each can provide evidence about motivation.

None is motivation itself.

This distinction sits at the heart of measurement. Many phenomena relevant to human change cannot be observed directly in their entirety. Researchers therefore need a way to connect abstract concepts to things that can actually be observed.

Construct, operational definition and indicator

A construct is a theoretical concept used to describe or organise a phenomenon researchers want to study. Stress, anxiety, motivation, identity and learning can all function as constructs.

Some constructs are latent: they are inferred through observable indicators rather than directly observed as complete variables.

To investigate a construct empirically, researchers must specify how it will be represented or measured. This is an operational definition.

If a study concerns motivation, for example, researchers might operationalise it through self-reported motivation, persistence on a difficult task or willingness to choose a demanding option.

The operational definition specifies what will count as observable evidence of the construct in that investigation. The resulting response, behaviour, score or recorded value can then function as an indicator.

A useful way to keep these distinctions visible is:

construct → operational definition → indicator → measurement → inference

This is a conceptual map rather than a universal technical taxonomy. Its purpose is to show that several steps can separate the phenomenon researchers want to understand from the conclusion eventually drawn about it.

Operationalisation creates empirical access to a construct.

It does not make the operation identical to the construct.

A measure is not the construct

Suppose anxiety is assessed with a questionnaire.

The questionnaire produces responses and perhaps a total score. That score can provide evidence about anxiety, but the score is not anxiety itself.

The same distinction applies elsewhere.

Performance on a cognitive task is not the entirety of the capacity the task is intended to assess. A physiological value is not automatically equivalent to a psychological state. An observed behaviour records what happened, not every process that produced it.

Measurement therefore involves representation and inference.

Researchers observe something measurable and use it as evidence about something they want to understand.

A measure represents a construct; it does not become the construct.

That inference may be supported by extensive theory and evidence. But treating the measure and the construct as interchangeable removes the very relationship that measurement needs to establish.

The same label can hide different measurements

Consider three studies of stress.

One uses a perceived-stress questionnaire. Another records exposure to specified demanding events. A third uses cortisol or cardiovascular reactivity as physiological indicators related to stress responses.

All three may provide evidence relevant to stress.

But they are not necessarily measuring the same aspect of it.

One focuses on perceived experience, another on environmental exposure, another on physiological response.

Shared terminology does not guarantee shared measurement.

This matters when research findings are compared. Two studies can use the same conceptual label while operationalising it differently.

That does not mean one operationalisation must be wrong. Complex constructs can have multiple legitimate manifestations, and different measures can capture different aspects of them.

Nor does it mean that every operationalisation is equally good.

The relevant question is whether the chosen indicators adequately support the particular interpretation being made.

Reliability and validity are different questions

Two concepts help evaluate that relationship: reliability and validity.

Reliability concerns the consistency or precision of measurement under relevant conditions. A highly unstable measurement procedure makes it harder to know what an observed value represents.

But consistency alone does not establish that the intended construct has been represented adequately.

A measure can produce highly consistent results while the interpretation made from those results remains poorly supported.

Validity concerns whether evidence and theory support the interpretation being made from a measurement for its intended use.

This makes validity more than a permanent stamp attached to an instrument.

Evidence supporting one interpretation in one context does not automatically justify every other interpretation, population or use.

The central question is:

What does this measurement legitimately allow us to infer?

Reliability matters to that answer.

But:

Reliability cannot substitute for validity.

Measurement has limits

Measurements are not perfectly transparent windows onto phenomena.

One reason is measurement error. Recorded values can contain uncertainty arising from the instrument, items, observer, occasion or other features of the measurement process.

Measures also differ in their sensitivity to relevant variation. If a measure cannot register the kind of difference or change being investigated, genuine variation may go undetected.

Indicators can also differ in how selectively they bear on a particular interpretation. Elevated heart rate, for example, can occur during fear, exercise, excitement, illness or heat. A change in heart rate therefore does not uniquely identify one psychological state.

This is the broad issue of measurement specificity intended here, rather than the technical definition of specificity used in diagnostic testing.

These limitations do not make measurement unreliable by definition.

They constrain what a particular observation can tell us.

Converging measurements can strengthen an interpretation

Sometimes different operationalisations provide compatible evidence about the same construct.

If people report greater motivation while also persisting longer on relevant tasks, the convergence may strengthen the interpretation because it depends less heavily on one operational choice.

This is one reason multiple methods can be useful in construct validation.

But convergence is not a universal requirement. A well-designed measure can sometimes provide adequate evidence for the question being asked, and different indicators need not move together if they represent different aspects of a construct.

Convergence strengthens an interpretation when the measures genuinely bear on the same claim.

It does not make every measure interchangeable.

What can a measurement legitimately tell us?

Measurement makes abstract constructs scientifically investigable by connecting them to observable indicators.

But each link matters.

Researchers decide how a construct will be operationalised. That choice determines which aspects become available for observation. The resulting measurements then provide evidence from which researchers infer something about the construct.

This gives the central criterion:

Measurement makes a construct empirically accessible through operational definitions and observable indicators, but conclusions are warranted only to the extent that those indicators reliably and validly represent the aspect of the construct being claimed.

Operational definitions therefore do two things at once.

They make empirical investigation possible.

And they establish boundaries around what the resulting evidence can support.

Ask how it was operationalised

When two studies say they measured stress, motivation, identity or learning, the shared label is not enough.

Ask:

How was the construct operationalised?

That question reveals what was actually observed, which aspect of the construct became measurable and what conclusions the evidence can reasonably support.

Scientific measurement does not give us direct access to every phenomenon in its entirety.

It gives us structured ways of representing phenomena through observable evidence.

Understanding that representation is part of understanding what the evidence means.

Sources and research record6 sources, with findings, strengths and limitations as entered

References

6 sources this piece rests on, as entered in the Library.

  1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

    Theoretical article · Psychological Bulletin, 52(4) · 281–302

    Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

    doi:10.1037/h0040957

  2. Messick, S. (1995) Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

    Theoretical article · American Psychologist, 50(9) · 741–749

    Reframes validity as the degree to which evidence and theory support the interpretation and use of scores, rather than as a permanent property of an instrument.

    doi:10.1037/0003-066X.50.9.741

  3. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

    Standards · American Educational Research Association, Washington, DC

    The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

    Read the source

  4. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

    Methodological article · Psychological Bulletin, 56(2) · 81–105

    Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

    doi:10.1037/h0046016

  5. Flake, J. K., Fried, E. I. (2020) Measurement schmeasurement: Questionable measurement practices and how to avoid them

    Methodological article · Advances in Methods and Practices in Psychological Science, 3(4) · 456–465

    Documents how unreported or unjustified operational choices make it unclear what a study measured, and argues that construct definition and operationalisation must be stated explicitly before results can be interpreted.

    doi:10.1177/2515245920952393

  6. Fried, E. I. (2017) The 52 symptoms of major depression: Lack of content overlap among seven common depression scales

    Empirical article · Journal of Affective Disorders, 208 · 191–197

    Shows that widely used scales sharing one construct label cover substantially different symptom content, demonstrating empirically that a shared term does not guarantee shared measurement.

    doi:10.1016/j.jad.2016.10.019

Further reading

Behind this page

The claims this concept makes, the evidence behind them, and the limits it accepts.

Evidence status

High confidence

Strongly supported, though resting on synthesis or principle rather than a single decisive body of evidence.

Claims

  1. A measure is an indicator of a construct, not the construct itself

    Established

    What this does not assert: Interpreting a score as evidence about the construct requires justification through a wider network of theory and evidence.

    1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

      Theoretical article · Psychological Bulletin, 52(4) · 281–302

      Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

      doi:10.1037/h0040957

  2. Measurement error and limited sensitivity constrain what an observation establishes

    Established

    What this does not assert: Recorded values carry uncertainty from the instrument, items, observer or occasion, and an insensitive measure can miss genuine variation.

    1. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

      Standards · American Educational Research Association, Washington, DC

      The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

      Read the source

  3. Convergence across differing methods can strengthen a construct interpretation

    Established

    What this does not assert: Convergence is informative because the methods carry different sources of error, and it is not a universal requirement.

    1. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

      Methodological article · Psychological Bulletin, 56(2) · 81–105

      Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

      doi:10.1037/h0046016

  4. An indicator can bear on several states at once

    High confidence

    What this does not assert: Elevated heart rate, for example, does not uniquely identify one psychological state; this is measurement specificity in the broad sense, not the diagnostic-testing definition.

    1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

      Theoretical article · Psychological Bulletin, 52(4) · 281–302

      Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

      doi:10.1037/h0040957

  5. Different operationalisations can capture different legitimate aspects of one construct

    Canonical inference

    What this does not assert: Plurality does not make every operationalisation equally adequate for the interpretation being made.

  6. Construct, operational definition, indicator, measurement and inference are distinct steps

    Canonical synthesis

    What this does not assert: This five-step map is The Shifting Point's explanatory schema for keeping the relationships visible, not a universally codified taxonomy in measurement science.

  7. Latent constructs are inferred through observable indicators

    Established

    What this does not assert: Being inferred rather than directly observed does not make a construct fictitious; it makes its measurement an inferential matter.

    1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

      Theoretical article · Psychological Bulletin, 52(4) · 281–302

      Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

      doi:10.1037/h0040957

  8. An operational definition specifies how a construct becomes empirically accessible

    Established

    What this does not assert: Operationalisation creates access to the construct; it does not exhaust the construct's meaning.

    1. Flake, J. K., Fried, E. I. (2020) Measurement schmeasurement: Questionable measurement practices and how to avoid them

      Methodological article · Advances in Methods and Practices in Psychological Science, 3(4) · 456–465

      Documents how unreported or unjustified operational choices make it unclear what a study measured, and argues that construct definition and operationalisation must be stated explicitly before results can be interpreted.

      doi:10.1177/2515245920952393

  9. Unstated operational choices make results difficult to interpret

    Established

    What this does not assert: Construct definition and operationalisation must be reported before a finding can be understood as evidence about the named construct.

    1. Flake, J. K., Fried, E. I. (2020) Measurement schmeasurement: Questionable measurement practices and how to avoid them

      Methodological article · Advances in Methods and Practices in Psychological Science, 3(4) · 456–465

      Documents how unreported or unjustified operational choices make it unclear what a study measured, and argues that construct definition and operationalisation must be stated explicitly before results can be interpreted.

      doi:10.1177/2515245920952393

  10. The same construct label can conceal substantially different measurements

    Established

    What this does not assert: Scales sharing a label can differ markedly in content, so findings using the same term are not automatically comparable.

    1. Fried, E. I. (2017) The 52 symptoms of major depression: Lack of content overlap among seven common depression scales

      Empirical article · Journal of Affective Disorders, 208 · 191–197

      Shows that widely used scales sharing one construct label cover substantially different symptom content, demonstrating empirically that a shared term does not guarantee shared measurement.

      doi:10.1016/j.jad.2016.10.019

    2. Flake, J. K., Fried, E. I. (2020) Measurement schmeasurement: Questionable measurement practices and how to avoid them

      Methodological article · Advances in Methods and Practices in Psychological Science, 3(4) · 456–465

      Documents how unreported or unjustified operational choices make it unclear what a study measured, and argues that construct definition and operationalisation must be stated explicitly before results can be interpreted.

      doi:10.1177/2515245920952393

  11. Reliability concerns the consistency or precision of measurement

    Established

    What this does not assert: Reliability is estimated under specified conditions and does not by itself justify an interpretation.

    1. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

      Standards · American Educational Research Association, Washington, DC

      The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

      Read the source

  12. Validity concerns whether evidence and theory support the intended interpretation and use

    Established

    What this does not assert: Validity is a property of interpretations for a purpose, not a permanent stamp attached to an instrument.

    1. Messick, S. (1995) Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

      Theoretical article · American Psychologist, 50(9) · 741–749

      Reframes validity as the degree to which evidence and theory support the interpretation and use of scores, rather than as a permanent property of an instrument.

      doi:10.1037/0003-066X.50.9.741

    2. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

      Standards · American Educational Research Association, Washington, DC

      The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

      Read the source

  13. Reliability cannot substitute for validity

    Established

    What this does not assert: A highly consistent measurement can still support a poorly justified construct interpretation.

    1. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

      Standards · American Educational Research Association, Washington, DC

      The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

      Read the source

    2. Messick, S. (1995) Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

      Theoretical article · American Psychologist, 50(9) · 741–749

      Reframes validity as the degree to which evidence and theory support the interpretation and use of scores, rather than as a permanent property of an instrument.

      doi:10.1037/0003-066X.50.9.741

  14. Validity evidence does not transfer automatically across contexts

    Established

    What this does not assert: Support for one interpretation, population or use does not establish every other interpretation, population or use.

    1. Messick, S. (1995) Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

      Theoretical article · American Psychologist, 50(9) · 741–749

      Reframes validity as the degree to which evidence and theory support the interpretation and use of scores, rather than as a permanent property of an instrument.

      doi:10.1037/0003-066X.50.9.741

    2. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

      Standards · American Educational Research Association, Washington, DC

      The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

      Read the source

Sources

  1. Cronbach, L. J., Meehl, P. E. (1955) Construct validity in psychological tests

    Theoretical article · Psychological Bulletin, 52(4) · 281–302

    Establishes that a test score is an indicator of a construct rather than the construct itself, and that interpreting the score requires justification through a wider network of theory and evidence.

    doi:10.1037/h0040957

  2. Messick, S. (1995) Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

    Theoretical article · American Psychologist, 50(9) · 741–749

    Reframes validity as the degree to which evidence and theory support the interpretation and use of scores, rather than as a permanent property of an instrument.

    doi:10.1037/0003-066X.50.9.741

  3. American Educational Research Association, American Psychological Association, National Council on Measurement in Education (2014) Standards for Educational and Psychological Testing

    Standards · American Educational Research Association, Washington, DC

    The standards-based framework for reliability and validity, treating reliability as precision of measurement and validity as the evidence supporting a specific intended interpretation and use.

    Read the source

  4. Campbell, D. T., Fiske, D. W. (1959) Convergent and discriminant validation by the multitrait-multimethod matrix

    Methodological article · Psychological Bulletin, 56(2) · 81–105

    Establishes convergence across methods as evidence for a construct, and shows that method-specific variance means agreement between measures is informative only when the methods differ in their sources of error.

    doi:10.1037/h0046016

  5. Flake, J. K., Fried, E. I. (2020) Measurement schmeasurement: Questionable measurement practices and how to avoid them

    Methodological article · Advances in Methods and Practices in Psychological Science, 3(4) · 456–465

    Documents how unreported or unjustified operational choices make it unclear what a study measured, and argues that construct definition and operationalisation must be stated explicitly before results can be interpreted.

    doi:10.1177/2515245920952393

  6. Fried, E. I. (2017) The 52 symptoms of major depression: Lack of content overlap among seven common depression scales

    Empirical article · Journal of Affective Disorders, 208 · 191–197

    Shows that widely used scales sharing one construct label cover substantially different symptom content, demonstrating empirically that a shared term does not guarantee shared measurement.

    doi:10.1016/j.jad.2016.10.019

What this opens up

What becomes readable once you have this.

Where to go from here

Next published piece

Correlation, Prediction, Causation and Mechanism

A variable can be associated with an outcome, help predict it, causally affect it, or operate through a particular mechanism. These are different scientific claims, and establishing one does not automatically establish the others.

Continue through the Library →See where this sits in the graph →

Back to the Library →