Concepts · Learning and Memory

Operant Learning

Operant learning occurs when experience with relationships between behaviour and its consequences changes future behaviour. What matters is not simply what happens after an action, but the contingency connecting behaviour to what follows.

By Yona Ole Lobulu ·

Concept9 min readFoundationalD5.4

Topic
Learning and Memory
Read first
One piece should be read before this one
Reading time
About 9 minutes of reading
Difficulty
Late reading: this page sits at the far end of the Library, after many other pieces

The question

What relationship existed between the behaviour and its consequences, and how did experience with that relationship alter future behaviour?

Definition

Operant learning is a form of associative learning in which experience with relationships between behaviour and its consequences changes the likelihood, selection or organisation of future behaviour.

A behaviour occurs. Something happens afterward. Later, the behaviour changes.

It is tempting to explain the sequence by saying that the consequence reinforced or punished the action. Sometimes that may be true. But the sequence alone does not establish it.

The event may have occurred by coincidence. It may have had little effect on later behaviour. The same outcome may have been available whether or not the action occurred. Or behaviour may have changed because other conditions changed at the same time.

The central question in operant learning is therefore not simply: what happened after the behaviour?

It is: what relationship existed between the behaviour and its consequences, and how did experience with that relationship alter future behaviour?

Operant learning is a form of associative learning in which experience with relationships between behaviour and its consequences changes the likelihood, selection or organisation of future behaviour.

The closely related term instrumental learning is often used for the same broad territory, although research traditions differ somewhat in emphasis. The shared phenomenon is learning through behaviour–consequence relationships.

What operant learning means

The simplest operant structure looks like this: behaviour or action → consequence or outcome.

But the arrow can be misleading if it is interpreted to mean: something happened after the action, therefore it changed the action.

At the broadest level, operant learning concerns functional relationships between behaviour and consequences. In some cases, stronger experimental evidence shows that organisms learn something more specific: which particular actions lead to which particular outcomes.

Those are different levels of inference.

If behaviour changes after an outcome, we have evidence that experience mattered. We do not automatically know that the learner acquired a detailed representation of an action–outcome relation, nor that one particular mechanism produced the change.

Operant learning concerns how experience with behaviour–consequence relationships changes future behaviour.
Figure — Operant learning as behaviour–consequence learning

EXPERIENCE — behaviour / action → consequence / outcome, under a particular contingency

↓ LEARNING — sensitivity to the behaviour–consequence relationship

↓ interacting with context · opportunity · current state · competing behaviour

↓ FUTURE BEHAVIOUR — likelihood · selection · organisation

Operant learning concerns how experience with behaviour–consequence relationships changes later behaviour. An event merely following an action is not enough to establish the learning relation. This is an analytic representation, not a claim that operant learning proceeds through a fixed sequence of internal stages.

What matters is the contingency

Timing can matter in operant learning. A consequence occurring soon after behaviour may be easier to relate to that behaviour than one arriving much later.

But temporal proximity is not the same as contingency.

Contiguity asks whether behaviour and consequence occurred close together. Contingency asks whether the probability or availability of the outcome differed depending on whether the behaviour occurred.

Suppose performing an action is followed by a particular outcome. If that outcome is much more likely after the action than when the action is absent, there is a meaningful action–outcome contingency.

Now change the arrangement. The same outcome begins occurring independently of the action.

The action and outcome may still sometimes appear close together, but their relationship has weakened because the outcome no longer depends as strongly on the action.

Researchers study this through contingency degradation: an outcome that was initially contingent on a response is also made available independently of that response. Instrumental behaviour can then decline or change selectively.

The important manipulation is not simply a change in the outcome itself. The relation between action and outcome has been weakened.

What follows behaviour matters differently when its occurrence depends on the behaviour.

This is one of the clearest demonstrations that operant learning cannot be reduced to: behaviour happened, then consequence happened.

Temporal proximity may facilitate learning, but immediacy alone does not define the operant relation.

Contingency should not be confused with causation either. A statistical relationship between an action and an outcome can make the action informative about what follows without settling every causal question about why that outcome occurred.

Learning specific consequences

Operant learning can contain more structure than a generic increase or decrease in one response.

An organism may learn that action A produces outcome X, and that action B produces outcome Y.

Later behaviour can remain sensitive to which outcome has historically followed which action.

One way researchers test this is by changing the value of an outcome after the action–outcome relation has already been learned. If performance of the corresponding action changes, that provides evidence that behaviour depended, at least partly, on information about the specific outcome associated with it.

This matters because a purely generic response-strengthening account would describe the learning only as: this response became more likely.

Outcome-sensitive behaviour shows that some instrumental learning preserves more information than that.

The conclusion should remain limited: some operant learning concerns which specific outcomes different actions produce.

That does not mean all operant behaviour is governed by detailed action–outcome representations. Nor does it tell us whether later behaviour should be classified as goal-directed or habitual. Those questions require additional distinctions developed downstream.

A consequence is not reinforcement

The word consequence should remain neutral.

Here, it refers broadly to an event or change in conditions that occurs in relation to behaviour and may become relevant to later behaviour. Calling something a consequence does not tell us what effect, if any, it had.

A consequence may be followed by more of a behaviour, less of it, little meaningful change, or changes that appear only under particular conditions.

A consequence is not reinforcement simply because it comes after behaviour.

This distinction is fundamental.

The same caution applies to the everyday word reward. If behaviour increases after an outcome, that does not by itself establish that pleasure caused the change. Operant learning can depend on contingency, prior learning, outcome properties and current conditions without being reducible to a simple rule that organisms repeat whatever feels good.

Reinforcement, Reward and Pleasure separates those concepts explicitly.

The opposite inference is equally risky. An event can be unpleasant or deliberately intended as punishment without its label telling us what it actually did to future behaviour.

Whether a consequence reduces later responding is an empirical question about behavioural change, not merely intention or subjective unpleasantness.

The full analysis belongs to What Punishment Changes.

Operant learning changes the organisation of behaviour

Operant learning is often described in terms of behaviour becoming more or less frequent.

That is important, but incomplete.

Learning through consequences can change whether behaviour occurs, which action is selected, when an action occurs, how responding is distributed across alternatives, and which conditions make one behaviour more likely than another.

Operant learning can change not only how much behaviour occurs, but which behaviour occurs under which conditions.

This is why behavioural probability should not be imagined as a fixed property of a person.

The expression of prior learning depends on current circumstances. Opportunity matters. So do available alternatives, current state, outcome availability and the environment in which the action occurs.

An action can produce a particular outcome only when surrounding conditions make that relation possible.

A temporary reduction in behaviour therefore does not necessarily mean earlier consequence learning has disappeared. Likewise, an increase in behaviour may reflect a changed opportunity or context rather than entirely new learning.

This Concept inherits that distinction from What Is Learning? without redefining it: learning contributes to current behaviour, but performance is shaped by more than learning alone.

Learning through consequences does not tell us how deliberate behaviour will be

Operant learning is sometimes pictured as explicit reasoning: if I do X, Y will happen, so I will choose X.

People can certainly reason this way.

But explicit deliberation is not the definition of operant learning.

Researchers can measure what people consciously report about an action–outcome relationship and separately measure whether their behaviour is sensitive to that relationship. These are different sources of evidence, and their correspondence varies across tasks and methods.

That means neither extreme is justified.

Operant learning should not be defined as necessarily conscious, but it should not be characterized as inherently unconscious either.

Learning through consequences also does not determine how deliberate every later instance of behaviour will be.

Behaviour learned through consequences is not automatically habitual behaviour.

Operant learning can contribute to behaviour that remains sensitive to goals and current outcomes, and it can also form part of the history from which more habitual control later develops. This Concept does not explain that transition.

That belongs to What Makes Behaviour Habitual? and Goal-Directed and Habitual Control.

Operant and classical learning differ by the relation being learned

Classical Conditioning established classical conditioning around predictive relationships among events: cue or event → outcome.

Operant learning centres on a different relation: action or behaviour → outcome.

The distinction is sometimes summarized as passive versus active learning. That is misleading.

Pavlovian learning can produce active approach, orientation and defensive behaviour. What distinguishes the two concepts is not whether the organism moves or acts.

Classical and operant learning differ primarily in the relationship being learned, not in whether the organism is active.

The distinction also does not imply isolated systems. Pavlovian cues can influence operantly learned actions, and the same behaviour can reflect several learning histories at once.

The two Concepts are therefore parallel analytic branches of the broader associative-learning architecture established in Associative Learning.

What consequence-dependent learning does—and does not—explain

Operant learning has a reciprocal temporal structure.

Behaviour can alter what happens next. Within the opportunities and constraints of the environment, those changes can affect which consequences occur or become available. Experience with those consequences can then alter later behaviour.

In that limited sense: behaviour changes conditions, and changed conditions influence later behaviour. The general causal principle is developed separately in Reciprocal Causation.

But this should not be inflated into a complete theory of behavioural feedback or human action.

Operant learning is one causal process among many. Complex behaviour can also depend on bodily state, perception, goals, social relationships, environmental constraints, prior learning and wider causal systems.

Consequence history matters without becoming the whole explanation.

This is also why the concept should not be turned into a simple behaviour-management formula. Knowing that consequences can participate in learning does not tell us, by itself, whether a particular consequence reinforces, punishes, motivates, suppresses, creates habit or produces durable change.

Those are separate empirical questions.

From consequences to reinforcement and punishment

Operant learning provides the general mechanism: experience with relationships between behaviour and consequences changes future behaviour.

The crucial element is the relation.

Something merely occurring after behaviour is not enough. Contingency matters. Some operant learning preserves information about specific outcomes. Learned behaviour remains dependent on the conditions under which it is expressed.

And the category consequence remains deliberately broader than the concepts built downstream from it.

The next question is therefore not whether consequences matter. It is what different behavioural effects of consequences should be called.

When a consequence is associated with increased future responding, how should that relationship be understood—and why is it not the same thing as reward or pleasure? That is Reinforcement, Reward and Pleasure.

When consequences are associated with reduced responding, a different set of distinctions becomes necessary. That is What Punishment Changes.

Sources and research record9 sources, with findings, strengths and limitations as entered

References

9 sources this piece rests on, as entered in the Library.

  1. Thorndike, E. L. (1911) Animal Intelligence: Experimental Studies

    Book · Macmillan, New York

    The founding experimental treatment of learning through consequences of action, and the origin of the functional analysis of effect.

    Read the source

  2. Skinner, B. F. (1938) The Behavior of Organisms: An Experimental Analysis

    Book · Appleton-Century-Crofts, New York

    The systematic statement of reinforcement as a functional relation between behaviour and consequence rather than a property of the consequence itself.

    Read the source

  3. Hammond, L. J. (1980) The Effect of Contingency upon the Appetitive Conditioning of Free-Operant Behavior

    Experimental article · Journal of the Experimental Analysis of Behavior, 34(3) · 297–304

    The classic contingency-degradation demonstration: responding declines when the same outcome also becomes available independently of the action.

    doi:10.1901/jeab.1980.34-297

  4. Adams, C. D., Dickinson, A. (1981) Instrumental Responding Following Reinforcer Devaluation

    Experimental article · Quarterly Journal of Experimental Psychology, 33B(2) · 109–121

    Devaluing an outcome after training changes the instrumental action that produced it — evidence that some instrumental learning preserves outcome information.

    doi:10.1080/14640748108400816

  5. Colwill, R. M., Rescorla, R. A. (1985) Postconditioning Devaluation of a Reinforcer Affects Instrumental Responding

    Experimental article · Journal of Experimental Psychology: Animal Behavior Processes, 11(1) · 120–132

    Outcome devaluation after training selectively changes responding, showing that different actions can carry information about different outcomes.

    doi:10.1037/0097-7403.11.1.120

  6. Dickinson, A. (1985) Actions and Habits: The Development of Behavioural Autonomy

    Theoretical review · Philosophical Transactions of the Royal Society B, 308(1135) · 67–78

    Why learning through consequences does not by itself make behaviour habitual, and how outcome sensitivity can decline with extended training.

    doi:10.1098/rstb.1985.0010

  7. Herrnstein, R. J. (1961) Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement

    Experimental article · Journal of the Experimental Analysis of Behavior, 4(3) · 267–272

    Behaviour is allocated across alternatives in relation to their consequences — operant learning organises choice, not only response frequency.

    doi:10.1901/jeab.1961.4-267

  8. Shanks, D. R., Dickinson, A. (1991) Instrumental Judgment and Performance under Variations in Action–Outcome Contingency and Contiguity

    Experimental article · Memory & Cognition, 19(4) · 353–360

    Explicit contingency judgments and instrumental performance are separable measures that respond differently to contingency and delay.

    doi:10.3758/BF03197143

  9. Estes, W. K. (1948) Discriminative Conditioning II: Effects of a Pavlovian Conditioned Stimulus upon a Subsequently Established Operant Response

    Experimental article · Journal of Experimental Psychology, 38(2) · 173–177

    Pavlovian cues can modulate instrumentally established behaviour — the classic demonstration that the two processes interact within one behaviour.

    doi:10.1037/h0057525

Further reading

Behind this page

The claims this concept makes, the evidence behind them, and the limits it accepts.

Evidence status

Established

Well supported by a substantial, converging empirical literature.

Claims

  1. Degrading the contingency between an action and its outcome reduces instrumental responding even when the outcome continues to occur

    Established

    What this does not assert: Demonstrated in free-operant appetitive designs; magnitude varies with schedule.

    1. Hammond, L. J. (1980) The Effect of Contingency upon the Appetitive Conditioning of Free-Operant Behavior

      Experimental article · Journal of the Experimental Analysis of Behavior, 34(3) · 297–304

      The classic contingency-degradation demonstration: responding declines when the same outcome also becomes available independently of the action.

      doi:10.1901/jeab.1980.34-297

  2. Operant and instrumental learning designate substantially overlapping scientific territory

    High confidence

    What this does not assert: Research traditions differ in emphasis and preferred measures.

    1. Skinner, B. F. (1938) The Behavior of Organisms: An Experimental Analysis

      Book · Appleton-Century-Crofts, New York

      The systematic statement of reinforcement as a functional relation between behaviour and consequence rather than a property of the consequence itself.

      Read the source

    2. Thorndike, E. L. (1911) Animal Intelligence: Experimental Studies

      Book · Macmillan, New York

      The founding experimental treatment of learning through consequences of action, and the origin of the functional analysis of effect.

      Read the source

  3. A consequence is not reinforcement simply because it follows behaviour

    Canonical inference

    What this does not assert: Reinforcement is a functional relation established by evidence, not by sequence.

  4. Unpleasantness or punitive intent does not establish that a consequence functioned as punishment

    Canonical inference

    What this does not assert: Whether responding decreased is an empirical question.

  5. Contingency is not causation

    Canonical inference

    What this does not assert: A dependency can make an action informative without settling why the outcome occurred.

  6. Operant learning is not defined as either conscious or unconscious

    Canonical inference

    What this does not assert: Deliberate reasoning about consequences occurs but is not definitional.

  7. Behaviour learned through consequences is not automatically habitual behaviour

    Canonical inference

    What this does not assert: Consequence learning can contribute to the history from which habit later develops.

  8. The classical–operant distinction concerns the relationship being learned, not passivity or activity

    Canonical inference

    What this does not assert: The processes are analytically distinct and can interact.

  9. Expression of operant learning depends on opportunity, context and current conditions

    Canonical inference

    What this does not assert: A change in behaviour need not mean a change in what was learned.

  10. Operant learning is one contributor to behaviour rather than a complete explanation of it

    Canonical inference

    What this does not assert: Bodily state, perception, goals, relationships and constraints also contribute.

  11. Instrumental behaviour can remain sensitive to the value of the specific outcome it has produced

    Established

    What this does not assert: Sensitivity depends on training conditions and devaluation method.

    1. Adams, C. D., Dickinson, A. (1981) Instrumental Responding Following Reinforcer Devaluation

      Experimental article · Quarterly Journal of Experimental Psychology, 33B(2) · 109–121

      Devaluing an outcome after training changes the instrumental action that produced it — evidence that some instrumental learning preserves outcome information.

      doi:10.1080/14640748108400816

    2. Colwill, R. M., Rescorla, R. A. (1985) Postconditioning Devaluation of a Reinforcer Affects Instrumental Responding

      Experimental article · Journal of Experimental Psychology: Animal Behavior Processes, 11(1) · 120–132

      Outcome devaluation after training selectively changes responding, showing that different actions can carry information about different outcomes.

      doi:10.1037/0097-7403.11.1.120

  12. Reducing the value of an outcome after training selectively changes the action that produced it

    Established

    What this does not assert: Selectivity is evidence about outcome information, not about a single mechanism.

    1. Colwill, R. M., Rescorla, R. A. (1985) Postconditioning Devaluation of a Reinforcer Affects Instrumental Responding

      Experimental article · Journal of Experimental Psychology: Animal Behavior Processes, 11(1) · 120–132

      Outcome devaluation after training selectively changes responding, showing that different actions can carry information about different outcomes.

      doi:10.1037/0097-7403.11.1.120

  13. Responding is allocated across alternatives in relation to the consequences those alternatives produce

    Established

    What this does not assert: Established for concurrent schedules; quantitative form varies by preparation.

    1. Herrnstein, R. J. (1961) Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement

      Experimental article · Journal of the Experimental Analysis of Behavior, 4(3) · 267–272

      Behaviour is allocated across alternatives in relation to their consequences — operant learning organises choice, not only response frequency.

      doi:10.1901/jeab.1961.4-267

  14. Pavlovian cues can modulate instrumentally established behaviour

    Established

    What this does not assert: Direction and size of the interaction depend on the preparation.

    1. Estes, W. K. (1948) Discriminative Conditioning II: Effects of a Pavlovian Conditioned Stimulus upon a Subsequently Established Operant Response

      Experimental article · Journal of Experimental Psychology, 38(2) · 173–177

      Pavlovian cues can modulate instrumentally established behaviour — the classic demonstration that the two processes interact within one behaviour.

      doi:10.1037/h0057525

  15. Extended training can reduce the sensitivity of instrumental behaviour to outcome value

    Established

    What this does not assert: The transition itself is owned downstream; here it only bounds the definition.

    1. Adams, C. D., Dickinson, A. (1981) Instrumental Responding Following Reinforcer Devaluation

      Experimental article · Quarterly Journal of Experimental Psychology, 33B(2) · 109–121

      Devaluing an outcome after training changes the instrumental action that produced it — evidence that some instrumental learning preserves outcome information.

      doi:10.1080/14640748108400816

    2. Dickinson, A. (1985) Actions and Habits: The Development of Behavioural Autonomy

      Theoretical review · Philosophical Transactions of the Royal Society B, 308(1135) · 67–78

      Why learning through consequences does not by itself make behaviour habitual, and how outcome sensitivity can decline with extended training.

      doi:10.1098/rstb.1985.0010

  16. Contiguity between behaviour and consequence facilitates operant learning without defining the operant relation

    High confidence

    What this does not assert: Delay effects are robust; immediacy is not the criterion of the relation.

    1. Shanks, D. R., Dickinson, A. (1991) Instrumental Judgment and Performance under Variations in Action–Outcome Contingency and Contiguity

      Experimental article · Memory & Cognition, 19(4) · 353–360

      Explicit contingency judgments and instrumental performance are separable measures that respond differently to contingency and delay.

      doi:10.3758/BF03197143

  17. Explicit judgments about an action–outcome contingency and behavioural sensitivity to it are separable measures

    High confidence

    What this does not assert: Correspondence varies across tasks and methods.

    1. Shanks, D. R., Dickinson, A. (1991) Instrumental Judgment and Performance under Variations in Action–Outcome Contingency and Contiguity

      Experimental article · Memory & Cognition, 19(4) · 353–360

      Explicit contingency judgments and instrumental performance are separable measures that respond differently to contingency and delay.

      doi:10.3758/BF03197143

  18. Operant learning changes the selection, timing and distribution of behaviour, not only its frequency

    High confidence

    What this does not assert: Frequency remains the most commonly reported measure.

    1. Herrnstein, R. J. (1961) Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement

      Experimental article · Journal of the Experimental Analysis of Behavior, 4(3) · 267–272

      Behaviour is allocated across alternatives in relation to their consequences — operant learning organises choice, not only response frequency.

      doi:10.1901/jeab.1961.4-267

Sources

  1. Thorndike, E. L. (1911) Animal Intelligence: Experimental Studies

    Book · Macmillan, New York

    The founding experimental treatment of learning through consequences of action, and the origin of the functional analysis of effect.

    Read the source

  2. Skinner, B. F. (1938) The Behavior of Organisms: An Experimental Analysis

    Book · Appleton-Century-Crofts, New York

    The systematic statement of reinforcement as a functional relation between behaviour and consequence rather than a property of the consequence itself.

    Read the source

  3. Hammond, L. J. (1980) The Effect of Contingency upon the Appetitive Conditioning of Free-Operant Behavior

    Experimental article · Journal of the Experimental Analysis of Behavior, 34(3) · 297–304

    The classic contingency-degradation demonstration: responding declines when the same outcome also becomes available independently of the action.

    doi:10.1901/jeab.1980.34-297

  4. Adams, C. D., Dickinson, A. (1981) Instrumental Responding Following Reinforcer Devaluation

    Experimental article · Quarterly Journal of Experimental Psychology, 33B(2) · 109–121

    Devaluing an outcome after training changes the instrumental action that produced it — evidence that some instrumental learning preserves outcome information.

    doi:10.1080/14640748108400816

  5. Colwill, R. M., Rescorla, R. A. (1985) Postconditioning Devaluation of a Reinforcer Affects Instrumental Responding

    Experimental article · Journal of Experimental Psychology: Animal Behavior Processes, 11(1) · 120–132

    Outcome devaluation after training selectively changes responding, showing that different actions can carry information about different outcomes.

    doi:10.1037/0097-7403.11.1.120

  6. Dickinson, A. (1985) Actions and Habits: The Development of Behavioural Autonomy

    Theoretical review · Philosophical Transactions of the Royal Society B, 308(1135) · 67–78

    Why learning through consequences does not by itself make behaviour habitual, and how outcome sensitivity can decline with extended training.

    doi:10.1098/rstb.1985.0010

  7. Herrnstein, R. J. (1961) Relative and Absolute Strength of Response as a Function of Frequency of Reinforcement

    Experimental article · Journal of the Experimental Analysis of Behavior, 4(3) · 267–272

    Behaviour is allocated across alternatives in relation to their consequences — operant learning organises choice, not only response frequency.

    doi:10.1901/jeab.1961.4-267

  8. Shanks, D. R., Dickinson, A. (1991) Instrumental Judgment and Performance under Variations in Action–Outcome Contingency and Contiguity

    Experimental article · Memory & Cognition, 19(4) · 353–360

    Explicit contingency judgments and instrumental performance are separable measures that respond differently to contingency and delay.

    doi:10.3758/BF03197143

  9. Estes, W. K. (1948) Discriminative Conditioning II: Effects of a Pavlovian Conditioned Stimulus upon a Subsequently Established Operant Response

    Experimental article · Journal of Experimental Psychology, 38(2) · 173–177

    Pavlovian cues can modulate instrumentally established behaviour — the classic demonstration that the two processes interact within one behaviour.

    doi:10.1037/h0057525

What this opens up

What becomes readable once you have this.

Where to go from here

Next published piece

Reinforcement, Reward and Pleasure

What changes behaviour, what carries value and what feels pleasurable are related questions—but they are not the same question.

Continue through the Library →See where this sits in the graph →

Back to the Library →