Models · Learning and Memory

Goal-Directed and Habitual Control

A model of how goal-sensitive and habitual influences can coexist, compete and cooperate in shaping learned behaviour.

By Yona Ole Lobulu ·

Model11 min readFoundationalD5.17

Topic
Learning and Memory
Read first
3 pieces should be read before this one
Reading time
About 11 minutes of reading
Difficulty
Late reading: this page sits at the far end of the Library, after many other pieces

The question

What evidence supports different contributions from goal-sensitive and habitual influences under these conditions?

Definition

A model organising two distinguishable influences on learned behaviour: goal-directed control, which remains sensitive to action–outcome relations and to the current value of the outcome, and habitual control, which describes learned context–response tendencies that make particular responses more likely. The two influences can coexist, compete or contribute at different levels of the same behaviour.

You leave work intending to stop somewhere before going home.

The destination matters. You know what you need to do there. But several minutes into the journey, you notice that you are following your usual route home.

It is tempting to say:

"Habit took over."

The familiar route may indeed have exerted habitual influence. The planned stop may still have exerted goal-sensitive influence. But the action cannot tell us, by itself, how much either contributed.

The Goal-Directed and Habitual Control Model starts from this possibility: both influences can be active at the same time.

The question is not simply whether behaviour was deliberate or habitual. It is: what evidence supports different contributions from goal-sensitive and habitual influences under these conditions?

Why choice or habit is too simple

Goal-directed and habitual control identify meaningful differences in behaviour.

The problem arises when those differences are treated as two exclusive owners of action. One controller supposedly makes deliberate choices until another takes over and produces a habit.

Real behaviour is rarely divided so cleanly.

A familiar action can remain sensitive to its consequences. A current goal can be implemented through a highly practised sequence. A learned response tendency can make one action more available without making alternatives impossible.

Sometimes both influences favour the same behaviour. Sometimes they compete. Sometimes they contribute at different levels of the same behavioural sequence.

The Model therefore represents relative influence, not exclusive ownership.

Goal-sensitive influence

Goal-directed control is defined more precisely than behaviour that merely appears purposeful.

It depends principally on two kinds of information: what an action is expected to produce, and how much that consequence currently matters.

This builds on Operant Learning, which establishes how actions and consequences become related through experience.

The Model's goal-sensitive pathway is: action–outcome knowledge and current outcome value contribute to goal-sensitive influence.

These are the principal informational inputs. Their expression also depends on whether the information can be retrieved, integrated and used under the current context and control conditions.

Three questions capture the pathway:

  • What will this action produce?
  • Does that outcome matter now?
  • Does the action still make a difference?

If behaviour adjusts when the answers change, that supports a goal-directed inference.

Sensitivity to current outcome value

Imagine someone who prepares the same drink every morning.

The sequence is familiar, fast and highly practised. But when they no longer want that drink, they immediately prepare another.

The behaviour remained sensitive to the expected outcome's current value. Familiarity did not make it independent of goals.

This connects to Valuation. The expected outcome may remain unchanged while its current significance changes. Valuation contributes to which consequences matter now, but it is not identical to motivation or reward.

Sensitivity to action effectiveness

Goal-sensitive behaviour also adjusts when an action stops producing the expected result.

A route closes. A workplace process changes. A response that previously produced an outcome no longer makes a difference.

If behaviour changes accordingly, this supports sensitivity to the action–outcome relationship.

Goal-sensitive behaviour does not require a slow, conscious calculation. A person can use learned consequence information quickly and without describing every step.

Nor does it guarantee rationality. Behaviour can remain sensitive to outcomes while relying on incomplete knowledge or pursuing an outcome that later proves harmful.

Serving a goal is not enough to identify goal-directed control. The relevant question is whether behaviour remains sensitive to action–outcome relations and current outcome value.

Habitual influence

Habitual control begins from a different form of learned influence.

Repetition in recurring situations can strengthen tendencies that make a response more available in those contexts. Different theories explain the underlying learning in different ways, but the central Model pathway is: a learned context–response tendency contributes to habitual influence.

The arrow represents increased influence or probability, not compulsory action.

The road outside work, the usual time of departure and a familiar junction may all support the well-learned route home. That route can become unusually available before the person has deliberately compared every alternative.

This builds on What Makes Behaviour Habitual? and How Context Triggers Habitual Action.

Habitual influence is often inferred when responding becomes less immediately sensitive to a change in the outcome's current value, or to a change in whether the action still produces that outcome.

Reduced sensitivity does not mean goals have disappeared. It means that current outcome information had less immediate influence over behaviour under the tested conditions.

Development and expression are different

Repetition in stable contexts can contribute to learned response tendencies. It does not prove habitual control.

A person may take the same route every day because it remains the best way home. They may train regularly because the current outcome remains strongly valued. Behavioural frequency describes what repeatedly occurs, not the process currently guiding it.

Formation, expression and maintenance must also remain separate.

A response tendency can be learned without being expressed in every context. A familiar context can retrieve an old tendency without making the response inevitable. A behaviour can stop appearing while its learning history remains.

Familiar, automatic and habitual are not interchangeable

A skilled action may be fast and efficient while remaining highly sensitive to current goals. It may be automatic in some respects without being habitually controlled.

This is the boundary established by Automaticity.

Awareness does not settle the distinction either. A person can consciously notice a familiar pull, and goal-sensitive action can occur without a fully conscious calculation.

Repetition, frequency, automaticity and persistence may all provide relevant information. None identifies habitual influence by itself.

How the influences work together

The two pathways were separated so their contributions could be described. They were not separated because behaviour must use only one.

Alignment

Suppose you intend to go home and the familiar route also leads home.

Current outcome information and the learned route tendency favour the same behaviour. The action may feel fluent because several influences support it.

The behaviour cannot reveal their weighting. Agreement makes the influences difficult to distinguish.

Competition

Suppose you intend to stop somewhere else while the familiar context supports the route home.

The influences now favour different actions. Following the usual route may be consistent with habitual influence, but it could also reflect a lost goal, divided attention or inadequate route planning.

Taking the alternative route does not prove habitual influence was absent. The familiar tendency may have been present but outweighed or redirected.

Different levels of action

A person may make a goal-sensitive decision to prepare dinner while relying on familiar sequences to carry it out.

A familiar sequence is not necessarily habitual. The example shows how a current goal can organise one level of behaviour while learned tendencies contribute to another. Establishing the control status of either level still requires evidence.

This hierarchical organisation differs from simple cooperation. In cooperation, both influences favour the same action. In hierarchy, they contribute at different levels of the behavioural structure.

The Goal-Directed and Habitual Control Model

The Model has two pathways and one shared set of conditions.

  • Action–outcome knowledge and current outcome value contribute to goal-sensitive influence.
  • A learned context–response tendency contributes to habitual influence.
  • Context and current control conditions affect both pathways.
  • Both influences contribute to behaviour.
  • Behaviour produces consequences that contribute to later learning.

Goal-sensitive influence draws on learned consequences and their current value. Habitual influence draws on learned response tendencies that become available in particular contexts.

The arrows represent proposed influence. They do not indicate compulsory responses, complete causation or directly measured strength.

Researchers do not observe the two influences themselves. They observe patterns of behaviour across defined conditions and manipulations, then infer their relative contribution.

Why their contribution changes

The relative contribution of each influence is not fixed.

Availability

Relevant goals and action–outcome relations must be retrieved and maintained. A goal that is not currently accessible cannot guide action in the same way as one that remains active.

Context can retrieve familiar responses, but it can also reactivate goals and signal which action–outcome relation applies. Context therefore affects both pathways.

Coordination

Behaviour may require maintaining one goal, inhibiting an immediately available response or switching to another plan.

These processes connect the Model to Executive Functions. Executive processes can support the expression of goal-sensitive behaviour, but they are not identical to goal-directed control and do not function as a central commander.

Allocation

Flexible consequence-based evaluation can require time and effort. Whether additional control is deployed may depend on its expected benefit.

Goals Compete helps explain why several valued actions may be available at once. Effort and Expected Value helps explain why control allocation depends partly on what the additional effort is expected to achieve.

Control mode also does not reveal motivation. As Motivation Is a Family of Processes establishes, initiation, persistence and allocation depend on multiple motivational processes. A behaviour can be strongly motivated while relying on familiar tendencies, or weakly pursued despite remaining goal-sensitive.

Goal-sensitive control supports flexibility. Habitual influence can support efficiency and stability. Either can help or hinder action depending on what the situation requires.

How researchers infer control

Researchers often use two experimental approaches.

Outcome devaluation

First, an action is learned to produce an outcome. The value of that outcome is then changed.

A food outcome may become less desirable after satiety. In a human task, participants may learn that one outcome no longer carries points or another form of value.

Researchers then test whether the associated action decreases.

Selective adjustment supports sensitivity to current outcome value and therefore a goal-directed contribution. Continued responding may support habitual influence if the devaluation was effective and the relevant action–outcome relation had been learned.

Contingency degradation

Another approach changes whether the action still makes a difference.

An outcome that previously depended on an action may begin occurring independently. The action is now less effective at producing it.

Reduced responding supports sensitivity to the changed contingency. Persistence can support a habitual-control inference under appropriate conditions.

What the result can and cannot show

A valid inference requires more than observing whether behaviour continued.

Researchers must establish that:

  • the outcome's value or contingency genuinely changed;
  • the relation was learned;
  • the test allowed that knowledge to be expressed;
  • task comprehension or another process does not explain the result.

Human tasks can also recruit instructions, working memory, inhibition and strategic reasoning.

The sequence is therefore: manipulate a relevant condition, observe behaviour, identify a pattern, then infer relative control contribution.

Devaluation and contingency-degradation tasks do not directly measure internal controllers. Reduced sensitivity can support habitual influence without identifying one unique mechanism.

Most importantly: failure to demonstrate goal-directed control does not automatically prove habitual control.

Sometimes the correct conclusion is that the task did not identify the controlling process clearly.

Control inference should therefore be graded and condition-specific. One task cannot classify a person as a goal-directed or habitual type.

The Model is not the system

Return to the route example.

At the behavioural level, the person followed the familiar route despite intending to stop elsewhere.

At the process level, the Model proposes possible contributions from current outcome information, a learned route tendency and the conditions affecting their expression.

At the computational level, researchers can represent flexible consequence-based choice and learned response policies in several ways. Model-based and model-free approaches offer useful partial parallels, but they are not exact synonyms for goal-directed and habitual control. Different algorithms can generate similar behavioural patterns.

At the biological level, neural studies identify distributed and overlapping contributions to learning, valuation, control and action selection. They do not reveal one goal-directed region and one habit region.

An arbitration model can mathematically represent changing reliance on strategies. It does not establish an internal arbitrator choosing between literal agents.

A successful model fit shows that one representation accounts for measured data under stated assumptions. It does not prove that the target system implements the Model literally.

This is the protection supplied by Scientific Models Are Tools: the model is not the target system.

The Model is useful because it organises evidence and exposes better questions. Its components are not the complete cognitive or biological machinery that produces action.

A better question about behaviour

The person intended to stop somewhere else but followed the usual route.

A current goal could have supported one action while a learned contextual tendency supported another. The observed route does not reveal how strongly either influence contributed.

This architecture becomes important for What Happens During Self-Control?, where familiar tendencies, competing goals and current control conditions become part of a wider behavioural conflict.

Goal-directed and habitual control are not rival owners of behaviour. They are explanatory influences whose relative contribution must be inferred from evidence.

Behind this page

The claims this model makes, the evidence behind them, and the limits it accepts.

Evidence status

Established

Well supported by a substantial, converging empirical literature.

Claims

  1. Goal-sensitive and habitual influences can both be active in the same behaviour at the same time

    Canonical inference

    What this does not assert: The founding premise of the Model.

  2. Behaviour that adjusts when outcome value changes supports a goal-directed inference

    Established

    What this does not assert: Support, not proof.

  3. A highly practised sequence can still change immediately when the outcome is no longer wanted

    Established

    What this does not assert: Familiarity did not make it independent of goals.

  4. An expected outcome may remain unchanged while its current significance changes

    Established

    What this does not assert: Inherited from D7.4, Valuation.

  5. Valuation is not identical to motivation or reward

    High confidence

    What this does not assert: Boundary preserved from D7.4.

  6. Goal-sensitive behaviour also adjusts when an action stops producing the expected result

    Established

    What this does not assert: Sensitivity to the action–outcome relation.

  7. Goal-sensitive behaviour does not require a slow, conscious calculation

    High confidence

    What this does not assert: Goal-direction is not consciousness.

  8. Goal-directed control does not guarantee rationality

    High confidence

    What this does not assert: Outcome sensitivity can rest on incomplete or harmful knowledge.

  9. Serving a goal is not sufficient to identify goal-directed control

    Canonical inference

    What this does not assert: Purposeful appearance is not the criterion.

  10. Repetition in recurring situations can strengthen tendencies that make a response more available in those contexts

    Established

    What this does not assert: The habitual pathway of the Model.

  11. Theories differ in how they explain the learning underlying habitual influence

    Contested

    What this does not assert: The Model does not adjudicate between them.

  12. An observed action cannot by itself reveal how much either influence contributed

    Canonical inference

    What this does not assert: Control status is inferred from conditions, not read off behaviour.

  13. Habitual influence is often inferred when responding becomes less immediately sensitive to changed outcome value or changed contingency

    Established

    What this does not assert: An inference from a pattern, not a measurement.

  14. Reduced sensitivity does not mean goals have disappeared

    Canonical inference

    What this does not assert: It concerns immediate influence under tested conditions.

  15. Repetition in stable contexts does not prove habitual control

    High confidence

    What this does not assert: Frequency describes occurrence, not the guiding process.

  16. Formation, expression and maintenance of a response tendency are separable

    Established

    What this does not assert: A tendency can be learned without being expressed.

  17. A familiar context can retrieve an old tendency without making the response inevitable

    Established

    What this does not assert: Inherited from D5.14.

  18. A behaviour can stop appearing while its learning history remains

    Established

    What this does not assert: Absence of behaviour is not erasure of learning.

  19. A skilled action can be fast and efficient while remaining highly sensitive to current goals

    Established

    What this does not assert: The boundary established by D5.15, Automaticity.

  20. Automaticity in some respects does not establish habitual control

    High confidence

    What this does not assert: Automatic and habitual are not interchangeable.

  21. Awareness does not settle the goal-directed / habitual distinction

    High confidence

    What this does not assert: A familiar pull can be consciously noticed.

  22. Repetition, frequency, automaticity and persistence each provide relevant but insufficient information

    Canonical inference

    What this does not assert: None identifies habitual influence by itself.

  23. Treating goal-directed and habitual control as two exclusive owners of action misdescribes ordinary behaviour

    High confidence

    What this does not assert: The Model represents relative influence, not ownership.

  24. When both influences favour the same action, their weighting cannot be distinguished

    Canonical inference

    What this does not assert: Alignment hides relative contribution.

  25. Following a familiar route against a current intention is consistent with habitual influence but does not establish it

    Canonical inference

    What this does not assert: A lost goal, divided attention or poor planning could also explain it.

  26. Taking the alternative route does not prove habitual influence was absent

    Canonical inference

    What this does not assert: The tendency may have been outweighed or redirected.

  27. A current goal can organise one level of behaviour while learned tendencies contribute to another

    High confidence

    What this does not assert: Hierarchical organisation differs from simple cooperation.

  28. A familiar sequence is not necessarily habitual

    Canonical inference

    What this does not assert: Control status of each level requires its own evidence.

  29. Context and current control conditions affect both pathways

    Established

    What this does not assert: Context is not exclusively a habit cue.

  30. Behaviour produces consequences that contribute to later learning

    Established

    What this does not assert: The Model is a loop, not a line.

  31. Researchers observe behavioural patterns across conditions rather than the two influences themselves

    Canonical inference

    What this does not assert: The influences are inferred entities.

  32. The relative contribution of each influence is not fixed

    Established

    What this does not assert: It varies with availability, coordination and allocation.

  33. A goal that is not currently accessible cannot guide action as one that remains active

    Established

    What this does not assert: Availability constrains goal-sensitive influence.

  34. A familiar action can remain sensitive to its consequences

    Established

    What this does not assert: Familiarity is not evidence of habitual control.

  35. Context can reactivate goals and signal which action–outcome relation applies

    High confidence

    What this does not assert: The same context can serve either pathway.

  36. Maintaining a goal, inhibiting an available response and switching plans connect the Model to executive processes

    Established

    What this does not assert: Inherited from D8.1.

  37. Executive processes are not identical to goal-directed control

    High confidence

    What this does not assert: And do not function as a central commander.

  38. Whether additional control is deployed may depend on its expected benefit

    Established

    What this does not assert: Inherited from D7.14, Effort and Expected Value.

  39. Several valued actions may be available at once

    Established

    What this does not assert: Inherited from D7.13, Goals Compete.

  40. Control mode does not reveal motivation

    High confidence

    What this does not assert: Inherited from D7.10; motivation is a family of processes.

  41. Goal-sensitive control supports flexibility while habitual influence can support efficiency and stability

    High confidence

    What this does not assert: Either can help or hinder depending on the situation.

  42. Outcome devaluation tests whether an action decreases after the value of its outcome changes

    Established

    What this does not assert: The principal goal-sensitivity manipulation.

  43. Continued responding after effective devaluation may support habitual influence

    Established

    What this does not assert: Only if the action–outcome relation had been learned.

  44. Contingency degradation changes whether the action still makes a difference to the outcome

    Established

    What this does not assert: The second principal manipulation.

  45. A current goal can be implemented through a highly practised sequence

    High confidence

    What this does not assert: Practised execution does not settle control status.

  46. A valid control inference requires that the manipulation worked, the relation was learned, and the test allowed expression

    Canonical inference

    What this does not assert: Plus the exclusion of task comprehension and other processes.

  47. Human tasks can recruit instructions, working memory, inhibition and strategic reasoning

    Established

    What this does not assert: Which complicates habit inferences in humans.

  48. Experimental habit induction in humans has repeatedly failed to produce clear habitual control

    Contested

    What this does not assert: A live methodological debate.

  49. Devaluation and contingency-degradation tasks do not directly measure internal controllers

    Canonical inference

    What this does not assert: They support graded inferences.

  50. Failure to demonstrate goal-directed control does not prove habitual control

    Canonical inference

    What this does not assert: Sometimes the task simply did not identify the process.

  51. One task cannot classify a person as a goal-directed or habitual type

    Canonical inference

    What this does not assert: Inference is condition-specific, not dispositional.

  52. Model-based and model-free approaches are partial parallels, not synonyms for goal-directed and habitual control

    Contested

    What this does not assert: Different algorithms can generate similar behavioural patterns.

  53. Neural studies identify distributed and overlapping contributions rather than one goal region and one habit region

    Established

    What this does not assert: No one-to-one mapping is available.

  54. An arbitration model can represent changing reliance on strategies without establishing an internal arbitrator

    High confidence

    What this does not assert: Mathematical description is not mechanism.

  55. A successful model fit shows that a representation accounts for data under stated assumptions

    Canonical inference

    What this does not assert: It does not prove literal implementation. Inherited from D2.12.

  56. A learned response tendency can make one action more available without making alternatives impossible

    Established

    What this does not assert: Influence, not compulsion.

  57. Goal-directed and habitual control are explanatory influences whose relative contribution must be inferred from evidence

    Canonical inference

    What this does not assert: The core thesis of this Model.

  58. Goal-directed control depends principally on what an action is expected to produce and how much that consequence currently matters

    Established

    What this does not assert: The two defining informational inputs.

  59. Actions and consequences become related through experience

    Established

    What this does not assert: Inherited from D5.4, Operant Learning.

  60. Expression of goal-sensitive influence depends on whether the relevant information can be retrieved, integrated and used under current conditions

    High confidence

    What this does not assert: Knowledge alone does not guarantee its use.

Read this first

This piece assumes them.

Helpful beforehand

Useful context, but not required.

Where to go from here

Next published piece

Skill Learning Changes Perception and Action

A skilled performer is not a novice running the same process faster. Experience can change what becomes relevant, when action can begin and what no longer has to be managed separately.

Continue through the Library →See where this sits in the graph →

Back to the Library →