Research · Perception, Attention and Belief

Predictive Processing: Scope and Limits

Predictive processing is best evaluated not as one theory to accept or reject, but as a family of claims whose evidential strength varies across computational, behavioural and neural levels.

By Yona Ole Lobulu ·

Research note15 min readFoundationalD6.13

Topic
Perception, Attention and Belief
Read first
2 pieces should be read before this one
Reading time
About 15 minutes of reading
Difficulty
Late reading: this page sits at the far end of the Library, after many other pieces

The question

Which predictive-processing claim is being tested, what alternatives does the evidence distinguish it from, and how much confidence does that particular result justify?

Definition

Predictive processing offers a powerful family of frameworks for understanding perception and cognition in terms of prediction, incoming information and the management of discrepancies between them. Its influence, however, does not mean that every version makes the same claims, that every compatible finding uniquely confirms the framework, or that computational descriptions automatically identify neural mechanisms. Evaluating predictive processing therefore requires separating its broad explanatory insights from stronger implementation claims, distinguishing related formulations from one another, and asking which observations discriminate it from plausible alternatives.

Predictive processing is compelling partly because of how much it seems able to connect.

The language of prediction, incoming information and discrepancy can be used to think about perception, expectation, learning, attention and action. It offers a way of seeing apparently different cognitive problems as variations on a common challenge: organisms must interpret incomplete information while anticipating what comes next.

That breadth can be scientifically valuable. A framework that connects phenomena can organise findings that would otherwise remain separate, suggest relationships between levels of explanation and generate new hypotheses.

But explanatory breadth also raises an evidential question.

If many different observations can be accommodated within predictive-processing language, then the fact that an observation fits the framework tells us less than it would if only a narrow set of outcomes were compatible with it.

The useful question is therefore not simply:

Is there evidence for predictive processing?

It is:

Which predictive-processing claim is being tested, what alternatives does the evidence distinguish it from, and how much confidence does that particular result justify?

That question leads to a more disciplined evaluation than either treating predictive processing as an established general theory of cognition or dismissing the entire framework because some of its stronger claims remain uncertain.

One Label, Different Claims

The first difficulty is terminological.

Predictive processing is not one uniformly specified theory, and the relationships among the labels used in this area are not fixed in exactly the same way by every researcher.

For evaluation, however, it is useful to distinguish several overlapping strands.

As introduced in Predictive Processing, predictive processing can be treated as a broad family of approaches in which predictions or expectations interact with incoming information and discrepancies between them contribute in some way to inference or updating.

Predictive coding often refers to more specific computational or neural architectures in which higher levels generate predictions about activity at lower levels and some form of mismatch or prediction error plays an important role in signalling what remains unexplained.

Bayesian approaches describe aspects of perception or cognition using probabilistic inference, allowing prior information and current evidence to contribute jointly to estimates under uncertainty. A system can be described successfully by a Bayesian model without that description uniquely specifying a predictive-coding neural implementation.

Free-energy formulations introduce a broader mathematical framework intended to connect perception, learning and action.

Active inference develops related formal ideas further into action and policy selection, adding commitments that go beyond the narrower claim that sensory processing makes use of prediction.

These descriptions are useful distinctions rather than a universally agreed taxonomy. Different authors draw the boundaries differently, and the traditions overlap historically, mathematically and conceptually.

The important point is therefore not to force them into a perfect hierarchy.

It is to avoid treating them as interchangeable.

Evidence that supports one predictive-coding architecture does not automatically establish a free-energy formulation. Evidence that behaviour is well captured by Bayesian inference does not automatically demonstrate a particular prediction-error circuit. Evidence for prediction-sensitive sensory processing does not by itself establish active inference.

And evidence against one classical predictive-coding implementation would not, by itself, decide every broader predictive-processing formulation.

Before asking whether predictive processing is supported, we first have to ask which claim within the family is actually at stake.

The Phenomenon Is Not the Explanation

A second distinction is even more fundamental.

Perception is constructive.

As established in Perception Is Constructive, perception is not a passive registration of sensory input. Context, prior information, learned regularities and the organisation of sensory signals can influence what a person perceives.

Changes in what someone reports experiencing can therefore establish that expectation or context shaped perception.

But that observation does not, by itself, identify the computational or neural mechanism responsible.

Suppose an experiment shows that expectation changes how people perceive an ambiguous stimulus.

That is an important empirical finding.

It tells us that expectation influences perceptual processing or perceptual judgement.

But several questions remain open.

Did the expectation alter the sensory representation itself? Did it change attention? Did it influence a later decision about the stimulus? Did it arise through associative learning, recurrent processing, hierarchical prediction-error signalling or some combination of mechanisms?

Predictive processing offers one family of explanations.

The phenomenon itself does not uniquely select that family.

This distinction matters because some of the broadest empirical claims in this area can be extremely well supported while remaining theoretically non-specific.

That expectations and context influence perception may deserve considerable confidence.

The stronger claim that these effects are produced by one particular predictive-processing architecture requires additional evidence.

The phenomenon is what needs explaining. The framework is a candidate explanation.

Constructive perception therefore does not logically entail predictive processing.

Compatible Is Not the Same as Confirmed

A finding can be compatible with a theory without showing that the theory is better supported than plausible alternatives.

Consider expectation suppression.

In many experiments, neural responses to expected stimuli are smaller than responses to unexpected stimuli. This pattern fits naturally within some predictive-coding accounts: when incoming information matches a prediction, less discrepancy may remain to be processed.

That is a plausible interpretation.

But smaller responses to expected stimuli are not unique to that mechanism.

A representation may become sharper. Repetition may produce adaptation or repetition suppression. Attention may change response magnitude. Learning and recurrent processing may produce related effects.

The observation:

expected stimulus → reduced neural response

therefore establishes more securely that expectation-sensitive processing exists than it establishes why that reduction occurs.

This does not make expectation suppression weak or irrelevant evidence.

It means that the evidential strength depends on the claim.

If the claim is:

expectations alter sensory responses,

the finding may provide strong support.

If the claim is:

those altered responses are specifically generated by a particular prediction-error mechanism,

the same result carries less weight unless competing explanations have been ruled out or made less plausible.

This is where Replication, Robustness and Scientific Confidence becomes essential.

An effect can replicate while its theoretical interpretation remains unsettled.

A replicated effect is not automatically a replicated explanation.

Replication increases confidence that something occurs. It does not, by itself, determine why it occurs.

To move beyond compatibility, we need experiments in which competing accounts do not all expect the same result.

What Stronger Evidence Looks Like

Some predictive-processing research has been designed around exactly this problem.

Consider a prediction-error account and a sharpening account of expectation.

For some measurements, both may predict that expected stimuli produce reduced overall neural responses. If that is the only observation available, the experiment cannot discriminate well between them.

A more informative test looks for conditions where the models diverge.

Researchers can manipulate the strength of a prediction, whether that prediction matches the incoming signal, or the detailed pattern of neural representation. Under those conditions, prediction-error and sharpening accounts may expect different outcomes.

In specific perceptual paradigms, studies using this logic have reported results that favour prediction-error accounts over the tested sharpening alternatives.

That evidence carries more weight for that particular comparison.

The reason is not that prediction-error terminology was used afterward.

It is that two plausible models were specified, they differed in what they expected to observe, and the experiment was able to discriminate between them.

The logic is:

competing accounts → different predictions → a test capable of separating them → evidence favouring one of the tested accounts

This is stronger than discovering a result first and then showing that predictive-processing language can accommodate it.

Even here, the conclusion must remain local.

A result favouring a prediction-error model over a sharpening model in one sensory paradigm supports that computational account in that context.

It does not establish predictive processing globally.

It does not establish every predictive-coding implementation.

It does not establish the free-energy principle.

And it does not yet tell us exactly how the computation is implemented biologically.

The more specific the evidential victory, the more specific the conclusion should remain.

Prediction Is Not Yet Mechanism

Predictive success and mechanistic explanation are different achievements.

Suppose a computational model predicts how perceptual judgements change when expectations are manipulated.

The prediction succeeds.

That matters.

But several further questions remain.

A model may successfully describe a pattern.

It may successfully predict what happens under new conditions.

A causal account requires evidence that manipulating relevant factors changes the effect in the way the account predicts.

A mechanistic account must go further and identify the operations through which those causal relationships are produced.

A neural implementation claim goes further again by proposing how biological systems realise those operations.

As established in Correlation, Prediction, Causation and Mechanism, these are not interchangeable achievements.

Nor should the variables in a successful formal model automatically be treated as literal components of the system.

A Bayesian model may represent uncertainty using probability distributions. A predictive model may contain quantities labelled prediction error or precision. A free-energy formulation introduces further mathematical variables.

These constructs may capture important structure in behavioural or neural data.

But Scientific Models Are Tools provides a crucial safeguard:

model ≠ target system.

A probability distribution in a model is not automatically a probability distribution explicitly represented by neurons. A fitted precision parameter is not automatically an identifiable neural precision signal. A prediction-error variable is not automatically a distinct population of neurons carrying that quantity in exactly the modelled form.

Those mappings are empirical hypotheses.

They require evidence of their own.

A model can therefore be useful, predictive and theoretically revealing while its biological implementation remains uncertain.

Successful prediction does not confer mechanistic reality on every variable the model contains.

When Flexibility Becomes an Evidential Problem

Broad frameworks often require auxiliary assumptions.

Predictive-processing models may differ in prior expectations; reliability or precision; hierarchical level; learning history; task context; the way prediction errors are weighted; or how different information sources interact.

There is nothing inherently problematic about this complexity. Cognition itself is complex, and serious models often need multiple interacting components.

The evidential problem arises when those assumptions remain unconstrained until after the result is known.

If a surprising outcome can always be accommodated by changing an unmeasured prior, shifting the relevant hierarchical level or invoking a different precision assignment only retrospectively, then compatibility provides limited evidence for the underlying framework.

The central question becomes:

What did the formulation rule out before the result was known?

This is the legitimate force behind criticisms of underspecification in predictive-processing theories.

But the stronger claim that predictive processing is inherently unfalsifiable goes too far.

Specific predictive-coding models can make risky predictions. Researchers have identified response patterns that would count against particular formulations, and experiments can be designed so that predictive accounts and competing models expect different outcomes.

Testability therefore varies across the family.

A broad organising framework may offer useful concepts while remaining too general to face a single decisive experiment. A narrower mechanistic model may make more specific commitments and therefore be easier to challenge experimentally.

The broader claim may currently deserve higher confidence because it makes fewer implementation-specific commitments.

The narrower mechanistic claim may deserve lower current confidence because its greater specificity requires additional evidence.

That does not make the narrower claim less scientific.

Quite the opposite: specifying enough to risk failure is what allows stronger empirical tests.

Post hoc explanation also has a legitimate role. An unexpected result can reveal a missing assumption, motivate a revised model or generate a new prediction.

But retrospective accommodation and prospective testing have different evidential status.

A post hoc explanation can begin the next test.

It should not be treated as though that test has already been passed.

The Harder Question of Neural Implementation

The evidential burden rises again when predictive processing is proposed as a particular neural architecture.

There is meaningful evidence in this direction.

Neural systems respond differently to expected and unexpected information in many experimental settings. Researchers have identified mismatch-sensitive responses in sensory systems. Experiments that separate repetition from expectation have found neural effects consistent with prediction-error computations. Manipulating cortical feedback can alter expectation-related mismatch responses, providing causal evidence that top-down pathways contribute to some of these phenomena.

Laminar and temporal studies have also reported patterns compatible with mechanisms proposed by specific hierarchical predictive-coding models.

These findings matter.

They show that prediction and expectation are not merely convenient metaphors detached from biological processing.

But they do not yet establish one universal predictive-coding circuit.

Unexpected stimuli can produce enhanced responses for several reasons. Adaptation, novelty, salience, attention and learning may all contribute depending on the paradigm.

This is why the terminology itself requires discipline.

A larger response to an unexpected stimulus is descriptively a mismatch or expectation-violation response.

Calling it a prediction-error signal adds a theoretical interpretation.

That interpretation becomes stronger when the experiment controls relevant alternatives or when the structure of the response matches independently specified predictions of a particular model.

Several different levels must therefore remain separate.

A formal prediction error is a quantity defined within a model.

A prediction-error computation is a proposed information-processing operation.

A prediction-error signal is an empirical response interpreted as reflecting that operation.

A prediction-error neuron or circuit is a still more specific biological implementation claim.

Evidence at one level does not automatically establish the others.

This distinction also protects the concept established in Uncertainty.

Uncertainty is not prediction error.

Uncertainty concerns what the available information leaves unresolved. Prediction error concerns a discrepancy between prediction and incoming information within particular predictive frameworks.

A system may remain uncertain even when a current prediction happens to match the input. A large mismatch does not automatically imply that uncertainty itself is large.

Research on neural implementation also remains active.

Some studies support hierarchical feedback, laminar organisation and expectation-sensitive effects predicted by particular coding models. Other syntheses have not found a simple large-scale segregation between prediction and error processing. Newer models suggest that prediction and prediction-error information may be represented more diffusely than architectures with sharply separated cell populations imply.

The appropriate conclusion is therefore graded.

There is strong evidence for expectation-sensitive, prediction-sensitive and mismatch-sensitive effects in sensory processing.

There is meaningful evidence consistent with mechanisms proposed by specific predictive-coding models in particular systems and tasks.

There is considerably less justification for claiming that one canonical prediction/error architecture has been established as the universal cortical implementation of predictive processing.

Those are different claims.

They should not inherit the same confidence.

Scope Does Not Transfer Evidence Automatically

The same discipline becomes more important as predictive-processing ideas expand beyond sensory perception.

Free-energy formulations and active inference connect predictive ideas to broader questions involving learning, action, policy selection and organism–environment interaction.

That wider scope may be theoretically productive.

But every extension adds commitments that require evidence of their own.

Evidence that cortical feedback contributes to an auditory mismatch response does not automatically establish an active-inference account of action.

Evidence favouring a prediction-error model over a sharpening model in visual or auditory perception does not establish every claim associated with the free-energy principle.

Evidence that behaviour is approximately Bayesian in one task does not establish a universal Bayesian neural architecture.

Support does not automatically transfer from one formulation or level of claim to another.

This is especially easy to forget when the theories fit together elegantly.

Theoretical coherence matters. It can reveal deep relationships and support model development.

But coherence is not discriminative empirical evidence.

As theoretical scope expands, the evidential burden expands with it.

What Should We Actually Conclude?

So how well supported is predictive processing?

There is no scientifically useful single confidence rating.

The answer changes with the claim.

At the broadest empirical level, there is high confidence that expectations, context and learned regularities influence sensory processing and perceptual experience. These findings are robust and important. They are also not uniquely proprietary to predictive processing.

At a narrower level, there is strong evidence that sensory systems contain expectation-sensitive and mismatch-sensitive processes. This moves closer to the territory emphasised by predictive accounts, while still leaving room for more than one computational interpretation.

Narrower still, specific prediction-error models have received meaningful and sometimes discriminative support in particular paradigms, including cases where they outperform tested alternatives. Those findings justify increased confidence in the specific accounts tested, not in every theory grouped under the predictive-processing label.

At the level of neural implementation, confidence should be more selective. Evidence supports top-down feedback, hierarchical processing and neural responses consistent with mechanisms proposed by predictive-coding models. But the claim that cortex universally implements one canonical prediction/error circuit architecture remains substantially more uncertain.

Broader free-energy and active-inference formulations introduce further commitments and therefore carry independent evidential burdens. Their status cannot simply be inherited from successful sensory prediction experiments.

This pattern is not a failure of predictive processing.

It is what serious evaluation of a broad scientific framework looks like.

Some broad phenomena can deserve high confidence precisely because they make fewer implementation-specific commitments and are shared across several theoretical traditions.

More specific mechanistic claims can carry lower current confidence while being scientifically valuable because they expose the framework to stronger tests.

A replicated phenomenon can coexist with uncertainty about its explanation.

A successful computational model can capture genuine structure without being a literal map of neural implementation.

A framework can organise a large body of evidence without every compatible finding uniquely confirming it.

Predictive processing should therefore neither be accepted because of its breadth nor rejected because stronger versions remain uncertain.

Its explanatory value must be assessed where each claim is made:

Which formulation?

Which prediction?

Compared with which alternative?

At which explanatory level?

Supported by what evidence?

And what result would make us revise the claim?

Those questions do not diminish predictive processing.

They are what allow its genuinely supported contributions to be distinguished from claims that remain promising, underspecified or unresolved.

The appropriate verdict is therefore neither that predictive processing is simply true nor that it is simply false.

It is a productive but heterogeneous theoretical family whose claims warrant different degrees of confidence—and whose scientific value is greatest when broad explanatory ideas are converted into specific accounts that can genuinely be distinguished from alternatives.

Behind this page

The claims this research note makes, the evidence behind them, and the limits it accepts.

Evidence status

High confidence

Strongly supported, though resting on synthesis or principle rather than a single decisive body of evidence.

Claims

  1. Predictive processing is a family of related approaches rather than one uniformly specified theory

    High confidence

    What this does not assert: Different authors draw the boundaries between formulations differently.

  2. In many experiments, neural responses to expected stimuli are smaller than responses to unexpected stimuli

    Established

    What this does not assert: Reported across several sensory paradigms.

  3. Reduced responses to expected stimuli are not unique to a prediction-error mechanism

    High confidence

    What this does not assert: Sharpening, adaptation, repetition suppression and attention can produce related effects.

  4. A finding compatible with predictive processing does not show that predictive processing is better supported than plausible alternatives

    Canonical inference

    What this does not assert: Compatibility is weaker evidence than discrimination.

  5. A replicated effect is not automatically a replicated explanation

    Canonical inference

    What this does not assert: Replication increases confidence that something occurs, not why.

  6. Experiments carry more evidential weight when competing accounts expect different outcomes

    High confidence

    What this does not assert: Discriminative design is what separates models.

  7. In specific perceptual paradigms, studies designed to discriminate models have reported results favouring prediction-error accounts over tested sharpening alternatives

    Established

    What this does not assert: The comparison is local to the paradigm and the alternatives tested.

  8. A discriminative result in one sensory paradigm does not establish predictive processing globally

    Canonical inference

    What this does not assert: The more specific the victory, the more specific the conclusion.

  9. Description, prediction, causal explanation, mechanism and neural implementation are different explanatory achievements

    High confidence

    What this does not assert: Success at one level does not confer success at another.

  10. Variables in a successful formal model should not automatically be treated as literal components of the system

    Canonical inference

    What this does not assert: Mappings from model to biology are empirical hypotheses.

  11. A fitted precision parameter is not automatically an identifiable neural precision signal

    Canonical inference

    What this does not assert: Nor is a prediction-error variable automatically a distinct neuronal population.

  12. Predictive coding refers to more specific computational or neural architectures than the broad predictive-processing family

    High confidence

    What this does not assert: Higher levels generate predictions and mismatch signals what remains unexplained.

  13. A model can be useful, predictive and theoretically revealing while its biological implementation remains uncertain

    High confidence

    What this does not assert: Predictive success does not confer mechanistic reality.

  14. Predictive-processing models differ in priors, precision, hierarchical level, learning history and task context

    High confidence

    What this does not assert: Complexity is not itself an objection.

  15. Compatibility provides limited evidence when auxiliary assumptions remain unconstrained until after the result is known

    Canonical inference

    What this does not assert: The question is what the formulation ruled out in advance.

  16. The claim that predictive processing is inherently unfalsifiable goes too far

    High confidence

    What this does not assert: Specific predictive-coding models can make risky predictions.

  17. Testability varies across the predictive-processing family

    Canonical inference

    What this does not assert: Broad frameworks resist decisive single experiments; narrower models do not.

  18. Retrospective accommodation and prospective testing have different evidential status

    Canonical inference

    What this does not assert: A post hoc explanation can begin the next test, not replace it.

  19. Neural systems respond differently to expected and unexpected information in many experimental settings

    Established

    What this does not assert: Mismatch-sensitive responses are widely reported in sensory systems.

  20. Manipulating cortical feedback can alter expectation-related mismatch responses

    Established

    What this does not assert: Provides causal evidence that top-down pathways contribute to some of these phenomena.

  21. Enhanced responses to unexpected stimuli can arise from adaptation, novelty, salience, attention or learning

    High confidence

    What this does not assert: Interpretation depends on the paradigm and its controls.

  22. A formal prediction error, a prediction-error computation, a prediction-error signal and a prediction-error circuit are different levels of claim

    Canonical inference

    What this does not assert: Evidence at one level does not automatically establish the others.

  23. A system can be described successfully by a Bayesian model without that description specifying a predictive-coding neural implementation

    Canonical inference

    What this does not assert: Levels of description are not interchangeable.

  24. Uncertainty is not prediction error

    Canonical inference

    What this does not assert: A system may remain uncertain when a prediction matches the input; a large mismatch need not mean large uncertainty.

  25. Some syntheses have not found a simple large-scale segregation between prediction and error processing in cortex

    High confidence

    What this does not assert: Newer models suggest more diffuse representation than sharply separated populations imply.

  26. The claim that cortex universally implements one canonical prediction/error circuit architecture remains substantially uncertain

    Canonical inference

    What this does not assert: Expectation-sensitive processing is far better supported than any universal architecture.

  27. Support does not automatically transfer from one formulation or level of claim to another

    Canonical inference

    What this does not assert: Every extension of scope adds its own evidential burden.

  28. Theoretical coherence is not discriminative empirical evidence

    Canonical inference

    What this does not assert: Elegance of fit between theories can obscure missing tests.

  29. There is no scientifically useful single confidence rating for predictive processing as a whole

    Canonical inference

    What this does not assert: Confidence must be assigned claim by claim and level by level.

  30. Free-energy formulations and active inference add commitments beyond the claim that sensory processing uses prediction

    High confidence

    What this does not assert: Active inference extends the formalism into action and policy selection.

  31. Evidence supporting one predictive-coding architecture does not automatically establish a free-energy formulation

    Canonical inference

    What this does not assert: Nor does it establish active inference.

  32. Evidence against one classical predictive-coding implementation would not by itself decide every broader predictive-processing formulation

    Canonical inference

    What this does not assert: Falsification is as local as confirmation.

  33. Context, prior information and learned regularities can influence what a person perceives

    Established

    What this does not assert: Established independently of any particular computational account.

  34. Demonstrating that expectation shapes perception does not identify the computational or neural mechanism responsible

    Canonical inference

    What this does not assert: Attention, decision processes and associative learning remain candidates.

  35. Constructive perception does not logically entail predictive processing

    Canonical inference

    What this does not assert: The phenomenon needs explaining; the framework is one candidate explanation.

Where to go from here

Next published piece

What Must a Theory of Emotion Explain?

Before emotion theories can be compared, the explanatory problem has to be clear: what must a scientifically adequate theory explain about emotional experience, bodily and neural processes, behaviour, context, variation and change through time?

Continue through the Library →See where this sits in the graph →

Back to the Library →