The question
How does predictive processing model perception and cognition as the interaction between prior expectations and incoming sensory information?
Definition
Predictive processing is a family of theoretical approaches that models perception and cognition in terms of predictions, incoming sensory information and the discrepancies between them. Rather than treating perception as a one-way progression from sensory input to representation, predictive-processing accounts propose recurrent interactions in which prior expectations help shape the interpretation of incoming signals while mismatches can drive updating. Its formulations differ in scope and empirical commitment, so predictive processing should be treated as a scientific model rather than as an established literal description of all neural or psychological function.
Perception does not begin from zero.
When you enter a familiar room, recognise an object in poor light or hear a familiar voice through noise, the perceptual system does not confront the incoming information as though nothing has ever been encountered before.
It already has structure.
Previous experience, current conditions and existing model states make some sensory patterns more expected than others.
Predictive-processing theories begin from that observation and propose a specific explanatory architecture: perception emerges through recurrent interaction between model-based expectations and incoming sensory information.
That proposal has become highly influential.
But it needs to be kept precise.
Predictive processing is not simply the claim that expectations affect perception.
It is not the claim that the nervous system ignores sensory evidence and sees whatever it expects.
And it is not one single, universally agreed algorithm.
It is better understood as a family of related scientific models organised around prediction, sensory constraint, mismatch and revision — one way of explaining the constructive perception established in Perception Is Constructive.
Predictive Processing Is a Family of Models
The phrase predictive processing is often used as though it referred to one theory.
In practice, it covers a family of related approaches.
The minimal shared architecture can be described through four ideas:
existing model-dependent structure generates expectations;
incoming sensory information constrains those expectations;
mismatch between expected and incoming information matters;
the current estimate or model can be revised.
Predictive processing is a family of related models rather than one universally agreed theory with one fixed implementation.
More specific formulations can add further commitments, including:
hierarchical predictive coding;
generative models;
Bayesian inference;
precision weighting;
explicit prediction-error minimisation;
free-energy formulations.
Those additions are important.
They should not all be treated as part of the minimal definition of the entire family.
The minimal predictive-processing family can be described through prediction, sensory constraint, mismatch and revision; hierarchy, precision and particular generative or Bayesian machinery belong to more specific formulations.
This distinction matters because a theoretical construct can easily begin to sound like an observed biological object.
Predictive-processing models may contain:
prediction signals;
prediction-error signals;
precision terms;
hierarchical levels;
generative models.
Those variables can organise explanation without automatically identifying distinct neural mechanisms.
That is the epistemic discipline established in Scientific Models Are Tools.
Prediction Does Not Mean Conscious Forecasting
The word prediction creates an immediate source of confusion.
In ordinary language, prediction usually means anticipating something that will happen later.
Predictive-processing models use the term more broadly.
A prediction can concern expected:
sensory patterns;
lower-level states;
latent states within the model;
relations among signals.
The relevant expectation may concern what sensory activity should look like now, given the model's current state.
A predictive-processing prediction need not be a conscious expectation or a forecast of a future event.
Suppose a partially visible object is encountered in dim light.
Existing model structure may make some interpretations more plausible than others.
Within predictive-processing terms, those expectations can influence the current perceptual estimate without the person consciously thinking: "I predict that this object is a chair."
Likewise, a prior within the model should not automatically be translated into an explicit belief.
Existing model structure may reflect:
previous experience;
learned regularities;
current contextual structure;
other constraints already represented in the system.
Not every prior need be learned, and not every prior is consciously accessible.
Predictive-processing terminology therefore operates at a different level from ordinary conscious expectation.
Predictions and Sensory Evidence Interact
Predictive processing is sometimes reduced to the slogan: perception is prediction.
That is too crude.
Model-based expectations matter.
So does sensory evidence.
Predictive processing is an interaction model, not a prior-dominance model.
Imagine seeing an indistinct shape where a chair normally stands.
Existing model structure may initially support "chair" as a plausible interpretation.
Then better sensory information becomes available and reveals a stack of boxes.
The incoming evidence constrains the current estimate.
Model-based expectations influence inference, while sensory evidence constrains and can revise it.
This interaction is central.
The framework does not require: prediction always overrides sensation.
Nor does it divide the process neatly into:
subjective top-down information;
objective bottom-up reality.
Incoming sensory information is already processed within an interacting nervous system.
The distinction is functional.
Existing model structure contributes one source of constraint — including, in relevant formulations, the surrounding conditions treated theory-generally in Context Alters Interpretation.
Current sensory information contributes another.
When they do not align, mismatch becomes important.
Prediction Error Can Change Inference and Learning
Prediction error is one of the most important concepts in predictive processing.
It is also one of the easiest to misunderstand.
At the broad conceptual level: prediction error refers to mismatch represented within the model between expected and incoming or lower-level information; the exact computation and implementation differ across formulations.
It does not necessarily mean:
conscious surprise;
knowing that something is wrong;
making a bad decision;
failing to forecast a future event.
Prediction error is a modelling construct for mismatch, not necessarily a consciously experienced mistake.
Mismatch can contribute to revision on different timescales.
Current inference
A prediction error can alter the system's present estimate of what is causing or generating sensory information.
The indistinct object first interpreted as a chair may be re-estimated as boxes.
More persistent model change
Under suitable conditions, repeated or informative mismatch may contribute to changes in:
parameters;
associations;
model structure;
future expectations.
Prediction error can contribute to revision on different timescales: rapidly altering the current estimate and, under suitable conditions, contributing to more persistent model change.
Neither consequence follows automatically from every mismatch.
A discrepancy may be:
weak;
noisy;
expected;
unreliable;
assigned little influence.
This is why prediction error should not simply be equated with learning.
A mismatch can contribute to updating without implying that every prediction error automatically produces learning.
Predictive-processing updating is one proposed model of some experience-dependent change.
It does not replace the broader concept of learning established elsewhere in the Library, in What Is Learning?.
Predictions and Errors Interact Across Levels
Many influential predictive-processing formulations are hierarchical.
In those models, different levels can represent structure at different:
scales;
timescales;
degrees of abstraction;
forms of latent organisation.
The levels interact recurrently.
One influential implementation is classic hierarchical predictive coding.
In that architecture, higher levels generate predictions about lower-level activity, while residual mismatch information is propagated in the opposite direction.
Classic hierarchical predictive-coding models often assign descending pathways to predictions and ascending pathways to residual prediction-error information.
That is an important implementation.
It is not the universal definition of predictive processing.
A simplified version looks like this:
higher-level model state
↓ prediction
expected lower-level activity
↕ compared with
incoming lower-level activity
↑ residual mismatch
prediction-error information
This is a classic predictive-coding example, not a universal predictive-processing circuit.
The broader family is defined by recurrent prediction-centred interaction, not by one fixed wiring diagram.
Likewise:
Top-down and bottom-up describe directions or roles in modelled information exchange; they are not synonyms for subjective and objective information.
Top-down does not necessarily mean conscious belief.
Bottom-up does not mean information untouched by prior neural processing.
Prediction errors may also occur at multiple levels rather than as one single global error signal.
And recurrence itself is not uniquely predictive-processing evidence.
Many neural theories contain recurrent processing.
Generative Models Turn Perception into an Inference Problem
Many influential predictive-processing formulations use generative models.
A generative model represents how latent or hidden states within the model relate to expected observations.
Generative models represent relations between latent states and the sensory observations expected from them.
Those latent states may be interpreted as candidate causes of sensory input.
But their presence in a model does not independently establish that they correspond one-to-one with real causal entities.
This is another application of Scientific Models Are Tools.
The model provides a representational architecture.
Ontological interpretation requires additional evidence.
Generative models allow perception to be framed as an inference problem.
Given:
current sensory information;
existing model structure;
which latent state best accounts for what is being observed?
Perceptual inference is modelled as an automatic estimation problem, not deliberate reasoning about sensory data.
No conscious analyst inside the brain is comparing hypotheses.
"Inference" names the role assigned by the model.
Some influential predictive-processing accounts formalise this process in Bayesian terms.
Bayesian methods provide a mathematical language for combining:
prior information;
sensory evidence;
uncertainty.
But that does not justify saying: the brain literally computes Bayes' theorem in one simple universal form.
A computational description and its neural implementation are separate questions.
A generative model should likewise not be imagined as an inner movie or miniature replica of the world.
It is a theoretical structure relating latent states to expected observations.
Precision Changes the Weight of Uncertain Information
Prediction errors need not exert equal influence.
A discrepancy carried by reliable information should matter differently from one carried by highly noisy information.
Influential precision-based predictive-processing models handle this through precision.
In precision-based predictive-processing models, prediction errors can be weighted by their expected precision—their modelled reliability relative to uncertainty.
In relevant formulations, greater expected precision corresponds to lower expected uncertainty and stronger weighting of the signal.
The key word is expected.
Precision concerns expected reliability within the model, not whether the information is objectively correct.
A model can assign high precision to information that is in fact misleading.
Precision is therefore not truth.
Nor is it necessarily conscious confidence.
Precision is a modelled estimate of reliability, not necessarily a conscious feeling of certainty.
It should also not be equated with attention.
Some predictive-processing formulations explain aspects of attention through precision weighting.
That is a substantive theoretical proposal.
It is not the Library's definition of attention, which remains theory-general in Attention Selects Information.
Likewise, D6.2 does not attempt to define uncertainty generally; that architecture belongs to Uncertainty.
It introduces uncertainty only insofar as some predictive-processing models use relative uncertainty to determine how much influence different error signals should have.
Prediction-Error Minimisation Is a Modelling Principle
Many influential predictive-processing formulations organise inference around reducing prediction error over time.
If model-generated expectations repeatedly fail to account for incoming information, current estimates or longer-term model parameters may change in ways that reduce future mismatch.
This is what prediction-error minimisation attempts to capture.
The phrase should not be read psychologically.
Prediction-error minimisation is a modelling principle, not a claim that organisms consciously try to eliminate every surprise.
Nor does it imply that prediction error reaches zero.
Real perceptual systems operate under changing and noisy conditions.
Residual discrepancy is unavoidable.
Error minimisation also does not guarantee truth.
Prediction-error minimisation should not be equated with truth optimisation.
A system may reduce mismatch relative to:
an imperfect model;
limited evidence;
noisy input;
incorrect assumptions.
This section should also be distinguished from the earlier discussion of prediction error.
Earlier, the question was: what can mismatch do?
Here, the question is: how do some influential models organise the overall dynamics?
Some formulations connect prediction-error minimisation to variational free-energy minimisation under particular modelling assumptions.
That relationship is important.
But the Free Energy Principle adds broader formal commitments and is not identical to the minimal predictive-processing architecture introduced here.
D6.2 therefore does not define predictive processing through free-energy minimisation.
Related Frameworks Are Not Interchangeable
Researchers do not use these labels with perfectly consistent boundaries.
For the Library, the following distinctions are used to prevent conceptually different commitments from being collapsed.
Predictive processing
The broad family organised around model-based prediction, sensory constraint, mismatch and revision.
Predictive coding
More specific computational or neural schemes that implement prediction and residual-error signalling, often through hierarchical message passing.
Bayesian brain approaches
Probabilistic frameworks that model perception or cognition as inference under uncertainty.
Some predictive-processing formulations are Bayesian.
The two labels are not identical.
Free-energy formulations
Specific variational formulations that relate inference and learning to free-energy minimisation under additional assumptions.
Active inference
A related extension that brings action and policy selection into the same broader generative/free-energy architecture.
These frameworks overlap, but they are not interchangeable labels.
The distinction matters because each additional framework introduces additional commitments.
Evidence compatible with prediction error does not automatically establish every claim associated with:
Bayesian brain theory;
the Free Energy Principle;
active inference.
The taxonomy here is therefore operational rather than universal.
It keeps the Library from treating neighbouring frameworks as synonyms when they are not.
Model Variables and Explanatory Reach Require Evidence
Predictive processing is powerful partly because its vocabulary can organise a wide range of problems involving:
perception;
learning;
uncertainty;
attention;
context;
broader cognition.
That reach is scientifically valuable.
But breadth of application does not automatically establish breadth of truth.
The ability to describe a phenomenon in predictive-processing language is not, by itself, evidence that predictive processing uniquely explains it.
Likewise, model quantities should not be reified automatically.
A useful variable in a predictive-processing model does not automatically identify a distinct neural mechanism.
Some proposed neural implementations assign prediction and prediction-error roles to different:
populations;
pathways;
cortical layers.
Those are testable mechanistic proposals.
They are not universal anatomical facts simply because the model contains those roles.
Evidence compatible with prediction-error signalling can support the framework without uniquely establishing the complete architecture.
Support becomes stronger when evidence discriminates predictions of the model from credible alternatives.
How strongly current evidence accomplishes that discrimination is a separate question.
That is the role of D6.13 — Predictive Processing: Scope and Limits.
D6.2 establishes the model.
D6.13 evaluates how far its evidential authority should extend.
A Powerful Model, Not a Licence to Explain Everything
Predictive processing offers a coherent way to explain why perception need not begin from zero.
Existing model structure generates expectations.
Sensory information constrains them.
Mismatch can alter the current estimate and, under some conditions, contribute to more persistent change.
More specific formulations add:
hierarchical predictive coding;
generative models;
Bayesian inference;
precision weighting;
prediction-error minimisation.
Those additions deepen the framework without all belonging to its minimal definition.
Predictive processing models perception and cognition as recurrent interaction between prior expectations and incoming information, with mismatches capable of updating predictions, while remaining a theoretical framework whose scope and interpretation require empirical constraint.
Its value lies in making those interactions explicit enough to investigate.
Predictive processing is most useful when treated as a powerful model to test, refine and constrain—not as a vocabulary that automatically explains whatever it can describe.