CACrown ArchivesThe cinema collection
Menu
Research dossier · General Reference

Posterior predictive distribution

in Bayesian statistics, the distribution of a new data point marginalized over the posterior

Cross-disciplinary reference desk with index cards, atlas, dictionary and catalogue
General referenceInterpretive dossier study · Crown Archives visual atlas
Record originEnglish Wikipedia
Text licenseCC BY-SA 4.0
Source revisionJul 17, 2026
Entity authorityQ7234227
Source-derived summary

In Bayesian statistics, the posterior predictive distribution is the distribution of possible unobserved values conditional on the observed values.

Given a set of N i.i.d. observations

X

=

{

x

1

,

,

x

N

}

{\displaystyle \mathbf {X} =\{x_{1},\dots ,x_{N}\}}

, a new value

x

~

{\displaystyle {\tilde {x}}}

will be drawn from a distribution that depends on a parameter

θ

Θ

{\displaystyle \theta \in \Theta }

, where

Θ

{\displaystyle \Theta }

is the parameter space.

p

(

x

~

|

θ

)

{\displaystyle p({\tilde {x}}|\theta )}

It may seem tempting to plug in a single best estimate

θ

^

{\displaystyle {\hat {\theta }}}

for

θ

{\displaystyle \theta }

, but this ignores uncertainty about

θ

{\displaystyle \theta }

, and because a source of uncertainty is ignored, the predictive distribution will be too narrow. Put another way, predictions of extreme values of

x

~

{\displaystyle {\tilde {x}}}

will have a lower probability than if the uncertainty in the parameters as given by their posterior distribution is accounted for.

A posterior predictive distribution accounts for uncertainty about

θ

{\displaystyle \theta }

. The posterior distribution of possible

θ

{\displaystyle \theta }

values depends on

X

{\displaystyle \mathbf {X} }

:

p

(

θ

|

X

)

{\displaystyle p(\theta |\mathbf {X} )}

And the posterior predictive distribution of

x

~

{\displaystyle {\tilde {x}}}

given

X

{\displaystyle \mathbf {X} }

is calculated by marginalizing the distribution of

x

~

{\displaystyle {\tilde {x}}}

given

θ

{\displaystyle \theta }

over the posterior distribution of

θ

{\displaystyle \theta }

given

X

{\displaystyle \mathbf {X} }

:

p

(

x

~

|

X

)

=

Θ

p

(

x

~

|

θ

)

p

(

θ

|

X

)

d

θ

{\displaystyle p({\tilde {x}}|\mathbf {X} )=\int _{\Theta }p({\tilde {x}}|\theta )\,p(\theta |\mathbf {X} )\operatorname {d} \!\theta }

Because it accounts for uncertainty about

θ

{\displaystyle \theta }

, the posterior predictive distribution will in general be wider than a predictive distribution which plugs in a single best estimate for

θ

{\displaystyle \theta }

.

Prior vs. posterior predictive distribution

The prior predictive distribution, in a Bayesian context, is the distribution of a data point marginalized over its prior distribution

G

{\displaystyle G}

. That is, if

x

~

F

(

x

~

|

θ

)

{\displaystyle {\tilde {x}}\sim F({\tilde {x}}|\theta )}

and

θ

G

(

θ

|

α

)

{\displaystyle \theta \sim G(\theta |\alpha )}

, then the prior predictive distribution is the corresponding distribution

H

(

x

~

|

α

)

{\displaystyle H({\tilde {x}}|\alpha )}

, where

p

H

(

x

~

|

α

)

=

θ

p

F

(

x

~

|

θ

)

p

G

(

θ

|

α

)

d

θ

{\displaystyle p_{H}({\tilde {x}}|\alpha )=\int _{\theta }p_{F}({\tilde {x}}|\theta )\,p_{G}(\theta |\alpha )\operatorname {d} \!\theta }

This is similar to the posterior predictive distribution except that the marginalization (or equivalently, expectation) is taken with respect to the prior distribution instead of the posterior distribution.

Editorial summary

The public source identifies “Posterior predictive distribution” as in Bayesian statistics, the distribution of a new data point marginalized over the posterior. This brief keeps that definition visible, then builds a research path around Posterior, predictive and distribution.

Editorial reviewA concise reference frame for defining the subject, testing terminology and identifying the institution closest to the evidence. The current 497-word lead offers orientation but no explicit four-digit date, so chronology should not be assumed. The selected authority fields contribute no independent date. Its value is orientation rather than verdict, with Posterior, predictive and distribution providing the first useful test.
Editorial analysis

Why this record matters

A short description can identify a subject without explaining its stakes. For “Posterior predictive distribution”, the useful work is to connect “in Bayesian statistics, the distribution of a new data point marginalized over the posterior” to the records capable of establishing context and consequence.

Evidence profile

Vocabulary and entity names are the principal evidence signals here, because they determine the precision of every later search. The source revision retrieved here is dated Jul 17, 2026. The linked authority identifier is Q7234227. None of the 0 selected statements returned an explicit reference.

Critical limits

The absence of detail may reflect summary conventions rather than a lack of surviving documentation. The source lead contains qualifying language; that uncertainty should survive quotation, summary and reuse. Authority statements aid reconciliation but still require their own references, qualifiers and ranks to be checked.

How to read it

Use the entry as an orientation point, then follow its citations and revision history. Names, dates and institutional relationships should be checked against the original record.

Best used for
  • Subject orientation
  • Search vocabulary
  • Locating named sources
Verify next

The closest primary source, responsible institution and strongest cited specialist reference.

Three-step research path

  1. Establish the record: confirm the title “Posterior predictive distribution”, its source revision and the description used here.
  2. Expand the search: follow Posterior predictive distribution primary sources, Posterior predictive distribution archive and Posterior research across catalogues and specialist indexes.
  3. Test the account: compare the strongest cited source with the responsible institution’s current record and note any disagreement.

Questions for further research

  1. Which source most directly establishes the central claim about “Posterior predictive distribution”?
  2. Which cited source is closest to the event, object or claim?
  3. Which institution is responsible for the underlying evidence?
Subject index

Search terms from this dossier

Source & attribution

This entry incorporates text from Posterior predictive distribution” on English Wikipedia. Contributors are listed in the page history. Text is available under the Creative Commons Attribution-ShareAlike 4.0 License. Selected authority identifiers and statements are retrieved from Wikidata under CC0; their references and qualifiers remain part of the verification path.