MMP.13:4.4 - Construct a posterior when conditional probability is needed
Supply a prior probability law Pi with its subject meaning. Reweight that law by the observed likelihood and normalize, preserving any point masses and continuous parts. For a set A to which the prior assigns a probability,
Pi_y(A) = integral_A L(theta;y) Pi(dtheta) / integral L(u;y) Pi(du).
Pi_y denotes the posterior law. Integration against Pi means combining values under the actual prior: a sum for discrete unknowns, integration against a density where one exists, and both contributions for a mixture. A function-valued unknown requires a specified prior law and likelihood on that space, not an assumed ordinary density.
The denominator must be finite and positive; an unnormalized expression alone does not establish a posterior probability law. For a real vector theta whose prior has a density pi(theta) with respect to ordinary volume dtheta, the posterior density is
pi(theta | y) = L(theta;y)*pi(theta) / integral L(u;y)*pi(u) du.
This density formula is a representation of the preceding law construction under that condition. A uniform prior depends on the parameterization, and an improper prior requires a separate argument that the posterior exists.
For example, give a failure probability p prior mass 1/2 at p=0 and a uniform distribution on [0,1] for the remaining 1/2. One failure-free Bernoulli observation has likelihood 1-p. The unnormalized atom has mass 1/2; the weighted continuous part has mass 1/4. Normalizing by 3/4 leaves posterior mass 2/3 at zero. The remaining 1/3 has conditional density 2*(1-p) on [0,1]. Using only the ordinary density would lose the atom.
Retain joint dependence when removing nuisance unknowns. Sum or integrate the joint posterior over them, or calculate g(theta) from joint posterior draws. Independently combining draws from marginal distributions changes the joint law unless independence is established.
Derive the posterior of q through that transformation. For a set B,
P(q in B | y) = integral 1{g(theta) in B} Pi_y(dtheta).
A credible set has its stated posterior probability under this model and prior. A posterior mean, median or quantile is a chosen summary of that law. In general, E[g(theta)|y] differs from g(E[theta|y]); :5.1 computes the difference.
A penalized optimum from MMP.12 can coincide with a posterior mode when its objective represents the chosen likelihood and prior. That optimum still does not supply the posterior spread. If a prior resolves an otherwise unidentified difference, preserve that source of the resolution in the returned result.