Bayesian Estimation & Credible Intervals
Bayesian updating gives you a whole posterior distribution. Eventually, though, someone asks "so what's your estimate, and how sure are you?" Summarizing the posterior as a number and a range is Bayesian estimation. The range, called the credible interval, has the interpretation people often wrongly give to a confidence interval.
Two summaries of a posterior
- A point estimate: usually the posterior mean (the balance point), but sometimes the median or the mode (the peak, called the MAP, maximum a posteriori).
- A credible interval: a range that contains a stated share of the posterior probability, say 95%. The most common version is the central interval, chopping 2.5% off each tail.
🎮 Read the Credible Interval Off the Posterior
Estimating a proportion with a flat prior. The purple curve is your posterior, and the shaded band is the central credible interval at the level you choose. Add data and watch it narrow.
What a credible interval means
A 95% credible interval means what it sounds like: given your data and prior, there is a 95% probability the parameter lies in this range. Compare that to a frequentist confidence interval, where the 95% refers to the long-run behavior of the procedure, not this particular interval. The Bayesian statement is the direct probability claim people instinctively (but wrongly) read into confidence intervals.
Central vs. highest-density intervals
The shaded band above is a central interval: equal probability trimmed from each tail. An alternative is the highest posterior density (HPD) interval: the shortest range containing the target probability. For a symmetric posterior they are identical. For a skewed one the HPD interval stays closer to the peak and is a little narrower. Both are legitimate: central intervals are simpler, and HPD intervals are the shortest possible.
Reading the region off a grid
The posterior above is a curve, so its interval comes from a quantile function. A grid posterior, the table of candidate values built in §3.7, has no quantile function and does not need one. The highest-density idea works directly on the rows.
Sort the hypotheses by posterior mass, biggest first, and take them one at a time until the running total clears your target. Whatever you took is the region. Because a Beta prior times a binomial likelihood is always single-peaked, the hypotheses that come out are always neighbors, so the region reads as a range rather than a scatter of separate values.
On the nine-row table from the previous lesson, where a spun bottle cap landed lip-up nine times in twelve:
| Taken in order | p = .7 | .8 | .6 | .9 | .5 |
|---|---|---|---|---|---|
| Posterior mass | .3110 | .3065 | .1841 | .1106 | .0697 |
| Running total | .3110 | .6175 | .8016 | .9122 | .9818 |
Four hypotheses hold 91.2%, short of the target, so the fifth is added and the region is p = .5 through .9, holding 98.2% of the posterior.
That overshoot is not an error, but reporting the region as exactly 95% would be. A discrete posterior can hold only certain totals, so it usually cannot hold exactly 95% of itself. Name the region and the mass it actually holds: "the smallest 95% region runs from .5 to .9 and holds 98.2% of the posterior." A finer grid reduces the overshoot. At nineteen candidates .05 apart, the smallest region clearing 95% runs from .50 to .90 and holds 95.8%. The continuous Beta(10, 4) posterior that both grids approximate has a 95% highest-density interval of [.486, .926]. You can see the nineteen-row version on the grid widget with nineteen hypotheses loaded.
When you report the region, give the level you achieved rather than the one you asked for, and say which rule produced it. Central and highest-density intervals are different, and a skewed posterior separates them. Beta(10, 4) is skewed: its central interval is [.462, .909] and its highest-density interval is [.486, .926], shifted right and about .008 narrower.
Watch it tighten
Increase the trials and the posterior narrows, so the credible interval shrinks. Raise the credibility from 90% to 99% and the band widens, because a higher probability needs a wider range. It is the same trade-off between precision and confidence that you saw with confidence intervals.
Why it matters: Bayesian results in trials, forecasts and industry experiments are reported with credible intervals. Part of the appeal of Bayesian methods is that a credible interval can be read directly as a probability statement about the parameter.
Problem 75 of the practice problems takes a Beta(8, 12) prior from last quarter and updates it on 33 completions out of 60 users. It asks for the posterior's mean, mode and standard deviation, then why that mean is below the observed completion rate. To watch the Beta distribution itself change shape as its two parameters move, the distribution playground has it alongside eight others.
Problem 85 asks for the smallest 95% region of a nine-row grid posterior and for the mass it actually holds, using the method this section describes.
Common questions
What is the difference between a credible interval and a confidence interval?
A 95% credible interval means "given the data (and prior), the parameter is 95% likely to be in here." A 95% confidence interval promises only that the procedure captures the true value 95% of the time across repeated studies. With flat priors and decent data the two are often numerically similar, but only the credible interval supports the direct probability statement.
How do I test a hypothesis with a credible interval?
Often you can simply read the probability off the posterior: P(θ > 0 | data) is an ordinary number you can quote, with no null hypothesis involved. When you do want a decision rule, the usual one is a ROPE, a region of practical equivalence. State in advance the range of values you would count as "no meaningful effect", say −0.1 to 0.1 on your outcome's scale, and compare it with the credible interval. If the interval falls entirely inside the ROPE, you can accept the null for practical purposes, which a p-value never allows. If it falls entirely outside, you reject the null. If it straddles the boundary, the data have not decided. Choosing the ROPE is a judgment about what matters in your field, so state it before you look, as you would a smallest effect size of interest.
Do Bayesian and frequentist results ever agree?
Very often. With flat or weak priors and reasonable sample sizes, credible and confidence intervals typically match to two decimals, because the likelihood dominates both. They differ when the prior carries real information or the data are thin, and their interpretation always differs. So the two frameworks usually agree, and a real disagreement tells you something: your prior is having a large effect on the result.