Section 3.7

Bayesian Thinking

The methods in this course so far have been frequentist: probability means long-run frequency, and parameters are fixed but unknown numbers we estimate. In Bayesian statistics, probability is a degree of belief, parameters get full probability distributions, and learning from data means applying one rule, again and again, to update that belief.

Bayes' rule, in words

Textbooks use the names Bayes' rule and Bayes' theorem for the same statement, and this lesson uses the shorter one. You start with a prior, what you believed before seeing the data. You collect data, whose likelihood says how well each possible parameter value explains what you saw. Multiply them and renormalize, and you get the posterior, your updated belief:

posterior ∝ likelihood × prior

That is the whole method. Today's posterior becomes tomorrow's prior, so beliefs change as evidence accumulates.

🎮 Prior → Data → Posterior

Estimating a proportion (say, a coin's chance of heads). Set your prior belief, then add data and watch the posterior (purple) form between your prior and what the data alone suggest.

Prior mean—
Data so far—
Posterior mean—

What the playground shows

Data overwhelms the prior. With little data the posterior stays close to your prior. Increase the trials and the posterior moves toward the observed rate and gets narrower. With enough evidence, your starting belief barely matters. A strong prior takes more data to move, and a flat (weak) prior lets the data dominate from the start.

  • A flat / weak prior says "I don't know much," so the posterior is driven almost entirely by the data, and often looks a lot like the frequentist answer.
  • A strong prior encodes real prior knowledge, useful when data is scarce, but it should be defensible, because it can shift your conclusions.

The same rule as a table

The playground draws Bayes' rule as three curves. An assignment almost always asks for it as a table, and the table is the version you can check by hand. Write down a short list of candidate values for the parameter and give each one a prior. For each candidate, work out how probable your data would be if it were the true value. Multiply the two columns together, and divide each product by the column's total.

Here is a concrete example. A bottle cap spun on its edge lands lip-up or lip-down at an unknown rate. Take nine candidate rates, p = .1 through .9, and give each the same prior weight of 1/9. A prior like that is called a non-informative prior, or a flat one. It makes the posterior depend only on the data, so it is the usual default in an assignment and the easiest choice to defend. It also has costs. It assumes that .1 and .7 were equally plausible for a spun cap, which they are not, and a list of nine rates has already ruled out every value in between. So even a non-informative prior is an assumption about the parameter.

Now spin the cap twelve times and count nine lip-up. The likelihood column is the binomial probability mass function read at each candidate:

P(9 of 12 | p) = C(12, 9) × p⁹ × (1 − p)³

which at p = .7 is 220 × .7⁹ × .3³ = .2397. Do that for each of the nine candidates to fill the column.

Hypothesis pPriorLikelihoodPrior × likelihoodPosterior
.1.11111.6e-71.8e-82.1e-7
.2.1111.00016.4e-6.0001
.3.1111.0015.0002.0019
.4.1111.0125.0014.0162
.5.1111.0537.0060.0697
.6.1111.1419.0158.1841
.7.1111.2397.0266.3110
.8.1111.2362.0262.3065
.9.1111.0852.0095.1106
total1.08561

Two quantities in the table have names that sound harder than they are.

The denominator is the total of the fourth column, .0856. It has two names, the marginal likelihood and the evidence. It is the probability of seeing 9 lip-up out of 12 at all, averaged over the nine hypotheses with their prior weights. Dividing each product by it makes the last column sum to 1, as a probability distribution over the nine candidates must.

The likelihood column is a function, though not the kind the notation suggests. P(9 of 12 | p) is the ordinary binomial formula from §1.8, which normally answers "how many successes will a known rate produce?" Here the data are fixed at 9 of 12 and p varies. Read that way, the same formula is called the likelihood function. It does not have to sum to 1 down the column, and it does not: .0001 + .0015 + .0125 and the rest come to .7708, a number with no meaning. So a likelihood is not a probability distribution over the parameter, and it cannot become one without a prior.

The posterior column gives the answer. The most probable single candidate is .7 at .311, with .8 almost level at .307: twelve spins cannot separate those two. Multiply each p by its posterior and add the nine products, and the posterior mean is .715. Everything at or below .4 has been all but eliminated, and .1 is down to two parts in ten million.

The table is a grid approximation. The nine rows are a grid laid across the parameter's range, and each row does by hand the multiplication that the smooth version does at every value at once. Add rows and the grid gets finer: nineteen candidates .05 apart put the peak on .75 and move the posterior mean to .7143. A grid works whenever there is no closed form, which is most of the time, provided there are few enough parameter values to list. The MCMC further down this page does the same job for parameter spaces too big to list.

🎮 Bayes by Hand

The table above as an interactive. Choose how many hypotheses to spread across the range, set the prior, and enter a count of successes out of trials. The four columns recompute on every change, the bars draw the posterior, and the shaded rows are the smallest set of hypotheses holding 95% of it. Nothing here is random.

Posterior over the hypotheses

Prior, likelihood, product, posterior

The shaded rows are the highest-posterior hypotheses, taken in order until their running total clears 95%. The next lesson reads a credible region off a grid the same way. The blue figure is the most probable single hypothesis.

Posterior mean—
Most probable—
95% region—
Its actual mass—
Evidence (column total)—
Gap from the Beta posterior—

Try two things. Drag the hypothesis count from nine up to twenty-one and watch the posterior mean settle: a coarse grid is an approximation, and you can see its error. Then switch the prior to Informed with the peak at .30, and the posterior mean falls from .715 to .542 on the same twelve spins. That prior counts as ten imaginary trials and moves the answer this far, so the reader of your report needs to see it stated.

Yesterday's posterior is today's prior

Data rarely arrive all at once, and the rule does not require them to. Spin the cap eight more times and get five lip-up. You can continue in two ways. Use the posterior from the first twelve as the prior for the next eight and run the table again. Or pool both samples, 14 of 20, and run the table once from the original flat prior.

Both routes give the same posterior, identical to the last digit the computer keeps. Press Combine a second sample and the two right-hand columns show the same nine numbers, .4026 at p = .7 by either route.

One line of algebra explains it. Updating twice multiplies the flat prior by p⁹(1 − p)³ and then by p⁵(1 − p)³; pooling multiplies it by p¹⁴(1 − p)⁶ in one go. Those are the same product. The binomial coefficients do differ between the routes, 220 × 56 against 38,760, but a coefficient is the same constant in every row, so it cancels in the division that normalizes the column.

So the order of the data does not matter, and neither does how they are divided into batches. A class exercise in which pairs pool into fours and fours into eights gives the same answer from either end. It also means a posterior summarizes everything the data have shown so far: once you have it, you can discard the counts and keep updating.

The conjugate shortcut

The grid does the arithmetic the long way. For this pairing of prior and likelihood there is a shortcut, and it is why the playground at the top of the page can draw an exact curve rather than a histogram.

A Beta prior with a binomial likelihood gives a Beta posterior, and the update is simple counting:

Beta(α, β) prior + k successes in n trials → Beta(α + k, β + n − k)

Here α (alpha) and β (beta) are the Beta distribution's two shape parameters. Add the successes to the first and the failures to the second. Nothing else changes. A prior family that produces a posterior in its own family this way is conjugate to the likelihood, and Beta is the conjugate prior for a proportion.

The flat prior over nine candidates is the grid's version of Beta(1, 1), which is level across the whole interval. Nine of twelve therefore gives Beta(1 + 9, 1 + 3) = Beta(10, 4), with mean 10/14 = .7143 and mode 9/12 = .75. Read that density at .1, .2 and so on up to .9, scale the nine numbers so they sum to 1, and you get the posterior column of the table above, digit for digit. The widget's last readout is the largest gap between the grid’s posterior column and that Beta density, and it reads .0000 at every setting, informed prior included.

Before computers, most practical Bayesian work used conjugate families. They were the cases where the integral in the denominator had a known answer, and a model without a conjugate prior could be written down but not used. MCMC removed that constraint, so applied Bayesian analysis is no longer limited to a small catalog of tractable cases. Conjugate priors are still useful for two reasons. They are exact, so they can check a sampler: if a chain cannot reproduce Beta(10, 4) from nine of twelve, the problem is the chain and not the mathematics. And they make a prior easy to interpret, because Beta(α, β) acts like α + β trials already run, of which α succeeded. That is the plainest way to tell a reader how much prior belief you brought.

The grid needs no conjugate pair. Swap in a likelihood with no conjugate prior and the five columns still work, so the grid generalizes and the conjugate shortcut checks it.

What "modern computing" actually does

The playground above draws an exact posterior because a proportion with a Beta prior is one of a small number of textbook cases where the multiplication has a closed-form answer. Realistic models almost never have one, and for most of the twentieth century that was the practical objection to Bayesian methods: the rule was simple, but the integral could not be computed.

MCMC (Markov chain Monte Carlo) changed that. It does not solve for the posterior. The computer takes a long random walk through parameter space, designed so that the time it spends in each region is proportional to that region's posterior probability. The points it visits are a sample from the posterior, which you can average, plot, or read quantiles off. MCMC replaces a formula with simulation, much as bootstrapping replaces one with resampling. Stan, JAGS, PyMC and the Bayesian modules inside JASP all run some version of MCMC, so a Bayesian analysis reports how many samples it drew and whether its chains agreed with each other.

Why people love (and argue about) Bayes

The Bayesian framework gives you what you usually want: a full distribution for the parameter. You can say "there's a 95% probability the rate is between X and Y", a statement that frequentist confidence intervals cannot make. The cost is the prior: it must be chosen, and different priors can give different answers when data are thin. Chosen openly and reported, though, a prior is also a way to use real prior knowledge.

Why it matters: Bayesian methods are now widely used in A/B testing, clinical trials, machine learning and election forecasting. They express uncertainty directly as probability, and modern computing has made them practical. The next lesson turns the posterior into concrete summaries with credible intervals.

Problem 74 in the practice problems runs the classic base-rate case by hand: a 95% sensitive test, a rare condition, and a positive result that leaves the patient at about 9%.

Problem 85 builds a nine-row grid from scratch on thirteen germinated seeds out of twenty, and Problem 86 runs the same count through the Beta shortcut and asks where the two would stop agreeing.

Common questions

What is the main difference between Bayesian and frequentist statistics?

The meaning of "probability". Frequentists treat parameters as fixed unknowns and put probability on data procedures ("5% of such intervals miss"). Bayesians put probability on parameter values themselves, as degrees of belief updated by evidence ("the rate is 95% likely between .55 and .72"). The Bayesian version answers the question people usually ask, but it requires you to specify a prior.

Can a hypothesis I gave zero prior weight ever come back?

It cannot. The posterior is proportional to prior times likelihood, so a row whose prior is 0 has a product of 0 whatever the data say, and it stays at 0 through every later update. Ruling a value out in the prior is permanent, and no amount of evidence reopens it. So leave a little weight on anything you are not certain is impossible, however unlikely it looks. The same applies in reverse: a hypothesis given a prior of 1 keeps a posterior of 1 whatever the data say. Bayes' rule rescales the beliefs you start with, and it cannot create one you did not have.

What is a Bayes factor?

The evidence ratio between two hypotheses: how much more probable the observed data is under H₁ than under H₀. BF = 10 means the data favors H₁ ten-to-one; BF = 1 means the data can't tell them apart. It's the Bayesian counterpart to significance testing, with two advantages over p-values: it can quantify evidence for a null, and it doesn't inflate with optional stopping.