Section 3.20

Signal Detection Theory

A radiologist looks at a chest X-ray. Somewhere in the gray there is either a tumor or nothing, and the image does not make clear which: diseased tissue sometimes looks clean, and healthy tissue sometimes looks suspicious. Signal detection theory is the statistics of decisions made on evidence that overlaps like that. It separates two things that every raw accuracy score mixes together: how well the observer can tell signal from noise (sensitivity), and how willing they are to say “yes” (response bias).

One question, four outcomes

The basic experiment is simple. On every trial something is either there (signal) or not (noise), the observer answers yes or no, and each trial is counted in one of four cells. A yes to a real signal is a hit; a no is a miss. A yes when nothing was there is a false alarm; a no is a correct rejection. This is the same 2×2 table you met as the confusion matrix, with a human in the role of the classifier.

Because each row must sum to 100%, the whole table comes down to two independent numbers. They are the hit rate H (the share of signal trials answered yes) and the false-alarm rate FA (the share of noise trials answered yes). A percent-correct score keeps only one number, and the number it loses is the more interesting one. An observer who answers yes on every trial gets a perfect 100% hit rate, along with a 100% false-alarm rate, and has detected nothing at all. Accuracy alone cannot tell that strategy apart from skill.

Two bells and one line

The model behind the theory is simple enough to draw. Every trial produces an internal evidence value in the observer's head: how much the X-ray looks like a tumor, how familiar a word feels. Noise trials produce evidence scattered around a low average. Signal trials produce evidence from the same bell-shaped distribution shifted to the right. The distance between the two bells, measured in standard deviation units, is d′ (“d-prime”): the observer's sensitivity. When the bells are far apart, signal and noise rarely produce similar evidence. When they overlap, the observer will often confuse them, and no strategy can prevent that.

The observer's decision rule is a single vertical line, the criterion: answer yes whenever the evidence is to its right. Placed low (to the left), the rule is liberal and produces many hits and many false alarms. Placed high, it is conservative and produces few of both. The statistic c measures the criterion's position as its distance from the neutral point halfway between the bells, so c = 0 is unbiased, positive is conservative, and negative is liberal.

Both parameters come from the z-scores you already know. Convert the two observed rates into normal quantiles:

d′ = z(H) − z(FA)      c = −[z(H) + z(FA)] / 2

A third index, β = ec·d′, expresses the same bias as a likelihood ratio. It says how many times more probable the evidence at the criterion must be under signal than under noise before the observer says yes. Try all of them in the playground below.

🎛️ The signal detection playground

Top: the noise and signal + noise evidence distributions. Drag the criterion line sideways (or focus the chart and use ← →); everything to its right gets a “yes.” The four shaded regions, the 2×2 grid, and the ROC point all follow. Sweep moves the criterion across the whole axis to trace the ROC curve.

Hit rate—
False alarms—
d′ = z(H) − z(FA)—
criterion c—
bias β—
AUC = Φ(d′/√2)—
respond “yes”
respond “no”
signal
present
Hit
—
Miss
—
noise
only
False alarm
—
Correct rejection
—

Skill or strategy: two ways to get more hits

The playground shows the central point of this lesson. There are exactly two ways to raise a hit rate. Pull the bells apart (raise d′), and hits rise while false alarms fall: the observer discriminates better. Or slide the criterion left (lower c), and hits rise while false alarms rise with them: perception has not improved, only the willingness to say yes. Only the first is skill. A memory drug that “improves recognition by 10%” has shown nothing until the false-alarm rate is reported too.

Neither parameter is more “real” than the other, though. Criterion placement is where costs, rewards, and base rates enter the decision. A smoke detector is engineered to be extremely liberal, because a false alarm costs a minute of annoyance while a miss costs a house. A criminal jury is instructed to be conservative for the mirror-image reason. Airport screeners, spam filters, and drowsy radiologists all operate at some point on this trade-off, and signal detection theory lets you say where, separately from how good they are.

Where to set the criterion

Given the costs, rewards and base rates, the best criterion can be calculated exactly. The optimal policy is to answer yes whenever the evidence is more likely under signal than under noise by some constant factor, and that factor is the β above:

βopt = [P(noise) / P(signal)] × [cost of a false alarm / cost of a miss]

Rare signals push β up, because most yeses would be wrong anyway, and expensive misses push it down. In the criterion units the playground uses, copt = ln(βopt) / d′. Try the examples below on the sliders, at their default d′ = 1.5. Equally common signal and noise trials with equally costly errors give βopt = 1 and copt = 0, the slider's home position. Make signal trials rare, one in ten, and βopt = 9 lifts the ideal criterion to c = +1.46, where hits fall to 24% and false alarms to 1.3%. That may look too cautious, but at a 10% base rate a neutral observer's yeses are wrong more often than not. A smoke alarm is at the other extreme: if a miss costs twenty times a false alarm, βopt = .05 and copt = −2.00.

The most useful case is one where both forces act at once. Suppose that in a screening program 5 people in 1,000 carry the disease, and a missed case costs a hundred times as much as a needless recall. The base-rate term (199) and the cost term (1/100) very nearly cancel: βopt = 1.99, copt = +0.46, a hit rate of 61% and a false-alarm rate of 11%. Neither of those two rates looks acceptable on its own, yet both follow from one line of arithmetic. It is the prior odds times the likelihood ratio from Bayesian thinking, applied to a yes/no decision.

Sweep the criterion and you draw the ROC

Every criterion position gives one (false-alarm rate, hit rate) pair: a single point in the square panel above. Sliding the criterion across the whole evidence axis moves that point along a curve, from the most conservative corner (never say yes: 0, 0) to the most liberal one (always say yes: 1, 1). That curve is the receiver operating characteristic. The “receiver” in the name is a 1950s radar operator judging blips, because the theory began in radar research. The ROC curve you met in the ML course is the same object, with a model's score threshold in place of the observer's criterion.

The curve's shape depends on d′ alone. Moving the criterion moves the operating point along the curve, and only a change in sensitivity lifts the whole curve toward the top-left corner. So the area under the ROC curve summarizes performance without being affected by bias. For the equal-variance model, AUC = Φ(d′/√2), so d′ = 1 gives 0.76 and d′ = 2 gives 0.92, the same numbers the ML lesson reports for its separation slider. Memory researchers use this all the time. Asking participants for confidence ratings, not just yes or no, gives several criteria at once, an empirical ROC, and a sensitivity estimate that a shift in criterion cannot change.

🕹️ Be the detector

Twenty quick trials. Each shows one evidence reading drawn from the model above at d′ = 1.5, half the trials signal and half noise, shuffled. Decide where to put your criterion and call each one; your own hits and false alarms then give your measured d′ and c.

Press Start.

From counts to d′

Real data arrive as four counts, and the arithmetic takes two z-transforms. Say a recognition-memory test used 20 old and 20 new words, and a participant produced 15 hits and 4 false alarms. Then H = .75 and FA = .20, so z(.75) = 0.674 and z(.20) = −0.842 (you can look up either one in the site's z, t, χ² and F tables). That gives d′ = 0.674 − (−0.842) = 1.52 and c = −(0.674 − 0.842)/2 = 0.08: good discrimination and close to neutral bias. The “Try it yourself” box below runs the same four lines in R or Python.

The 0-and-1 problem. A perfect hit rate or a zero false-alarm rate breaks the formula, because z(0) and z(1) are infinite. The standard fixes move the extreme proportions slightly inward. This lesson (including the game above) uses the log-linear rule: add 0.5 to every count and 1 to every trial total before converting, for all participants, extreme or not. The common alternative, the 1/(2N) rule, replaces only the 0s and 1s. Either is acceptable, as long as you apply one consistently and name it in your methods section.

When the two bells are not the same width

All of the above assumes the two evidence distributions have the same spread. Only then can a single d′ stand for sensitivity. Memory data have contradicted that assumption for decades. The way to check it is the zROC: take the several (false-alarm rate, hit rate) pairs a confidence-rating experiment gives you, convert both to z-scores, and plot z(H) against z(FA). Equal variance predicts a straight line of slope 1. Old/new recognition studies consistently find a slope nearer 0.8, meaning the studied-item distribution is about 1.25 times wider than the new-item one. There is a simple reason: studied items differ in how well they were learned, while unstudied items have nothing to differ in.

A slope other than 1 causes a problem. Plain d′ = z(H) − z(FA) then depends on the observer's criterion as well as their sensitivity. Put the signal distribution at 1.5 with an SD of 1.25 against standard-normal noise, which gives a zROC of slope .80 and intercept 1.20. Then compute the usual statistics at six criterion positions:

CriterionHit rateFalse alarmsPlain d′c
liberal94.5%69.1%1.10−1.05
 88.5%50.0%1.20−0.60
 78.8%30.9%1.30−0.15
 65.5%15.9%1.40+0.30
 50.0%6.7%1.50+0.75
conservative34.5%2.3%1.60+1.20

The observer and the true sensitivity are the same in every row, yet plain d′ gives six different answers, because only the willingness to say yes changed. So two groups that differ in criterion will look as though they differ in sensitivity. That is the confound d′ is meant to remove.

Fixing it takes more than one point on the ROC, and that is the practical reason to collect confidence ratings. Given the whole zROC, the summary that remains valid under unequal variance is da, the separation divided by the root-mean-square of the two spreads:

da = (μS − μN) / √[(σN² + σS²) / 2]

For the observer in the table that is 1.33, with an area Az = Φ(da/√2) = .826, where Φ (phi) is the normal CDF. The six equal-variance readings above would have given areas anywhere from .782 to .871 for that same curve. Check the zROC slope before you quote a d′, and when it is not close to 1, report da or Az and say which.

The same two-number logic applies to recognition-memory experiments (old/new judgments), audiology and vision screening, radiologists' diagnostic performance, eyewitness lineups, and quality inspection. It also connects directly to the previous lesson: the cumulative-Gaussian psychometric function is the curve a signal detection observer produces as the stimulus grows stronger. Section 3.19 fits that response curve, and this lesson gives the decision theory behind it.

Why it matters: whenever a person or a model makes yes/no calls on ambiguous evidence, percent correct mixes strategy with skill. Reporting d′ and c (or an ROC and its area) keeps them separate. A training program that sharpens perception and one that only coaches people to say yes more often can produce identical hit rates and completely different signal detection profiles.

Problem 76 of the practice problems computes d′, c, β and AUC from raw hit and false-alarm counts, including the participant who scored worse on percent correct while being the more sensitive observer.

Common questions

What is a good d′ value?

Zero means the observer cannot tell signal from noise at all, and values grow without a fixed ceiling. The ROC identity AUC = Φ(d′/√2) gives useful reference points: d′ = 1 corresponds to getting a two-alternative comparison right about 76% of the time, and d′ = 2 to about 92%. Beyond 3, performance is so close to perfect that hit and false-alarm rates approach 1 and 0, where the estimate itself becomes unstable. What counts as good depends on the task. A d′ of 1 is respectable for faint stimuli near threshold and alarming for a tumor-versus-clean judgment.

My task was two-alternative forced choice, not yes/no. Does d′ still work?

Yes, with a different formula, and the result is less affected by bias. The observer compares two intervals and picks one, so there is no yes/no criterion to place, and any preference for one interval is usually small. That makes proportion correct close to a pure measure of sensitivity on its own. Convert it with d′ = √2 · z(Pc): 76% correct is d′ = 1 and 92% is d′ = 2, the same two reference points as the AUC identity, because it is the same formula. The conversion does not let you mix designs, though. A 2AFC d′ and a yes/no d′ from the same observer estimate the same underlying sensitivity, but they come from different tasks with different amounts of measurement noise. Report the design beside the number, and do not pool the two.

What should I do when a hit or false-alarm rate is exactly 0 or 1?

The z-transform sends those proportions to infinity, so d′ cannot be computed from them directly. Two standard corrections exist. The log-linear rule adds 0.5 to every count and 1 to every trial total before converting, and it is applied to all observers, not just the extreme ones. The 1/(2N) rule replaces only the 0s and 1s, with 1/(2N) in place of 0 and 1 − 1/(2N) in place of 1. Both pull extreme estimates toward the middle. The log-linear version is less biased in simulations and is what this lesson uses. Whichever you pick, apply it to every observer and name it in your methods section.