Psychometric Functions & the PSE
Psychophysics turns perception into numbers: how bright a light must be before you notice it, how much a tone must rise before it sounds louder, how long an interval must last before it feels long. Every such measurement relies on the psychometric function, an S-shaped curve that links a stimulus you control to the responses an observer gives. The analysis consists of fitting that curve and reading two numbers off it, with a method you already met in logistic regression.
Psychophysics and the method of constant stimuli
Psychophysics measures the link between a physical stimulus and the perception it produces. Gustav Fechner began this program in 1860, and that work is often credited as the start of experimental psychology. Its most transparent procedure is the method of constant stimuli. You fix a ladder of stimulus levels spanning the interesting range, from a level almost everyone judges one way to a level almost everyone judges the other. You present each level many times in random order and record the proportion of one response at each. Plotted against stimulus level, those proportions trace the psychometric function. (Adaptive procedures such as staircases set each stimulus according to the last response; the last section of this lesson measures what that gains. Learn constant stimuli first, because it shows the entire curve.)
That one procedure answers a wide range of questions. In vision: the faintest contrast at which a pattern becomes visible, or which of two lines looks longer. In hearing: the quietest detectable tone, the measurement behind the audiogram in every hearing clinic, or which of two tones is louder. In touch: the smallest gap at which two pressed points stop feeling like one. In the lifted-weight experiments that produced Weber's law: which of two weights feels heavier. In the perception of time: whether an interval feels short or long. Detection tasks ask whether a stimulus is present at all; discrimination tasks ask which of two is more intense. Both yield the same shape, a proportion climbing smoothly from near 0 to near 1 as the stimulus grows.
The psychometric function
The psychometric function maps stimulus level to response probability. The analysis describes its smooth rise with a simple curve and summarizes that curve with a few parameters.
Three families do most of the work in practice. The logistic is the same S-curve that logistic regression uses. The cumulative Gaussian comes from classic signal detection theory: it is what you get if the observer's internal reading of the stimulus is the true level plus normal noise. The Weibull is asymmetric on the stimulus axis and is widely used for detection tasks. Each has two parameters, a location and a spread, and each is fit by maximum likelihood on the raw trial counts. Each trial is a Bernoulli outcome whose probability depends on the stimulus, so the fit uses the same binomial likelihood as a GLM.
To make this concrete, the rest of the lesson stays with one task: the temporal bisection procedure, a standard way to study how people perceive duration. The observer first learns two anchor durations, a short one (200 ms) and a long one (800 ms). Probe intervals of in-between lengths are then presented one at a time, and for each one the observer judges whether it felt closer to short or to long. The share of “long” responses at each probe duration is the proportion the method of constant stimuli collects, and the interactive below fits a curve to it.
🎮 The Bisection Experiment
Each dot is the share of “long” responses at one probe duration; the curves are maximum-likelihood fits, with dashed drops at each condition's PSE. The simulated observer is always Gaussian; the selector changes only the fitted model. Sliders re-analyze the same frozen trials; “New trials” draws fresh ones.
Two numbers: PSE and JND
The point of subjective equality (PSE) is the stimulus level at which the two responses are equally likely, the 50% point of the fitted curve. In bisection it answers a psychological question: which duration feels like the exact midpoint between the anchors? The answer is not necessarily the arithmetic middle. Timing studies often find bisection points near the geometric mean of the anchors: for 200 and 800 ms that is √(200 × 800) = 400 ms, not 500. This is one of the classic signs that our sense of time is compressive.
The just-noticeable difference (JND) measures precision: half the distance between the curve's 25% and 75% points. A steep curve means a small JND and a sharp sense of duration; a shallow curve means a large one. The two numbers vary independently. One observer can be biased but precise (shifted PSE, steep slope), another unbiased but noisy. Dividing the JND by the PSE gives the Weber fraction, which lets you compare precision across different stimulus ranges.
The Weber fraction is useful because of Weber's law. The lifted-weight experiments mentioned above found that the just-noticeable difference grows in proportion to the baseline, so ΔI / I stays roughly constant as I changes. If you can just feel one gram added to 100 g, at 400 g you need four grams. The law also explains the geometric-mean bisection point above. If equal ratios feel equally different, then 400 ms is halfway between 200 and 800 ms, because 400 is double 200 and 800 is double 400. Like most laws in perception it is an approximation. It fails near the faintest stimulus an observer can detect at all, where the Weber fraction climbs steeply, and the usual correction adds a small constant to the denominator: ΔI / (I + a).
Shift or slope: two different findings
Now manipulate something. Play the probe tones louder, make a visual stimulus brighter, add a distracting task. A condition can change the curve in two distinct ways. A PSE shift slides the whole curve sideways: more intense stimuli tend to be judged longer, so a brighter probe moves the curve left and the PSE drops while the slope stays the same. A slope change makes the curve flatter or steeper without moving its midpoint: the observer's timing has become noisier or more precise, and the JND grows or shrinks accordingly.
Which of the two happens is usually the finding. “The manipulation changed how long things seem” is a PSE result; “the manipulation changed how well people can time at all” is a JND result. The comparison condition in the interactive lets you set each one independently and see what the fits report.
Does the choice of function matter? Less than students fear, though more for some families than others. The two symmetric curves, logistic and cumulative Gaussian, both place their 50% point where the data are most informative, so their PSEs agree to within a millisecond or two. The asymmetric Weibull can be several milliseconds off, because it is a detection function anchored at zero stimulus, and its lopsided shape fits symmetric bisection data a little poorly. The spread parameter differs more: a logistic fit to Gaussian data reports a JND about 10% smaller. The families differ most in the tails, where lapses occur. Pick the family your subfield uses, report it, and keep it constant across conditions.
When the curve cannot start at zero
Every curve so far climbs from 0 to 1, which is right for bisection: an observer who never feels a probe as long says “short” every time. Plenty of psychophysics does not work that way. In a two-alternative forced-choice task (2AFC) the observer gets two intervals and reports which one held the stimulus, so pure guessing already gives 50% correct. The curve's floor is chance, not zero. By convention the threshold is read at 75%, halfway between chance and perfect, and not at the 50% point used for the PSE here.
That floor is the guess rate γ (gamma). The psychometric model used in practice includes it explicitly, along with a lapse rate λ (lambda) for the stray errors at the easy end. The curve itself is written ψ (psi):
ψ(x) = γ + (1 − γ − λ) · F(x)
with F one of the three families above. A plain two-parameter curve fitted to forced-choice data gives badly wrong answers. Simulate this lesson's own observer as a 2AFC task, with a true 75% threshold at 500 ms and 40 trials at each of the seven levels. The 50% point of an uncorrected fit comes out at 280 ms, 220 ms below the truth, and the JND is inflated from 54 ms to 160 ms. Fixing γ at .5 recovers the threshold to within a millisecond.
Lapses do less damage, and mostly to the JND. Give the same observer a 4% lapse rate, and a two-parameter fit still finds the PSE at 500 ms but reports a JND of 73 ms against a true 54 ms, a third too wide. A curve fixed at 0 and 1 can only reach those stray errors by being flatter. Estimate the lapse rate along with the curve, and the JND comes back to 53 ms. The 2AFC floor and the lapse rate are also a reason to keep the signal detection reading of a psychometric function in mind. A forced-choice proportion correct converts directly to d′ = √2 · z(Pc), so 75% correct is a sensitivity of 0.95. A yes/no proportion correct mixes sensitivity with wherever the observer put their criterion.
Fitting it in practice
For a simple two-parameter fit you can use logistic regression. Model each trial's response from the stimulus level, and both numbers follow from the coefficients: PSE = −b₀/b₁ and JND = ln 3 / b₁. To compare conditions, add the condition and its interaction with stimulus level: the condition term tests a PSE shift, and the interaction tests a slope change. Real laboratory data usually need one more parameter: lapse rates, the small probabilities of stray finger errors that keep the curve from quite reaching 0 and 1. Purpose-built tools such as psignifit and quickpsy estimate them along with the curve; the “Try it yourself” box below shows both routes.
In a full study you fit the function for each participant and condition. Each participant's PSE or JND then becomes a single data point in a familiar analysis: a paired t-test across two conditions, or a mixed model for larger designs.
How adaptive procedures find the threshold
The ladder of levels has one obvious weakness. You have to choose it before you know where the observer's threshold is, and levels far from the threshold tell you very little. Take this lesson's default observer, with PSE 500 ms and σ = 80 ms. At 200 ms and at 800 ms the answer is the same on roughly 9,999 trials in 10,000. Add up the Fisher information each level contributes. The 160 trials at 200, 300, 700 and 800 ms provide 7% of everything the experiment learns about the PSE, while the 40 trials at 500 ms alone provide 44%.
A staircase lets the observer's own responses decide where the stimulus goes. Start anywhere, and after each response move the stimulus one step toward the answer they did not give: if they said “long”, shorten the next probe; if they said “short”, lengthen it. The stimulus moves toward the level where the two answers are equally likely, then oscillates around it. Each change of direction is a reversal, and the step size usually starts large and halves at the first few reversals. This 1-up-1-down rule converges on the 50% point, which for a bisection task is the PSE itself. Detection experiments usually want a threshold higher up the curve. Levitt's transformed up-down rules reach it by requiring a run of correct answers before stepping down: k in a row to step down, one error to step up. That converges where pk = ½, so 2-down-1-up targets 70.7% and 3-down-1-up 79.4%. Simulating both on a forced-choice observer gives .705 and .790.
Does it save trials? The answer is different for placing trials and for estimating the PSE. With the same budget of 280 trials as the ladder above, a staircase on this lesson's observer puts 97% of its trials within 80 ms of the true PSE. The ladder puts only one level in seven there. But the textbook estimator loses most of that advantage. Averaging the last eight reversals recovers the PSE with a standard deviation of 24 ms, against the ladder's 11 ms. It also stops improving after about forty trials, because it always uses just eight numbers, however long you run. Fit the psychometric function to all 280 staircase trials and that standard deviation drops to 6 ms. The fixed ladder needs roughly 790 trials, 2.8 times as many, to match it.
So an adaptive procedure has two advantages. The one people usually quote, and the weaker of the two, is that it concentrates trials where the information is. The other is that you do not need to know the answer in advance. Run the same simulation on an observer whose true PSE is 525 ms with a sharp σ of 25 ms. Only one level of the fixed ladder is then anywhere near the rising part of the curve. Constant stimuli returns 505 ms, off by twenty, while the staircase gives 525.0. The advantage still depends on the analysis: average the reversals and you lose it, fit the curve and you keep it.
Why it matters: psychometric functions are the standard tool of psychophysics, and they appear wherever someone makes a binary judgment about a graded stimulus: perception labs, audiology clinics fitting hearing thresholds, vision screening, animal timing studies. If you can fit one S-curve and read off its PSE and JND, you can quantify what an observer perceives and separate bias from sensitivity with numbers.
Problem 77 of the practice problems reads a psychometric function off its logistic coefficients, the point of subjective equality and the just-noticeable difference, then compares how a quieter probe shifts both.
Common questions
Can I pool everyone's trials into one curve?
Fit each observer separately and average the parameters, not the other way round. Pooling trials fits one curve to a mixture of people, and a mixture of shifted curves is always shallower than the curves it is made of. For the cumulative Gaussian the group curve's spread comes out at exactly √(σ² + τ²), where σ is a typical observer's own spread and τ is the SD of their PSEs. Simulate 24 observers with a true JND of 54 ms whose PSEs scatter with an SD of 60 ms, and the pooled curve reports 67 ms, about a quarter too wide, while the PSE itself comes back correct. So pooling distorts precision, which is usually what the study was measuring. When observers contribute unequal numbers of trials, fit a mixed model to the raw trials to combine them correctly.
Which function should I fit: logistic, cumulative Gaussian, or Weibull?
The two symmetric families (logistic and cumulative Gaussian) put the PSE in essentially the same place, and the asymmetric Weibull is usually within a few milliseconds. The spread parameter, and with it the JND, varies more, and the tails differ most. The cumulative Gaussian has the cleanest signal-detection interpretation (internal noise is normal), the logistic is computationally convenient and matches logistic regression, and the Weibull suits detection tasks where performance is anchored at zero stimulus. If none is clearly better for your task, follow the convention in your subfield. Report the family and keep it constant across conditions.
Do I need a lapse rate in my psychometric model?
For real observers, usually yes. People blink, press the wrong key and lose attention, so a few errors appear even at the easiest stimulus levels. A two-parameter fit has to flatten the whole curve to accommodate those trials, and that biases the slope and so the JND. A small lapse parameter, either fixed at something like 0.02 or estimated with an upper bound, absorbs those errors. Tools built for psychophysics (psignifit, quickpsy) include lapse and guess rates by default, so move to them once the basic fit makes sense.