Type I and Type II errors
A test can be wrong in two opposite ways. Type I = false alarm: reject a true claim (convict the innocent). Type II = missed it: accept a false claim (free the guilty). The fact that ties it together: the significance level itself.
False alarm vs missed it
| Accept H₀ | Reject H₀ | |
|---|---|---|
| H₀ true | correct | Type I (false alarm) |
| H₀ false | Type II (missed it) | correct |
= the significance level. — computable only with the true alternative value, working in that model. The asymmetry matters: rejecting rules out a Type II; accepting it rules out a Type I.
P(Type I) lives in the H₀ model; P(Type II) lives in the TRUE H₁ model. You can never make a Type II if you rejected, nor a Type I if you accepted.
Worked example
Type I from a binomial test
16 coin tosses, suspect bias towards tails. , at 5%.
but , so the critical region is . The Type I probability is the achieved size of that region.
Not exactly 5%: a discrete binomial only reaches the nearest achievable tail ≤ 5%. In context, a Type I error here means “deciding the coin is biased when it is actually fair.”
Worked example
Type I = the critical region's probability (real paper)
Back to the cube test, , at 5%. We found , so 8 is not in the region; , so the region is .
.
(9709/62/O/N/23 Q3) — once the rejection region is fixed, P(Type I) is just its total probability under H₀, never the nominal 5%.
Type II needs the true model
A Type II probability only exists once you are told what is actually true. Swap to that distribution and find the chance the test still lands in the acceptance region.
Worked example
Type I AND Type II from two normal models
A protein level screens for a condition. Healthy people have ; those with it have . The doctor flags anyone with .
- Type I (healthy, flagged): use . .
- Type II (has it, missed): use . .
A continuous variable (protein level) needs NO continuity correction. The two probabilities use different curves — switching models is the whole skill.
Worked example
Reverse Type II: how good must the test be?
Same condition model . The doctor wants a cut-off (flag when ) so the test misses a true case with probability under 0.03. Find the range of .
- A miss is under the true model, so require .
- Standardise: , and .
- .
A lower cut-off catches more true cases (smaller Type II) — but that same move lets more healthy people past the line, so Type I climbs. You are choosing where on the trade-off to sit.
Worked example
The balanced cut-off: equal Type I and Type II
Both models again, healthy and with the condition. Find the cut-off where .
- Type I (healthy past ): . Type II (condition below ): .
- Set equal: . By symmetry of , the arguments are negatives: .
- Cross-multiply: .
This single value is the balance point the slider below has to land on — push the line either side and one error grows while the other shrinks.
Drag the cut-off below: push it right and false alarms (Type I) shrink while misses (Type II) balloon. You trade them — you cannot kill both at once.
P(Type I) = 0.010·P(Type II) = 0.238·raising the bar misses more
Why cutting one error grows the other
Both errors are areas cut by the same critical line: slide it to shrink the false-alarm tail (Type I) and the missed-detection area (Type II) grows — you trade them, not remove both. The honest way to cut both is a larger , which narrows both curves so they overlap less.
Examiner note
Common mistake
Your turn— tap to reveal the worked answer (Type II under the true alternative)
Sunflower heights are . A friend claims singing raises the mean; the gardener will believe it if one plant exceeds m. If the true mean of sung-to plants is m (same SD), find .
Use the true model : .