Hypotheses, significance and exact tests
A sample can differ from a claim just by chance. A hypothesis test asks whether the difference is too extreme to explain comfortably under that claim.
Start with one question, not six definitions
We write for the claim being tested and for the possible change. A tail is an extreme end of a probability distribution.
One question, four moves
Assume H₀
Use its parameter as the model.
Choose the tail
H₁ gives the direction.
Find probability
Observed result + more extreme.
Make a decision
Compare with α; write in context.
Small p-value = evidence against H₀, not proof.
Imagine that a die is claimed to be fair, but the number 6 appears unusually often. First pretend the claim is true. Then calculate how likely it would be to obtain this many sixes, or even more, by chance.
- is the starting model. It contains the equality, such as .
- is the change being investigated. Its sign tells you which extreme results count as evidence.
- The test statistic is a quantity calculated from the sample and used in the decision, such as an observed count, sample mean or z-value.
- The tail probability is calculated as if H₀ were true. It is not the probability that H₀ itself is true.
Key idea
Write the model, direction and decision precisely
Definition
1 markWhat is meant by the significance level of a hypothesis test?
Model answer: The chosen upper bound for the probability of rejecting H₀ when H₀ is true.
Definition
1 markWhat is meant by the critical region?
Model answer: The set of values of the test statistic for which H₀ is rejected.
- The significance level is chosen before the data are inspected. Common values are 5%, 2.5% and 1%.
- If the relevant tail probability is at most , reject . If it is larger, fail to reject .
- Results inside the critical region lead to rejection of . Its boundary is a critical value.
- If the result is not significant, write “fail to reject H₀” or “there is insufficient evidence”. Do not claim that H₀ has been proved.
- Two decision routes are equivalent: compare the tail probability with , or check whether the observed test statistic lies in the critical region.
Examiner note
For an exact test, add the whole relevant tail
Lower tail means 0, 1 and 2
P(X ≤ 2) = P(0) + P(1) + P(2) = 0.1028
Compare the whole tail with α = 0.05.
Worked example
Current-paper exact binomial test
A claim says that the probability of a success is . In 30 independent trials there are 2 successes. Test at the 5% level whether the probability is lower.
“Lower” means the observed value 2 and every smaller value:
Since , fail to reject .
Conclusion: There is insufficient evidence that the success probability is lower than .
(9709/62/M/J/25 Q8(a))
- Calculator lower tail: use cumulative probability .
- Calculator upper tail: use , not .
- For a Poisson rate, first scale the null mean to the full observation interval. The same tail rule then applies.
Examiner note
Common mistake
Poisson bridge: if the claimed rate is 4 per hour and 20 events are observed in 3 hours, then under H₀ and . The upper-tail result is significant at 5%.
A fair coin is tossed 16 times and gives 12 heads. Test at the 5% level whether it is biased towards heads. You may use for .
Show worked answer
Conclusion: There is sufficient evidence that the coin is biased towards heads.
