MathsFurther Statistics 1 › The quality of tests

The quality of tests

Every test can fail two ways: raise a false alarm, or miss a real change. The size measures the first, the power measures the second, and shrinking either one inflates the other.

Year FMEDEXCEL 9FM0 FS1

Builds on Hypothesis tests for Poisson and geometric models and Hypothesis testing with the binomial.

IN THIS TOPIC

  • Distinguish Type I and Type II errors and compute their probabilities.
  • Find the size of a test from its critical region.
  • Evaluate and interpret the power function at particular alternatives.

WHAT YOU PROBABLY THINK

A test at the 5% level is wrong 5% of the time.

Two ways to be wrong

A Type I error rejects a true H₀: a false alarm. Its probability is the size of the test, which is the actual probability of landing in the critical region when H₀ holds, and for a discrete distribution it usually falls a little below the nominal level. A Type II error fails to reject H₀ when some alternative is true: a miss.

The 5% figure describes only the first kind of failure, and only when H₀ is true, so the opening claim is doubly wrong: it ignores Type II errors entirely and misreads a conditional probability as an overall error rate.

One critical region, two risks: Type I under H₀ and Type II under the alternative p = 0.15H₀: p = 0.3p = 0.15Type I: 0.0355Type II: 0.5951critical region X ≤ 2
FIG. 1The two error types on one picture: the critical region carries the Type I risk under H₀, and the same region leaves a Type II risk under the alternative.

WORKED EXAMPLE

Size and Type II probability

For X ~ B(20, p), H₀: p = 0.3 against H₁: p < 0.3, with critical region X ≤ 2. Find the size, and the probability of a Type II error when p = 0.15.

Size = P(X ≤ 2 | p = 0.3) = 0.0355: below the nominal 5%, because the discrete jump from X = 2 to X = 3 overshoots.

When p = 0.15: P(X ≤ 2) = 0.4049, so the test rejects with probability 0.4049.

The Type II probability is 1 − 0.4049 = 0.5951: this test misses a drop to 0.15 more often than it catches it.

The power function

The power at a particular alternative is the probability of correctly rejecting H₀ there, so power = 1 − P(Type II error). Plotted against the parameter, it gives the power function: low near H₀, where the truth is hard to distinguish from the null, and rising towards 1 as the alternative moves further away.

The power function for X ~ B(20, p) with critical region X ≤ 2: barely above the size near p = 0.3, near-certain by p = 0.050.050.150.310.920.400.0355: the sizepower rises as the truth moves away from H₀
FIG. 2The power function for the same test: barely above the size near p = 0.3, climbing past 0.9 by p = 0.05, since a large change is easy to detect.

WORKED EXAMPLE

Reading the power function

For the test above, the power is 0.0355 at p = 0.3, 0.2061 at p = 0.2, 0.4049 at p = 0.15 and 0.9245 at p = 0.05. Comment.

At p = 0.3 the power equals the size, as it must: H₀ is true there, so 'correct rejection' means a false alarm.

Detection improves steadily as p falls, but even at p = 0.15, half the true departures go unnoticed.

Only for large drops does the test become reliable. Widening the critical region would raise power everywhere, at the cost of a larger size: the two cannot both be reduced with the sample size fixed.

YOUR TURN

Comparing two tests

Test A has size 0.05 and power 0.62 at a given alternative; test B has size 0.01 and power 0.44 at the same alternative. Which is preferable, and on what grounds?

Show the working

Neither dominates: A detects the alternative more often, B raises fewer false alarms.

The choice depends on the costs. Where a missed change is expensive, prefer A; where a false alarm triggers something costly, prefer B.

Raising the sample size is the only way to improve both at once.

THE EXAM BIT

  • Compute the size from the critical region, not from the nominal level; for discrete tests they differ.
  • Type II probabilities need a specific alternative value; without one, the question is incomplete.
  • Power = 1 − P(Type II error), evaluated at the stated alternative.
  • When comparing tests, mention both size and power; a test cannot be judged on either alone.

CHECK YOURSELF

A test has size 0.043 and, at a particular alternative, a probability of 0.68 of failing to reject H₀. State the probability of a Type I error and the power at that alternative.

Show a hint

Size is the Type I probability; power is one minus the Type II probability.

Show the answer

P

(

T

y

p

e

I

e

r

r

o

r

)

=

0

.

0

4

3

.

P

o

w

e

r

=

1

0

.

6

8

=

0

.

3

2

:

t

h

i

s

t

e

s

t

m

i

s

s

e

s

t

h

e

a

l

t

e

r

n

a

t

i

v

e

a

b

o

u

t

t

w

o

t

h

i

r

d

s

o

f

t

h

e

t

i

m

e

.

Type I rejects a true H₀ with probability equal to the test's size; Type II misses a real change.

Power = 1 − P(Type II error), rising as the alternative moves away from H₀; size and power trade off.

WORKBOOK

Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.

CHECK YOUR PROGRESS

Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device only.

  • Distinguish Type I and Type II errors and compute their probabilities.
  • Find the size of a test from its critical region.
  • Evaluate and interpret the power function at particular alternatives.

Open the full revision checklist to see every objective in the course in one place.

No animated video for this topic yet; these notes stand alone.