Mathematical Methods · Units 3 & 4

Sampling Distributions of Proportions

Understand the sampling distribution of the sample proportion the easy way, with plain English intuition, its mean and standard deviation, worked examples and an auto marked practice test. VCE Maths Methods Units 3 and 4.

Learn

Imagine the whole country has already decided, deep down, what fraction of people will vote yes. You will never know that exact fraction, so you ring a sample of people and work out the fraction in your sample instead. Ring a different sample and you get a slightly different fraction. The sampling distribution is the picture of all those possible fractions, and the surprising part is how predictable that picture is. It sits right on the truth and its spread shrinks in a way you can calculate exactly.

0.40.50.60.70.8 0.20.40.60.81
The sampling distribution of the sample proportion is approximately a bell curve, centred on the true population proportion p, with a spread that shrinks as the sample size grows.

The sample proportion and why it moves

You want to know the population proportion pp, the fraction of the whole population with some feature. You cannot measure everyone, so you take a random sample of size nn, count the XX people in the sample who have the feature, and report the fraction:

p^=Xn\hat{p} = \frac{X}{n}

This p^\hat{p} is the sample proportion, and it is your estimate of pp. The key idea is that p^\hat{p} is not a fixed number. Before you collect the sample it is random, because XX is random. A different random sample gives a different p^\hat{p}. So p^\hat{p} has its own distribution, called the sampling distribution, and our whole job is to describe its centre and its spread.

The centre: the mean of p hat

Here is the reassuring fact. Although each p^\hat{p} misses the truth a little, the sample proportion does not lean high or lean low. On average it lands exactly on the population proportion:

E(p^)=pE(\hat{p}) = p

That is what makes p^\hat{p} a fair estimate of pp. There is no built in bias pulling your samples above or below the truth, so the whole distribution is centred on the number you are trying to find. The catch is that any single sample still misses, which is why the spread matters just as much as the centre.

The spread: the standard deviation of p hat

The count XX in a sample of size nn is a binomial variable, so p^=X/n\hat{p} = X/n inherits its spread from the binomial. Dividing by nn shrinks that spread, and the standard deviation of the sample proportion works out to

sd(p^)=p(1−p)n\mathrm{sd}(\hat{p}) = \sqrt{\frac{p(1-p)}{n}}

Read the formula as a story. The p(1−p)p(1-p) on top is largest when p=0.5p = 0.5, so a population split fifty fifty is the hardest to pin down. The nn underneath is inside a square root, so a bigger sample shrinks the spread, just more slowly than you might expect. To halve the spread you do not double the sample, you have to multiply it by four.

The shape: an approximately normal curve

For a large enough sample the sampling distribution of p^\hat{p} is not just centred and spread in a known way, it is also a familiar shape. It is approximately a normal curve:

p^≈N ⁣(p, p(1−p)n)\hat{p} \approx N\!\left(p,\ \frac{p(1-p)}{n}\right)

That means once you have the mean pp and the standard deviation p(1−p)/n\sqrt{p(1-p)/n}, you can answer probability questions about p^\hat{p} exactly the way you handle any normal distribution. Questions like “what is the chance a sample gives a proportion below 0.150.15” become standard normal calculations. This bell shape is also the foundation for confidence intervals, which is the next topic.

How to actually do it

Every sampling distribution question runs on the same short recipe.

  1. Identify the population proportion pp and the sample size nn from the wording.
  2. The mean of p^\hat{p} is just pp.
  3. The standard deviation of p^\hat{p} is p(1−p)/n\sqrt{p(1-p)/n}.
  4. If the question asks for a probability, treat p^\hat{p} as normal with that mean and standard deviation.

Two traps catch students every year. The first is mixing up the count and the proportion. Always divide by nn, so the answer is a fraction like 0.320.32, never the raw count. The second is rounding too early. Carry full accuracy through the square root and only round the final answer to the number of decimal places asked for.

Lock it in with active recall

Cover the answer and say each one out loud before you flip. Rate yourself honestly — the cards you find hard come back sooner, the ones you know are spaced further out.

Active recall

Answer from memory first, then flip. Rate yourself and each card returns on a spaced schedule (1 → 3 → 7 → 16 days).

What is the sample proportion p^\hat{p}, and how does it differ from pp?
What is the mean of the sampling distribution of p^\hat{p}?
Write the standard deviation of p^\hat{p}.
If you multiply the sample size nn by 44 (with pp fixed), what happens to sd(p^)\mathrm{sd}(\hat{p})?
For a large sample, what is the approximate distribution of p^\hat{p}?
Recall · The binomial distribution
Why is the count XX in a sample binomial, and how does p^\hat{p} relate to it?
Recall · The normal distribution
How do you find a probability like Pr⁡(p^<0.15)\Pr(\hat{p} < 0.15)?

See this recipe in action in the Worked Examples tab, then test yourself in Try It.

Worked examples

Worked Example 1Luggage labelled heavy, from a real exam

On a particular day a random sample of 3535 pieces of luggage is selected. Let P^\hat{P} be the proportion of luggage labelled heavy in random samples of size 3535, where the population proportion is p=0.234p = 0.234. (i) Find Pr⁡(P^>0.2)\Pr(\hat{P} > 0.2), correct to three decimal places. (ii) Find the probability that P^\hat{P} lies within one standard deviation of its mean, correct to three decimal places. Do not use a normal approximation.

  1. 1

    (i) Since P^=X/35\hat{P} = X/35, the event P^>0.2\hat{P} > 0.2 means X>0.2×35=7X > 0.2 \times 35 = 7, that is X≥8X \ge 8, where X∼Bi(35,0.234)X \sim \mathrm{Bi}(35, 0.234).

    Pr⁡(X≥8)=1−Pr⁡(X≤7)≈0.595\Pr(X \ge 8) = 1 - \Pr(X \le 7) \approx 0.595
  2. 2

    (ii) The mean is p=0.234p = 0.234 and sd(P^)=p(1−p)35≈0.0716\mathrm{sd}(\hat{P}) = \sqrt{\dfrac{p(1-p)}{35}} \approx 0.0716. Within one standard deviation of the mean corresponds to XX from 66 to 1010.

    Pr⁡(6≤X≤10)≈0.684\Pr(6 \le X \le 10) \approx 0.684
  3. 3

    The examiner report notes the common errors were using the wrong inequality in (i) and using a normal approximation, or the wrong standard deviation, in (ii).

    (i) ≈0.595,(ii) ≈0.684\text{(i)}\ \approx 0.595, \qquad \text{(ii)}\ \approx 0.684
Answer
(i) Pr⁡(P^>0.2)≈0.595,(ii) ≈0.684\text{(i)}\ \Pr(\hat{P} > 0.2) \approx 0.595, \qquad \text{(ii)}\ \approx 0.684

VCAA 2024 Mathematical Methods Exam 2, Section B Q4d

Worked Example 2Smallest sample size for a target spread, from a real exam

A sample of nn people gives a sample proportion of p^=0.6\hat{p} = 0.6. Find the smallest value of nn for which the standard deviation of the sample proportion is 250\dfrac{\sqrt{2}}{50}.

  1. 1

    Set the standard deviation of the sample proportion equal to the target value, with p^=0.6\hat{p} = 0.6.

    p^(1−p^)n=0.6×0.4n=250\sqrt{\frac{\hat{p}(1-\hat{p})}{n}} = \sqrt{\frac{0.6 \times 0.4}{n}} = \frac{\sqrt{2}}{50}
  2. 2

    Square both sides to clear the square roots.

    0.24n=22500\frac{0.24}{n} = \frac{2}{2500}
  3. 3

    Solve for nn. The examiner report notes that students who set up the standard-deviation equation correctly generally reached the answer.

    n=0.24×25002=300n = \frac{0.24 \times 2500}{2} = 300
Answer
n=300n = 300

VCAA 2024 Mathematical Methods Exam 1, Q5c.ii

Worked Example 3Mean and standard deviation of p hat

In a large population, the true proportion who own a pet is p=0.3p = 0.3. Random samples of size n=50n = 50 are taken. Find the mean and the standard deviation of the sample proportion p^\hat{p}, with the standard deviation correct to four decimal places.

  1. 1

    The mean of p^\hat{p} is always the population proportion pp.

    E(p^)=p=0.3E(\hat{p}) = p = 0.3
  2. 2

    The standard deviation of p^\hat{p} is p(1−p)/n\sqrt{p(1-p)/n}. Substitute p=0.3p = 0.3 and n=50n = 50.

    sd(p^)=0.3×0.750=0.0042\mathrm{sd}(\hat{p}) = \sqrt{\frac{0.3 \times 0.7}{50}} = \sqrt{0.0042}
  3. 3

    Evaluate the square root and round at the very end.

    0.0042≈0.0648\sqrt{0.0042} \approx 0.0648
Answer
E(p^)=0.3,sd(p^)≈0.0648E(\hat{p}) = 0.3, \quad \mathrm{sd}(\hat{p}) \approx 0.0648
Worked Example 4A bigger sample sharpens the spread

For a population proportion p=0.6p = 0.6, random samples of size n=400n = 400 are taken. Find the standard deviation of the sample proportion p^\hat{p}, correct to four decimal places.

  1. 1

    Write the standard deviation formula with p=0.6p = 0.6 and n=400n = 400.

    sd(p^)=0.6×0.4400=0.0006\mathrm{sd}(\hat{p}) = \sqrt{\frac{0.6 \times 0.4}{400}} = \sqrt{0.0006}
  2. 2

    Evaluate.

    0.0006≈0.0245\sqrt{0.0006} \approx 0.0245
  3. 3

    Notice the spread is small because the sample is large. The same pp with n=50n = 50 would give a spread about three times wider.

    sd(p^)≈0.0245\mathrm{sd}(\hat{p}) \approx 0.0245
Answer
sd(p^)≈0.0245\mathrm{sd}(\hat{p}) \approx 0.0245
Worked Example 5Using the normal shape of p hat

A population has proportion p=0.2p = 0.2. A random sample of size n=100n = 100 is taken, so p^\hat{p} is approximately normal. Find Pr⁡(p^<0.15)\Pr(\hat{p} < 0.15), correct to four decimal places.

  1. 1

    The mean is p=0.2p = 0.2. Find the standard deviation with p=0.2p = 0.2 and n=100n = 100.

    sd(p^)=0.2×0.8100=0.0016=0.04\mathrm{sd}(\hat{p}) = \sqrt{\frac{0.2 \times 0.8}{100}} = \sqrt{0.0016} = 0.04
  2. 2

    So p^\hat{p} is approximately normal with mean 0.20.2 and standard deviation 0.040.04.

    p^≈N(0.2, 0.042)\hat{p} \approx N(0.2,\ 0.04^2)
  3. 3

    Find the probability below 0.150.15 using the normal distribution.

    Pr⁡(p^<0.15)≈0.1056\Pr(\hat{p} < 0.15) \approx 0.1056
Answer
Pr⁡(p^<0.15)≈0.1056\Pr(\hat{p} < 0.15) \approx 0.1056

Practice questions

Practice test

Try it yourself

7 questions, 10 marks

Choose your answers, then submit to see your score and the full worked solutions. Multiple choice is marked for you, just like Exam 2 Section A.

Q1.The probability a driver is late on any given working day is 0.087040.08704. Let P^\hat{P} be the proportion of days the driver is late in a five-day working week. Find Pr⁡(0.4≤P^≤0.6)\Pr(0.4 \le \hat{P} \le 0.6), correct to four decimal places.

2marks

Work this on paper. The worked solution appears once you submit.

VCAA 2025 Mathematical Methods Exam 2, Section B Q3biii

Q2.A random sample of n=250n = 250 people contains 8080 who use public transport. The value of the sample proportion p^\hat{p} is:

1mark
Need a hint?
The sample proportion is the count of successes divided by the sample size, so work out 80÷25080 \div 250.

Q3.Random samples of size nn are taken from a population with proportion pp. The mean of the sample proportion p^\hat{p} is:

1mark
Need a hint?
The sampling distribution is centred on the truth, so ask which option describes the centre rather than the spread.

Q4.For a population proportion p=0.25p = 0.25 and samples of size n=64n = 64, the standard deviation of p^\hat{p} is closest to:

1mark
Need a hint?
Use p(1−p)/n\sqrt{p(1-p)/n} with p=0.25p = 0.25 and n=64n = 64, and remember the final square root.

Q5.Keeping the population proportion pp fixed, the sample size nn is multiplied by 44. The standard deviation of p^\hat{p} is:

1mark
Need a hint?
In p(1−p)/n\sqrt{p(1-p)/n} the nn sits under a square root, so think about what 4\sqrt{4} does to the spread.

Q6.A population has proportion p=0.5p = 0.5 and samples of size n=100n = 100 are taken, so p^\hat{p} is approximately normal. Which statement is correct?

1mark
Need a hint?
The mean is just pp; work out the standard deviation as p(1−p)/n\sqrt{p(1-p)/n} with p=0.5p = 0.5 and n=100n = 100.

Q7.A population has proportion p=0.5p = 0.5. Random samples of size n=100n = 100 are taken. Find the mean and the standard deviation of the sample proportion p^\hat{p}. Show your working.

3marks

Work this on paper. The worked solution appears once you submit.

Frequently asked questions

What is the difference between p and p hat?
The plain p is the true proportion for the whole population, a fixed number you usually never know. The p hat is the proportion you actually measure in one sample, and it changes from sample to sample. You use p hat as your best estimate of p.
Why does a bigger sample give a smaller standard deviation?
Because the sample size sits underneath a square root in the spread formula. More people in the sample means the random ups and downs average out more, so your estimate jumps around less. To halve the spread you need four times the sample, not just double it.
When can I treat the sample proportion as normal?
When the sample is reasonably large, the sampling distribution of the sample proportion is close enough to a bell curve to use the normal distribution. That lets you answer probability questions about the sample proportion the same way you would for any normal variable.