MATH 1401 · Sections 7.4 & 7.6 · The Central Limit Theorem for Proportions & Assessing Normality

The Central Limit Theorem for Proportions & Assessing Normality

Sections 7.4 & 7.6

Today's Plan

You can already:

Find probabilities with a normal distribution

Find a value from a normal distribution with a given percentile

Find probabilities involving a sample mean using the Central Limit Theorem

Today:

Use the Central Limit Theorem for a sample proportion, p^

Find probabilities for a sample proportion

Use a dotplot to decide whether a sample came from an approximately normal population

Example

Review Problem: Hours of Sleep

Adults sleep an average of μ=6.8 hours per night with standard deviation σ=1.4 hours. Assume sleep times are normally distributed.

Question A: What is the probability that a randomly selected adult sleeps more than 7.5 hours?

Question B: A researcher samples 35 adults. What is the probability that the sample mean sleep time is more than 7.5 hours?

Think first

Both questions use the same cutoff, 7.5 hours. One is about one adult, the other is about the mean of 35 adults. Which standard deviation does each one use? And which probability will be smaller, by a little or by a lot?

A young man asleep on a sofa under a knitted blanket.
Example

Review Problem: Question A

What is the probability that a randomly selected adult sleeps more than 7.5 hours?

Recall

μ=6.8 hours

σ=1.4 hours

Solution:

One adult's sleep time is normally distributed with μ=6.8 and σ=1.4.

Find the area under the normal curve:

Statistics Calculator
The Statistics Calculator's Normal tool set to Area to the right of x, with mean 6.8, standard deviation 1.4 and x 7.5. The result is 0.30854.
TI-84 PlusA TI-84 Plus home screen showing normalcdf(7.5,99999999,6.8,1.4) and its result, 0.3085375322.

The probability that a randomly selected adult sleeps more than 7.5 hours is approximately 0.3085.

Example

Review Problem: Question B

A researcher samples 35 adults. What is the probability that the sample mean sleep time is more than 7.5 hours?

Recall

μ=6.8 hours

σ=1.4 hours

n=35

Solution:

Since n=35>30, we can apply the Central Limit Theorem.

Compute μx¯ and the standard error σx¯:

μx¯=μ=6.8andσx¯=σn=1.435=0.237

Find the area under the normal curve:

Statistics Calculator
The Statistics Calculator's Normal tool set to Area to the right of x, with mean 6.8, standard deviation 0.237 and x 7.5. The result is 0.0015706.
TI-84 PlusA TI-84 Plus home screen showing normalcdf(7.5,99999999,6.8,0.237) and its result, 0.0015705914.

The probability that the sample mean sleep time is more than 7.5 hours is approximately 0.0016.

The Key Difference

Individual Adult

High variability: σ=1.4 hours

Probability one adult sleeps more than 7.5 hours: P(X>7.5)=0.3085

Individual values are naturally spread out across the population

Sample of 35 Adults

Low variability: σx¯=1.435=0.237 hours

Probability the sample mean of 35 adults is more than 7.5 hours: P(x¯>7.5)=0.0016

Sample means cluster tightly around the population mean

Think about it

Has Support Grown?

A 2019 city study found that 49% of residents supported adding bike lanes.

This year you survey 200 residents. 110 of them support it. That's 55%.

Has support grown? Or could a sample of 200 come out at 55% even if nothing has changed?

Everything we've done with the Central Limit Theorem so far has been about sample means. 55% isn't a sample mean. It's a sample proportion.

Two cyclists in helmets riding in a painted bike lane on a city street.

Proportions: A Different Kind of Question

MEANS (x¯)
A woman asleep in bed.

Average hours of sleep

A teenager looking at a phone at a table.

Average time on TikTok

A white delivery van parked at the curb, with a hand truck and a cardboard box beside it.

Average delivery time

PROPORTIONS (p^)

Proportion of people who sleep less than 6 hours

Proportion of teenagers who use TikTok daily

Proportion of on-time deliveries

In a population, the proportion who have a certain characteristic is called the population proportion, denoted p. In a simple random sample of n individuals, if x of them have the characteristic, the sample proportion is p^=xn.

Sampling Distribution of p^

Just as different samples produce different values of x¯ . . .

. . . different samples will also produce different values of the sample proportion p^.

Each sample: 1,000 people. 672 have the characteristic, so p^=6721,000=0.672

Because p^ varies from sample to sample, it has a probability distribution of its own, called the sampling distribution of p^.

What Does the Sampling Distribution of p^ Look Like?

Population: 57% of people can correctly identify a deepfake video, so p=0.57; 43% can't

Sample size n=80

Think first

Before you draw, predict where the sample proportions will center and what shape they will make.

Samples0
Mean of the sample proportions
Standard deviation of the sample proportions

Draw more samples

The sample proportions pile up in the shape of a normal curve.

The Central Limit Theorem for MeansProportions

RECALL

x¯p^ is approximately normal with mean μx¯=μμp^=p
and standard error σx¯=σnσp^=p(1−p)n as long as
n>30 or the population is approximately normaln·p and n·(1−p) are both at least 10

Checking the Theorem Against the Machine

The machine drew samples of n=80 from a population with p=0.57.

Theorem μp^0.57
Mean of the sample proportions
Theorem σp^0.0554
Standard deviation of the sample proportions

Center. The theorem says μp^=p=0.57.

Spread. The theorem says the standard error is σp^=p(1−p)n=0.57(0.43)80=0.0554.

Shape. The theorem says approximately normal, since n·p=80(0.57)=45.6 and n·(1−p)=80(0.43)=34.4 are both at least 10. The machine's sample proportions piled up in the shape of a normal curve.

The theorem and the machine agree on all three.

Why Both Conditions Matter

1. Drag n to the right until both conditions are met.

2. Leave n there, and move p close to 0 or close to 1.

3. Bring p back to the middle, then lower n.

When p is close to 0 or close to 1, n has to be larger. When n is small, p has to be near the middle.

Check your understanding

Does the Central Limit Theorem Apply?

Think first

For each one, compute n·p and n·(1−p).

(a)A simple random sample of size 60 will be drawn from a population with proportion p=0.12.

(a)NO. n·p=60(0.12)=7.2, which is less than 10.

(b)A simple random sample of size 24 will be drawn from a population with proportion p=0.55.

(b)YES. n·p=24(0.55)=13.2 and n·(1−p)=24(0.45)=10.8 are both at least 10.

(c)A simple random sample of size 10,000 will be drawn from a population with proportion p=0.002.

(c)YES. n·p=10,000(0.002)=20 and n·(1−p)=10,000(0.998)=9,980 are both at least 10.

Computing Probabilities for p^

Check that n·p≥10 and n·(1−p)≥10

Find μp^=p and the standard error σp^=p(1−p)n

Same normal curve work as for sample means: find the area under the normal curve

Example

Example 1: Deepfake Detection

A 2024 meta-analysis of 56 studies found that only 57% of people can correctly identify a deepfake video. A media literacy organization tests 80 people. What is the probability that fewer than half correctly identify the deepfake?

Solution:

We have n=80 and p=0.57

n·p=80(0.57)=45.6 ✓

n·(1−p)=80(0.43)=34.4 ✓

A man in glasses studies a woman's face on a computer monitor.
Example

Example 1: Deepfake Detection

What is the probability that fewer than half correctly identify the deepfake?

Recall

n=80

p=0.57

Compute:

μp^=p=0.57andσp^=p(1−p)n=0.57(0.43)80=0.0554

Find the area under the normal curve:

Statistics Calculator
The Statistics Calculator's Normal tool set to Area to the left of x, with mean 0.57, standard deviation 0.0554 and x 0.5. The result is 0.1032.
TI-84 PlusA TI-84 Plus home screen showing normalcdf(-99999,0.5,0.57,0.0554) and its result, 0.103198032.

The probability that fewer than half correctly identify the deepfake is approximately 0.1032.

Example

Example 2: Bike Lane Support

A 2019 city study found that 49% of residents supported adding bike lanes. You believe support has grown since then. You survey 200 residents and find 55% now support it. If support hasn't changed, what is the probability of getting a sample proportion at least this high?

Solution:

We have n=200 and p=0.49

n·p=200(0.49)=98 ✓

n·(1−p)=200(0.51)=102 ✓

Two cyclists in helmets riding in a painted bike lane on a city street.
Example

Example 2: Bike Lane Support

If support hasn't changed, what is the probability of getting a sample proportion at least this high?

Recall

n=200

p=0.49

Compute:

μp^=p=0.49andσp^=p(1−p)n=0.49(0.51)200=0.0353

The standard error uses p=0.49, the value we're assuming, not the sample's p^=0.55.

Find the area under the normal curve:

Statistics Calculator
The Statistics Calculator's Normal tool set to Area to the right of x, with mean 0.49, standard deviation 0.0353 and x 0.55. The result is 0.044592.
TI-84 PlusA TI-84 Plus home screen showing normalcdf(0.55,9999,0.49,0.0353) and its result, 0.044592081.

If support hasn't changed, the probability of a sample proportion of 0.55 or more is approximately 0.0446.

Unusual? 0.0446 is less than 0.05, so if support hasn't changed, a sample proportion this high would be unusual.

Example

Example 3: Who Wants Ice Cream!

A Harris poll found that 27% of Americans prefer chocolate ice cream. A food blogger thinks Gen Z prefers chocolate more than the general population. She surveys 100 college students and finds 30% prefer chocolate. What is the probability of getting a sample proportion at least this high if Gen Z is the same as everyone else?

Solution:

We have n=100 and p=0.27

n·p=100(0.27)=27 ✓

n·(1−p)=100(0.73)=73 ✓

Three ice cream cones in a row, a chocolate scoop with fudge sauce in front.
Example

Example 3: Who Wants Ice Cream!

What is the probability of getting a sample proportion at least this high if Gen Z is the same as everyone else?

Recall

n=100

p=0.27

Compute:

μp^=p=0.27andσp^=p(1−p)n=0.27(0.73)100=0.0444

Find the area under the normal curve:

Statistics Calculator
The Statistics Calculator's Normal tool set to Area to the right of x, with mean 0.27, standard deviation 0.0444 and x 0.30. The result is 0.24962.
TI-84 PlusA TI-84 Plus home screen showing normalcdf(0.30,9999,0.27,0.0444) and its result, 0.2496232196.

If Gen Z is the same as everyone else, the probability of a sample proportion of 0.30 or more is approximately 0.2496.

Unusual? 0.2496 is not less than 0.05, so if Gen Z is the same as everyone else, a sample proportion this high would not be unusual.

Assessing Normality

Many statistical procedures require sampling from an approximately normal population. Often we don't know whether this condition is met, so the only way to assess normality is to examine the sample.

Not Seeking Perfection

We're not trying to determine if the population is exactly normal

Sample Size Matters

Assessing normality matters more for small samples than for large ones

No Hard Rules

Context and judgment matter. Hard and fast rules don't work well in practice

When We Reject Normality

We reject the assumption of normality if the sample has any of these features:

1Outliers Present

Sample contains extreme values far from the rest

2Severe Skewness

Distribution shows large degree of asymmetry

3Multiple Modes

More than one distinct peak in the data

Example

Example 4: Oven Temperatures

An oven is set to 360°F, and the temperature when the thermostat turns off is recorded. A sample of 7 readings:

358 363 361 355 367 352 368

Is it reasonable to treat this as a sample from an approximately normal population? Explain.

Outliers? No.

Severe skewness? No.

Multiple modes? No.

Solution:

The dotplot shows no outliers, no strong skewness, and no evidence of multiple modes. We can treat this as a sample from an approximately normal population.

A white oven door with a steel handle and a glass window.
Example

Example 5: Heart Rate Monitoring

A fitness company tests a new smartwatch's heart rate monitor. Six readings from a simple random sample:

68 71 79 98 67 75

Is it reasonable to treat this as a sample from an approximately normal population? Explain.

Outliers? Yes: 98 is far from the rest.

Solution:

The dotplot shows that 98 is an outlier. We should not treat this as a sample from an approximately normal population.

A runner's wrist wearing a slim fitness tracker.
Check your understanding

Approximately Normal or Not?

For each sample, is it reasonable to treat it as a sample from an approximately normal population?

Think first

Check each dotplot for outliers, severe skewness and multiple modes.

(a)Minutes on hold for 10 callers

(a)NO. The dotplot is strongly skewed to the right.

(b)Commute times, in minutes, for 12 employees

(b)NO. The dotplot shows two separate clusters, so there is more than one mode.

(c)Weights, in grams, of 8 chocolate bars

(c)YES. The dotplot shows no outliers, no strong skewness, and no evidence of multiple modes.

You Are Ready For:

7.4 & 7.6: Central Limit Theorem for Proportions & Assessing Normality

Check that n·p and n·(1−p) are both at least 10

Find the mean of p^ and its standard error: p and p(1−p)n

Find probabilities for a sample proportion

Use a dotplot to check a sample for outliers, severe skewness and multiple modes

1 / 28 → ← advance · Z zoom · B blank · H hides this bar