MATH 1401 · Dr. Barry Monk

Measures of Spread (Continued) and Measures of Position

Sections 3.2 & 3.3

We pick up Section 3.2 where the last lecture left off.

Section 3.2 · Objective 3 of 4

Compute the standard deviation of a population and a sample

Definition

Standard Deviation

Since variance is calculated using squared deviations, its units are the squared units of the data. However, it’s often more practical to use a measure of spread with the same units as the data. This is achieved by taking the square root of the variance, resulting in the standard deviation.

Sample Standard Deviation s=s2
Population Standard Deviation σ=σ2
Example

Recall that the variance for the San Francisco monthly temperatures was σ2 = 14.083 squared degrees.

JanFebMarAprMayJunJulAugSepOctNovDec
San Francisco515455565860606163625852
The standard deviation is
σ=14.083=3.753
Calculator

The 1-Var Stats command in the TI-84 Plus calculator returns both the sample standard deviation and the population standard deviation.

Data:578101113
TI-84 Plus
Online calculator

The Standard Deviation is Not Resistant

Like the mean, it is significantly affected by extreme values.

Mean 42.9 Standard deviation 24.9 Median 40.0
Drag any value. Watch which measure stays put.
Section 3.2 · Objective 4 of 4

Use The Empirical Rule to summarize data that are unimodal and approximately symmetric

The Empirical Rule

When a data set has a bell-shaped histogram, it is often possible to use the standard deviation to provide an approximate description of the data using a rule known as The Empirical Rule.

Approximately 68% of the data will be within one standard deviation of the mean.

μσ to μ+σ40 to 60

Approximately 95% of the data will be within two standard deviations of the mean.

μ2σ to μ+2σ30 to 70

All, or almost all, of the data will be within three standard deviations of the mean.

μ3σ to μ+3σ20 to 80
Example

Example: The Empirical Rule

Projections for the population percentage aged 65 and over for each state is given. Describe the population using the Empirical Rule.

14.114.314.417.812.014.912.613.712.813.813.712.413.8
14.113.314.316.08.111.514.110.212.413.415.612.813.9
12.314.115.313.013.610.512.413.513.910.711.514.312.7
13.112.212.415.012.613.613.715.514.69.012.214.0
Solution:

Note that the histogram is bell-shaped, so the Empirical Rule applies. We find the mean and standard deviation.

Mean:μ=13.249 Standard Deviation:σ=1.683
Online calculator
Example

Example: The Empirical Rule

We compute six quantities:

μσ=13.251.683=11.57 μ+σ=13.25+1.683=14.93

Approximately 68% of the data values are between 11.57 and 14.93.

μ2σ=13.252(1.683)=9.88 μ+2σ=13.25+2(1.683)=16.62

Approximately 95% of the data values are between 9.88 and 16.62.

μ3σ=13.253(1.683)=8.20 μ+3σ=13.25+3(1.683)=18.30

Almost all of the data values are between 8.20 and 18.30.

Recap

Section 3.2 — Measures of Spread

Standard Deviation

The square root of the variance, so it carries the same units as the data.

s=s2 σ=σ2
The Empirical Rule

For a bell-shaped data set, approximately 68% of the data lies within one standard deviation of the mean, approximately 95% within two, and all, or almost all, within three.

Measures of Position

Section 3.3
Section 3.3 · Objective 1 of 4

Compute and interpret z-scores

Definition

z-Score

Let x be a value from a population with mean μ and standard deviation σ. The z-score of x is:

z=xμσ
Interpretation:

The z-score of an individual data value tells how many standard deviations that value is from its population mean.

For example, a value one standard deviation above the mean has a z-score of z=1 and a value two standard deviations below the mean has a z-score of z=2.

A man and a woman standing outdoors in tall grass under a pale overcast sky. The man is nearer the camera on the left and the woman a little behind him on the right; he is noticeably the taller of the two.

Example:

A study shows the average height for U.S. men is 69.4 inches (σ=3.1) and for women is 63.8 inches (σ=2.8). Who is taller relative to their gender: a 73-inch man or a 68-inch woman?

ZMan’s Height=xμσ=7369.43.1=1.16 ZWoman’s Height=xμσ=6863.82.8=1.50

The woman is taller, relative to the population of women’s heights.

Section 3.3 · Graphical interpretation

The Two Heights on One Scale

Adult men
μ=69.4
Adult women
μ=63.8
One shared scale
z

Finding a Value with a Given z-Score

To find a value from a given z-score, multiply the z-score by the standard deviation, then add the result to the mean.

x=μ+z·σ
Example:

Data from the National Health and Examination Survey shows that for U.S. adults, the mean cholesterol level is μ=200 mg/dL (σ=44). What is the cholesterol level for someone with a z-score of 1.5?

Solution:
x=μ+z·σ=200+1.5(44)=266

The Empirical Rule & z-Scores

A long airport concourse in low evening light, windows down one side and gate signs overhead. Two travellers with wheeled luggage walk away from the camera in the middle distance.
Check Your Understanding

Check Your Understanding

The Wall Street Journal reports that the average maximum distance from check-in to the gate at top U.S. airports is μ=0.79 miles (σ=0.42 miles).

1

The maximum distance at the Seattle-Tacoma International Airport is 0.57 miles. What is the z-score for this distance?

z=0.570.790.42=0.52 2

What distance would have a z-score of 1?

x=0.79+1(0.42)=1.21miles
Section 3.3 · Objective 2 of 4

Compute the quartiles of a data set

Quartiles

We’ve learned how to compute the mean and median of a data set as measures of the center. Sometimes, it is useful to compute measures of position other than the center. Quartiles divide a data set into four approximately equal pieces.

Q1

The first quartile, Q1, separates the lowest 25% of the data from the highest 75%.

Q2

The second quartile, Q2, divides the data so that 50% is below and 50% is above, essentially representing the median.

Q3

The third quartile, Q3, separates the lowest 75% of the data from the highest 25%.

An American goldfinch perched on a thin branch, in profile.

Suppose…

The third quartile of the lifespan of a certain type of Goldfinch is approximately 72 months.

This means…
75%

That 75% of this type of Goldfinch do not live for more than 72 months.

25%

Or 25% of them live for more than 72 months.

Example

Example: Computing Quartiles

The table below shows the February annual rainfall in Los Angeles over several years (in inches). Calculate the quartiles.

0.000.030.080.160.170.200.290.560.700.790.830.921.221.301.48
1.641.721.902.372.843.063.123.213.293.543.573.583.714.134.17
4.274.374.644.894.945.545.596.106.617.968.878.9111.0212.7513.68
First quartileThird quartile
Online calculator

Visualizing the Quartiles

Percentiles

Percentiles divide a data set into hundredths.

p and (100p)50 and 50

For any number p between 1 and 99, the pth percentile separates the lowest p50% of the data from the highest (100p)50%.

Visualizing Percentiles

A smiling student in a classroom holding up a completed exam paper towards the camera, with classmates at desks behind her.
Check Your Understanding

Check Your Understanding

After getting her recent statistics exam score, Aminah boasted to her friends, “I scored in the 10th percentile of the class!”

Does Aminah have a reason to brag?

No. Scoring in the 10th percentile means 90% of the class scored above her.

Section 3.3 · Objective 3 of 4

Compute the five-number summary for a data set

Definition

Five-Number Summary

The five-number summary of a data set consists of the median, the first quartile, the third quartile, the smallest value, and the largest value. These values are generally arranged in order.

Minimum First Quartile Median Third Quartile Maximum
Calculator

Five-Number Summary with Technology

The five-number summary is often part of the output when using technology.

TI-84 Plus
Online calculator
Section 3.3 · Objective 4 of 4

Understand the effects of outliers

Definition

Outliers

An outlier is a value significantly larger or smaller than most in a data set.

Some outliers arise from errors, like a misplaced decimal point, causing the value to stand out.

Other outliers are accurate and reflect the presence of extreme values in the population.

Definition

Interquartile Range

One method for detecting outliers involves a measure called the Interquartile Range (IQR). The interquartile range is found by subtracting the first quartile from the third quartile: IQR=Q3Q1.

Any data value beyond the outlier boundaries may be an outlier.

Lower Outlier Boundary Q11.5(IQR)
Upper Outlier Boundary Q3+1.5(IQR)
An empty classroom. Rows of black tablet-arm chairs stand unoccupied in front of tall daylit windows, with no students present.
Example

Example: Outliers

Online calculator

The table shows the number of students absent from school each day in January. Identify any outliers in the data.

6567715751494441594942
56457744424546100595351
Solution:
IQR=Q3Q1=5945=14
Lower Boundary:Q11.5(IQR)=451.5(14)=24
Upper Boundary:Q3+1.5(IQR)=59+1.5(14)=80

The value 100 is an outlier because it is greater than 80.

Recap

Section 3.3 — Measures of Position

z-Scores

The z-score of an individual data value tells how many standard deviations that value is from its population mean.

z=xμσ
Quartiles & Percentiles

Quartiles divide a data set into four approximately equal pieces; percentiles divide it into a hundred.

Q1Q2Q3
Five-Number Summary

The five-number summary of a data set consists of the median, the first quartile, the third quartile, the smallest value, and the largest value.

Outliers

One method for detecting outliers involves a measure called the Interquartile Range (IQR).

IQR=Q3Q1

You Are Ready For:

3.1 & 3.2 HW: Measures of Center and Spread

3.3 HW: Measures of Position

1 / 36 → ← advance · Z zoom · B blank · H hides this bar