MATH 1401 · Dr. Barry Monk

Measures of Center

Section 3.1
Section 3.1 · Objective 1 of 3

Compute the mean and median of a data set

Mean of a Data Set

The mean of a data set represents its center. It’s like the balance point where all the data values, viewed as weights, would even out.

Balance point 81.2
Balanced. The beam is level — this is the mean.

The beam levels at 81.2, and that is the mean of the five scores.

Definition

Notation - Population Versus Sample

A population includes all individuals of interest, while a sample is a smaller group taken from the population. The mean is calculated the same way for both, but the notation is different.

Notation:
Population Mean:μ Sample Mean:x¯ Sum of Values:x
Definition

Computing the Mean

Given a list of numbers x1,x2,x3,, the sample mean is computed as:

x¯=xn

… and the population mean is computed as:

μ=xN

Median of a Data Set

The median is a measure of center that divides an ordered data set in half, with half the values below it and half above it. The method to calculate the median varies depending on whether the number of observations is even or odd.

If n is odd, the median is the middle number.

If n is even, the median is the average of the two middle numbers.

A student alone at a classroom desk, writing on a paper with a pen, in an otherwise empty room.
Example

Example: Computing the Mean and Median

During a semester, a student took five exams. The population of exam scores is 78, 83, 92, 68, and 85. Find the mean and median of the exam scores.

Mean
x=78+83+92+68+85=406
N=5
μ=xN=4065
μ=81.2
Median

Arrange the data values in increasing order.

68 78 83 85 92

The median is the middle number, 83.

A quiet hospital room with made-up beds, a drip stand and a chair by the window — a recovery ward, not an operating theatre.
Example

Example: Median with Even Number of Values

Eight patients undergo a new surgical procedure, and the number of days spent in recovery for each is as follows. Find the median number of days in recovery.

20 15 12 27 13 19 13 21
Solution:

Arrange the data values in increasing order.

12 13 13 15 19 20 21 27

The median is the average of the two middle numbers.

15+192=17
Calculator

1-Var Stats Command

The 1-Var Stats command in the TI-84 Plus calculator displays a list of the most common parameters and statistics for a given data set. This command is accessed by pressing STAT and then highlighting the CALC menu.

Calculator

Example: Mean and Median with 1-Var Stats

The same five exam scores, 78, 83, 92, 68, and 85, are entered in list L1. Then 1-Var Stats is selected from the STAT ▸ CALC menu.

The data in L1

1-Var Stats output

Scrolled down for the median

The calculator gives the mean as 81.2 and the median as 83, the same values found by hand. These five scores are the entire population, so the mean it reports is the population mean.

Interactive

The Complete Statistics Calculator

barrymonk.com/stats-calculator
1Paste the data
A panel headed Paste data with the five exam scores typed into a text box, above buttons reading Add to lists and Cancel.
2Data in L1
The DATA rail. Tabs for L1, L2 and L3; L1 is selected and holds 78, 83, 92, 68 and 85 in rows 1 to 5, with rows 6 to 10 empty.
Interactive

Running 1-Var Stats in the Calculator

3Choose Summarize
The Summarize form. A single menu labeled Data list is set to L1, above a note reading One list in, the full 1-Var summary out.
4The output
The 1-Var Stats result for L1. Center: mean x-bar 81.2, median 83, no mode. Spread: sample standard deviation 8.927, population standard deviation 7.985, sample variance 79.700, population variance 63.760, range 24, IQR 15.5. Position: minimum 68, Q1 73, median 83, Q3 88.5, maximum 92. Beside them a dot plot, a histogram and a boxplot of the same five scores.
Section 3.1 · Objective 2 of 3

Compare properties of the mean and median

The Mean is More Influenced by Extreme Values

Both the mean and median are measures of center. However, the mean is more influenced by extreme values than the median.

A statistic is considered resistant if it is not significantly affected by extreme values. The median is resistant, while the mean is not.

A large watermelon at one end of a wooden table beside six similar-sized apples in a row. One value is far larger than the rest.
Several scratch lottery tickets fanned out on a windowsill.
Example

Example: The Median is Resistant

Five families had annual incomes of $25,000, $31,000, $34,000, $44,000, and $56,000. The family earning $25,000 won the lottery, and their income jumped to $1,025,000.

Before the lottery win, the mean and median are as follows.

Mean = $38,000Median = $34,000

After the lottery win, the mean and median are as follows.

Mean = $238,000Median = $44,000

The extreme income of $1,025,000 significantly raised the mean from $38,000 to $238,000, while the median only increased from $34,000 to $44,000, showing the median’s resistance to extreme values.

Mean, Median, and the Shape of a Data Set

The mean and median measure the center of a data set in different ways.

Mean = Median · approximately symmetric
Symmetric

The mean and median are equal.

Skewed to the right

There are large values in the right tail. The mean is often greater than the median.

Skewed to the left

The mean is often less than the median.

Check Your Understanding

Skewed to the right, skewed to the left, or approximately symmetric?

1

A café tracks customers served in its first hour each day. The mean is 11 and the median is 4.

Skewed to the right
2

A hobby club records weekly hours members spend on projects. The mean is 4 hours and the median is 12 hours.

Skewed to the left
3

A study finds car owners drive a mean of 801 miles per year with a median of 798 miles.

Approximately symmetric
Section 3.1 · Objective 3 of 3

Find the mode of a data set

A glass bowl of brightly colored gumballs with more scattered on the surface beside it. Red is the most common color.
Definition

Mode of a Data Set

A value that is sometimes classified as a measure of center is the mode.

The mode of a data set is the value that appears most frequently.

If two or more values are tied for the most frequent, they are all considered to be modes.

If the values all have the same frequency, we say that the data set has no mode.

Four children of different ages crowded together, laughing.
Example

Example: Mode of a Data Set

Ten students were asked how many siblings they had. The results, arranged in order, were 0, 1, 1, 1, 1, 2, 2, 3, 3, 6. Find the mode of this data set.

0 1 1 1 1 2 2 3 3 6
Solution:

The value that appears most frequently is 1. Therefore, the mode is 1.

Mode – Measure of Center?

The mode is sometimes classified as a measure of center. However, this isn’t really accurate. The mode can be the largest value in a data set, or the smallest, or anywhere in between.

Mode for Qualitative Data

Means and medians apply only to quantitative data, while the mode can be computed for both qualitative and quantitative data.

Example:

Following is a list of the makes of all the cars rented by an automobile rental company on a particular day. Which make of car is the mode?

The most frequent category is “Toyota,” which appears six times.

HondaToyotaToyotaHondaFord
ChevroletNissanFordChevroletChevrolet
HondaDodgeFordFordToyota
ChevroletToyotaToyotaToyotaNissan
Recap · Section 3.1

Three measures of center

Mean

The mean of a data set represents its center. It’s like the balance point where all the data values, viewed as weights, would even out.

Median

The median is a measure of center that divides an ordered data set in half, with half the values below it and half above it.

The median is resistant, while the mean is not.

Mode

The mode of a data set is the value that appears most frequently.

MATH 1401 · Dr. Barry Monk

Measures of Spread

Section 3.2

San Francisco vs. St. Louis

Suppose you were deciding between living in San Francisco and St. Louis. One thing you might consider is weather. The table presents the average monthly temperatures, in degrees Fahrenheit, for both cities.

A row of San Francisco’s painted Victorian houses with the downtown skyline behind them, in clear daylight. The Gateway Arch above the St. Louis skyline at dusk, reflected in the Mississippi.
JanFebMarAprMayJunJulAugSepOctNovDec
San Francisco515455565860606163625852
St. Louis303544576675797870594535

The mean temperatures are similar: 57.5° for San Francisco and 56.1° for St. Louis. However, St. Louis has more temperature variation.

The mean shows the center, but we also need to measure the spread to fully describe the data set.

A halved watermelon beside a scatter of apples on a wooden surface, the watermelon many times the size of any apple.
Section 3.2 · Objective 1 of 2

Compute the range of a data set

Range of a Data Set

The range is a simple way to measure spread. The range of a data set is the difference between the largest value and the smallest value.

JanFebMarAprMayJunJulAugSepOctNovDec
San Francisco515455565860606163625852
St. Louis303544576675797870594535

The range of the San Francisco temperatures is: 63 − 51 = 12.

The range of the St. Louis temperatures is: 79 − 30 = 49.

Although the range is easy to compute, it is not often used in practice because it only involves two values from the data set.

Section 3.2 · Objective 2 of 2

Compute the variance of a population and a sample

Five darts clustered in and around the bullseye of a dartboard.

The Variance

When a data set has little spread, most values are close to the mean. With more spread, values are farther from the mean. Variance measures how far, on average, data values are from the mean.

Variance is calculated differently for populations and samples.

Definition

Population Variance

Population Variance σ2=(xμ)2N
Example

Example: Population Variance

Compute the population variance for the San Francisco temperatures.

Step 1: Compute the population mean μ. μ=xN=69012=57.5
Step 2: For each population value compute the difference between the value and the mean:
JanFebMarAprMayJunJulAugSepOctNovDec
x515455565860606163625852
(xμ)−6.5−3.5−2.5−1.50.52.52.53.55.54.50.5−5.5
Example

Example: Population Variance

JanFebMarAprMayJunJulAugSepOctNovDec
x515455565860606163625852
(xμ)−6.5−3.5−2.5−1.50.52.52.53.55.54.50.5−5.5
(xμ)242.2512.256.252.250.256.256.2512.2530.2520.250.2530.25
Step 3: Square the differences.
Step 4: Add the squared differences: (xμ)2=42.25+12.25+6.25+2.25+0.25+6.25+6.25+12.25+30.25+20.25+0.25+30.25=169
Step 5: Divide the sum by the population size: σ2=(xμ)2N=16912=14.083.
Definition

Sample Variance

When computing sample variance, the sample mean is used for deviations. Deviations using the sample mean are typically smaller than those using the population mean.

If we divided by n when calculating sample variance, it would generally be smaller than the population variance. To correct this, we divide the sum of squared deviations by n1 instead.

Sample Variance s2=(xx¯)2n1
Example

Example: Sample Variance

A new type of battery is being tested for laptop computers. The lifetimes, in hours, of six batteries, are 3, 4, 6, 5, 4, 2. Find the sample variance of the lifetimes.

Solution:

We find the sample mean to be 4. The sample variance is:

s2=(xx¯)2n1=[(34)2+(44)2+(64)2+(54)2+(44)2+(24)2]/(61)=10/5=2.
Recap · Section 3.2

Two measures of spread

Range

The range of a data set is the difference between the largest value and the smallest value.

Population variance

Variance measures how far, on average, data values are from the mean.

σ2=(xμ)2N
Sample variance

To correct this, we divide the sum of squared deviations by n1 instead.

s2=(xx¯)2n1

You Are Ready For:

3.1 HW: Measures of Center Some of 3.2 HW: Measures of Spread
1 / 9 → ← advance · Z zoom · B blank · H hides this bar