Part 3: Centre and spread
mean, median, mode, standard deviation, variance, range, quartiles, interquartile range, outliers, skewness
Part 2 let you decide which summaries a column can support. This part builds those summaries.
A manager asking “how long do customers wait?” wants one number. This part produces that number. It also produces the measure of spread that qualifies the centre, and it explains how the shape of the distribution determines which pair to report. The standard deviation is calculated by hand once, using five values, so that you can interpret and defend the result later.
Measures of centre
A measure of centre is a single value representing a whole column. Three measures are common, and the choice among them depends on the shape of the data.
Mean
The arithmetic mean adds all values and divides by the number of values. For a sample of \(n\) values \(x_1, x_2, \dots, x_n\):
\[\bar{x} = \frac{\sum x}{n}\]
The symbol \(\sum\) means “add up all of these”. For Waiting_Time_Mins in the teaching dataset the 200 values add to 9,091 minutes, so
\[\bar{x} = \frac{9{,}091}{200} = 45.455 \text{ minutes}\]
which you would report as 45.5 minutes. Reporting 45.455 would claim a precision that the underlying measurements do not support, because each value was rounded to the nearest minute.
The mean uses every value. That is useful, and it also makes the mean sensitive to extreme values: a single unusual case can move it a long way.
Worked example 1.9
Mean in a staffing decision
Five shifts record 28, 39, 22, 51 and 34 calls.
Calculate the mean, then decide whether it is the right basis for a staffing rule.
\[\bar{x} = \frac{28 + 39 + 22 + 51 + 34}{5} = \frac{174}{5} = 34.8 \text{ calls}\]
The mean is 34.8 calls. As a summary of the month, that figure is correct.
It is the wrong basis for the staffing rule. Staffing for 34.8 calls leaves the 51-call shift understaffed. The rule should use the upper end of the observed range.
This illustrates a common reporting failure. The statistic is correct, yet it answers an unasked question.
Median
The median is the middle value when the data are ordered. It is the value with half the data at or below it and half at or above.
To find it, sort the values and take the one in position \((n+1)/2\). With an even count, that position falls between two values. Take the mean of the two values.
For the 200 waiting times, \((200+1)/2 = 100.5\), so the median lies between the 100th and 101st values in the sorted list. Both are 46, so
\[\text{median} = \frac{46 + 46}{2} = 46 \text{ minutes}\]
For the small set 22, 28, 34, 39, and 51 the position is \((5+1)/2 = 3\), so the median is the third value, 34.
The median depends on positions. However extreme the outer values are, each affects the median only through its position. This resistance to extremes is the main advantage of the median.
Mode
The mode is the most frequent value. It is the only measure of centre available for nominal data, and the right one when the question is “what happens most often”.
You read the mode from a frequency table, so for a categorical variable you already have it. The table in Table 1.10 gives the mode of Ticket_Type at a glance: Standard, held by 124 of 200 customers.
For Waiting_Time_Mins there are two modes. Both 42 minutes and 47 minutes occur 10 times, which is more often than any other value. A distribution with two modes is called bimodal. Report both, because two peaks often indicate that two different groups have been combined in one column.
Some columns have no useful mode. Satisfaction holds 143 distinct values across 200 customers, and most occur once, so asking which value is most frequent tells you little. This is typical for a finely measured continuous variable. The mode is mainly useful for categorical data.
Choosing between them
| Situation | Use | Because |
|---|---|---|
| Roughly symmetric, no extremes | Mean | Uses all the data |
| Skewed, or has extreme values | Median | Resists the pull of the tail |
| Nominal data | Mode | Nothing else is defined |
| Ordinal data | Median | Order is meaningful. Equal gaps are not assumed |
| The question is “most common” | Mode | The mode directly answers the question |
Part A. A department’s eight salaries are 28, 30, 31, 32, 33, 35, 36, and 180 thousand.
Calculate the mean and the median. State which you would report to staff asking about typical pay, and say what the other one would tell them.
Part B. The same department records each employee’s contract type: 5 permanent, 2 fixed-term, 1 consultant.
Give the measure of centre for contract type, state its value, and say why the other two measures are unavailable here.
Part A, mean. The total is 405, so \(\bar{x} = 405 \div 8 = 50.6\) thousand.
Part A, median. With \(n = 8\) the position is \((8+1)/2 = 4.5\), so take the mean of the 4th and 5th values: \((32 + 33) \div 2 = 32.5\) thousand.
Report the median. The single salary of 180 pulls the mean above every other salary in the department, so the mean lies above all salaries except the highest. Seven of the eight staff earn less than the mean. Reporting 50.6 to staff asking about typical pay would be technically correct and would give them a misleading picture of the department.
Part B. Contract type is nominal, so the mode is the measure, and its value is permanent, held by 5 of the 8 staff.
The median requires ordered categories, and permanent, fixed-term, and consultant carry no natural order. The mean requires numbers to add, and contract categories are text labels, not quantities. The mode is therefore the appropriate measure for a nominal variable.
Measures of spread
A measure of centre without a measure of spread gives only part of the summary, and often the less useful part. Spread tells a manager how much the cases differ from one another. It also indicates how much uncertainty the average conceals.
Range
The range is the largest value minus the smallest value.
\[\text{range} = \text{maximum} - \text{minimum}\]
For the waiting times, \(90 - 12 = 78\) minutes. It is simple to calculate, but it depends entirely on the two most extreme cases. A single unusual value can therefore determine it. It is fragile on its own and useful mainly alongside other measures.
Deviation, variance, and standard deviation
The deviation of a value is its distance from the mean, \(x - \bar{x}\). A positive deviation lies above the mean, and a negative deviation lies below.
Averaging the deviations gives zero because the deviations always sum to zero. This property follows from the definition of the mean as the point at which deviations balance. For this reason, the deviations are squared before averaging.
The variance is the sum of the squared deviations divided by \(n - 1\):
\[s^2 = \frac{\sum (x - \bar{x})^2}{n - 1}\]
The standard deviation is its square root, which returns the measure to the original units and makes it interpretable:
\[s = \sqrt{\frac{\sum (x - \bar{x})^2}{n - 1}}\]
Variance is expressed in squared minutes, which is difficult to interpret directly. Standard deviation is expressed in minutes, which matches the original variable. Taking the square root returns the measure to the original units.
Worked example 1.10
Calculating a standard deviation by hand
Take the first five waiting times in the teaching dataset: 28, 39, 37, 48 and 61 minutes.
Calculate the variance and the standard deviation by hand, and say what the result tells a manager.
Step 1, the mean.
\[\bar{x} = \frac{28 + 39 + 37 + 48 + 61}{5} = \frac{213}{5} = 42.6\]
Step 2, the deviations and their squares.
| \(x\) | \(x - \bar{x}\) | \((x - \bar{x})^2\) |
|---|---|---|
| 28 | -14.6 | 213.16 |
| 39 | -3.6 | 12.96 |
| 37 | -5.6 | 31.36 |
| 48 | 5.4 | 29.16 |
| 61 | 18.4 | 338.56 |
| Sum | 0.0 | 625.20 |
The deviation column sums to zero, which provides an arithmetic check on the mean. If the sum differs from zero, the mean is incorrect and later calculations will also be incorrect.
Step 3, the variance.
\[s^2 = \frac{625.20}{5 - 1} = \frac{625.20}{4} = 156.3\]
Step 4, the standard deviation.
\[s = \sqrt{156.3} = 12.50 \text{ minutes}\]
Interpretation. These five customers waited 42.6 minutes on average, and a typical customer differed from that average by about 12.5 minutes.
Across all 200 customers, the same calculation gives \(s^2 = 141.74\) and \(s = 11.91\) minutes, so these five are reasonably representative of the file.
You will do this once by hand and let software do it thereafter. Do it once and the standard deviation stops being a button and becomes a quantity you can argue about.
What a standard deviation means in practice
A standard deviation of 12 minutes means waiting times typically differ from the average by about 12 minutes. A standard deviation of 2 minutes with the same mean describes a more consistent service, and a customer would notice the difference long before a manager read the report.
As a rough guide for roughly symmetric data, most cases fall within one standard deviation of the mean. In the teaching dataset, 133 of the 200 waits, or 66.5 per cent, fall between 33.5 and 57.4 minutes, which is the mean plus or minus one standard deviation. Chapter 2 makes this precise with the empirical rule.
Why the sample divides by n minus 1
The sample variance divides by \(n - 1\). The explanation for the minus one is brief.
Deviations are measured from the sample mean, and the sample mean sits, by construction, at the centre of that particular sample. Deviations from it are therefore slightly smaller than deviations from the unknown population mean would be. Dividing by \(n\) would carry that shortfall into the estimate and understate the true spread.
Dividing by \(n - 1\) corrects it. In worked example 1.10, dividing by 5 instead of 4 gives a standard deviation of 11.18 compared with 12.50, roughly 11 per cent smaller. The correction matters for small samples, and its effect becomes smaller as the sample grows. The two formulas converge for large datasets, so the issue mainly affects small samples.
Now that spread has its own symbols, here is the full notation table that Table 1.3 opened. Greek letters describe populations. Latin letters describe samples.
| Quantity | Population | Sample |
|---|---|---|
| Mean | \(\mu = \dfrac{\sum x}{N}\) | \(\bar{x} = \dfrac{\sum x}{n}\) |
| Variance | \(\sigma^2 = \dfrac{\sum (x - \mu)^2}{N}\) | \(s^2 = \dfrac{\sum (x - \bar{x})^2}{n - 1}\) |
| Standard deviation | \(\sigma = \sqrt{\sigma^2}\) | \(s = \sqrt{s^2}\) |
| Proportion | \(\pi\) | \(p\) |
| Size | \(N\) | \(n\) |
Comparing spread across different units
A standard deviation of 11.91 minutes is on a different scale from one of 256.84 USD. A direct comparison of the two therefore provides little information.
The coefficient of variation expresses the standard deviation as a proportion of the mean, which removes the units:
\[CV = \frac{s}{\bar{x}} \times 100\%\]
For the teaching dataset:
| Variable | \(\bar{x}\) | \(s\) | \(CV\) |
|---|---|---|---|
Age |
35.35 years | 7.35 | 20.8% |
Waiting_Time_Mins |
45.46 minutes | 11.91 | 26.2% |
Spending |
505.65 USD | 256.84 | 50.8% |
Spending varies far more, relative to its own average, than either age or waiting time. A manager can use this fact: a forecast of next month’s average spend carries much more uncertainty than a forecast of next month’s average wait.
The coefficient of variation requires a true zero, so it applies to ratio variables only. Applying it to temperature in Celsius produces a number that changes when you switch to Fahrenheit. This shows that the measure is inappropriate for interval variables.
Worked example 1.11
Two teams with the same mean
Two service teams both resolve calls in 24 minutes on average. Team A has a standard deviation of 3 minutes. Team B has a standard deviation of 14 minutes.
A report states that the two teams perform identically. Is the report right?
No. The means match, but the customer experience differs.
Team A is predictable: almost every call lands between about 18 and 30 minutes. Team B is unpredictable. Calls range from very fast to very slow, and a customer’s experience depends on chance.
The means are equal, but the risk differs. Reporting only the average implies that the teams perform identically, which is false.
Two branches both average 300 transactions a day. Branch A has a standard deviation of 20. Branch B has a standard deviation of 95.
Calculate the coefficient of variation for each. Explain what the difference means for staffing, and say which branch you would visit first.
Coefficient of variation. Branch A: \(20 \div 300 = 6.7\%\). Branch B: \(95 \div 300 = 31.7\%\). Branch B’s daily volume is almost five times as variable, relative to its own average.
Both branches average 300 transactions, so on volume alone they look identical.
Branch A, with a standard deviation of 20, is predictable. Most days fall roughly between 260 and 340, so a fixed staffing level works.
Branch B, with a standard deviation of 95, ranges from very quiet to very busy. A fixed level either wastes staff on quiet days or fails customers on busy ones, so it needs flexible staffing and a way to forecast the busy days.
Visit Branch B first, because its variation is the problem the average conceals.
Percentiles, quartiles, and the interquartile range
A percentile marks the position below which a given share of the data falls. The 90th percentile is the value that 90 per cent of cases fall at or below.
Quartiles split the ordered data into four equal parts:
- \(Q_1\), the first quartile, is the 25th percentile
- \(Q_2\), the second quartile, is the median
- \(Q_3\), the third quartile, is the 75th percentile
For Waiting_Time_Mins, the software reports \(Q_1 = 37.75\), \(Q_2 = 46\) and \(Q_3 = 53.25\) minutes.
Read as a sentence: a quarter of customers waited under 38 minutes, half waited under 46 minutes, and a quarter waited over 53 minutes.
The interquartile range is the spread of the middle half:
\[IQR = Q_3 - Q_1\]
Here the value is \(53.25 - 37.75 = 15.5\) minutes. Because it discards the outer quarters, it resists extremes more effectively than the range. The range for the same column is 78 minutes, set by two customers out of 200.
Quartiles fall between data points, so software must interpolate. Nine conventions exist for this interpolation.
For \(Q_1\) of the 200 waiting times, the values in sorted positions 50 and 51 are 37 and 38.
The convention used by JASP, R, and most statistical software puts \(Q_1\) at position \(1 + (n-1) \times 0.25 = 50.75\), giving \(37 + 0.75 \times 1 = 37.75\).
A common textbook convention puts it at position \((n+1)/4 = 50.25\), giving \(37 + 0.25 \times 1 = 37.25\).
Each convention is defensible, and the gap here is half a minute. The difference rarely changes a decision, but it explains why two packages can disagree on the same file. Report which tool produced your figures.
Boxplots and the outlier rule
A boxplot draws the quartiles. The box runs from \(Q_1\) to \(Q_3\), a line inside the box marks the median, and whiskers extend to the furthest values that are still within reach of the box.
“Within reach” has a definition. The standard rule places fences at
\[\text{lower fence} = Q_1 - 1.5 \times IQR \qquad \text{upper fence} = Q_3 + 1.5 \times IQR\]
and flags anything beyond them as a potential outlier.
For the waiting times:
\[\text{lower fence} = 37.75 - 1.5 \times 15.5 = 14.5 \text{ minutes}\]
\[\text{upper fence} = 53.25 + 1.5 \times 15.5 = 76.5 \text{ minutes}\]
Four of the 200 customers fall outside: two waited 12 and 14 minutes, and two waited 77 and 90 minutes. Those four points appear individually on the boxplot, beyond the ends of the whiskers.
The word “outlier” here means “look at this one again”. If you are managing complaints, the customer who waited 90 minutes may be the most urgent case in the file.
Software reports \(Q_1\) as 22 minutes, the median as 31, and \(Q_3\) as 41 for a different branch.
- Calculate the interquartile range.
- Calculate the upper fence.
- State what the IQR tells a manager.
- Explain why the IQR is preferable to the range as a summary of spread when a few customers waited over two hours.
- $IQR = 41 - 22 = $ 19 minutes.
- Upper fence $= 41 + 1.5 = 41 + 28.5 = $ 69.5 minutes. Any wait beyond 69.5 minutes is flagged for a second look.
- The middle half of customers waited between 22 and 41 minutes, a spread of 19 minutes. This describes the typical experience and is the figure to quote alongside the median.
- The range would be driven entirely by the small number of customers who waited over two hours. It would describe those cases. It would fail to describe the service as a whole. Those cases matter and should be reported. They should not define the normal wait. The IQR discards the outer quarters, so it describes what is typical. The fences describe what is unusual.
Shape, skewness, and outliers
Comparing the mean with the median reveals the shape before you draw a single graph.
| Relationship | Shape | What it means |
|---|---|---|
| Mean close to the median | Roughly symmetric | The average describes the typical case well |
| Mean well above the median | Right-skewed | A tail of high values pulls the mean up |
| Mean well below the median | Left-skewed | A tail of low values pulls the mean down |
Apply it to the teaching dataset.
| Variable | Mean | Median | Difference | Shape |
|---|---|---|---|---|
Waiting_Time_Mins |
45.46 | 46.00 | -0.54 | Close to symmetric |
Satisfaction |
3.81 | 3.86 | -0.04 | Close to symmetric |
Spending |
505.65 | 494.28 | +11.37 | Mildly right-skewed |
Waiting time here is close to symmetric, although waiting times in operational data often skew to the right. Right skew is common in business data. Waiting times, spending, and incomes have a floor at zero and often lack a fixed ceiling, so the tail runs to the right. In this branch, waits sit in a fairly narrow band with a few long values. The description should follow the data.
Spending shows the mild right skew you would expect from money, with the mean 11 USD above the median.
An outlier is a value far from the rest. It may be an error, a special case, or an important observation. Investigate it and report what you found. Deleting it silently is inappropriate, because a reader who sees only the remaining cases evaluates an edited file.
Worked example 1.12
Skewness in online orders
Order values at a different company have a mean of 480 and a median of 395.
Identify the shape of the distribution and state which figure belongs in a report on typical spending and which belongs in a revenue forecast.
The mean exceeds the median by 85, so the distribution is right-skewed. Most orders are modest and a few large ones pull the average up.
Reporting 480 as the typical order overstates what most customers spend. The median, 395, is the defensible figure for “typical”.
The mean, 480, is the right figure for forecasting total revenue, because total revenue is the mean multiplied by the number of orders and the large orders count fully towards it.
Both are correct, and they answer different questions. The report should name the question it is answering.
A branch reports these figures for waiting time: mean 52 minutes, median 41 minutes, \(Q_1\) 33, \(Q_3\) 58, minimum 9, maximum 186.
- What shape is the distribution, and how do you know?
- Calculate the IQR and the upper fence.
- Which measure of centre would you quote to a customer asking about a typical wait?
- The branch manager proposes removing the 186-minute case as “obviously an error”. What would you advise?
- Right-skewed. The mean sits 11 minutes above the median, which happens when a tail of long waits pulls the average up. The gap between the median and the maximum is 145 minutes. The gap between the median and the minimum is 32 minutes. This supports the same conclusion.
- \(IQR = 58 - 33 = 25\) minutes. Upper fence \(= 58 + 1.5 \times 25 = 95.5\) minutes. The 186-minute wait is far beyond it.
- The median, 41 minutes. Half of customers waited no longer than this value. The mean is inflated by waits that most customers avoided.
- Ask what evidence identifies it as an error. A very busy afternoon could produce the same value in the column. A 186-minute wait is possible in a service where customers queue. If the record is a data entry error, correct it and record the correction. If it is real, it may be one of the most important rows in the file, because that customer is likely to complain. In either case, record the decision in the report. A reader who sees only the remaining cases evaluates an edited file.
Data quality and ethical description
Every choice in this chapter can mislead while every sentence remains literally true.
Examples include choosing the mean over the median because it is larger, truncating an axis, reporting a percentage while leaving out its denominator, and presenting a convenience sample as though it represented everyone. Each choice may appear defensible in isolation, and each choice can move a reader toward a conclusion the evidence withholds.
One question provides a test: would you make the same choice if the number pointed the other way? If the honest summary is the one that happens to suit you, the choice is defensible. If you would have chosen differently, the choice is selective.
A practical habit follows from it. Record the choices you made in one sentence in the report: which measure of centre and why, how any bands were drawn, what happened to unusual cases, and what the sample covers. One sentence converts a claim into a statement a reader can check.
Chapter 8 returns to this topic in detail. The issue begins here, because the opportunity to mislead begins with these choices.
Section summary
The mean, median, and mode answer different questions, and the shape of the data decides which one to report. Spread is an essential part of the summary. Equal means with different standard deviations describe different risks. Dividing by \(n - 1\) corrects for measuring deviations from the sample mean. The IQR and the 1.5 multiplier separate typical values from unusual values. Comparing the mean with the median reveals skew without a graph.