Mean, Median, Mode (and Range)

How to find the mean, median, mode and range of a data set, with one worked example for all four, what to do with no mode or two modes, and a memory trick.

Mean, Median or Mode: Which One Answers Your Question

You have a homework problem that asks for the mean, median and mode of one dataset, and you need all three, plus the range, in one pass. The answer is to sort the numbers first, then work through three different definitions on the same sorted list; the mean is the arithmetic average, the median is the physical middle, and the mode is the most frequent value. For the set {4, 2, 8, 2, 6}, the mean is 4.4, the median is 4, and the mode is 2, because 2 appears twice. A single worked example walks that exact sequence, then a practice set with answers lets you check each step.

Mean Median Mode Range: How to Find Each

You have a homework problem that asks for the mean, median and mode of one dataset, and you need all three, plus the range, in one pass. The answer is to sort the numbers first, then work through three different definitions on the same sorted list; the mean is the arithmetic average, the median is the physical middle, and the mode is the most frequent value. For the set {4, 2, 8, 2, 6}, the mean is 4.4, the median is 4, and the mode is 2, because 2 appears twice. A single worked example walks that exact sequence, then a practice set with answers lets you check each step.

How to Find Mean and Median: The Step-by-Step

Sort First, Then Calculate

The mean is found by adding every number in the dataset, then dividing by the count of numbers. The median is found by sorting the numbers from smallest to largest, then locating the middle value; if the count is even, average the two middle values. For example, to find the mean of {5, 1, 9, 3, 7}, add them to get 25, then divide by 5 to get 5. To find the median, sort to {1, 3, 5, 7, 9}, and the middle value is 5.

When the dataset has an even count, like {2, 4, 6, 8}, the median is the average of 4 and 6, which is 5. This works even if the two middle values are equal; the median is then that repeated value. For grouped data in a frequency table, find the cumulative frequency and locate the class where the cumulative frequency first reaches half the total; that class contains the median. Within that class, interpolate to estimate the exact median value, but know that this is an estimate, not an actual data point.

Outliers and the Median Advantage

The most common mistake is skipping the sort. If you average the two middle numbers from an unsorted list, you get a wrong answer. Always sort first, then apply the rule. For a dataset with a large outlier, like {1, 2, 3, 100}, the mean is 26.5 but the median is 2.5; the median better represents the typical value because the outlier skews the mean. This is why the median is preferred for income data, where a few high earners distort the average.

Mode Definition and When It Matters

The mode definition is simple: it is the value that appears most frequently in a dataset. Unlike the mean or median, the mode can be used with nominal data, such as favourite colour or product category, because it requires no numerical order. A dataset can have one mode, two modes (bimodal), three or more modes (multimodal), or no mode at all if every value occurs equally often.

For the dataset {1, 1, 2, 3, 3, 3}, the mode is 3 because it occurs three times. If two values tie for the highest frequency, like {1, 1, 2, 2, 3}, the dataset is bimodal with modes 1 and 2. Do not average the modes; report both. For a frequency table, the mode is the value with the highest frequency, and the modal class is the interval with the highest frequency in grouped data.

The mode is the only measure of central tendency that can be a non-numeric value, like 'red' or 'SUV'. However, it is often unstable with small samples; a single change in data can shift the mode dramatically. For continuous data, no value repeats exactly, so the mode is often reported as the midpoint of the modal class. When a distribution is skewed, the mode sits at the peak of the curve, while the median is the middle and the mean is pulled toward the tail.

Mean Median Mode Relation in Skewed Data

In a right-skewed distribution, also called positively skewed, the mean is greater than the median, which is greater than the mode. For example, with data like {1, 2, 2, 3, 10}, the mean is 3.6, the median is 2, and the mode is 2. The mean is dragged up by the extreme value 10, while the median and mode stay near the bulk of the data. In a left-skewed distribution, the pattern reverses: the mean is less than the median, which is less than the mode.

For symmetric data, the mean equals the median equals the mode, assuming a single-peaked distribution. This is true for a normal bell curve. The relationship between the three measures is a quick check on the shape of your data: if mean and median differ substantially, the data is likely skewed. If they are close, the data is roughly symmetric.

This relation matters in practice. When a government reports median household income, it is choosing the median over the mean because income distributions are right-skewed; the mean would be inflated by top earners. The median gives the middle household's income, which is more representative of typical earnings. But the mean is still necessary for calculating total income, since you cannot sum medians to get a total.

When the Median Fails: Open Classes and Bimodal Gaps

Open-Ended Classes and Bimodal Distributions

The median is robust, but it is not a cure-all. For grouped data with an open-ended class, like '100 or more', you cannot compute an exact median because the upper boundary is unknown. You must either assume a plausible upper limit, which is a guess, or report the median of the closed classes only. Statistical agencies often top-code such values, replacing them with a cap, which is an approximation.

For a bimodal distribution, the median may fall in a low-frequency gap between the two peaks. Consider data {1, 1, 2, 2, 9, 9, 10, 10}; the median is (2+9)/2 = 5.5, which is not representative of either cluster. Reporting just the median hides the fact that there are two distinct groups. In such cases, present the two modes and the gap, not a single central value.

Spreadsheet Caveats

When you use a spreadsheet, Excel's MEDIAN function ignores text and logical values, but it includes numbers stored as numbers. Numbers entered as text are ignored, which can silently change your result. To avoid this, ensure numeric columns are formatted as numbers, not text. Google Sheets behaves the same way. If you need a conditional median, use an array formula like =MEDIAN(IF(range=criteria, values)) and press Ctrl+Shift+Enter; this is not a built-in PivotTable option, so you must build the array yourself.

Which Traveller Should Use the Median and Which Should Not

The median suits anyone summarising skewed data: a job hunter comparing salaries, a home buyer looking at neighbourhood prices, or a policymaker reading income reports. If you are a student in a first statistics course, the median is a homework staple, but you must also learn the mean for sums. A data analyst using spreadsheets should default to the median for skewed columns, but switch to the mean when the total matters, such as summing sales.

The median is not for everyone. If you need the total cost, revenue, or count, the mean is required; you cannot add medians to get a meaningful total. For nominal data like product categories, the median is invalid; use the mode instead. For time-series trend analysis, the median is not a standard summary; use a moving average or regression. And if your data is symmetric and clean, the mean and median are nearly identical, so the choice is cosmetic, but the mean is easier to compute in a spreadsheet.

When the normal route is closed, such as a dataset with an open-ended class or a calculator that does not state its quartile method, do not guess. Use the median of the closed classes, or check the calculator's manual for the method. If you are comparing two calculators and they differ, the difference is the quartile interpolation method, not an arithmetic error. The median's breakdown point is 50%, meaning half the data can be arbitrarily contaminated without changing it; the mean has a breakdown point of 0%, so a single extreme value ruins it. That is the strongest argument for the median, and the reason it survives contact with real, messy data.

Frequently Asked Questions

What is the mean of the dataset {4, 2, 8, 2, 6}?

The mean is 4.4. This is found by adding every number in the dataset (4+2+8+2+6 = 22) and then dividing by the count of numbers, which is 5, giving 22/5 = 4.4.

How do you find the median when a dataset has an even number of values?

When the dataset has an even count, like {2, 4, 6, 8}, the median is the average of the two middle values. For that set, the median is the average of 4 and 6, which is 5.

Can the mode be used for non-numeric data?

Yes, the mode is the only measure of central tendency that can be a non-numeric value, like 'red' or 'SUV'. It can be used with nominal data, such as favourite colour or product category, because it requires no numerical order.

What is the relationship between mean, median, and mode in a right-skewed distribution?

In a right-skewed distribution, the mean is greater than the median, which is greater than the mode. For example, with data like {1, 2, 2, 3, 10}, the mean is 3.6, the median is 2, and the mode is 2.

Why might the median fail to represent a bimodal distribution?

For a bimodal distribution, the median may fall in a low-frequency gap between the two peaks. For data {1, 1, 2, 2, 9, 9, 10, 10}, the median is (2+9)/2 = 5.5, which is not representative of either cluster.

What is the breakdown point of the median compared to the mean?

The median's breakdown point is 50%, meaning half the data can be arbitrarily contaminated without changing it. The mean has a breakdown point of 0%, so a single extreme value ruins it.