Median vs Mean: Which Should You Use?

When the median beats the mean: skewed data, outliers, incomes and house prices. How the gap between them shows skew, and when the mean is better.

The Median Is Not the Average of the Two Middle Numbers

The most common mistake about the median vs mean question is assuming the median is just the average of the two middle numbers in an even-sized dataset, without first sorting the data. That is wrong. The median is the middle value of a dataset when it is ordered from smallest to largest, and for an even count, it is the average of the two middle values, but only after sorting. The mean, by contrast, is the arithmetic average, the balancing point of all values. When data is symmetric and free of outliers, the two land close together. When data is skewed, they diverge.

Here is what is true: the median resists the pull of extreme values. A single outlier can yank the mean upward or downward, but the median shifts only if that outlier changes which value sits in the middle. For a reader deciding which average to report, the choice depends on the data's shape. If the data has a long tail in one direction, the median is the honest summary of the typical case. If the data is bell-shaped and balanced, the mean is the more powerful tool because it supports further calculations like variance and standard deviation.

Quick Answer: Which Average Should You Report?

Report the median when your data is skewed or has outliers. It tells you the middle value of the sorted dataset, unaffected by extreme observations. Report the mean when the data is symmetric and has no significant outliers. It uses every value and supports more advanced statistics. For a quick decision, plot the data: if the tail extends to the right, the mean exceeds the median, and the median is the safer choice. If the tail extends to the left, the mean falls below the median, and the median still wins for describing the center. If the plot looks symmetric, either works, but the mean is the standard for further analysis.

The failure case is when you report the mean on skewed data without noting the skew. A salary dataset with a few executives earning millions will show a mean far above the median. A reader who sees only the mean assumes the typical worker earns far more than they do. That is misleading. When in doubt, report both, but lead with the median for skewed data, and say why.

How Outliers Move the Mean but Not the Median

An outlier is a value far from the rest of the data. Its effect on the mean is dramatic. Consider a dataset of five house prices: a low price, three mid-range prices, and a very high price. The mean becomes a number no house resembles. The median, after sorting, is the third value, which is the typical price. The outlier moved the mean significantly, but the median did not budge. The outlier only changes the median if it crosses the middle position. That is the core of the skewed data median argument: the median is completely unaffected by the magnitude of an outlier, only by its position in the sorted order.

This is not a minor technicality. When you use the median, you choose a measure that ignores the exact value of extreme observations. That is why the median is the robust statistic of choice for income, home prices, and any dataset where a few large values could mislead. The mean is not useless. It is the only measure that uses every data point. That makes it sensitive to outliers but more efficient when the data is clean. The practical rule: if you see a value that looks like a mistake or a rare event, the median is your friend. If you see a value that is genuinely representative, the mean is fine.

Reading Skew from Mean vs Median

The relationship between the mean and the median tells you the shape of your distribution without a plot. In a right-skewed distribution, where the tail extends toward higher values, the mean is pulled in that direction. The mean exceeds the median. In a left-skewed distribution, with a tail toward lower values, the mean falls below the median. This is the standard relation taught in introductory statistics. For example, OpenStax Introductory Statistics, section 2.6, states that in a right-skewed distribution, the mean is greater than the median, and in a left-skewed distribution, the mean is less than the median.

When the mean and median are close, the data is approximately symmetric. Either measure is acceptable. But the difference between them is a diagnostic tool. If you compute both and the mean is much higher than the median, you have high outliers or a right tail. Investigate. If the mean is much lower, you have low outliers. The size of the gap is not standardized, but a gap larger than about 10% of the median's value often signals meaningful skew. This is how you read skew from mean vs median. It separates a careful analyst from someone who just runs a spreadsheet function.

Why Incomes and Home Prices Are Reported as Medians

Governments and real estate agencies report the median household income and the median home price for a simple reason: these distributions are right-skewed. The mean would be misleading. The U.S. Census Bureau states the rationale directly in its publications on income and poverty: because the median household income is not affected by the exact income of high-income households, it is a better measure of the typical income than the average household income. The same logic applies to home prices. A handful of mansions can inflate the average, but the median shows the price of the middle home.

When you read that the median household income in a county is a certain amount, it means half of households earn more and half earn less. The mean might be higher because of a few billionaires. That number does not describe the typical household. For a business analyst or a policy maker, the median is the only defensible choice for such data. The failure case is when a report uses the mean for income or price data without acknowledging the skew. It misleads readers into thinking the typical value is higher than it is. If you see a mean income reported for a region with known inequality, treat it with suspicion and ask for the median.

When the Mean Is Better

The mean is the better measure when your data is symmetric and free of outliers. It uses every observation and supports further statistical calculations. For example, if you are measuring the average test score in a class where no student scored dramatically higher or lower, the mean is the standard choice. It is also the only measure that can be used in formulas for variance, standard deviation, and regression. The median is less efficient in symmetric data, meaning it has more sampling variability for the same sample size.

The mean is also necessary when you need a total. If you are summing expenses or projecting total revenue, you cannot sum medians. The median of a set of costs does not tell you the total cost. The mean does, when multiplied by the count. The rule is not that the mean is always wrong. It is that the mean is wrong for skewed data. For a clean, symmetric dataset, report the mean, and pair it with a measure of spread like the standard deviation. For anything with a long tail, use the median and the interquartile range.

Worked Comparison: A Single Outlier Changes the Story

Here is a worked example with an outlier to show the difference. Take five wait times at a clinic in minutes: 5, 10, 12, 15, and 90. The mean is 26.4 minutes. The median, after sorting, is the third value, 12 minutes. The outlier, the 90-minute visit, nearly doubled the mean. The median says the typical wait is 12 minutes. If you report the mean to a patient, they expect to wait over 25 minutes. If you report the median, they know most visits are under 15 minutes. The median is the honest answer here.

Now consider the reverse: a dataset with no outliers, such as 10, 11, 12, 13, 14. The mean is 12, and the median is also 12. They agree. Either is fine. The lesson is that the median vs mean question is not about which is more accurate in the abstract. It is about which describes the center of your specific data. For skewed data, the median is the robust choice. For symmetric data, the mean is more precise. Always compute both. Plot the data if you can. Let the shape of the distribution make the decision for you.

What the Median Tells You That the Mean Cannot

The median tells you the middle value of a sorted dataset. It is a direct statement about position: half the data falls below it, half above. The mean tells you the balancing point, where the sum of deviations above equals the sum below. In a symmetric distribution, these coincide. In skewed data, the median is the only one that describes the typical observation. The median is valid for ordinal, interval, and ratio data, but not for nominal data, where categories have no order. You cannot take the median of a list of colors or zip codes.

The median also has a useful property with grouped data. When you have a frequency table with open-ended classes, such as "a high amount or more," the median can still be estimated from the cumulative frequencies. The mean requires an assumed value for the upper bound. This is why the median is preferred in official statistics, where the top income bracket is often capped or open-ended. The practical takeaway: if your data has a natural order and a skewed shape, the median will not lie to you about the center.

How to Use the Median in Spreadsheets and Reports

In a spreadsheet, the MEDIAN function ignores text and logical values. It includes numbers only, which is what you want when cleaning messy data. The AVERAGE function includes everything. A stray label or an error can distort it. When you have a column of salaries with a few blanks, use =MEDIAN(A2:A100) and =AVERAGE(A2:A100) side by side. If they differ wildly, you have outliers. The median is the one to report. This is the median vs average decision made concrete for a data analyst.

The failure case is when a spreadsheet user copies data with a text string like "N/A" in a cell. The AVERAGE function returns an error or a distorted value. The MEDIAN function will ignore that text, but only if it is truly text. A number stored as text may be included depending on the software. Always verify the data type. For a business report, pair the median with the interquartile range, the range from the 25th to the 75th percentile, to show the spread around the middle. The median alone does not tell you about variability. Use both.

When the Median Fails: Bimodal Data and Other Edge Cases

The median is not a universal answer. For bimodal data, where two distinct clusters exist, the median may fall in a low-frequency gap between the peaks. It becomes unrepresentative of either group. For example, test scores that cluster at 40 and 90, with nothing in between, will have a median around 65, a value no student scored. In such cases, report the two modes or a histogram, not a single center. The median is also not helpful for nominal data, where categories have no order. You cannot even compute it.

Another edge case is the tie convention for even sample sizes. When n is even, the median is the average of the two middle values. Some software uses the lower value for the 50th percentile, depending on the quantile method. This is a real source of confusion when two calculators give different answers for the same data. If you are hand-calculating, always sort first, then average the two middle values. For a frequency table, use the median class, the class where the cumulative frequency crosses n divided by 2, and interpolate if needed. These details matter because a median computed wrong is worse than none.

Practical Rules for Your Next Dataset

Here is the rule of thumb that ends the median vs mean debate: if the data has extreme values or is skewed, use the median. If the data is symmetric with no outliers, use the mean. For income, home prices, or any real-world data with a long tail, default to the median. For test scores, measurements, or any controlled experiment, the mean is usually fine. Never report one without checking the other. The difference between them is the first sign of skew.

  • Check the data's shape with a quick histogram or box plot before choosing.
  • Compute both the mean and median; a large gap signals outliers.
  • Report the median with the IQR for skewed data, and the mean with the standard deviation for symmetric data.

If you are writing a report, state which measure you used and why. A single sentence, such as "The median is reported because the distribution is right-skewed," prevents misunderstanding. The mean is a poor choice when a few values dominate. The median is a poor choice when the data is bimodal. Know which case you are in, and you will never mislead your reader again.

Median vs Mean FAQ

What is the most common mistake people make when calculating the median for an even-sized dataset?

The most common mistake is assuming the median is just the average of the two middle numbers without first sorting the data. The correct method is to sort the dataset from smallest to largest first, then average the two middle values.

How can you tell if a distribution is right-skewed by comparing the mean and median?

In a right-skewed distribution, where the tail extends toward higher values, the mean exceeds the median. This is the standard relation taught in introductory statistics, such as in OpenStax Introductory Statistics, section 2.6.

Why do governments and real estate agencies report the median for incomes and home prices instead of the mean?

These distributions are right-skewed, so the mean would be misleading. The U.S. Census Bureau states that because the median household income is not affected by the exact income of high-income households, it is a better measure of the typical income than the average.

In the worked example with five wait times (5, 10, 12, 15, and 90 minutes), what are the mean and median?

The mean is 26.4 minutes, and the median, after sorting, is the third value, 12 minutes. The outlier nearly doubled the mean, while the median remained at the typical value.

When is the mean a better measure to report than the median?

The mean is better when your data is symmetric and free of outliers, such as measuring average test scores with no extreme values. It uses every observation and is the only measure that can be used in formulas for variance, standard deviation, and regression.

What is a failure case for the median with bimodal data?

For bimodal data with two distinct clusters, such as test scores clustering at 40 and 90 with nothing in between, the median may fall around 65, a value no student scored. In such cases, the median becomes unrepresentative of either group.