How to Find the Median of Grouped Data

Find the median from a grouped frequency table with the formula L + ((n/2 − CF)/f) × h. Identify the median class, then work through a full example.

When Data Is Grouped

You rarely meet the median of grouped data in the raw. A spreadsheet hands you individual incomes, and you sort them, find the middle, and you are done. But the moment someone collapses those numbers into classes like “40,000-49,999,” the exact values vanish. What you have left is a frequency distribution, a table of ranges and counts, and the median of that grouped data is no longer a value you can point to. It is a number you must estimate.

Grouped data exists because raw values are bulky, or because the source itself only published ranges. Census tables, salary surveys, and exam score reports all arrive pre-binned. The median of grouped data answers the same question the plain median answers: what value splits the ordered set in half? But because the individual observations are gone, you cannot sort what you do not have. You work with the class boundaries, the class width, and the cumulative frequency to reconstruct where the middle falls.

The estimate is not a guess. It is an interpolation, a straight-line assumption that the values inside the middle class spread evenly across its width. That assumption is wrong in detail and right on average, which is why the result is called an estimate, not a measurement. The formula you are about to use is the standard one from NCERT Class 10 Mathematics chapter 13, and it appears in nearly identical form in OpenStax Introductory Statistics. Both texts agree on the mechanics, even if their notation differs slightly.

The Grouped Data Median Formula and Each Symbol

Formula Breakdown

The grouped data median formula is written as:

Median = l + [(n/2 - cf) / f] × h

Each symbol does one job. The l is the lower boundary of the median class, the class that contains the middle observation. The n is the total number of observations, which you get by summing every frequency in the table. The cf is the cumulative frequency of the class just before the median class, the running total of all observations that fall below it. The f is the frequency of the median class itself, how many observations sit inside that one interval. The h is the class width, the difference between the upper and lower boundary of the median class.

Read the formula as a sequence of steps. First find n/2, the position of the middle observation in the ordered data. Then subtract cf, the number of observations that came before the median class. What remains is how far into the median class the middle observation sits. Divide that by f, the class frequency, which converts the count into a fraction of the class. Multiply by h, the class width, which turns that fraction into a distance along the number line. Add l, the lower boundary, and you have the estimated median.

OpenStax writes the same idea as L + (n/2 - f_cum)/f × w. The letters differ, L for lower boundary, w for width, but the arithmetic is identical. The formula assumes the data inside the median class is distributed uniformly. That is the interpolation assumption, and it is the only reason the calculation is possible. Without it, the median of grouped data would be indeterminate, because the middle observation could be anywhere inside that class.

Finding the Median Class

Using Cumulative Frequency

The median class is the class interval whose cumulative frequency first reaches or exceeds n/2. You cannot find it by looking at the frequencies alone. A class with a huge frequency might sit early in the distribution, and a class with a small frequency might sit late. The cumulative frequency is the running total, and it is the only column that tells you where the middle observation lives.

Build the cumulative frequency column by adding each class's frequency to the total of all previous classes. The first class has a cumulative frequency equal to its own frequency. The second adds its frequency to the first. The third adds to the second, and so on. When the cumulative frequency reaches or passes n/2, that class is the median class. If n is even, n/2 falls on a boundary between two observations, and the median class is still the one whose cumulative frequency first reaches that value.

A common error is to pick the class with the highest frequency, confusing the mode class with the median class. They are different things. The mode class is the most common interval. The median class is the one that contains the middle of the ordered distribution. In a skewed distribution, they can be far apart. The median class is always found through the cumulative frequency, never through the frequency column alone.

For an ungrouped frequency table, the rule is simpler. If n is odd, the median is the value where the cumulative frequency reaches (n+1)/2. If n is even, it is the average of the two middle values.

Worked Example: Class Intervals and Frequencies
Class IntervalFrequency (f)Cumulative Frequency (cf)
10–1944
20–29610
30–39818
40–49523
50–59225

Worked Example: Step-by-Step

Take the table above. Total frequency n is 25, so n/2 is 12.5. The cumulative frequency column reads 4, then 10, then 18. The first class whose cumulative frequency reaches or exceeds 12.5 is the third class, 30-39. That is the median class. Its lower boundary l is 29.5, because the class runs from 29.5 to 39.5 if you assume continuous data. The cumulative frequency before this class, cf, is 10. The frequency of the median class, f, is 8. The class width h is 10.

Plug the numbers into the formula: Median = 29.5 + [(12.5 - 10) / 8] × 10. The numerator inside the brackets is 2.5. Divide by 8 to get 0.3125. Multiply by 10 to get 3.125. Add 29.5 to get 32.625. The estimated median is 32.6, rounded to one decimal place.

Notice the median class boundaries. The table says 30-39, but the lower boundary is 29.5, not 30. This is because class intervals in grouped data are continuous, and the boundary sits halfway between 29 and 30. If the table used 30-39.9 instead, the lower boundary would be 30.0. The formula is sensitive to this, so check the table's convention before you start. The value 32.6 means that, based on the assumption of uniform distribution, half the observations fall below 32.6 and half above.

Why It Is an Estimate

The result 32.6 is not a value that exists in your data. No single observation in the original dataset necessarily equals 32.6. The median of grouped data is an estimate because the raw values inside the median class are lost. The formula assumes they spread evenly across the class width, but in reality they could cluster at the lower end, pile up in the middle, or bunch near the upper boundary. The interpolation smooths over that unknown distribution and reports a single number.

The estimate gets worse as the class width grows. A narrow class like 30-34 leaves less room for error than a wide class like 30-59. The estimate also suffers if the data inside the median class is heavily skewed. If most values in that class sit near the lower boundary, the true median is lower than the estimate. If they sit near the upper boundary, the estimate undershoots. The formula cannot detect this because the individual values are gone.

This is why the median of grouped data is reported with a caveat: it approximates the median of the underlying raw data. For a quick sense of the center, it is good enough. For a precise figure, you need the original values or a dataset that has not been binned. The failure case is when someone treats the estimate as exact, then builds further calculations on it. The error compounds.

Ogive Method (Graphical)

You can also find the median of grouped data by drawing, and the drawing is called an ogive, a cumulative frequency curve. Plot the cumulative frequency against the upper boundary of each class. The points are (upper boundary, cumulative frequency). Connect them with a smooth curve or straight line segments. The median is the value on the horizontal axis where the cumulative frequency equals n/2.

Draw a horizontal line from n/2 on the vertical axis to the curve, then drop a vertical line to the horizontal axis. The point where the vertical line lands is the median. The ogive gives the same answer as the formula, because both use the same cumulative frequency and the same interpolation assumption. The graph is a visual way to see the middle rather than calculate it.

The ogive method is worth learning because it shows you the shape of the distribution. A steep ogive means a dense class, a shallow one means a sparse class. The median sits where the cumulative frequency crosses the halfway line, and you can see whether that happens in a steep or shallow section. The method fails gracefully when the data has open-ended classes, like “100 or more,” because the upper boundary is unknown. In that case, the ogive cannot be drawn without an assumption, and the formula needs the same assumption. The practical rule: use the ogive to check your arithmetic, not to replace it.

Median of Frequency Distribution: Common Pitfalls

Three Mistakes to Avoid

The median of frequency distribution trips up newcomers in three predictable ways. First, they forget to sort the data before finding the middle. With raw data, the median of {5, 1, 3} is 3, not 5, and the error is obvious. With grouped data, the classes are already ordered, so the sorting step is built in, but the cumulative frequency column must be computed correctly. One arithmetic slip in the running total sends the median class to the wrong row.

Second, they confuse the median with the mean. The mean of grouped data uses class midpoints as stand-ins for the actual values, which biases the result. The median uses only the position of the middle observation and the width of the median class. For skewed data, the mean and median diverge, and the median is the one that resists the pull of a long tail. A single extreme value changes the mean but not the median, because the median only cares about the middle, not the edges.

Third, they interpolate the median for open-ended classes without stating the assumption. If the last class is “100 or more,” you do not know the upper boundary, so you cannot compute the class width. The formula requires h, and h is undefined. The standard workaround is to assume the open class has the same width as the previous class, but that is a guess. State it as a guess, or the estimate carries false confidence.

Unequal Class Widths and Other Edge Cases

Not every frequency distribution uses equal class widths. You might have classes like 0-9, 10-19, 20-49, and 50-99. The median formula still works, but you must use the class width h of the median class itself, not an average width. The median class selection, based on cumulative frequency reaching n/2, is unaffected by unequal widths. The lower boundary l and the width h come from the median class alone.

The tricky part is when the median class has a different width than its neighbors. Suppose the median class is 20-49, width 30, and the previous class is 10-19, width 10. The formula uses h = 30 for the median class. The interpolation assumes the 30-unit span distributes the observations uniformly, which is a stronger assumption when the class is wide. The estimate is still valid, but the confidence in it drops.

Open-ended classes are the other edge case. If the first class is “under 10” or the last is “100 or more,” the formula needs a boundary that does not exist. The convention is to assume the open class mirrors the width of the adjacent class, but this is a stated assumption, not a fact. If the data is truly open-ended, consider reporting the median class instead of the median, or use a different measure entirely.

Median of Grouped Data in Spreadsheets and Tools

Spreadsheets rarely compute the median of grouped data directly, because they expect raw values. Excel's MEDIAN function takes a range of cells, ignores text, logical values, and empty cells, and returns the middle value. It does not understand a frequency table. If you have a frequency table, you must expand it into a list of individual values, which defeats the purpose of grouping. The workaround is to use the formula above by hand, or to build a helper column that repeats each class midpoint according to its frequency.

Excel's MEDIAN is not a built-in PivotTable value field option. If you need a median inside a PivotTable, you must use Power Pivot's DAX MEDIAN function or add a helper column to the source data. Google Sheets has a similar limitation: its MEDIAN function does not support an array of arrays directly, so you must use FLATTEN or an array literal to combine multiple ranges.

The deeper problem is quartile method disagreement. Different tools use different methods for computing quartiles, and the median is the 50th percentile. Hyndman and Fan's 1996 paper catalogs nine methods, labeled R-1 through R-9. Excel's QUARTILE.INC uses R-7, the default in most statistical software. QUARTILE.EXC uses a different method, R-6. For the median itself, all methods agree, because the middle value is unambiguous. But the quartiles around it do not agree.

The Weighted Median and When It Applies

The median of grouped data is sometimes confused with the weighted median, but they are different tools. The weighted median assigns a weight to each observation, and the median is the value where the cumulative weight reaches half the total. This is useful when some observations represent more people or more events than others. The weighted median method is the right tool when your frequencies are not simple counts but importance weights.

For most grouped data, the frequencies are counts, and the plain median formula applies. The weighted median is for surveys with sampling weights, where each respondent represents a different number of people in the population. If your frequency table comes from a weighted survey, do not use the grouped data median formula. Use the weighted median instead. The distinction matters because the two can give different answers, and the weighted version is correct for survey data.

Practical Guidance for the Reader

Your Workflow

When you face a frequency table, your first move is to compute the cumulative frequency column. Do this before you look at any formula. Find n/2, then scan the cumulative frequency column from top to bottom until you hit a value that is greater than or equal to n/2. That row is your median class. Write down its lower boundary, its frequency, and its width. Then plug them in.

If you are using a calculator or spreadsheet, check the quartile method it uses. Most calculators do not disclose it, and the difference shows up in the quartiles, not the median. The median of grouped data is safe because the middle is the middle. But if a calculator gives you a median that seems off, verify the cumulative frequency column first. That is where the errors live.

If your data has an open-ended class, stop and decide whether the median is the right measure. The mean is worse, because it needs a midpoint for the open class, and the midpoint is pure invention. The median needs a width, which is also an assumption. If you can, get the raw values or a less aggregated table. If you cannot, state your assumption and move on.

Common Questions

What is the median of grouped data?

It is an estimate of the median calculated from class intervals, using the formula l + [(n/2 - cf)/f] × h, where the median class is found via cumulative frequency.

How do I find the median class?

Compute cumulative frequencies and select the first class where the cumulative frequency is greater than or equal to n/2.

Why is the median of grouped data not exact?

Because the raw values inside the median class are unknown, the formula assumes uniform distribution across the class width.

Can I use the median formula for open-ended classes?

Only with an assumption about the missing boundary, typically that the open class matches the width of the adjacent class.

Does Excel calculate the median of grouped data?

No, Excel's MEDIAN function needs raw values. For grouped data, use the interpolation formula or expand the frequency table.

What is the difference between median and mean for grouped data?

The mean uses class midpoints and is pulled by extreme values; the median uses only the middle position and class width, resisting outliers.

Which quartile method should I use for the median?

All methods agree on the median itself; disagreement appears in quartiles. Use R-7 (Excel's QUARTILE.INC) unless told otherwise.