How to Calculate a Weighted Median
A weighted median is the value where cumulative weight reaches half the total. The method, an example, the exact-50% tie, and weighted mean vs median.
What a Weighted Median Is
The ordinary median is the middle value of a sorted list. When every observation carries the same weight, that is all you need. But real data rarely works that way. A single household may represent many people in a survey. One test question may be worth three times another. A weighted median is the value that splits a weighted dataset in half, so that half of the total weight lies below it and half lies above it. It is the median you use when the observations do not each count once.
Here is the core difference from the plain median: you sort the values, but you sort them by their weights, not by how many rows exist. The weighted median is the value where the cumulative weight crosses 50% of the total. If that sounds abstract, the step-by-step method below makes it concrete. The weighted median procedure has only four steps.
Step-by-Step Method for a Weighted Median
Four Steps to Find the Weighted Median
First, list every distinct value in your dataset. Second, assign each value its weight. The weight is the importance, the frequency, or the number of observations that value represents. Third, sort the values from smallest to largest, keeping each value paired with its weight. Fourth, add the weights in that sorted order until the cumulative sum reaches or passes half of the total weight. The value at which that happens is the weighted median.
Handling the Tie Zone
If the cumulative weight lands exactly on 50% at a particular value, you are in the tie zone. The convention is to take the arithmetic mean of that value and the next larger value. If the cumulative weight jumps from 45% to 55% at a single value, that value is the weighted median, no averaging needed.
The most common mistake is treating weights as if they were frequencies in a plain median. A frequency table lists how many times each value occurs, and the plain median finds the middle row. A weighted median uses the same cumulative logic but allows the weights to be anything: survey weights, monetary amounts, or importance scores. If your weights are all equal to 1, the weighted median collapses to the ordinary median.
Weighted Median Example With a Frequency Table
Here is a worked example. A small business pays five employees, but one is an executive with an outsized salary. The salaries are: in a band around $30,000 to $45,000, and one at $200,000.In that case, each salary has equal weight and the median stays $40,000.
Now change the setup. Suppose you have survey data where each response represents a different number of people in a region. The values are 10, 20, 30, and 40. The weights are 5, 3, 2, and 1, respectively. The total weight is 11. Half of 11 is 5.5. Sorted values with their weights: 10 (weight 5), 20 (weight 3), 30 (weight 2), 40 (weight 1). Cumulative weight after the first value is 5, which is below 5.5. Add the second value: cumulative weight is 8, which is above 5.5. The weighted median is 20.
The table below shows the same calculation in a form you can reuse for any weighted dataset.
The Exact-50% Tie Convention
When the cumulative weight lands exactly on half of the total, you have a tie. The standard convention, traced to Edgeworth's 1888 paper in the Philosophical Magazine and cited in Hyndman and Fan's 1996 classification of quantile estimators, is to average the value at that point with the next larger value. This matches the type 2 definition in their table. For an odd number of observations, the median is the value at position (n+1)/2, the type 1 definition. For an even number, the median is the mean of the values at positions n/2 and (n/2)+1.
Why average instead of picking the lower value? Because the median is meant to split the probability mass in half. Averaging the two central values is the only way to ensure that exactly half the weight lies on each side when the tie falls between two distinct values. If you picked the lower value, you would be biased low; if you picked the higher, biased high. The average is the neutral choice.
Some software, particularly older quantile functions, uses a different tie rule and returns the lower value for the 50th percentile. That is why your calculator and your friend's calculator can disagree on the same dataset. The difference is not arithmetic error; it is a different convention. For the weighted median, the tie convention is the same as for the plain median: average the two central weighted values when the cumulative weight lands exactly on half.
Weighted Median vs Weighted Mean
The weighted mean multiplies each value by its weight, sums those products, and divides by the total weight. The weighted median finds the value that splits the cumulative weight in half. For symmetric distributions, they are close or identical. For skewed data, they diverge. A single extreme value, an outlier, pulls the weighted mean toward it but leaves the weighted median untouched. That is the practical argument for the weighted median when your data has a heavy tail.
Consider income data. A few billionaires inflate the weighted mean income, making it look like the average person earns far more than they do. The weighted median income, by contrast, is the income of the household in the middle of the income distribution. It is robust to the outlier. The mean is still useful for totals: total income, total expenditure, total weight. But if you are describing the center of a skewed distribution, the weighted median is the honest summary.
The relationship between mean and median also signals skew. If the weighted mean exceeds the weighted median, the distribution is right-skewed, meaning a long tail of large values. If the weighted mean is below the weighted median, the distribution is left-skewed, with a tail of small values. This is not a mathematical theorem about all datasets; it is a property that holds for most real-world skewed data. The weighted median does not tell you about the shape, only about the middle.
Weighted Median in Excel and Other Spreadsheets
How to Compute It Without a Built-in Function
Excel does not have a dedicated weighted median function. The MEDIAN function, which ignores text and logical values but includes numbers, works only for unweighted data. To compute a weighted median in a spreadsheet, you need to construct the cumulative weight column yourself and then use a lookup or an array formula.
Put your values in column A, your weights in column B. In column C, compute the cumulative weight with a formula that adds the current weight to the sum of all previous weights. In column D, compute the cumulative weight as a percentage of the total weight. Then use INDEX and MATCH to find the first value where that percentage is at or above 0.5. That value is the weighted median. If you need the tie convention, check whether the cumulative weight lands exactly on 0.5; if so, average the value at that row with the one below it.
Quartile Methods and Grouped Data
For quartiles, Excel's QUARTILE.INC uses method R-7, the default in R, while QUARTILE.EXC uses method R-6. These give different quartile values for the same data. The weighted median, being the 50th percentile, is the same under both methods only when the cumulative weight crosses 50% at a single value. In the tie case, the methods can differ. If you are using a spreadsheet for a homework assignment, state which quartile method you used; your instructor will want to know.
For grouped data, the median class is the class where the cumulative frequency reaches or exceeds n/2. The weighted median for grouped data uses the same logic but with weights. The formula is the standard one for grouped data medians.
Weighted Median Formula for Different Data Shapes
The weighted median calculation changes slightly depending on whether you have raw values or a frequency table. For raw values with weights, you sort the pairs by value, compute the cumulative weight, and find the crossing point. For a frequency table, the values are the class midpoints and the weights are the frequencies. The median class is the class where the cumulative frequency reaches or passes half the total. Within that class, you interpolate linearly to estimate the median value.
That interpolation is where most hand calculations go wrong. You cannot simply take the class midpoint and call it the median. You need the lower boundary of the median class, the cumulative frequency of the class before it, the frequency of the median class, and the class width. The formula is: L + ((n/2 - F) / f) * w, where L is the lower boundary, n is the total weight, F is the cumulative weight before the median class, f is the weight of the median class, and w is the class width.
This grouped-data median formula is standard and appears in every introductory statistics textbook. It is an approximation, because it assumes the values within the median class are evenly distributed. If the class has an open end, such as "100 or more," the formula cannot be used without assuming an upper bound. That is a limitation, not a flaw; the median for grouped data is always an estimate.
Common Mistakes and How to Avoid Them
The most common mistake is sorting the weights instead of the values. The weights stay attached to their values; you sort the values and carry the weights along. If you sort the weights independently, you destroy the pairing and the result is meaningless. The second mistake is confusing the weighted median with the weighted mean. The weighted mean is a balance point, the weighted median is a midpoint. They are different statistics and answer different questions.
A third mistake is using QUARTILE.INC when the data has an even count and expecting the same result as QUARTILE.EXC. The methods differ by definition. A fourth mistake is applying the median to nominal data, categories like colors or names that have no order. The median requires order; without order, the concept does not apply. A fifth mistake is using the median to describe a bimodal distribution without noting the two peaks. The median may sit in a low-frequency gap between the modes, giving a misleading impression of the center.
In a spreadsheet, the most common failure is copying a formula that references the wrong range. Shifting rows or columns changes the result silently. Always double-check the range in the cumulative weight column before trusting the output. For hand calculations, write the cumulative weights in a separate column and check your arithmetic by confirming the final cumulative weight equals the total weight.