The Number in the Middle Is Lying to You
The average hides where outcomes actually concentrate. Learn why outliers drive disproportionate results, and what to do about it in your analysis.
Why Outliers Matter More Than Averages
Most analytical frameworks are built around the average. Managers track it, dashboards display it, and quarterly reviews live and die by it. The average is comfortable. It fits on a slide. It tells a coherent story. The problem is that in most domains that actually matter, the average is nearly useless as a guide to what is happening and almost completely useless as a guide to what will happen next.
This post is not an argument against statistics. It is an argument about which statistics deserve your attention, and why the ones that seem inconvenient are often the ones doing the most work.
Key Takeaways
- In most real-world distributions, outcomes are not evenly spread. A small number of extreme cases account for a disproportionate share of total results.
- According to research by Hendrik Bessembinder at Arizona State University, the best-performing 4% of listed U.S. companies explain the entire net gain of the U.S. stock market since 1926. The other 96% collectively matched Treasury bills.
- Removing outliers from your data to make analysis "cleaner" often removes the most important signal in that data.
- Averages describe the middle of a distribution. Outliers describe where value is actually created, destroyed, or transformed.
- The operational shift worth making is not to stop measuring averages, but to stop treating them as the destination and start treating them as the starting point.
The Classic Idea: Averages Are How We Make Sense of Complexity
The mean is genuinely useful. That is worth saying plainly, because the rest of this post is going to argue against over-relying on it.
If you are managing a supply chain and need to estimate typical delivery time, the average gets you close enough. If you are teaching a statistics class and want to describe a population's height, the mean plus a standard deviation tells you most of what you need. The average exists because the world produces more data than any human can hold in their head, and central tendency is the compression algorithm we invented to deal with that problem.
In statistics, the mean is described as a measure of central tendency, alongside the median and the mode. Each one tries to answer the same question: what is the typical case? The assumption embedded in that question is that the typical case is the one that matters most.
For a narrow set of problems, that assumption holds. Most people's height, most coin flips, most manufacturing tolerances, these follow something close to a normal distribution. The center is genuinely representative. The outliers are genuinely noise.
But the domains where the average is actually reliable are much narrower than most analysts assume.
The number in the middle is a compression.
An average can be mathematically correct and strategically false. It tells you where the center lands after the extremes have already done the important work.
The middle-value machine
Change the distribution and amplify its tail. Watch the mean stop describing almost everyone in the dataset.
In a bounded, roughly normal distribution, the center describes the population reasonably well.
The average changes meaning by domain
In bounded systems the center can be a useful guide. In multiplicative systems the tail often determines the outcome.
Center is useful
Natural limits keep observations clustered. The average resembles an actual person.
Tail creates wealth
A tiny fraction of firms accounts for essentially all long-horizon net gains.
Top tier dominates
The average customer may describe neither the high-value segment nor everyone else.
One win carries the fund
A single extreme return can exceed the rest of the portfolio combined.
Edges expose needs
Designing for an extreme user can produce a better system for everyone.
Where the outcome actually lives
These are not average-driven systems. Their aggregate result is concentrated in a small minority.
The remaining 96% collectively matched one-month Treasury bills.
A fraction of more than 64,000 firms accounted for all net global wealth creation.
Stop cleaning the signal out
The goal is not to worship every anomaly. It is to investigate unusual observations before compressing or deleting them.
Flag
Keep the unusual point visible before deciding what it means.
Verify
Trace it to a real event, behavior, error, or measurement failure.
Segment
Separate populations whose behavior makes one shared mean misleading.
Study
Treat top performers and extreme users as research questions.
Then summarize
Use the average as context after understanding the distribution.
What Everyone Gets Wrong: Most Interesting Distributions Are Not Normal
Here is where the reasoning tends to collapse.
The normal distribution (the famous bell curve) works beautifully when the thing being measured has a natural ceiling and floor, and when no single observation can dramatically shift the whole population. Human height qualifies. Stock market returns do not. Revenue per customer does not. Employee performance does not. Scientific discoveries do not.
In all of those cases, the distribution is skewed. Often severely. A small number of extreme outcomes pull the total so far from the middle that the average becomes a kind of fiction, technically correct but practically misleading.
The clearest evidence of this comes from research by Hendrik Bessembinder, a finance professor at Arizona State University's W. P. Carey School of Business, who evaluated lifetime returns for every U.S. common stock traded since 1926. His findings are worth sitting with for a moment: the best-performing 4% of listed companies account for the entire net gain of the U.S. stock market over that period. The remaining 96% of stocks, collectively, performed no better than one-month Treasury bills.
Read that again. Not underperformed. Not underperformed by a little. Matched Treasury bills. The bottom 96% of the market, in aggregate, generated zero excess return.
Globally, the concentration is even sharper. Bessembinder's subsequent research covering more than 64,000 global stocks found that the top 2.4% of firms account for all $75.7 trillion in net global stock market wealth creation from 1991 to 2020.
```htmlAverages describe the field. Outliers create the result.
Across markets, customers, investments, and design, a small number of extreme cases can determine the outcome of the whole system.
| Domain | What the Average Suggests | What the Outliers Reveal |
|---|---|---|
| U.S. Stock Market | Most stocks contribute to long-term returns | Top 4% of stocks drive all net gains since 1926 (Bessembinder, ASU) |
| Global Stock Market | Equity markets broadly create wealth | Top 2.4% of firms account for all $75.7T in global wealth creation, 1991–2020 |
| Customer Revenue | Revenue is distributed across the customer base | Roughly 20% of customers typically generate 80% of revenue (Pareto Principle) |
| Venture Capital | Diversified portfolios spread and reduce risk | One or two outlier investments routinely drive the returns of an entire fund |
| Product Design | Build for the typical user’s needs | Designing for extreme users often creates products that work better for everyone |
The stock market example is not an anomaly. It is a pattern.
Venture capital operates the same way. A single investment in a company that becomes dominant can return more than the rest of a portfolio combined. The math only works if you accept, upfront, that most investments will not perform and that you are looking for the one that will. The average expected return of a VC portfolio is almost irrelevant to how you should think about constructing one.
Customer revenue follows the same logic. The Pareto Principle, the observation that roughly 80% of outcomes derive from 20% of inputs, keeps showing up across industries because it reflects something real about how skewed distributions behave. Your top customers are not just "better" versions of your average customers. They may be operating in a completely different mode.
This does not mean the average is worthless. It means that if the average is your only lens, you are looking at the center of a distribution while most of the action is happening at the edges.
What Changed: Data Is Now Revealing the Edges More Clearly
The argument for paying attention to outliers is not new. What has changed is how visible those outliers have become, and how much it costs to ignore them.
As organizations collect more granular behavioral data, they are increasingly able to see patterns that used to get washed out in aggregate reporting. A spike in failed logins is not random noise. A single customer logging 300 interactions in a day is not a data error. A production dip from one machine is not irrelevant. Each of these is an outlier relative to the average, and each of them can be the earliest visible signal of something significant: fraud, a power user, equipment failure, a new market.
For decades, the standard analytical response to these signals was to clean them out. Outliers skew the mean. They stretch the y-axis. They make the dashboard look messier than the leadership team wants it to look. So they were removed, smoothed, or averaged away.
The problem is that removing them also removes the information. Sigma Computing, an analytics platform, put it directly: "If you're always building dashboards that flatten out these moments, you're only seeing the middle of the distribution, not the events that pull things forward or throw them off course."
Machine learning has complicated this further. Anomaly detection has become a first-class discipline precisely because models trained only on average behavior are often blind to the cases where something meaningful is happening. Fraud detection does not work if you train a model on normal transactions. Predictive maintenance does not work if you normalize away the unusual readings. The outlier is the signal.
There is also a design argument worth including here. In the early 2000s, a team designed a self-stabilizing spoon for people with Parkinson's disease, a condition affecting roughly 0.9% of the global population. From a purely statistical perspective, designing for that population made no sense. The mean user does not have Parkinson's.
But the product, designed for an extreme edge case, turned out to be useful for a much broader population. Designing for the outlier created something better for the average user, not despite the extreme constraint but because of it. This pattern repeats in product development often enough that it has its own name: inclusive design.
What This Means Operationally: Stop Cleaning the Signal Out of Your Data
None of this is an argument for treating every anomaly as meaningful. Some outliers are genuinely noise. A data entry error is not a business insight. A one-time spike caused by a system glitch is not a trend. The goal is not to worship every unusual data point but to stop automatically discarding them before you understand what they are.
The operational shift looks like this:
Flag before you delete. When something falls outside the expected range, mark it and keep it rather than removing it from the dataset. You can always decide it is noise later. You cannot recover it after deletion.
Segment your analysis. If the top 20% of your customers behave differently from the other 80%, analyzing them together produces a mean that describes neither group accurately. Separate the distributions. Ask what is true for each segment on its own terms.
Look at the shape, not just the center. Standard deviation tells you something about variability, but skewness tells you something about direction. A right-skewed distribution means most observations are below the mean and a few are dramatically above it. That shape has different strategic implications than a normal distribution with the same mean.
Treat your best performers as a research project. Whether those are customers, employees, products, or investment positions, the outliers at the top of your distribution are telling you something the average cannot. What are they doing differently? What conditions produced them? Can those conditions be replicated?
Track what you almost removed. Some of the most valuable analyses come from re-examining data points that were nearly discarded. A handful of early adopters who used a product in unexpected ways, a small region that outperformed every model's prediction, a single employee who processed twice the volume of anyone else. These are not statistical inconveniences. They are case studies.
This does not mean abandoning averages. Averages are a starting point. They tell you where the middle of the distribution is, which is useful context for identifying what counts as an outlier in the first place. The mistake is treating the middle as the destination rather than as a reference point.
There is something uncomfortable in all of this. If a small fraction of stocks account for essentially all stock market wealth creation since 1926, then most of what the market does on any given day is, at a very long horizon, irrelevant. If a small fraction of customers drive most of your revenue, then most of your acquisition activity is, at best, supporting infrastructure for the outliers you are trying to find and keep.
That is a strange thing to accept. It suggests that a large portion of analytical effort is spent describing a middle that does not ultimately drive outcomes, while the actual value concentrates in the tails that are often most difficult to predict, identify, and retain.
Not every outlier perfectly supports this thesis, and that is the honest part. Some extreme observations are flukes. Some power users churn. Some top-performing stocks eventually revert. The tail distributions that make outliers so powerful also make them unpredictable. You cannot assume every anomaly is a future winner any more than you should assume every anomaly is a data error.
What you can do is build analytical systems that keep outliers visible rather than compressing them into an average, and make sure the decisions you are treating as settled are not simply an artifact of looking at the wrong part of the distribution.
Stop Treating the Average as the Answer
The average is a compression. A useful compression, but a compression nonetheless. Most of what it discards, especially in skewed distributions, is the information that determines where outcomes actually land.
The practical version of this is not complicated. Look at distributions, not just means. Segment before you summarize. Flag anomalies before deleting them. Treat top performers as the research question, not the baseline. And be honest about how much of your current analytical infrastructure is built around describing the center of a distribution that does most of its important work at the edges.
Bessembinder's research on stock market returns is striking not because it is counterintuitive, but because it is so thoroughly documented. Ninety-plus years of data, tens of thousands of stocks, and the result is the same: the net gain comes from a fraction of the population that the average systematically obscures.
The market is not unusual in this regard. It is just unusually well-documented.
Frequently Asked Questions
What is the difference between an average and an outlier?
An average (or mean) is a single number representing the center of a dataset, calculated by summing all values and dividing by the count of observations. An outlier is a data point that sits far from the rest of the distribution. The two concepts are related because outliers can significantly distort the average, which is one reason averages are often unreliable in skewed distributions.
Why do outliers matter more than averages in financial markets?
Research by Hendrik Bessembinder at Arizona State University found that the top 4% of U.S. listed companies account for the entire net gain of the U.S. stock market since 1926. The remaining 96% of stocks, collectively, generated returns equivalent to Treasury bills. This means the average stock performance is not representative of where value is actually created, and investment strategies built around average expectations miss where the return concentration actually sits.
Does the Pareto Principle (80/20 rule) apply to outliers?
The Pareto Principle is one expression of skewed distributions, the observation that roughly 80% of outcomes derive from 20% of inputs. Outliers are the extreme version of this pattern: the top performers in a distribution that do not just outperform by a little but by an order of magnitude. The 80/20 relationship is more moderate than the concentration seen in stock market wealth creation (where 4% drives 100% of net gains), but both reflect the same underlying phenomenon.
Should businesses ever remove outliers from their data?
Yes, but with caution. Outliers caused by genuine data entry errors, system glitches, or measurement failures can distort analysis without providing meaningful signal. The practical rule is to investigate before deleting. If the outlier reflects real behavior by a real user or system, removing it also removes information. If it is a technical artifact, it can be excluded. The default should be to flag and retain, not to filter first.
How do you identify whether an outlier is a signal or noise?
There is no single test. Useful questions include: Can the observation be traced to a real event or behavior? Does it appear in multiple data sources consistently? Does it recur across different time periods? Is it isolated to one channel or variable? If an anomaly appears repeatedly, survives cross-referencing, and corresponds to a plausible real-world event, it is more likely signal than noise. Visualization tools, anomaly detection algorithms, and segmentation analysis all help distinguish the two.
What is a power law distribution and how does it relate to outliers?
A power law distribution is one where a small number of observations are dramatically larger than the rest, and there is no strong central tendency pulling values toward a typical case. Wealth, market returns, and city population sizes all follow approximate power laws. In these distributions, the average is heavily influenced by the extreme values and describes almost no individual observation accurately. Outlier thinking is particularly important in power law domains because the tail of the distribution determines the aggregate outcome.
How should teams change their reporting to account for outliers?
Start by adding distribution views (histograms, box plots, percentile breakdowns) alongside summary statistics. Segment key metrics by performance tier rather than reporting a single mean. Flag observations that exceed a defined threshold for separate review rather than removing them before analysis reaches decision-makers. Over time, build the habit of asking "what is true for our top performers" as a distinct question from "what is true on average."
An independent voice that will raise an eyebrow.
The Off Label is marketing strategy in action. We go further than what's on the surface. Every play, brief, strategy, and trend published here is proof of how we connect dots and turn ideas into an advantage.
Browse Full Foundations Archive →Published from the Charleston, South Carolina strategy lab. Synthesizing marketing behavior into actionable strategy for New York City and the world's creative hubs.
© 2026 The Off Label. All rights reserved. Content on this site may not be reproduced without prior permission.
NYC / LDN / CDMX / CHS
