Skip to content
Mathematics & Probability

Data Dredging

Model #0777Category: Mathematics & ProbabilityDepth to apply:

By Updated 3 sources

4 min read
Mathematics & Probability
Section 1

Core Idea

Data dredging — also called data fishing or data snooping — is the practice of searching through large datasets for statistically significant patterns without a prior hypothesis. When you test enough variables against each other, some will correlate by pure chance. At a 5% significance level, one in twenty random comparisons will appear significant. The problem isn't looking at data; it's mistaking patterns found after the fact for genuine causal relationships. Data dredging produces "discoveries" that don't replicate because they were never real — they were artefacts of multiple testing on noisy data. The antidote is pre-registered hypotheses and out-of-sample validation.

Get Faster Than Normal by email

Ideas from founders and companies.

Free newsletter. Unsubscribe anytime.

Or open the full subscribe page.

Section 2

How to See It

Marketing
You're seeing it when a marketing team runs dozens of A/B test variants, finds one with p < 0.05, and declares victory — without adjusting for the fact that testing 20 variants makes one false positive likely.
Research
You're seeing it when a study claims a surprising result — like a specific food reducing cancer risk — but the researchers tested 200 food items and only reported the one that crossed the significance threshold.
Business
You're seeing it when an analyst slices revenue data by 30 different demographic variables, finds one segment with unusual growth, and presents it as a strategic insight without acknowledging the multiple comparisons problem.
Section 3

How to Use It

Before analyzing data, state your hypothesis. If you're exploring, label the results as exploratory — not confirmatory. When multiple comparisons are tested, apply statistical corrections (Bonferroni, false discovery rate) or validate findings on a holdout dataset. Be especially skeptical of surprising findings that emerged from undirected analysis — the more surprising the result, the more likely it's a false positive from dredging.
Decision filter
"Did I predict this finding before looking at the data — or did I find it by searching the data until something appeared significant?"
As a founder
When your data team surfaces an unexpected insight — a segment outperforming, a feature correlating with retention — ask whether the hypothesis was stated before the analysis. If not, validate on a separate sample before investing resources. The cost of acting on a dredged finding is a misallocated sprint or a failed campaign.
Section 5

Founders & Leaders

Richard FeynmanNobel Prize-winning physicist; educator
Feynman's famous "Cargo Cult Science" lecture warned against the tendency to find patterns that aren't real and mistake them for discoveries. He insisted on a principle he called "bending over backwards" — actively trying to disprove your own findings rather than confirming them. This intellectual honesty is the direct antidote to data dredging. For founders awash in metrics, Feynman's lesson is to distrust the exciting finding from undirected analysis. State your hypothesis first, test it cleanly, and try to prove yourself wrong before betting the company on a number.
Section 7

Connected Models

Pairs-with
P-hacking
P-hacking manipulates analysis until a significant result appears. Data dredging is the broader practice — searching through data without a hypothesis. P-hacking is the specific technique of tweaking methods to manufacture significance.
Exploits
Confirmation Bias
Confirmation bias makes us seek data that supports what we already believe. Data dredging provides the raw material — with enough variables, confirmation bias will always find something to latch onto.
Creates
Correlation vs Causation
Correlation vs causation is the error of assuming that co-occurrence implies cause. Data dredging generates spurious correlations at scale — each one tempts the analyst to infer a causal relationship that doesn't exist.
Section 8

One Key Quote

"The first principle is that you must not fool yourself — and you are the easiest person to fool."
Richard Feynman
Section 11

Summary & Further Reading

Data dredging finds "significant" patterns in noisy data by testing enough variables without a prior hypothesis. The antidote is pre-registered hypotheses, multiple-testing corrections, and out-of-sample validation.
01
Book
On intellectual honesty, avoiding self-deception, and the discipline of rigorous scientific thinking.
02
Book
On the cognitive biases that make humans prone to seeing patterns in noise and trusting them as real.
03
Book
On the most common statistical errors in research and practice, including data dredging and multiple comparisons problems.

Why this matters next

Frequently asked questions

What is Data Dredging?

Data Dredging is a mental model used for better thinking and decision-making.

How do you apply Data Dredging?

To apply Data Dredging, identify situations where this framework is relevant, then use it as a lens to evaluate your options and decisions. The model is most useful when combined with other complementary mental models.

What category does Data Dredging fall under?

Data Dredging falls under the Mathematics & Probability category of mental models. Other models in this category can be found on the Mathematics & Probability hub page.

Why is Data Dredging important?

Data Dredging is important because it provides a structured way to think about problems that would otherwise be approached with intuition alone. Understanding this model helps you avoid common reasoning errors and make better decisions.

Continue exploring

Get Faster Than Normal by email

Ideas from founders and companies.

Free newsletter. Unsubscribe anytime.

Or open the full subscribe page.

Popular Mental Models