AboutHow we built thisSponsorshipShop
SearchSubscribeDecision ToolsBusiness ModelsFrameworksReading Lists
Privacy PolicyTerms of UseCookie PolicyRefund PolicyAccessibilityDisclaimer

© 2026 Faster Than Normal. All rights reserved.

Faster Than Normal
DecisionsPeopleBusinessesNewsletterSubscribe
Start reading →
  1. Home
  2. Mental models
  3. Data Dredging
Mathematics & Probability

Data Dredging

Model #0777Category: Mathematics & ProbabilityDepth to apply:
4 min read

On this page

  • Core Idea
  • How to See It
  • How to Use It
  • Founders & Leaders
  • Connected Models
  • One Key Quote
  • Summary & Further Reading

Contents

  1. 1. Core Idea
  2. 2. How to See It
  3. 3. How to Use It
  4. 4. Founders & Leaders
  5. 5. Connected Models
  6. 6. One Key Quote
  7. 7. Summary & Further Reading
·Mathematics & Probability
Section 1

Core Idea

Data dredging — also called data fishing or data snooping — is the practice of searching through large datasets for statistically significant patterns without a prior hypothesis. When you test enough variables against each other, some will correlate by pure chance. At a 5% significance level, one in twenty random comparisons will appear significant. The problem isn't looking at data; it's mistaking patterns found after the fact for genuine causal relationships. Data dredging produces "discoveries" that don't replicate because they were never real — they were artefacts of multiple testing on noisy data. The antidote is pre-registered hypotheses and out-of-sample validation.
Section 2

How to See It

Marketing
You're seeing it when a marketing team runs dozens of A/B test variants, finds one with p < 0.05, and declares victory — without adjusting for the fact that testing 20 variants makes one false positive likely.
Research
You're seeing it when a study claims a surprising result — like a specific food reducing cancer risk — but the researchers tested 200 food items and only reported the one that crossed the significance threshold.
Business
You're seeing it when an analyst slices revenue data by 30 different demographic variables, finds one segment with unusual growth, and presents it as a strategic insight without acknowledging the multiple comparisons problem.
Section 3

How to Use It

Before analyzing data, state your hypothesis. If you're exploring, label the results as exploratory — not confirmatory. When multiple comparisons are tested, apply statistical corrections (Bonferroni, false discovery rate) or validate findings on a holdout dataset. Be especially skeptical of surprising findings that emerged from undirected analysis — the more surprising the result, the more likely it's a false positive from dredging.
Decision filter
"Did I predict this finding before looking at the data — or did I find it by searching the data until something appeared significant?"
As a founder
When your data team surfaces an unexpected insight — a segment outperforming, a feature correlating with retention — ask whether the hypothesis was stated before the analysis. If not, validate on a separate sample before investing resources. The cost of acting on a dredged finding is a misallocated sprint or a failed campaign.
Section 5

Founders & Leaders

Richard FeynmanNobel Prize-winning physicist; educator
Feynman's famous "Cargo Cult Science" lecture warned against the tendency to find patterns that aren't real and mistake them for discoveries. He insisted on a principle he called "bending over backwards" — actively trying to disprove your own findings rather than confirming them. This intellectual honesty is the direct antidote to data dredging. For founders awash in metrics, Feynman's lesson is to distrust the exciting finding from undirected analysis. State your hypothesis first, test it cleanly, and try to prove yourself wrong before betting the company on a number.
Section 7

Connected Models

Pairs-with
P-hacking
P-hacking manipulates analysis until a significant result appears. Data dredging is the broader practice — searching through data without a hypothesis. P-hacking is the specific technique of tweaking methods to manufacture significance.
Exploits
Confirmation Bias
Confirmation bias makes us seek data that supports what we already believe. Data dredging provides the raw material — with enough variables, confirmation bias will always find something to latch onto.
Creates
Correlation vs Causation
Correlation vs causation is the error of assuming that co-occurrence implies cause. Data dredging generates spurious correlations at scale — each one tempts the analyst to infer a causal relationship that doesn't exist.
Section 8

One Key Quote

"The first principle is that you must not fool yourself — and you are the easiest person to fool."
— Richard Feynman
Section 11

Summary & Further Reading

Data dredging finds "significant" patterns in noisy data by testing enough variables without a prior hypothesis. The antidote is pre-registered hypotheses, multiple-testing corrections, and out-of-sample validation.
01
Surely You're Joking, Mr. Feynman! — Richard Feynman (1985)
Book
On intellectual honesty, avoiding self-deception, and the discipline of rigorous scientific thinking.
02
Thinking, Fast and Slow — Daniel Kahneman (2011)
Book
On the cognitive biases that make humans prone to seeing patterns in noise and trusting them as real.
03
Statistics Done Wrong — Alex Reinhart (2015)
Book
On the most common statistical errors in research and practice, including data dredging and multiple comparisons problems.

Why this matters next

mental modelsConfirmation Bias

Data Dredging applied the Confirmation Bias mental model

mental modelsFalse Positives & False Negatives

Data Dredging applied the False Positives & False Negatives mental model

mental modelsScale

Data Dredging applied the Scale mental model

mental modelsCost

Data Dredging applied the Cost mental model

mental modelsP-hacking

Data Dredging applied the P-hacking mental model

mental modelsCorrelation vs Causation

Data Dredging applied the Correlation vs Causation mental model

Frequently asked questions

What is Data Dredging?+

Data Dredging is a mental model used for better thinking and decision-making.

How do you apply Data Dredging?+

To apply Data Dredging, identify situations where this framework is relevant, then use it as a lens to evaluate your options and decisions. The model is most useful when combined with other complementary mental models.

What category does Data Dredging fall under?+

Data Dredging falls under the Mathematics & Probability category of mental models. Other models in this category can be found on the Mathematics & Probability hub page.

Why is Data Dredging important?+

Data Dredging is important because it provides a structured way to think about problems that would otherwise be approached with intuition alone. Understanding this model helps you avoid common reasoning errors and make better decisions.

Continue exploring

BT

Mental model

Bayes Theorem

A mathematical framework for updating beliefs based on new evidence, proportiona

BT

Mental model

Black Swan Theory

Rare, unpredictable events with extreme impact that are retrospectively rational

CO

Mental model

Compounding

Small consistent gains accumulate exponentially over time — the most powerful fo

CC

Mental model

Correlation vs Causation

The critical distinction between two variables that move together and one actual

ER

Mental model

Ergodicity

The distinction between ensemble averages and time averages — what works across

EG

Mental model

Exponential Growth

Growth that accelerates proportionally to its current size, producing deceptivel

More like this, in your inbox

I send a newsletter every week — free, no spam, unsubscribe anytime.

Or open the full subscribe page.

On this page

  • Core Idea
  • How to See It
  • How to Use It
  • Founders & Leaders
  • Connected Models
  • One Key Quote
  • Summary & Further Reading

Popular Mental Models

First Principles ThinkingOccam's RazorCircle of CompetenceInversionConfirmation BiasSecond-Order ThinkingDunning-Kruger EffectSurvivorship BiasPareto PrincipleOpportunity Cost