AboutHow we built thisSponsorshipShop
SearchSubscribeDecision ToolsBusiness ModelsFrameworksReading Lists
Privacy PolicyTerms of UseCookie PolicyRefund PolicyAccessibilityDisclaimer

© 2026 Faster Than Normal. All rights reserved.

Faster Than Normal
DecisionsPeopleBusinessesNewsletterSubscribe
Start reading →
  1. Home
  2. Mental models
  3. Field Testing
Business & Strategy

Field Testing

Model #0198Category: Business & StrategySource: Drew HoustonDepth to apply:
37 min read

On this page

  • The Core Idea
  • How to See It
  • How to Use It
  • The Mechanism
  • Founders & Leaders in Action
  • Visual Explanation
  • Connected Models
  • One Key Quote
  • Analyst's Take
  • Test Yourself
  • Top Resources

Contents

  1. 1. The Core Idea
  2. 2. How to See It
  3. 3. How to Use It
  4. 4. The Mechanism
  5. 5. Founders & Leaders in Action
  6. 6. Visual Explanation
  7. 7. Connected Models
  8. 8. One Key Quote
  9. 9. Analyst's Take
  10. 10. Test Yourself
  11. 11. Top Resources
·Business & Strategy
Section 1

The Core Idea

Focus groups loved New Coke. Two hundred thousand blind taste tests confirmed that consumers preferred the new formula to both old Coke and Pepsi. Coca-Cola launched with total confidence in April 1985. Within seventy-nine days, they brought back the original formula and absorbed one of the most expensive product failures in corporate history. The taste tests were not wrong — in a controlled setting, people genuinely preferred the sweeter formula. The taste tests were irrelevant. They measured preference in a vacuum. They could not measure brand attachment, nostalgic identity, the psychological ownership people felt over "their" Coke, or the revolt that erupts when you change something woven into cultural fabric. The lab said yes. The field said no. The field was right. Field testing — validating your product or idea with real users in real conditions, not in a laboratory or conference room — is the discipline of closing the gap between what people say they will do and what they actually do. That gap has killed more products than bad engineering ever will.
Drew Houston understood the gap in 2007. He had built the earliest version of Dropbox but could not convince investors that consumers wanted cloud file storage — the concept felt abstract, the market already had competitors, and the problem seemed solved by USB drives. So Houston made a three-minute demonstration video showing exactly how Dropbox worked: drag a file into a folder, watch it appear on another computer. He did not commission a survey. He did not convene a focus group. He posted the video on Hacker News — a community dense with his target users — and measured whether people would act. Within twenty-four hours, 75,000 people had signed up for the beta waitlist. The video was a field test: a real product concept, shown to real potential users, in the real environment where they made decisions, measured by the only metric that matters — whether they moved. No amount of market research could have generated that signal. The signal required real people making a real commitment — even one as small as entering an email address — in the context where they actually lived.
Nick Swinmurn ran the same logic with shoes. In 1999, he wanted to know whether people would buy footwear online — a proposition that every retail expert and every focus group would have dismissed. People need to try on shoes. They need to feel the material. They need to check the fit. Swinmurn ignored the objections and ran the test. He walked into shoe stores around San Francisco, photographed the inventory, posted the photos on a basic website, and waited. When someone placed an order, he went back to the store, bought the shoes at full retail price, and shipped them himself. Zero inventory. Zero warehouse. Zero risk beyond his time. The field test measured one variable: would real people spend real money on shoes they had not tried on? They did. The company Swinmurn built on that signal was Zappos. Amazon acquired it for $1.2 billion in 2009.
The power of field testing is ecological validity — a concept from experimental psychology that describes how well findings from one setting predict behaviour in another. Laboratory experiments maximise internal validity by controlling variables. Field tests sacrifice some control for massive gains in ecological validity — the fidelity of the signal to real-world conditions. For business decisions, this trade-off is almost always worth making. The variables that determine whether a product succeeds — willingness to pay, switching costs, competitive alternatives, contextual friction, the chasm between stated preference and revealed behaviour — are precisely the variables that disappear inside a conference room. A focus group can tell you whether people find your concept interesting. Only a field test can tell you whether they will reach for their wallet.
Section 2

How to See It

Field testing shows up wherever someone chose to measure real behaviour instead of collecting opinions. The signal is action — purchases, sign-ups, retention, repeat usage — not sentiment. The most reliable field tests measure behaviour that costs the user something: money, time, reputation, or effort. Free opinions are cheap to give. Expensive actions are honest signals.
The cost hierarchy matters: a credit card transaction is the strongest signal. A time investment (installing software, creating an account) is the next strongest. An email address is the weakest actionable signal — but still orders of magnitude more reliable than a survey response, because even entering an email requires the user to evaluate the proposition and decide it is worth the inevitable follow-up emails.
You're seeing field testing when a founder sells the product before it fully exists — taking pre-orders, posting a demo video, running a concierge service — and uses the behavioural response to decide whether to build.
Startups
You're seeing field testing when a founder validates demand with zero infrastructure. Pebble raised $10.3 million on Kickstarter in 2012 before manufacturing a single watch. The campaign was not a fundraiser. It was a field test: would real people commit real money to a smartwatch? Ten million dollars answered the question with a clarity that no survey could match. Buffer tested demand with a landing page that described pricing tiers for a product that did not exist — and measured how many visitors clicked "Sign Up." The visitors were not in a focus group. They were on the internet, distracted, with a hundred competing demands on their attention. The ones who clicked were the real signal.
Product
You're seeing field testing when a product team ships a feature to 1% of users and measures behaviour before rolling it out to everyone. Netflix tests thumbnail images, recommendation algorithms, and UI layouts on small user cohorts before scaling changes across 230 million subscribers. Each test is a field test — real users, real content, real viewing decisions, measured in minutes watched rather than survey responses about which layout they prefer. The distinction matters: asking users "which thumbnail do you like better?" produces one answer. Measuring which thumbnail they actually click on produces a different — and more profitable — answer.
Growth
You're seeing field testing when a startup validates demand by fulfilling orders manually before building automation. DoorDash's founders delivered food themselves in Palo Alto before building a logistics platform. They were not building a delivery business. They were running a field test: would restaurant customers pay for delivery, and could unit economics work at a local level? The answer came from real orders they fulfilled with their own cars — not from the market sizing deck they could have built instead.
Investing
You're seeing field testing when an investor asks a founder "how many paying customers do you have?" instead of "how big is the market?" The first question demands a field test result. The second accepts a research output. The gap in predictive value between the two is the gap between what people do and what analysts project they might do. Revenue from paying customers is field-tested truth. A TAM slide is a hypothesis dressed in a suit.
Section 3

How to Use It

Field testing is the fastest path from hypothesis to truth. The discipline is not in the execution — any founder can post a landing page or sell a product manually. The discipline is in the willingness to let real behaviour override your beliefs. The field test that contradicts your thesis is more valuable than the one that confirms it, because it saves you from building something nobody will buy.
Decision filter
"Before committing resources to build, ask: have we tested whether real people will pay real money under real conditions? If we have only measured interest, intention, or enthusiasm, we have measured nothing."
As a founder
Your first field test should cost almost nothing and reveal almost everything. The Zappos method — photograph existing products, post them online, fulfil manually — works for any category where you can simulate the buying experience without building infrastructure. The Dropbox method — demonstrate the product concept and measure sign-up intent — works for categories where the product cannot yet exist. The DoorDash method — personally deliver the service and measure unit economics — works for marketplace and logistics businesses. The common thread: strip every variable except the one you are testing. Swinmurn did not test logistics, warehouse design, or customer service protocols. He tested whether people would buy shoes online. Period.
The second discipline is speed. A field test should take days or weeks, not months. If your field test requires three months of preparation, you are building, not testing. Shorten the loop. The founders who extract the most value from field testing run many small tests rather than one elaborate one — because each test either validates a hypothesis or kills it, and both outcomes are progress.
As an investor
The strongest signal in an early-stage pitch is field test evidence. Revenue from real customers is a field test. A waitlist of email addresses is not — email addresses are free to give and the conversion rate from stated intent to actual purchase is routinely below 10%. When evaluating a pre-product startup, ask what the founder has done to test demand in real conditions. The founder who says "we've talked to fifty potential customers and they're all excited" has collected stated preference data. The founder who says "we've sold twelve units at $200 each to people we found on LinkedIn" has run a field test. The difference in predictive value is the difference between opinion and evidence.
Look also for founders who let field test results reshape their plans. Hastings killed Netflix's per-rental model because the field test — real subscriber behaviour — showed subscriptions produced better retention. The willingness to follow the data, even when it contradicts the original thesis, is the investor's strongest quality signal.
As a decision-maker
Use field tests to resolve internal debates that data cannot settle and politics will not survive. When the product team believes customers want feature A and the sales team believes they want feature B, do not commission a survey. Build the cheapest possible version of both, ship them to a small cohort, and measure which one users actually engage with. The field test converts a political negotiation — who has more influence in the room? — into an empirical question — which feature produces the behaviour we want? Google famously tested forty-one shades of blue for a link colour and measured click-through rates. The field test settled a debate that opinion never would have.
The second application: use field tests to de-risk major investments. Before committing to a new market, new product line, or major feature, run a field test that simulates the core value proposition at minimal cost. Amazon tested same-day delivery in a single zip code before rolling it out nationally. The field test cost a fraction of the infrastructure investment and revealed operational constraints that the planning models had missed entirely.
Common misapplication: Treating positive feedback as a field test result. "Everyone we talked to said they love the concept" is not field testing. It is social desirability bias in action — people telling you what you want to hear because disagreeing feels socially costly. Field testing requires a behavioural commitment: money spent, time invested, usage observed. Enthusiasm without action is noise.
Second misapplication: Over-building before testing. A founder who spends six months building a polished product and then "tests" it with beta users has not run a field test — they have run a product launch with a euphemistic name. Zappos tested demand with zero inventory. Dropbox tested demand with a video. Buffer tested demand with a landing page for a product that did not exist. The less you build before testing, the more you learn per dollar spent.
Third misapplication: Testing the wrong variable. A field test that measures whether people like your product is nearly useless. A field test that measures whether people will pay for it, return to use it, or recommend it to others is invaluable. The variable must connect to the economic engine of the business. Zappos did not test whether people enjoyed browsing shoes online. It tested whether they would complete a purchase. Variable selection determines whether the test produces actionable intelligence or interesting trivia.
Section 4

The Mechanism

Section 5

Founders & Leaders in Action

The founders below did not arrive at field testing through methodology or frameworks. They arrived at it because they had no alternative. Neither had the budget for market research, elaborate prototypes, or controlled testing environments. Both were forced into the only laboratory available — the real world — and both discovered that the constraints produced a signal quality that no amount of money could have purchased.
What connects them is a feedback loop between making and measuring that operated at the speed of weeks, not quarters. Knight tested shoes at track meets and incorporated runner feedback into the next batch. Hastings tested business models with real subscribers and killed the ones that did not retain. Both treated every customer interaction as data, every transaction as an experiment, and every failure as a signal rather than a verdict.
Phil KnightCo-founder, Nike (then Blue Ribbon Sports), 1964–2004
Knight's company was born at the interface between a track and a garage. His partner Bill Bowerman was the legendary track coach at the University of Oregon, and Bowerman's obsession was lighter, faster running shoes. He experimented with designs in his garage workshop and tested them immediately on real runners at real track meets — not prototypes evaluated by a focus group but shoes worn by competitive athletes on actual tracks, in actual races, under actual competitive pressure. The waffle trainer — the shoe that made Nike — was born when Bowerman poured liquid urethane into his wife's waffle iron to create a lighter sole with better traction. He tested the resulting shoe on Oregon runners the following week. Knight, meanwhile, sold shoes out of the back of his Plymouth Valiant at those same meets, watching which designs runners reached for, which they returned, and which they ran personal bests in. Every transaction was a field test. Every race was a data point.
The feedback loop between Bowerman's garage and the track was Nike's first competitive advantage. A shoe that failed on the track was redesigned within days, not months. A shoe that produced personal bests was immediately moved into production. By the time established manufacturers like Adidas recognised Knight's threat, Nike had run more field iterations in five years than its competitors had run in twenty. Nike did not commission a consumer research study until it was already a national brand — because the field had already told Knight everything he needed to know.
Reed HastingsCo-founder & CEO, Netflix, 1997–2023
Netflix's founding story is a sequence of field tests, each smaller and cheaper than the decision it informed. In 1997, Hastings and Marc Randolph wanted to know whether DVDs could survive being mailed through the US Postal Service. So Hastings mailed a CD to his own house in Santa Cruz in a standard envelope. It arrived intact. Field test passed — cost: the price of a stamp. They then tested whether consumers would rent DVDs by mail by launching a bare website and fulfilling orders manually. The test revealed demand — but also that the pay-per-rental model produced high churn because late fees created the same friction that made Blockbuster frustrating. So Hastings ran another field test: a flat-rate subscription model with no late fees and no due dates. The subscription customers stayed. The per-rental customers churned. The data — from real customers making real choices with real money — killed the rental model and created the subscription model that would eventually destroy Blockbuster.
The pattern repeated at every inflection point. When Netflix considered streaming, Hastings did not commission a study on whether consumers would watch movies on their computers. He added streaming as a free feature for existing DVD subscribers and measured whether they used it. They did — and the usage data, not a strategic thesis, justified the multi-billion-dollar pivot that transformed Netflix from a mail-order DVD service into a $150 billion streaming platform. Hastings has described Netflix's culture as one of "informed risk-taking" — but the information always came from field tests, not from analysis.
Each decision point that could have consumed months of board-level deliberation was resolved by a field test that took weeks. The subscription model was not a strategic insight born in a conference room. It was a field test result born from watching which customers stuck around and which did not.
Section 6

Visual Explanation

FIELD TESTING — THE INTENTION-ACTION GAPWhat people say they'll do is not what they do. Test the difference.STATED INTENTREVEALED BEHAVIOURSURVEY"Would you buy this?" — 80% say yesFOCUS GROUP"I'd definitely use this" — enthusiastic nodsMARKET RESEARCH"$50B addressable market" — impressive deckFIELD TEST"Here's the product. Buy it now."Real money. Real friction. Real alternatives.3% convert — this is the truth.The other 77% were being polite.Politeness does not generate revenue.FIELD TESTS THAT BUILT COMPANIESDROPBOX (2007)3-min demo video → 75,000 beta signups overnightZAPPOS (1999)Store photos + manual fulfillment → $1.2B exitThe credit card is the only focus group that doesn't lie.
Field Testing — The gap between what people say they'll do (stated intent) and what they actually do (revealed behaviour). Field testing measures the side that matters.
The diagram maps the core tension in product validation: the methods that feel most rigorous — surveys, focus groups, market research — produce the least reliable signal for predicting real-world success, while the method that feels most primitive (show people the product and see if they buy it) produces the only signal that matters. The left column represents the tools most companies default to when validating a new concept. Each generates encouraging data. Surveys produce optimistic percentages. Focus groups produce enthusiastic reactions. Market research produces impressive TAM figures. None measure whether real people will make real trade-offs in real conditions.
The right column is the field test. One box against three, because the field test does not need variety — it needs one thing: a mechanism for real people to demonstrate real behaviour with real consequences. The 3% conversion rate is illustrative but grounded: the gap between stated purchase intent and actual purchase behaviour typically falls in the range of 10:1 to 30:1 depending on category, price point, and switching cost. The bottom examples anchor the abstraction in two of the most consequential field tests in technology history. Neither founder built much before testing. Both generated signals that billions in market research could not have replicated.
The visual asymmetry between three left-column boxes and one right-column box is deliberate. Companies invest disproportionately in stated-intent methods because they feel thorough and produce reassuring numbers. The single field test box is smaller, cheaper to execute, and produces a number that is often uncomfortable — but it is the only number that predicts revenue.
Section 7

Connected Models

Field testing sits at the intersection of experimental method and entrepreneurial practice. It borrows the hypothesis-testing rigour of the scientific method, applies it through the minimum viable product, and produces the signal that A/B testing later optimises at scale. The connected models below trace how field testing both draws from and feeds into the broader ecosystem of decision-making under uncertainty.
Understanding these connections reveals why field testing is not a standalone technique but a node in a network of models that collectively govern how businesses discover truth under uncertainty. The reinforcing connections show how field testing gains power from scientific rigour, MVP methodology, and the willingness to do unscalable work. The tension connections reveal where the model's strengths create trade-offs — particularly the irreducible tension between ecological validity and experimental control.
Reinforces
MVP
The minimum viable product is the instrument of the field test. Dropbox's three-minute video was an MVP — the smallest possible artefact that could generate a demand signal. Zappos's photographed shoe inventory was an MVP — the cheapest possible storefront that could simulate a real purchase. The MVP is not a product strategy. It is a field testing strategy: build the minimum needed to produce a real behavioural signal, and let the signal determine whether to build more. The field test defines what "viable" means — not "functional enough to demonstrate" but "real enough to trigger genuine behaviour." The reinforcement is tight: every good MVP is a field test, and every well-designed field test produces an MVP as its instrument.
Reinforces
Scientific Method
Field testing is the scientific method applied to business hypotheses. The structure is identical: form a hypothesis ("people will buy shoes online"), design an experiment (photograph shoes, post them, measure purchases), collect data (real orders from real customers), and update the hypothesis based on results. The reinforcement is mutual: the scientific method provides the rigour that prevents field tests from becoming anecdotal fishing expeditions, and field testing provides the ecological validity that laboratory experiments sacrifice. A field test without a clear hypothesis is aimless exploration. A hypothesis without a field test is an untested belief.
Reinforces
Do Things That Don't Scale
Paul Graham's principle that startups should do things that don't scale is a field testing manifesto. Swinmurn buying shoes at retail and shipping them individually did not scale. The DoorDash founders delivering food personally did not scale. Airbnb's founders photographing apartments themselves did not scale. None of these activities needed to scale — they needed to produce a signal about whether the underlying demand was real. Unscalable actions generate the highest-fidelity field test data because they occur at the exact interface between the product and the customer, with no abstraction layer diluting the signal. The paradox: the least scalable approach to validation produces the most scalable insight, because the demand signal — once confirmed — justifies building the infrastructure that makes scaling possible.
Leads-to
[A/B Testing](/mental-models/a-b-testing)
Field testing in real conditions is the ancestor of digital A/B testing. The field test establishes whether demand exists at all. A/B testing optimises the product for the demand the field test discovered. Dropbox's field test ("does anyone want cloud storage?") preceded years of A/B testing ("which referral incentive maximises sign-ups?"). The progression is natural: field testing produces the initial signal, and A/B testing refines the signal into a growth engine. Skipping the field test and jumping straight to A/B testing optimises a product that may be solving the wrong problem — the equivalent of polishing a car that no one wants to drive.
Tension
[Hypothesis](/mental-models/hypothesis)
A strong hypothesis focuses the field test — but too strong a hypothesis blinds the tester to unexpected findings. Swinmurn's hypothesis was "people will buy shoes online." But his field test revealed something he was not looking for: the critical importance of free returns. Customers would buy shoes online only if they could return them without friction. The return policy insight was more valuable than the original demand validation, and it emerged because Swinmurn's test was open enough to surface unexpected data. The tension: hypotheses provide direction, but the field test's greatest value often lies in what it reveals beyond the hypothesis — the surprises, the edge cases, the unanticipated behaviours that reshape the entire business model.
Tension
Randomized Controlled Experiment
Randomized controlled experiments maximise internal validity by isolating variables in controlled conditions. Field tests maximise ecological validity by measuring behaviour in real conditions where variables interact naturally. The tension is irreducible: you cannot simultaneously control all variables and observe natural behaviour. In medicine, where isolating causal mechanisms is paramount, the RCT dominates. In business, where the goal is predicting real-world adoption rather than identifying causal mechanisms, the field test produces more actionable intelligence. The strongest validation strategies use both: field tests to establish whether demand exists in the wild, and controlled experiments to isolate which specific variables drive the demand. The sequence matters — field-test first to confirm the phenomenon exists, then design controlled experiments to understand why it exists and how to amplify it.
Section 8

One Key Quote

"There are no facts inside the building, so get the hell outside."
— Steve Blank, The Four Steps to the Epiphany (2005)
Blank's instruction is the entire philosophy of field testing compressed into an imperative. Inside the building, you have hypotheses, assumptions, market sizing models, competitive analyses, and the collective opinions of people who have every incentive to believe the product will work. Outside the building, you have customers who will either buy or walk away. The distinction is between belief and evidence, and Blank's formulation insists that the only evidence that matters is the evidence you collect in the field.
The deeper insight is in the word "facts." Blank does not say there are no ideas inside the building, or no opinions, or no data. He says there are no facts. A fact, in Blank's usage, is a piece of evidence generated by real customer behaviour — not a projection, not a survey response, not an analyst's estimate. The sixty-slide pitch deck contains zero facts by this definition. The conversation where a potential customer says "I'd definitely buy that" contains zero facts. The moment a customer enters their credit card number — that is a fact. Field testing is the practice of manufacturing facts before committing irreversible resources. Every dollar spent before the first fact is a bet. Every dollar spent after the first fact is an investment.
The quote also encodes a bias correction. Founders spend their days inside the building — surrounded by teammates who share the vision, investors who funded the vision, and advisors who validated the vision. The building is an echo chamber by construction. "Getting outside" is not a logistical instruction. It is a cognitive one: escape the environment where confirmation bias operates at maximum strength and enter the environment where the only feedback mechanism is customer behaviour. The building tells you what you want to hear. The field tells you what is true.
Section 9

Analyst's Take

Faster Than Normal — Editorial View
Field testing is the single most important discipline in early-stage building — and the one founders skip most often. The reason is psychological, not strategic. Building feels productive. Testing feels like stalling. An engineer who spends six months building a product can point to code, features, and architecture as evidence of progress. A founder who spends two weeks running a field test has a landing page, forty-seven sign-ups, and an uncomfortable truth about whether anyone cares. The field test produces a smaller output and a more valuable one — but the emotional calculus favours building because building generates the feeling of momentum even when the direction is wrong.
The pattern that separates great founders from good ones: the willingness to let field test results kill their favourite ideas. Bezos killed the Fire Phone not because the technology failed but because the field test — actual consumer adoption — revealed that the demand thesis was wrong. Hastings killed Netflix's per-rental model not because it was unprofitable but because the field test — real subscriber retention data — showed that subscriptions produced dramatically better economics. The field test is the moment of truth, and the founder's response to an unfavourable result is the highest-leverage decision in the company's life. Kill the idea and redirect to what the field revealed. Or ignore the data and build toward a market that does not exist. The first path is painful and productive. The second is comfortable and fatal.
The most expensive mistake in product development is not a failed product. It is a product that should have been field-tested but wasn't. Every month of building without field validation accumulates assumptions that compound like debt. By the time the product launches, the assumption stack is so tall that a single failed assumption near the base collapses everything above it. Field testing is the practice of testing assumptions sequentially, starting with the most critical — does anyone want this? — and proceeding only when the answer is yes. The companies that skip this discipline do not save time. They spend the same amount of time — or more — and learn the same lesson. They just learn it after the money is gone.
The AI-era version of field testing is already diverging from the startup playbook. Companies building AI products face a novel asymmetry: the cost of building a prototype has collapsed (you can wrap an LLM in a weekend), but the cost of field-testing with real users has not (real workflow integration, real data sensitivity, real enterprise procurement). The result is an explosion of demos that look impressive and a drought of products that have survived contact with real conditions. The companies that win in AI will be the ones that field-test with real users doing real work — not the ones that demo best on stage. The Dropbox principle holds: the video gets attention, but the sign-ups tell the truth.
The operational test I apply to any product claim: show me the field test. "Customers love us" — show me the retention curve. "We've found product-market fit" — show me the organic growth rate. "This feature will drive conversion" — show me the A/B test with real users. Every assertion that is not backed by field-tested evidence is an assumption. Some assumptions are reasonable. But treating assumptions as facts is the structural error that field testing was invented to prevent.
Section 10

Test Yourself

The scenarios below test whether you can distinguish genuine field testing — measuring real behaviour in real conditions — from its common imitations: market research, customer interviews, and stated-preference surveys. The diagnostic is behavioural commitment. A field test measures what people do when it costs them something. Everything else measures what people say when it costs them nothing.
The most common analytical error is accepting positive sentiment as field test evidence. "People love it" is not a field test result — it is a description of sentiment, and sentiment does not predict behaviour with any reliability. The second error is dismissing field tests with small sample sizes. Twelve paying customers is a small number. It is also infinitely more signal than zero paying customers, and the qualitative data from those twelve transactions — why they bought, what they hesitated on, whether they returned — is often more valuable than a thousand survey responses.

Is this mental model at work here?

Scenario 1

A food startup surveys 500 people at a farmers market: 'Would you pay $12 for a jar of organic, small-batch hot sauce?' 73% say yes. The founder uses this data to raise a seed round and build a production facility.

Scenario 2

A SaaS startup builds a landing page describing a project management tool that does not yet exist. The page includes pricing tiers ($19/month, $49/month, $99/month) and a 'Start Free Trial' button. Clicking the button leads to a waitlist sign-up form that says 'We're launching soon — enter your email to get early access.' In two weeks, 340 people click the button and 280 enter their email.

Scenario 3

Before launching Airbnb nationally, Brian Chesky and Joe Gebbia personally travelled to New York — their most promising market — photographed hosts' apartments themselves (replacing amateur photos with professional-quality images), and measured whether bookings increased for apartments with upgraded photography.

Section 11

Top Resources

The literature on field testing spans experimental methodology, lean startup practice, and the memoirs of founders who built companies by testing in the real world before committing to build. Start with Blank for the philosophy, Ries for the methodology, and Knight for the visceral experience of field testing shoes at track meets in the 1960s. The strongest foundation combines the theoretical framework (why field testing works) with the operational memoir (what it feels like to do it under conditions of genuine uncertainty).
01
The Lean Startup — Eric Ries (2011)
Book
The foundational text on validated learning through real-world experimentation. Ries formalises the build-measure-learn loop and introduces the minimum viable product as the vehicle for testing business hypotheses with real customers. The book's core argument — that startups should treat every product decision as a hypothesis to be validated through the cheapest possible experiment — is the methodological backbone of field testing applied to technology companies and beyond.
02
The Four Steps to the Epiphany — Steve Blank (2005)
Book
Blank's customer development methodology provides the philosophical framework for field testing in the startup context. His insistence that founders "get out of the building" and test hypotheses directly with customers — before writing a business plan, before building a product, before raising capital — established the priority of field evidence over market research. The four-step process (customer discovery, customer validation, customer creation, company building) places field testing at the foundation of everything that follows.
03
Shoe Dog — Phil Knight (2016)
Memoir
Knight's memoir is an inadvertent masterclass in field testing. Every chapter describes a product tested in real conditions with real athletes — from the early Tiger shoes imported from Japan and sold at track meets, to Bowerman's hand-crafted prototypes tested on University of Oregon runners, to the waffle trainer field-tested on actual running surfaces. The book demonstrates field testing as a lived practice rather than a methodology, providing the operational texture that frameworks cannot capture.
04
The Mom Test — Rob Fitzpatrick (2013)
Book
Fitzpatrick's guide to customer conversations addresses the specific failure mode that field testing prevents: collecting positive feedback that tells you nothing. The book's central principle — never ask customers whether they would buy your product; instead, observe whether they have already spent money solving the problem — is a field testing discipline applied to the interview format. Essential reading for anyone who confuses enthusiasm with demand.
05
Do Things That Don't [Scale](/mental-models/scale) — Paul Graham (2013)
Essay
Graham's essay provides the strategic rationale for the unscalable field tests that produce the highest-fidelity demand signals. His argument that startups should recruit users manually, deliver the product personally, and do whatever it takes to generate early traction maps directly onto the field testing discipline: test demand through direct, unscalable interaction before investing in the scalable infrastructure. The essay reframes "things that don't scale" as the highest-signal experiments in a startup's early life — and provides the philosophical permission that many founders need to stop building and start testing.
06
Testing Business Ideas — David J. Bland & Alexander Osterwalder (2019)
Book
The most comprehensive catalogue of field testing experiments available. Bland and Osterwalder document forty-four experiments organised by the type of hypothesis being tested — desirability, feasibility, and viability — providing specific instructions, cost estimates, and evidence strength ratings for each. The book operationalises field testing from a philosophical principle into a playbook of specific techniques, making it the best tactical complement to Blank's strategic framework.

Why this matters next

mental modelsConfirmation Bias

Field Testing applied the Confirmation Bias mental model

mental modelsLeverage

Field Testing applied the Leverage mental model

mental modelsMomentum

Field Testing applied the Momentum mental model

mental modelsScientific Method

Field Testing applied the Scientific Method mental model

mental modelsInertia

Field Testing applied the Inertia mental model

mental modelsRevealed Preference

Field Testing applied the Revealed Preference mental model

Frequently asked questions

What is Field Testing?+

Field Testing is a mental model used for better thinking and decision-making.

How do you apply Field Testing?+

To apply Field Testing, identify situations where this framework is relevant, then use it as a lens to evaluate your options and decisions. The model is most useful when combined with other complementary mental models.

What category does Field Testing fall under?+

Field Testing falls under the Business & Strategy category of mental models. Other models in this category can be found on the Business & Strategy hub page.

Why is Field Testing important?+

Field Testing is important because it provides a structured way to think about problems that would otherwise be approached with intuition alone. Understanding this model helps you avoid common reasoning errors and make better decisions.

Where does Field Testing come from?+

Field Testing is discussed in the tradition of Drew Houston.

Continue exploring

1L

Mental model

100 People Love

Build something 100 people love rather than something 1 million people kind of l

7H

Mental model

7 Powers (Hamilton Helmer)

The seven — and only seven — sources of durable competitive advantage: Scale Eco

BA

Mental model

BATNA

Best Alternative to a Negotiated Agreement — your power in any negotiation comes

CL

Mental model

Competition is for Losers

Peter Thiel's thesis: every moment spent competing is a moment not spent buildin

DC

Mental model

Disagree and Commit

Once a decision is made after genuine debate, everyone — including dissenters —

DI

Mental model

Disruptive Innovation

Inferior products that start in niche markets and improve until they overtake in

More like this, in your inbox

I send a newsletter every week — free, no spam, unsubscribe anytime.

Or open the full subscribe page.

On this page

  • The Core Idea
  • How to See It
  • How to Use It
  • The Mechanism
  • Founders & Leaders in Action
  • Visual Explanation
  • Connected Models
  • One Key Quote
  • Analyst's Take
  • Test Yourself
  • Top Resources

Popular Mental Models

First Principles ThinkingOccam's RazorCircle of CompetenceInversionConfirmation BiasSecond-Order ThinkingDunning-Kruger EffectSurvivorship BiasPareto PrincipleOpportunity Cost