Customer Validation: An Evidence Ladder for Startups
Move from problem evidence to commitment, paid pilots, repeat use, and scalable economics with a practical validation sequence and honest limits for every signal.

Customer validation is the process of testing whether a startup can produce a repeatable buying and delivery pattern—not merely whether people say the idea is interesting.
In Steve Blank’s Customer Development model, Customer Validation follows Customer Discovery and focuses on developing a sales model that can be repeated and scaled. His original four-stage explanation also makes clear that teams iterate between stages.
Validation is therefore a chain of evidence. Each test supports a narrower claim. No single survey, waitlist, pilot, or sale proves the whole business.
Customer discovery vs. customer validation
| Customer discovery | Customer validation |
|---|---|
| Tests who has the problem and how it appears in real life | Tests whether a specific offer can be sold and delivered repeatedly |
| Investigates workflows, triggers, alternatives, and decision context | Investigates commitment, buying process, value realization, retention, and economics |
| Produces a sharper problem and customer hypothesis | Produces a more credible sales and delivery model |
| Often uses interviews, observation, and early prototypes | Often uses real offers, pilots, sales, usage, and renewal evidence |
Discovery and validation overlap. A failed offer can reveal that the original problem or segment was misunderstood and send the team back to discovery.
Start with the Customer Development overview if you need the full four-stage context.
The customer validation evidence ladder
The ladder below moves from weaker signals of interest toward stronger behavior. It is not a mandatory sequence; choose tests that match the business and risk.
Level 1: Existing evidence and market context
Review market data, public workflows, alternatives, job postings, procurement documents, search behavior, community discussions, and competitor positioning.
Supports: The problem and category may exist.
Does not support: Your chosen customer will buy your offer.
Level 2: Problem interviews and observation
Ask qualified customers about recent events, current alternatives, cost, urgency, and who makes decisions. Observe the workflow where possible.
Supports: The problem recurs in a defined context and has observable consequences.
Does not support: The customer prefers or will pay for your solution.
Use these customer discovery questions to focus on behavior rather than hypotheticals.
Level 3: Solution interaction
Show a prototype, workflow, sample output, demo, or manual version. Give the customer a real task rather than a guided tour.
Supports: The proposed mechanism is understandable and may help complete the job.
Does not support: The customer can adopt it in their real environment or obtain approval to buy.
Level 4: Costly commitment
Ask for a behavior that costs something meaningful: time from a decision-maker, internal data, security review, a letter of intent with substance, a deposit, or another relevant commitment.
Supports: The problem and offer are important enough to justify movement.
Does not support: The full buying and delivery process will complete.
Match the commitment to the business. A consumer product and an enterprise security tool require different steps.
Level 5: Paid test or pre-sale
Offer a transparent paid pilot, paid diagnostic, pre-order, or limited service. Define scope, timing, success measures, cancellation, and refund treatment.
Supports: At least some buyers will exchange money for the proposed value under the tested conditions.
Does not support: Acquisition, delivery, retention, and margin will scale.
The pre-selling guide covers honest pre-build offers. For B2B, use the paid-pilot examples to design a bounded evaluation.
Level 6: Value realization and repeat use
Measure whether customers complete the core workflow, reach the intended outcome, return where repeat use is expected, and require a sustainable amount of support.
Supports: The offer can deliver value in the tested environment.
Does not support: A scalable market or channel by itself.
Level 7: Renewal, expansion, and referrals
Look for customers choosing to continue, expand, introduce peers, or incorporate the offer into normal work.
Supports: Value persists beyond the novelty or pilot period.
Does not support: The same behavior will appear across all segments or channels.
Level 8: Repeatable acquisition and economics
Test whether qualified customers can be sourced, sold, onboarded, served, and retained with understandable economics outside the founder’s closest network.
Supports: A candidate repeatable business model.
Does not support: Guaranteed scale. Competition, costs, product quality, and market conditions can still change.
How to design a validation test
State one risky assumption
Operations leaders at independent logistics companies will pay for a weekly exception review because unresolved shipment exceptions create measurable customer penalties.
This contains a customer, behavior, offer, and reason.
Choose the behavior that would matter
Possible evidence:
- access to recent exception records;
- involvement from the budget owner;
- a paid diagnostic;
- agreement to a defined pilot; and
- continued use after the first review.
Define the test and boundary
Specify:
- eligible customer profile;
- offer and price;
- what is delivered;
- duration;
- success and guardrail measures;
- maximum time and cost; and
- the review decision.
Do not copy a universal conversion threshold. Use a sample and decision rule appropriate to the risk, sales cycle, and available evidence.
Capture disconfirming evidence
Record why qualified buyers decline, fail to activate, stop using, or refuse a next step. A validation system that stores only success stories is a sales archive, not research.
Build a one-page validation scorecard
A useful scorecard forces a decision; it does not collect every metric the company could track. Before a test starts, write the signal, target, source, review date, and action attached to a green, amber, or red result. Set targets for your category, price, buying process, and available sample rather than borrowing a universal benchmark.
The current 100 Tasks Task 25 applies that rule to a landing-page demand test: define the traffic experiments, record the decision gate before launch, and compare upstream attention with downstream commitment. A test becomes decision evidence only when its audience, offer, sample, measurement, and threshold were appropriate and documented before the result.
The nine-signal structure below is adapted from the 100 Tasks AI customer-validation framework. Use only the rows that test a real risk in your business, but do not drop a difficult signal merely because it is harder to collect.
| Signal | What to record | Evidence source | What it helps you decide |
|---|---|---|---|
| Product-loss response | How an appropriately engaged cohort would react if the product or proposed solution disappeared, plus the cohort definition and sample size | Survey of customers or prospects with enough experience to judge | Whether the solution creates meaningful desire rather than polite interest |
| Paid commitments | Deposits, pre-orders, paid pilots, or other credible commercial commitments under clearly stated terms | Payment records, signed pilot agreements, or purchase documentation | Whether interest survives a real commitment |
| Problem-language match | Whether qualified customers independently describe the problem, trigger, and stakes in similar language | Interview transcripts and observation notes | Whether the segment shares a coherent problem |
| Workaround investment | Money, time, headcount, risk, or manual effort already spent on alternatives | Interviews, workflow reviews, and existing-tool evidence | Whether the problem is costly enough to motivate change |
| Landing-page intent | The defined action taken by relevant visitors, separated by traffic source and offer | Landing-page analytics and follow-up responses | Whether the message and next step produce behavior from this audience |
| Acquisition economics | The cost and effort required to produce a qualified opportunity or customer under the tested motion | Channel spend, founder time, pipeline records, and resulting revenue | Whether the route to market could support the business model |
| Unprompted referrals | Introductions or peer recommendations offered without an incentive or scripted request | Interview and sales notes | Whether customers recognize other people with the same urgent problem |
| Founder-energy check | Whether sustained contact with the problem increases or reduces the founder’s willingness to keep working on it | Dated founder note reviewed with an adviser or accountability partner | Whether long-term founder commitment still fits the evidence |
| Pre-committed kill criteria | The previously written conditions for continuing, narrowing, changing, or stopping | The original assumption and decision document | Whether the team is honoring the rule it set before seeing results |
After the review, add the actual result, status, date, and next test to each active row. Keep the source beside the score so another person can audit it. An unsupported green status is not evidence, and a target rewritten after the result defeats the purpose of the scorecard.
A complete validation sequence
Imagine a founder building software for independent accounting firms.
Assumption
Firm managers struggle to see which client requests are blocking month-end work and will pay to reduce follow-up.
Test 1: Workflow review
The founder reviews recent month-end cases with managers and maps where requests stall.
Decision: The pain is real, but it concentrates in firms using one common client portal.
Test 2: Manual concierge service
The founder manually consolidates outstanding requests and produces a daily priority list for two firms.
Decision: Managers use the list, but staff ignore it unless ownership is assigned inside the existing portal.
Test 3: Paid pilot
The offer becomes a four-week paid pilot that integrates one portal export, assigns request owners, and measures follow-up time and work completed before the deadline.
Decision: One firm completes and renews; another stops because data export is unreliable. The founder has evidence of value and a major technical constraint—not broad product-market fit.
Test 4: Repeatable acquisition
The founder targets similar firms using the supported portal and compares referral, partner, and direct outreach channels.
Decision: A portal consultant partnership produces better-qualified opportunities, but the economics and renewal pattern still need more cycles.
The sequence works because every test changes the next offer.
What common validation metrics really mean
| Metric | Useful interpretation | Common overclaim |
|---|---|---|
| Landing-page conversion | Message and offer produced a defined action from this traffic source | “The market is validated” |
| Waitlist size | People shared contact information under stated conditions | “These people will pay” |
| Pilot acceptance | A buyer will test under negotiated scope | “We have repeatable sales” |
| Activation | A user reached a defined first value event | “They will retain” |
| Retention | A defined cohort continued the target behavior | “The business is profitable” |
| Renewal | Customers chose to continue at the tested terms | “The channel scales” |
For landing-page experiments, the startup landing-page examples can help match the CTA to the claim being tested.
Customer validation failure modes
Interviewing only friends
Warm contacts can help refine language, but social pressure can distort feedback. Recruit people who match the target context and do not owe the founder encouragement.
Leading with the solution
Pitching too early makes the customer evaluate your explanation rather than reveal their workflow.
Treating free use as willingness to pay
Free use can test usability or value. Payment tests a different constraint. Label the evidence correctly.
Ignoring implementation
A buyer may want the outcome but be unable to pass security, integration, procurement, or change-management requirements. Those are part of the product and sales model.
Moving the goal after the result
Write the decision rule before the test. If new evidence makes it inappropriate, document the change and why.
Scaling before retention
Acquisition can hide a leaky product. Follow customers through delivery and repeat use before increasing spend.
The validation decision
At the end of a cycle, choose one:
- Continue: The evidence supports the assumption and the next test should increase realism.
- Change: A specific customer, problem, offer, channel, price, or delivery belief needs revision.
- Stop: The evidence or downside limit no longer justifies investment.
- Inconclusive: The test did not expose the behavior and needs redesign.
If you have no existing audience, the no-audience validation guide explains how to recruit a small relevant sample without pretending reach is validation.
Customer validation is complete only in a narrow, temporary sense. You have earned the next investment when behavior supports it and the remaining uncertainty is visible.

Martin Bell
Founder of 100 Tasks. Martin Bell has launched or supported 120+ startups and turned Rocket Internet venture-building discipline into a step-by-step system used by 25,000+ founders and startups.


