Martin BellMartin Bell10 Min ReadUpdated Jul 14, 2026

Customer Validation: An Evidence Ladder for Startups

Move from problem evidence to commitment, paid pilots, repeat use, and scalable economics with a practical validation sequence and honest limits for every signal.

Customer Validation: The Secret to Irresistible Products

Customer validation is the process of testing whether a startup can produce a repeatable buying and delivery pattern—not merely whether people say the idea is interesting.

In Steve Blank’s Customer Development model, Customer Validation follows Customer Discovery and focuses on developing a sales model that can be repeated and scaled. His original four-stage explanation also makes clear that teams iterate between stages.

Validation is therefore a chain of evidence. Each test supports a narrower claim. No single survey, waitlist, pilot, or sale proves the whole business.

Customer discovery vs. customer validation

Customer discoveryCustomer validation
Tests who has the problem and how it appears in real lifeTests whether a specific offer can be sold and delivered repeatedly
Investigates workflows, triggers, alternatives, and decision contextInvestigates commitment, buying process, value realization, retention, and economics
Produces a sharper problem and customer hypothesisProduces a more credible sales and delivery model
Often uses interviews, observation, and early prototypesOften uses real offers, pilots, sales, usage, and renewal evidence

Discovery and validation overlap. A failed offer can reveal that the original problem or segment was misunderstood and send the team back to discovery.

Start with the Customer Development overview if you need the full four-stage context.

The customer validation evidence ladder

The ladder below moves from weaker signals of interest toward stronger behavior. It is not a mandatory sequence; choose tests that match the business and risk.

Level 1: Existing evidence and market context

Review market data, public workflows, alternatives, job postings, procurement documents, search behavior, community discussions, and competitor positioning.

Supports: The problem and category may exist.

Does not support: Your chosen customer will buy your offer.

Level 2: Problem interviews and observation

Ask qualified customers about recent events, current alternatives, cost, urgency, and who makes decisions. Observe the workflow where possible.

Supports: The problem recurs in a defined context and has observable consequences.

Does not support: The customer prefers or will pay for your solution.

Use these customer discovery questions to focus on behavior rather than hypotheticals.

Level 3: Solution interaction

Show a prototype, workflow, sample output, demo, or manual version. Give the customer a real task rather than a guided tour.

Supports: The proposed mechanism is understandable and may help complete the job.

Does not support: The customer can adopt it in their real environment or obtain approval to buy.

Level 4: Costly commitment

Ask for a behavior that costs something meaningful: time from a decision-maker, internal data, security review, a letter of intent with substance, a deposit, or another relevant commitment.

Supports: The problem and offer are important enough to justify movement.

Does not support: The full buying and delivery process will complete.

Match the commitment to the business. A consumer product and an enterprise security tool require different steps.

Level 5: Paid test or pre-sale

Offer a transparent paid pilot, paid diagnostic, pre-order, or limited service. Define scope, timing, success measures, cancellation, and refund treatment.

Supports: At least some buyers will exchange money for the proposed value under the tested conditions.

Does not support: Acquisition, delivery, retention, and margin will scale.

The pre-selling guide covers honest pre-build offers. For B2B, use the paid-pilot examples to design a bounded evaluation.

Level 6: Value realization and repeat use

Measure whether customers complete the core workflow, reach the intended outcome, return where repeat use is expected, and require a sustainable amount of support.

Supports: The offer can deliver value in the tested environment.

Does not support: A scalable market or channel by itself.

Level 7: Renewal, expansion, and referrals

Look for customers choosing to continue, expand, introduce peers, or incorporate the offer into normal work.

Supports: Value persists beyond the novelty or pilot period.

Does not support: The same behavior will appear across all segments or channels.

Level 8: Repeatable acquisition and economics

Test whether qualified customers can be sourced, sold, onboarded, served, and retained with understandable economics outside the founder’s closest network.

Supports: A candidate repeatable business model.

Does not support: Guaranteed scale. Competition, costs, product quality, and market conditions can still change.

How to design a validation test

State one risky assumption

Operations leaders at independent logistics companies will pay for a weekly exception review because unresolved shipment exceptions create measurable customer penalties.

This contains a customer, behavior, offer, and reason.

Choose the behavior that would matter

Possible evidence:

  • access to recent exception records;
  • involvement from the budget owner;
  • a paid diagnostic;
  • agreement to a defined pilot; and
  • continued use after the first review.

Define the test and boundary

Specify:

  • eligible customer profile;
  • offer and price;
  • what is delivered;
  • duration;
  • success and guardrail measures;
  • maximum time and cost; and
  • the review decision.

Do not copy a universal conversion threshold. Use a sample and decision rule appropriate to the risk, sales cycle, and available evidence.

Capture disconfirming evidence

Record why qualified buyers decline, fail to activate, stop using, or refuse a next step. A validation system that stores only success stories is a sales archive, not research.

Build a one-page validation scorecard

A useful scorecard forces a decision; it does not collect every metric the company could track. Before a test starts, write the signal, target, source, review date, and action attached to a green, amber, or red result. Set targets for your category, price, buying process, and available sample rather than borrowing a universal benchmark.

The current 100 Tasks Task 25 applies that rule to a landing-page demand test: define the traffic experiments, record the decision gate before launch, and compare upstream attention with downstream commitment. A test becomes decision evidence only when its audience, offer, sample, measurement, and threshold were appropriate and documented before the result.

The nine-signal structure below is adapted from the 100 Tasks AI customer-validation framework. Use only the rows that test a real risk in your business, but do not drop a difficult signal merely because it is harder to collect.

SignalWhat to recordEvidence sourceWhat it helps you decide
Product-loss responseHow an appropriately engaged cohort would react if the product or proposed solution disappeared, plus the cohort definition and sample sizeSurvey of customers or prospects with enough experience to judgeWhether the solution creates meaningful desire rather than polite interest
Paid commitmentsDeposits, pre-orders, paid pilots, or other credible commercial commitments under clearly stated termsPayment records, signed pilot agreements, or purchase documentationWhether interest survives a real commitment
Problem-language matchWhether qualified customers independently describe the problem, trigger, and stakes in similar languageInterview transcripts and observation notesWhether the segment shares a coherent problem
Workaround investmentMoney, time, headcount, risk, or manual effort already spent on alternativesInterviews, workflow reviews, and existing-tool evidenceWhether the problem is costly enough to motivate change
Landing-page intentThe defined action taken by relevant visitors, separated by traffic source and offerLanding-page analytics and follow-up responsesWhether the message and next step produce behavior from this audience
Acquisition economicsThe cost and effort required to produce a qualified opportunity or customer under the tested motionChannel spend, founder time, pipeline records, and resulting revenueWhether the route to market could support the business model
Unprompted referralsIntroductions or peer recommendations offered without an incentive or scripted requestInterview and sales notesWhether customers recognize other people with the same urgent problem
Founder-energy checkWhether sustained contact with the problem increases or reduces the founder’s willingness to keep working on itDated founder note reviewed with an adviser or accountability partnerWhether long-term founder commitment still fits the evidence
Pre-committed kill criteriaThe previously written conditions for continuing, narrowing, changing, or stoppingThe original assumption and decision documentWhether the team is honoring the rule it set before seeing results

After the review, add the actual result, status, date, and next test to each active row. Keep the source beside the score so another person can audit it. An unsupported green status is not evidence, and a target rewritten after the result defeats the purpose of the scorecard.

A complete validation sequence

Imagine a founder building software for independent accounting firms.

Assumption

Firm managers struggle to see which client requests are blocking month-end work and will pay to reduce follow-up.

Test 1: Workflow review

The founder reviews recent month-end cases with managers and maps where requests stall.

Decision: The pain is real, but it concentrates in firms using one common client portal.

Test 2: Manual concierge service

The founder manually consolidates outstanding requests and produces a daily priority list for two firms.

Decision: Managers use the list, but staff ignore it unless ownership is assigned inside the existing portal.

Test 3: Paid pilot

The offer becomes a four-week paid pilot that integrates one portal export, assigns request owners, and measures follow-up time and work completed before the deadline.

Decision: One firm completes and renews; another stops because data export is unreliable. The founder has evidence of value and a major technical constraint—not broad product-market fit.

Test 4: Repeatable acquisition

The founder targets similar firms using the supported portal and compares referral, partner, and direct outreach channels.

Decision: A portal consultant partnership produces better-qualified opportunities, but the economics and renewal pattern still need more cycles.

The sequence works because every test changes the next offer.

What common validation metrics really mean

MetricUseful interpretationCommon overclaim
Landing-page conversionMessage and offer produced a defined action from this traffic source“The market is validated”
Waitlist sizePeople shared contact information under stated conditions“These people will pay”
Pilot acceptanceA buyer will test under negotiated scope“We have repeatable sales”
ActivationA user reached a defined first value event“They will retain”
RetentionA defined cohort continued the target behavior“The business is profitable”
RenewalCustomers chose to continue at the tested terms“The channel scales”

For landing-page experiments, the startup landing-page examples can help match the CTA to the claim being tested.

Customer validation failure modes

Interviewing only friends

Warm contacts can help refine language, but social pressure can distort feedback. Recruit people who match the target context and do not owe the founder encouragement.

Leading with the solution

Pitching too early makes the customer evaluate your explanation rather than reveal their workflow.

Treating free use as willingness to pay

Free use can test usability or value. Payment tests a different constraint. Label the evidence correctly.

Ignoring implementation

A buyer may want the outcome but be unable to pass security, integration, procurement, or change-management requirements. Those are part of the product and sales model.

Moving the goal after the result

Write the decision rule before the test. If new evidence makes it inappropriate, document the change and why.

Scaling before retention

Acquisition can hide a leaky product. Follow customers through delivery and repeat use before increasing spend.

The validation decision

At the end of a cycle, choose one:

  • Continue: The evidence supports the assumption and the next test should increase realism.
  • Change: A specific customer, problem, offer, channel, price, or delivery belief needs revision.
  • Stop: The evidence or downside limit no longer justifies investment.
  • Inconclusive: The test did not expose the behavior and needs redesign.

If you have no existing audience, the no-audience validation guide explains how to recruit a small relevant sample without pretending reach is validation.

Customer validation is complete only in a narrow, temporary sense. You have earned the next investment when behavior supports it and the remaining uncertainty is visible.

Martin Bell

Martin Bell

Founder of 100 Tasks. Martin Bell has launched or supported 120+ startups and turned Rocket Internet venture-building discipline into a step-by-step system used by 25,000+ founders and startups.

Proven 100-Task Roadmap

Building A Startup Is Agonizing. Use The Proven 100-Task Roadmap.

Most founders are overworked, under-resourced, and forced to build without the operating sequence. 100 Tasks AI turns Martin Bell's 120+ launch process into a 100-task checklist, AI co-founder, Powersheets, and dashboard so you can launch and scale 3-5x faster.

Rocket InternetDeliverooDelivery HeroZalandoTEDDeloitteKPMGFinancial TimesThe Wall Street Journal
Start For $1
Martin Bell speaking on stage