Key takeaways
- Model verdicts, not business outcomes. 17.5% of assessment records received go, 76.0% caution and 6.4% no-go. These are not observed startup success or failure rates.
- Early-stage risk is the most frequent primary label. It appears in 36.9% of tagged analyses, followed by go-to-market at 25.7%. These are model-assigned risks, not verified causes of failure.
- Competition labels coincide with higher scores. Records flagged for competitive pressure average 64.5, the highest major risk-group mean. Competitor signals also enter the scoring process, so this is not independent evidence of demand.
- The median assessment score is 54. The mean is 56.3, and fewer than 0.2% score above 90. Scores are not calibrated business success probabilities.
- Selected rescans had higher scores. Across 734 re-scored iterations, the reported mean change was +8.9 points and 75.2% increased. Selection and model-version differences prevent a causal interpretation.
The median model assessment score is 54 out of 100. These startup validation benchmarks describe anonymized data from 6,000+ completed Preuve analyses. They are not observed startup failure or success rates, and a score is not a probability that your business will succeed.
This is the third edition, with an observation cutoff of July 30, 2026. The first edition covered 1,000 analyses; the April 2026 edition used 4,000+. The September 12, 2026 editorial correction clarifies interpretation and sourcing. It preserves the July numbers and adds no new observations.
The Preuve AI Score
The Preuve AI Score is a model assessment on a 0-100 scale. In this benchmark, 70 or above is go, 40-69 is caution and below 40 is no-go. These bands describe model judgments, not observed outcomes or calibrated success probabilities. Current product capabilities do not establish which model version scored each historical record.
Overall Viability Distribution
The average viability score across the benchmark set is 56.3 out of 100. The median is 54. The distribution is still centered in the 50-59 range, where a third of assessment records land.
| Score range | % of benchmark set |
|---|---|
| 0-9 | 0.2% |
| 10-19 | 0.3% |
| 20-29 | 1.8% |
| 30-39 | 4.1% |
| 40-49 | 23.8% |
| 50-59 (peak) | 33.4% |
| 60-69 | 18.9% |
| 70-79 | 15.4% |
| 80-89 | 2.0% |
| 90-100 | 0.2% |
76.0% of assessment records fall between 40 and 69, the model caution band. That label can help you identify questions to investigate, but it does not measure whether the business will survive.
Fewer than 0.2% of assessment records score above 90. The published edition reports 9 records above 90 and 2 at 100. The 90-100 table bucket includes scores of exactly 90, unlike the strictly-above-90 figure. Neither statistic measures later business success.

Eight months, one distribution
The July edition reports these monthly average scores: December 57.0, January 53.1, February 57.3, March 56.3, April 55.7, May 58.4, June 56.0, July 55.3.
The monthly means are descriptive. Changes in submission mix, tier mix and scoring models can offset or reinforce each other. Without fixed-input comparisons and model-version information, these means do not establish score comparability between January and July.
Verdict Breakdown: Go, Caution, No-Go
Every analysis ends with one of three verdicts based on the composite viability score:
17.5%
Go
Avg score: 74.5
76.0%
Caution
Avg score: 54.3
6.4%
No-Go
Avg score: 30.1
The no-go share rose from 4.8% in April to 6.4% in July. This compares published model verdict shares across editions; it cannot separate a change in submitted records from a change in scoring.
The published mean score difference between the go and caution groups is 20.2 points. Because the groups are defined by score, this gap is descriptive, not an independent test of what makes businesses succeed.

Which risks does the model flag most often?
A primary risk label was captured on 4,995 assessment records. This tagged-analysis subset is the denominator for the risk shares below, not all completed analyses. Missing-label coverage for the cleaned comparison set is not disclosed.
| Top risk | % of tagged analyses | Avg score |
|---|---|---|
| Early-stage risk | 36.9% | 51.7 |
| Go-to-market | 25.7% | 57.9 |
| Regulatory risk | 12.0% | 57.1 |
| Competitive pressure | 11.5% | 64.5 |
| Funding requirements | 6.6% | 59.3 |
| Production challenges | 3.5% | 58.0 |
| Team execution | 2.2% | 46.6 |
Minor categories like technical complexity, legal risk, and brand risk account for the remaining 1.7% of tagged analyses.
Early-stage risk is the most frequent primary label, at 36.9%, compared with 25.6% in the April edition. Its mean score is 51.7, the second lowest major risk-group mean. The label reflects insufficiently specified concepts. The aggregates do not establish why its frequency changed.
Go-to-market ranks second at 25.7% of tagged analyses, roughly 1 in 4. The label concerns distribution and reaching a buyer. The benchmark does not verify whether those businesses later acquired customers.
The competition paradox
Competition ranks fourth at 11.5%, with a mean score of 64.5. The agents returned three or more named competitors in more than 98% of analyses; 16 reports returned zero. Competitor signals are inputs to scoring, so their association with scores is partly built into the measure. A missing competitor result can reflect retrieval limits and does not establish an empty market or absent demand. The zero-competitor subgroup is small.
What does the selected rescan comparison show?
The July edition reports 734 re-scored iterations. These are assessment records, not a count of unique founders or independently verified changes to a business. Some follow-ups may involve AI-assisted revisions.
734
re-scored iterations
+8.9
average score gain
75.2%
improved their score
The reported mean change is +8.9 points, with 75.2% of iterations receiving higher scores. This is a selected, uncontrolled comparison. Selection into rescanning, changed inputs, tier or model-version differences and repeat chains may affect the result. It does not prove that rescanning caused score gains, that founders completed homework, or that business outcomes improved.
How do free and paid assessment scores compare?
This comparison separates free quick scans from paid deep assessments. Both groups contain assessment records, with different selection and scoring processes.
Free Scan
55.7
Average score
Deep Analysis (paid)
66.6
Average score
Paid assessment records average 66.6 versus 55.7 for free records, a published 10.9-point gap. The April edition reported a 13.0-point gap. Selection, submission detail and different scoring recipes may contribute; the comparison does not measure founder seriousness, preparation or the effect of paying. Per-tier denominators and a model-version breakdown are not disclosed.
If you use a paid report, treat it as research to check against customer evidence. Naming your buyer, researching alternatives and testing a price are useful next steps, but this comparison does not estimate their effect.

What separates a 50 from a 75?
I built Preuve to organize research questions. The aggregate comparisons below suggest places to investigate; they do not measure which actions move a score from 50 to 75.
1. Named competitors with pricing
Records with zero named competitors averaged 42.9; records with three or more averaged 56.3, a reported 13.4-point gap. The zero-result subgroup contains only 16 reports, and competitor evidence also affects the score. This is not independent evidence of market demand.
2. Demand evidence from outside the founder's network
Reddit threads, review complaints, Google Trends upticks, community forum posts. External demand signals show that strangers already care about the problem. Without them, the analysis can only score what the founder claims.
3. A specific buyer, not "everyone"
"Small business owners" is not a target market. "Solo accountants in the US who still use spreadsheets for client onboarding" is. Narrow targeting makes every other validation signal sharper: competitors get easier to name, demand gets easier to measure, and pricing gets a real reference point instead of a guess.
A more detailed submission may give a model more material to assess. It still needs to be checked against reality. These aggregates do not isolate the effect of preparation, and higher scores are not proof of a better business.
How to read this benchmark
The set is self-selected and includes repeated assessment records. It does not represent all startups or unique founders. The scores of people who never use Preuve are unknown, so no direction of population bias can be inferred.
A score of 54 is the benchmark median. Use it to locate your assessment in this historical distribution, then investigate the assumptions in your own report. The rescan comparison measures score differences, not a proven intervention.
Methodology
The unit is an assessment record. The July edition starts from anonymized data from 6,000+ completed analyses through July 30, 2026 and describes a cleaned, like-for-like comparison set.
The published exclusions are the founder's own test and demo scans, a promotional batch scored by a lighter model, agency scans and investor packages. The remaining scope is free scans, paid deep analyses and follow-up improved iterations.
Risk shares use the published subset of 4,995 records with a stored primary risk label. Rescan figures use 734 selected re-scored iterations, not unique founders. Per-tier and monthly statistics have their own subsets.
All published data is aggregate and anonymized. No individual ideas, founder names or company names are disclosed.
These are model assessment scores and risk labels, not observed startup failure or success rates or calibrated outcome probabilities. No survival follow-up, failure event definition or outcome-validation study is disclosed.
The set is self-selected and includes repeated assessment records. It does not represent all startups. Selection, different tier recipes and model-version changes limit comparisons; aggregate monthly means cannot establish scoring stability.
The exact cleaned-set size and exclusion counts are not disclosed. Monthly and tier denominators, missing-label coverage, pair eligibility, repeat-chain handling, starting timestamp and timezone, and the model-version mix are not fully documented in the public edition.
Rescan differences are uncontrolled associations. They do not establish that rescanning caused improvement, that founders independently completed homework or that business prospects improved. Competitor evidence also enters scoring, so score associations are not independent proof of demand.
Source queries and July aggregate exports are not available with this publication. The numbers cannot be independently reproduced from the materials here. The September editorial correction does not recompute or newly validate the July study.
Percentages are independently rounded. Rounded buckets, verdict shares and risk shares may not sum to 100%, and adding rounded buckets may differ from a separately rounded total. Do not reconstruct exact counts from these percentages.
For the full validation methodology, read how Preuve AI validates startup ideas. For the scoring formula, see how viability scores work.
Cite or download the published aggregates
This is a transcription of selected published aggregates from the July 2026 edition, not raw data, not a new computation and not a verified source export. It preserves the published values without recalculation.
Vincent, Preuve AI. Startup Validation Benchmarks 2026: 6,000+ Analyses. 3rd edition, observation cutoff July 30, 2026; editorial revision September 12, 2026. https://preuve.ai/blog/startup-validation-benchmarks-2026.
Download published aggregates as JSON. Each metric includes its unit and subset. The file also includes the cutoff, methodology and limitations above.
FAQ
What share of assessments received a go verdict?
In the July 30, 2026 benchmark, 17.5% of assessment records scored 70 or above, 76.0% scored 40-69 and 6.4% scored below 40. These model verdict shares are not observed startup success or failure rates. Percentages are independently rounded.
What is the most common risk label?
Early-stage risk appears in 36.9% of the tagged-analysis subset, followed by go-to-market at 25.7%, regulatory risk at 12.0% and competitive pressure at 11.5%. Missing risk labels are excluded from this distribution; the labels do not establish why businesses fail.
What does a startup viability score of 70 mean?
A score of 70 or above falls in the model go band. The benchmark median is 54 and the mean is 56.3; fewer than 0.2% score above 90. A score is not a calibrated probability of startup success or proof that a business is ready to launch.
Does this benchmark measure startup failure rates?
No. It describes anonymized Preuve assessment records, not observed business outcomes. There is no disclosed survival follow-up, failure event definition or outcome-validation study. The 6.4% no-go share cannot be used as a startup failure rate; the April edition reported 4.8% in that model band.
Do rescored iterations have higher scores?
The published July comparison reports +8.9 points on average and 75.2% with higher scores across 734 re-scored iterations, not unique founders. This selected, uncontrolled comparison does not establish that rescanning caused improvement or that business prospects improved. Pair eligibility, repeat-chain handling and model-version comparability are not fully documented.
How were these benchmarks prepared?
The July edition describes anonymized data from 6,000+ completed analyses through July 30, 2026 and a cleaned, like-for-like comparison set. Founder test scans, a lighter-model promotional batch, agency scans and investor packages are excluded. The exact cleaned-set size and exclusion counts are not disclosed. September 12, 2026 is an editorial correction, not a new study or recomputation.
Vincent
5 years in B2B growth, building Preuve AI in public. 82% of ideas it scores aren't ready, the point is finding out in 10 minutes, not 3 months.
Follow on X →Building is expensive. Validation is free.
Run your idea through 10 AI agents before you write a line of code. Every claim source-linked.





