Methodology · Sources

The sources behind every Preuve report

Preuve AI doesn't rate your idea from training data. Ten agents pull live evidence from 280+ recurring publications across 10 source categories. Every claim in your report links back to one.

10source categories
280+recurring pubs
7,300+citations · 410 reports

Last updated: July 2026 · View as Markdown

How we use these sources

Validation runs as a multi-agent pipeline. Each agent owns a domain (competitors, market, sentiment) and is allowed to call only the source families relevant to its job. The verdict is the convergence point - not a single LLM's opinion.

  1. Step 1

    10 agents query in parallel

    Competitors, market sizing, demand signals, community sentiment. Each agent owns its slice.

  2. Step 2

    Cross-validated across models

    When runs disagree, the verdict reruns until two models agree.

  3. Step 3

    Every claim links back

    No number in a report without a source URL behind it. You can fact-check every line.

How the Preuve AI Score works

The Preuve AI Score is a single 0-100 viability rating. Each idea is scored on six factors, every one checked against the same 60+ live sources with a source link on every claim, then cross-checked across models. The weighting is calibrated against thousands of scored ideas so the number stays honest: only about 17.5% reach launch-ready (70+), where most AI validators cluster at 75-85.

  1. Factor 1

    Market size & growth

    Bottom-up TAM/SAM/SOM and whether the market is expanding or shrinking, sized from market-research and funding sources.

  2. Factor 2

    Real demand signals

    Whether people actively search for, complain about, or pay to solve this, read from community and search sources, not a hunch.

  3. Factor 3

    Competitive density

    How crowded the space is and whether incumbents already solve it well, mapped from real competitors with funding and pricing.

  4. Factor 4

    Moat durability

    How easily the idea is copied, and whether anything (distribution, data, timing) makes it defensible.

  5. Factor 5

    Business model viability

    Whether the unit economics can work: pricing power, acquisition cost reality, and a path to margin.

  6. Factor 6

    Regulatory & execution risk

    Legal, compliance, and operational blockers that quietly kill ideas, flagged from filings and regional sources.

The exact weights and thresholds are proprietary and stay unpublished. What is not proprietary, and what most tools skip: every factor is scored against live, cited evidence you can click and check yourself.

Source families

Audit of 410 recent paid Preuve reports, July 2026. Source URLs were extracted and filtered to remove infrastructure and internal domains. Each row shows how often a family appeared as a cited source.

Heavy use

≥ 267 of 410 · 3 families
01

Market research firms

355/410

TAM, CAGR, segmentation, and forecast cited from published industry reports.

Grand ViewMordor IntelligenceStatistaIBISWorld
02

Regulatory filings

307/410

Audited revenue, risk factors, and direct competitor mentions from public filings.

SEC EDGAREUR-LexCompanies House
03

Tech & business press

291/410

Funding announcements, product launches, and pivot reporting from named outlets.

TechCrunchCNBCFortuneThe Information

Common

164-266 of 410 · 6 families
04

Community & social

247/410

Unfiltered pain points, demand language, and founder threads pulled from public posts.

RedditXYouTubeHacker News
05

Professional & hiring

244/410

Org charts, hiring velocity, and team composition signal growth and product direction.

LinkedInGlassdoorWellfound
06

Funding databases

242/410

Round size, lead investors, valuation, and competitor cap-table data.

CrunchbasePitchBookTracxn
07

PR & funding wires

202/410

First-party press releases, syndicated funding alerts, and milestone announcements.

PR NewswireBusinessWireGlobeNewswire
08

Review sites & marketplaces

174/410

Verified user reviews, ratings distributions, and unmet-need clusters by product.

G2CapterraTrustpilotProduct Hunt
09

Regional & international press

171/410

Local market dynamics and emerging-market signals beyond US/EU mainstream coverage.

Economic TimesEU-StartupsMaddynessNikkei Asia

Specialty

< 164 of 410 · 1 family
10

Developer & code

26/410

Open-source traction, dependent counts, and technical depth of competing teams.

GitHubdev.toStack Overflow

Common questions

What sources does Preuve AI use?
Ten families: regulatory filings (SEC EDGAR), professional and hiring data (LinkedIn), community signals (Reddit, Hacker News), market research firms (Grand View, Statista), tech and business press, funding databases, regional press, PR wires, review sites, and developer sources. Each agent calls the families relevant to its slice of the report.
How many sources does a typical Preuve report cite?
A paid report cites 40+ distinct domains, with a median of 70 clickable source links, drawn from across ten source families. The July 2026 audit of 410 paid reports surfaced over 7,300 distinct source URLs across those families, with market research firms (the most-cited family) appearing in 355 of the 410 reports and regulatory filings in 307.
Is Preuve AI's data live, or trained on a snapshot?
Live. Each scan queries the web at run time, not a training-data snapshot. Every report stamps the moment it ran.
How does Preuve AI verify what it cites?
Every number in a report links to a source URL. Independent models cross-check verdicts; runs rerun on disagreement.
How does the Preuve AI Score (0-100) work?
The Preuve AI Score is a composite 0-100 viability rating built from six factors: market size and growth, real demand signals, competitive density, moat durability, business model viability, and regulatory and execution risk. Each factor is scored against 60+ live sources with a source link on every claim, then cross-checked across models. The weighting is calibrated against thousands of scored ideas, so only about 17.5% reach launch-ready (70+); the exact weights and thresholds are proprietary.
What does "60+ live sources" mean across the site?
Per scan, the ten agents query 60+ live data sources (categories like Google Trends, Reddit, Crunchbase, news feeds, G2). What lands in the report is larger: a paid report cites 40+ distinct domains, with a median of 70 clickable source links. The July 2026 audit of 410 paid reports found those queries surfaced 280+ recurring publications grouped into the 10 families above.