---
title: "Sources - How Preuve AI cites live evidence"
slug: sources
description: "The 10 source families behind every Preuve AI report. Audit of 410 paid reports."
canonical: https://preuve.ai/sources
audit_date: 2026-07-23
reports_audited: 410
source_families: 10
---

# The sources behind every Preuve report

> **TL;DR:** Preuve AI cites live evidence from 280+ recurring publications across 10 source families. A July 2026 audit of 410 paid reports surfaced over 7,300 distinct source URLs, with market research firms (the most-cited family) appearing in 355 of the 410 reports. Every claim in every paid report links back to a source URL.

Preuve AI doesn't rate your idea from training data. Ten agents pull live evidence from 280+ recurring publications across 10 source categories. Every claim in your report links back to one.

## Audit summary

- 10 source families
- 280+ recurring publications
- 7,300+ distinct source URLs
- 410 paid Preuve reports audited (July 2026)

## How we use these sources

Validation runs as a multi-agent pipeline. Each agent owns a domain (competitors, market, sentiment) and is allowed to call only the source families relevant to its job. The verdict is the convergence point, not a single LLM's opinion.

### Step 1: 10 agents query in parallel
Competitors, market sizing, demand signals, community sentiment. Each agent owns its slice.

### Step 2: Cross-validated across models
When runs disagree, the verdict reruns until two models agree.

### Step 3: Every claim links back
No number in a report without a source URL behind it. You can fact-check every line.

## Source families

Audit of 410 recent paid Preuve reports, July 2026. Source URLs were extracted from customer-visible report sections, then matched against an explicit editorial allowlist - infrastructure, internal caches, the websites of the companies being analysed, and recommended SaaS tooling are all excluded, so every count below is a lower bound. Each row shows how often a family appeared as a cited source.

### Heavy use (>= 267 of 410 - 3 families)

1. **Market research firms** - 355 of 410. TAM, CAGR, segmentation, and forecast cited from published industry reports. Examples: Grand View, Mordor Intelligence, Statista, IBISWorld.
2. **Regulatory filings** - 307 of 410. Audited revenue, risk factors, and direct competitor mentions from public filings. Examples: SEC EDGAR, EUR-Lex, Companies House.
3. **Tech & business press** - 291 of 410. Funding announcements, product launches, and pivot reporting from named outlets. Examples: TechCrunch, CNBC, Fortune, The Information.

### Common (164-266 of 410 - 6 families)

4. **Community & social** - 247 of 410. Unfiltered pain points, demand language, and founder threads pulled from public posts. Examples: Reddit, X, YouTube, Hacker News.
5. **Professional & hiring** - 244 of 410. Org charts, hiring velocity, and team composition signal growth and product direction. Examples: LinkedIn, Glassdoor, Wellfound.
6. **Funding databases** - 242 of 410. Round size, lead investors, valuation, and competitor cap-table data. Examples: Crunchbase, PitchBook, Tracxn.
7. **PR & funding wires** - 202 of 410. First-party press releases, syndicated funding alerts, and milestone announcements. Examples: PR Newswire, BusinessWire, GlobeNewswire.
8. **Review sites & marketplaces** - 174 of 410. Verified user reviews, ratings distributions, and unmet-need clusters by product. Examples: G2, Capterra, Trustpilot, Product Hunt.
9. **Regional & international press** - 171 of 410. Local market dynamics and emerging-market signals beyond US/EU mainstream coverage. Examples: Economic Times, EU-Startups, Maddyness, Nikkei Asia.

### Specialty (< 164 of 410 - 1 family)

10. **Developer & code** - 26 of 410. Open-source traction, dependent counts, and technical depth of competing teams. Examples: GitHub, dev.to, Stack Overflow.

## Real citation examples

A hand-picked sample from the July 2026 audit. These are deep source URLs, not domain homepages, selected to show the kind of evidence reports link back to.

- **EDGAR S-1 filing archive** (Regulatory filings, SEC): Public-company risk factors, business model details, and audited disclosure. URL: https://www.sec.gov/Archives/edgar/data/1400118/000110465923074215/tm237052-9_s1.htm
- **Business software market report** (Market research firms, Grand View Research): Market sizing, segmentation, and growth-rate context. URL: https://www.grandviewresearch.com/industry-analysis/business-software-market
- **Vibe-coding founder discussion** (Community & social, Hacker News): Demand language, objections, and unscripted founder pain points. URL: https://news.ycombinator.com/item?id=44739556
- **Glean funding coverage** (Tech & business press, TechCrunch): Funding momentum, valuation context, and competitive signal. URL: https://techcrunch.com/2025/06/10/enterprise-ai-startup-glean-lands-a-7-2b-valuation/
- **Enterprise AI launch announcement** (PR & funding wires, PR Newswire): Launch timing, funding claims, and first-party positioning. URL: https://www.prnewswire.com/news-releases/woz-raises-6m-to-build-enterprise-grade-ai-apps-that-businesses-can-trust-302584904.html
- **Validator review marketplace** (Review sites & marketplaces, G2): Review density, customer language, and category alternatives. URL: https://www.g2.com/products/venturusai/reviews

## Common questions

### What sources does Preuve AI use?
Ten families: regulatory filings (SEC EDGAR), professional and hiring data (LinkedIn), community signals (Reddit, Hacker News), market research firms (Grand View, Statista), tech and business press, funding databases, regional press, PR wires, review sites, and developer sources. Each agent calls the families relevant to its slice of the report.

### How many sources does a typical Preuve report cite?
A paid report cites 40+ distinct domains, with a median of 70 clickable source links, drawn from across ten source families. The July 2026 audit of 410 paid reports surfaced over 7,300 distinct source URLs across those families, with market research firms (the most-cited family) appearing in 355 of the 410 reports and regulatory filings in 307.

### Is Preuve AI's data live, or trained on a snapshot?
Live. Each scan queries the web at run time, not a training-data snapshot. Every report stamps the moment it ran.

### How does Preuve AI verify what it cites?
Every number in a report links to a source URL. Independent models cross-check verdicts; runs rerun on disagreement.

### How does the Preuve AI Score (0-100) work?
The Preuve AI Score is a composite 0-100 viability rating built from six factors: market size and growth, real demand signals, competitive density, moat durability, business model viability, and regulatory and execution risk. Each factor is scored against 60+ live sources with a source link on every claim, then cross-checked across models. The weighting is calibrated against thousands of scored ideas, so only about 17.5% reach launch-ready (70+); the exact weights and thresholds are proprietary.

### What does "60+ live sources" mean across the site?
Per scan, the ten agents query 60+ live data sources (categories like Google Trends, Reddit, Crunchbase, news feeds, G2). What lands in the report is larger: a paid report cites 40+ distinct domains, with a median of 70 clickable source links. The July 2026 audit of 410 paid reports found those queries surfaced 280+ recurring publications grouped into the 10 families above.

## Canonical

- HTML: https://preuve.ai/sources
- Markdown: https://preuve.ai/sources.md
- Audit date: 2026-07-23
