Are AI Generated Startup Ideas Any Good?

Share
AI generated startup ideas as grey cards sealed under a glass dome, with one teal card outside in the open

Key takeaways

  • Multi-agent ideation solved the wrong half. Tools like the ADHD skill fan out parallel agents and return genuine novelty at volume, but every score attached to an idea is written by the same family of model that wrote the idea.
  • A critic agent is not evidence. Pruning weak branches makes the list shorter and more internally consistent. It adds no information from outside the model about whether anyone wants the thing.
  • Novelty scoring points at the wrong markets. Ideas flagged for competitive pressure average 65.4, the highest score of any risk group, based on anonymized data from 4,000+ ideas. An empty market usually means no demand.
  • Volume does not raise the hit rate. Only 18.3% of 4,000+ analyzed ideas earned a go verdict and fewer than 0.2% broke 90. Thirty-six generated ideas is a longer list, not a better one.

There is a free skill for Claude Code called ADHD that fans a single prompt out across five independent agents, has a critic agent rank what comes back, and returns the survivors. Point it at business ideas and it will hand you roughly 36 of them in a few minutes, each scored on novelty, viability and fit. The ideas are genuinely good. That makes the scores attached to them more dangerous rather than less, and if you are holding that list right now, those scores are the only part you were planning to act on.

Will your idea survive the market?

Preuve AI runs 10 agents against live market data and links every claim to a source. Free analysis in 60 seconds.

Are AI generated startup ideas any good?

The ideas are good raw material, but the scores are not decisions. That gap is what this whole piece is about.

Idea quality really has improved. Parallel divergent ideation escapes the beige output a single prompt returns, because the branches never see each other and cannot converge on the same obvious answer. Ask for online business ideas under a 500 dollar budget with agencies, courses and dropshipping ruled out, and you get things like scavenging shut-down SaaS products for a working codebase and an abandoned niche, or mining one-star app reviews for the features nobody ever shipped. Those are not lazy-prompt ideas.

What has not improved is the part that tells you which one to build. Across anonymized data from 4,000+ ideas, only 18.3% earned a go verdict and fewer than 0.2% broke 90. The median was 55. Most ideas are not terrible, they are unfinished, and generating more of them does not touch the unfinished part.

What the ADHD skill actually does

It is worth being precise about the mechanism, because the mechanism is genuinely clever and the failure is not where most people assume.

You install it into Claude Code and invoke it as a slash command. From there it does two things. First it splits your prompt across several parallel branches, each pushed into a different frame so they explore different industries and angles independently. Second it adds a critic agent, effectively a manager, that reads everything the branches return, ranks it, discards the branches that underperformed and passes you only what survived. It can run that judgment across several criteria at once.

Both halves are real improvements over one model answering one prompt. The fan-out fixes convergence, and the critic saves you from wading through a wall of undifferentiated output on your own. For breaking a blank page, that structure beats anything a single prompt does. I built a free startup idea generator for the same moment.

The trouble starts when the ranking gets read as a verdict.

Why does the critic agent not solve the problem?

Because the critic is a model reading model output, with no new information entering the loop between the generation and the judgment.

Think about what the manager agent can actually observe. It sees text, and it compares that text against other text produced minutes earlier by its own siblings. It has no search volume, no competitor pricing pages, no customer interview, no record of anyone ever paying for anything adjacent. So the branch it throws out is just the one that argued its case less convincingly.

A matte card facing a mirror, showing an AI critic agent that only ever sees its own output
The critic reads output from its own siblings, so the most it can confirm is that the batch agrees with itself.

That is a real signal about writing quality and internal coherence, just not one that says anything about demand. Pruning makes the list shorter and more internally consistent. That is not the same as making it true. The dangerous part is that the confidence climbs anyway, and it climbs faster than the accuracy does.

Skip weeks of manual research

Get complete market research, sourced proof, competitor map, and pricing data for your idea instantly.

Why a novelty score points you at the wrong markets

Scoring on novelty is worse than neutral. It actively selects for the failure mode, and there is data on exactly how.

In anonymized data from 4,000+ ideas, the ones flagged for competitive pressure averaged 65.4, the highest score of any risk group. A crowded market is the strongest green flag in the dataset. The top killer was having no go-to-market plan, at 29.4%, roughly twice the rate of a crowded market at 14.5%.

Now read that against a scoring rubric that rewards an idea for being unlike everything else in the batch. It steers you toward empty markets and calls the emptiness an opportunity. Most of the time an empty market is a market where people tried and nobody paid. I broke the full distribution down in my startup validation benchmarks, and the pattern survives every cut of the data.

There is a second-order version of this too. When the UK site Startups.co.uk tested five AI platforms on a 500 pound budget, the name ChatGPT invented turned out to belong to an existing company in Sacramento, Copilot returned ideas built on brands that already exist, and Stratup.ai proposed hardware-heavy plant monitoring that a 500 pound budget cannot fund. A model optimizing for plausibility will hand you a trademark collision and a number that does not add up, with total confidence.

Split it by question and the boundary gets obvious. There is a set of things a generated score genuinely settles, and a set it cannot touch no matter how many agents vote.

Question before you buildAI score?What actually answers it
Is the idea internally coherent?YesThe critic agent reading the branch
Is it different from the other 35?YesThe novelty ranking, which is what it is built for
Is anyone searching for it?NoSearch volume, forum threads, a subreddit arguing about it
Is anyone already selling it?NoA named competitor list, the group averaging 65.4
Has anyone paid for it?NoA paid tool, or an agency charging for the manual version
Can you actually reach the buyer?NoA go-to-market plan, the gap in 29.4% of ideas

What is the Outside Test?

The Outside Test is the rule I apply to anything a generator hands me: every claim that decides whether you build has to come from outside the model that produced the idea. Three checks. Start with the cheapest.

1

Someone else is searching for it. Demand has to exist without you creating it. Search volume, a subreddit arguing about the problem, or forum threads where people describe the workaround they built by hand. If nobody is looking, you are funding the education of an entire market.

2

Someone else is already selling it. Name the competitors. You are not copying their playbook, you are using them as proof the market clears. The 65.4 number says this is the check founders get backwards most often, reading a competitor list as a reason to stop.

3

Someone else has already paid. Not stated interest, actual money changing hands. An existing paid tool, an agency charging for the manual version, or a budget line the buyer already defends. Interest costs the person giving it nothing, and it is worth about that much.

A novelty score fails all three by construction. It cannot see search behavior or enumerate real competitors, and it has never once watched anyone reach for a card.

Three outside checks for AI generated startup ideas: search demand, existing sellers, proof of payment
Run these in order, because each check costs you more than the one before it.

How to use the ADHD skill without trusting its scores

Keep the fan-out. The ranking is the part to throw away, and that costs you nothing, because the two halves come apart cleanly.

Use the parallel branches for what they are unmatched at, which is generating options you would not have reached alone. Then treat all 36 outputs as equally unproven, because with respect to demand they are. The scores tell you which idea argued best, and you should not confuse the best-argued idea with the best one.

From there, cut before you research. Discard anything you could not explain to a specific named buyer in one sentence, which usually halves the list. Run the Outside Test by hand on the three that remain. If a competitor search comes back empty, that is a red flag, not an opening.

Then get a read from something that is not the model that generated the idea. That is the whole reason I built Preuve the way I did, with 10 parallel AI agents pulling from 50+ live data sources instead of one model reasoning from memory. Same fan-out instinct as the ADHD skill, pointed at what the market already shows instead of at what a model can invent. I wrote up the full pipeline in how AI actually validates a startup idea. The free Reality Check scan runs in under a minute, which is less time than you will spend arguing with yourself about idea number seven.

If you would rather start from ideas that already carry demand signals instead of generating cold, I keep a ranked list in AI startup ideas for 2026. Either path works. The failure mode is not picking the wrong one, it is trusting a number the idea generated about itself.

FAQ

Are AI generated startup ideas any good?

The ideas are good raw material. The scores attached to them are not decisions. Multi-agent tools now produce real novelty, well past the generic output of a single prompt, but every rating comes from the same family of model that generated the idea, so it reflects how plausible the idea sounds rather than whether anyone wants it.

What is the ADHD skill for Claude Code?

It is a free open-source skill that adds parallel divergent ideation to a coding agent. Instead of one model answering one prompt, it fans the prompt out across several independent branches that cannot see each other, then a critic agent scores the results, prunes the weak branches and returns the survivors. It can score across several criteria at once, such as novelty, viability and fit.

Does the critic agent make the ideas more reliable?

It makes them more consistent, not more reliable. The critic is a model reading model output with no new information between the generation and the judgment. It filters for coherence and internal quality, which is useful, but coherence is not demand. Nothing in the loop consults a customer.

Is a crowded market a reason to reject an AI generated idea?

No, and this is the most common mistake. In anonymized data from 4,000+ ideas, the ones flagged for competitive pressure averaged 65.4, the highest score of any risk group. Competitors prove someone is paying. An idea nobody else pursues is more often a sign of no demand than an untapped opening.

How many AI generated ideas are actually worth pursuing?

Fewer than the list length suggests. Across 4,000+ analyzed ideas, 18.3% earned a go verdict and 76.9% landed in the caution zone between 40 and 69. Generating 36 ideas instead of 5 gives you more to sort through, not a higher proportion of good ones.

Vincent

Vincent

Founder of Preuve AI · Last updated Jul 24, 2026

5 years in B2B growth, building Preuve AI in public. 82% of ideas it scores aren't ready, the point is finding out in 8 minutes, not 3 months.

Follow on X →

Building is expensive. Validation is free.

Run your idea through 10 AI agents before you write a line of code. Every claim source-linked.