Before you waste 3 months building — know what to kill, validate, and build.
What you get: at most five ideas, each with the documented problem behind it and a link you can open, a 0–100 score for what one person with your skills, time, budget and reach can actually sell and support, one first move you can do this week, and the point where you should walk away. In the example run we publish, 69 candidates came back as 5. A thin market comes back as zero.
You could ask a chatbot. It will answer from memory, differently every time, and it will not go and look. Point this at a market — or at nothing at all — and it searches for candidates itself, drops the ones with no documented problem behind them, and ranks what survives against your situation. No calls, no coaching, no flattery.
SoloPipe — the best idea the run we publish came back with, scored for one person.
Next move
Ship a narrow landing page targeting the single-user pipeline wedge and test organic community response before writing code.
Real output, unedited. The whole run is at /sample-discovery.
One real sweep of 30 markets, 2026-06-04 — look at how much it throws away.
ideas found
made the shortlist
markets searched
most we shortlist per market
Why not just ask ChatGPT or Claude?
You can prompt a chatbot to be harsh. You can't prompt it to go and look.
It goes and looks — a chat window can't
A chatbot answers from memory, so it invents plausible competitors and plausible complaints. This runs live searches, reads the pages it is allowed to read, and links what it kept. The published example made 60 searches, read 24 pages, and carries 41 sources you can open right now, before you pay anything.
It throws almost everything away
One run considers up to 150 candidates and keeps at most 5. On the published run, 36 of 69 were killed and each one says which check killed it. A chat gives you back about as many ideas as you asked for.
Every run inherits the last one
Each sweep writes the problems it verified into one shared record, and the next sweep checks against it. A chat starts from zero every time you open it.
The rule doesn't move
Ask a chat twice and it grades you against whatever it feels like that time. Here the checklist, the weights and the threshold are fixed in code — the AI rates each area and never sets the total, so it cannot talk itself into a passing number. Its ratings still vary between runs, like any model's; the standard they're measured against doesn't.
Show me the receipts
“Real data” is a claim anyone can make. We link the source.
Every Discovery pain carries a clickable source — a fetched page or a search snippet — or is honestly marked not verified. On the run we publish, 32 of 56 problems carry a link you can open before you ever pay, and the page says which ones don't. Browse the source-linked pain library by topic, the live /trending feed, and the /library of finalist examples, plus the Evidence Base inside every run. Rivals advertise “real data, not opinions” with numbers nobody can verify. We ship the actual link.
As of 2026-06-10, a leading validator scored our own idea 72/100 — “PROMISING 🚀.”
We ran the same one-liner through our own tools: Kill My Idea → KILL, 15/100, with 10 specific reasons. Single Idea Check → KILL, 0/100 (KF-07 failed: “remove the LLM, what's left?”).
Dated and screenshotted — a real run, on a real day.
Don't just check ideas. Discover them.
Validators score whatever you bring them. WhittleOS goes and finds the candidates: it searches live sources across a vertical, mines documented pains, kills the noise, and hands you a ranked shortlist — or honestly returns none when the market is weak. You bring yourself, we bring the ideas.
- Live sourcing. Real search hits and fetched pages — not model memory. Three different kinds of source — a search API, our own page reader, and public APIs — so no single provider can switch us off, and every run lists the sources it actually read.
- Backed by real problems. If there's no proof people actually have the problem, we drop the idea. A weak market honestly comes back with nothing — that's the filter working, not a bug.
- Compounding evidence. Every run writes what it verified into one shared record, and the next run checks against it — across all users and markets.
- Matched to you. Every idea that passes the checks is scored against your founder profile — skills, time, budget, and reach.
- Enough to act on. Every idea on your shortlist comes with the evidence behind it, a build / test / drop call, the one thing to do next, and the point where you should walk away.
One sweep of 30 markets: 3,430 ideas found, 93 made a shortlist — about 3 per market, capped at 5. That is a 2.7% survival rate. The gate is supposed to hurt.
How it works
Three steps to an honest verdict.
Bring an idea — or none
Describe one idea in four fields, or sweep a whole niche and let the pipeline source candidates for you.
It’s scored for solo fit, by a fixed checklist
Deal-breaker checks first, then seven weighted areas — judged on what one person can build, sell, and support. The AI rates each area; the total and the gate are computed in code, so it can't hand itself a passing number.
Get a verdict you can act on
Build it, test it, or drop it — with the reasons, the biggest unknown, and one next move you can do without a single sales call.
Behind the price
What actually runs when you press Run.
A planner maps your market into sub-markets. We run live searches through three different kinds of source — a search API, our own page reader (which respects each site's robots rules and never follows a link the model invented), and public APIs — and every run lists the sources it actually read. Every candidate has to be tied to a problem somebody actually described, or it is dropped. What is left goes through a fixed list of deal-breaker checks in batches, the best of those get a full scoring pass, and at most 5 come back. One market takes 10–40 minutes, and it runs in the background — we email you when it's done.
The run we publish did exactly that:
- markets planned
- 20
- searches run
- 60
- results returned
- 93
- full pages read
- 24
- sources you can open
- 41
- ideas considered
- 69
- killed, with the reason
- 36
- shortlisted
- 5
Every one of those numbers is on that page, and each of the 36 says which check killed it.
See a real Discovery runWhat you can do
One OS, the whole decision.
Discover
Searches a whole market for you → a ranked shortlist, each idea with the evidence behind it.
Check one idea
A single Build it / Test it first / Drop it verdict with reasons.
Compare
Rank several ideas against your situation.
Validate
A 1–2 week async validation plan, and the point where you should walk away. No calls, ever.
Scope the MVP
Turn a validated idea into a narrow v0.1 build.
Kill My Idea
A deliberately adversarial teardown before you commit.
Stop building the wrong thing.
Check your own idea in under a minute — or give the pipeline half an hour to go find, screen and score ideas for you. Saying no early is the whole point.
Pay for the decision, not the dopamine.

