The post-launch research review: did your evidence hold up?
Every research conclusion that reached the roadmap was a prediction. Ninety days later you can grade them — and the grading improves the next pass more than the product.
The short answer: after about thirty real conversations or ninety days of usage, list every research conclusion that reached the roadmap as a prediction, score each one true, false or unsettled, and diagnose the false ones by their source pattern rather than by their topic. The output that matters is a change to how you sample, not a change to the product.
Launch retrospectives are common and almost always cover execution: what shipped late, what broke, what the launch channel returned. The layer underneath — whether the research that decided what to build was actually right — is rarely examined, which is why the same sampling mistakes recur across products in the same company for years.
Every conclusion was a prediction
Research documents are written in the present tense, which disguises what they are doing. “Practitioners find reconciliation the most painful step” reads as an observation and functions as a forecast: it predicts that when you build for reconciliation, people will care. That forecast can be graded.
Start by extracting them literally. Go through the research artefact and the roadmap it produced, and write each load-bearing claim as a sentence that could be false. Vague claims resist this and that resistance is itself informative — a conclusion that cannot be phrased as a falsifiable prediction was never doing analytical work.
| Prediction type | Evidence that settles it | Common verdict |
|---|---|---|
| This step is the painful one | What people ask about unprompted | Often right about the step, wrong about the intensity |
| This segment will pay | Conversions by segment | Frequently wrong; a neighbouring segment converts |
| This objection will dominate | Objections in real calls | Right on content, wrong on ranking |
| They will switch from tool X | Migration behaviour | Usually over-optimistic |
| This is table stakes | Whether its absence blocks deals | Often over-built |
| The workflow ends here | Where users export or hand off | The end is usually further downstream |
The pattern in that last column is remarkably stable across teams: research is generally good at identifying where the pain is and poor at judging how much of it converts into willingness to pay. Knowing that about your own method is worth more than any individual verdict, and it is the same calibration problem behind the market-size sanity check.
Score honestly, in three buckets
Use true, false and unsettled — and be disciplined about the third. Most teams either force a verdict on everything or quietly file inconvenient predictions as unsettled. A prediction is genuinely unsettled only when you have not yet met the situation that would test it; if you have met it and the answer was uncomfortable, it is false.
Split compound claims before scoring. “Operations leads find this painful and will pay for a fix” is two predictions, and the interesting case is when the first is true and the second is false. That combination is extremely common and gets lost entirely if the pair is scored as one.
The person who wrote a conclusion is systematically its most generous reader, and the generosity is invisible from the inside. Put the person who built against the research in the room; they usually remember which claims turned out to be uncomfortable long before anyone wrote it down.
Diagnose by source pattern, not by topic
The valuable half of the review is asking why the false predictions were false. Do not ask what you got wrong about reconciliation; ask what the wrong predictions had in common in how they were evidenced.
Four patterns account for most of it. The claim rested on one voice with unusual authority. It came from a single channel whose audience was not your market. It was inferred from an adjacent industry and never tested in yours — the risk flagged in thin-niche research. Or it came from promotional content that described an aspiration rather than a practice, which is the trap examined in separating hype from signal.
Each of those has a specific remedy in the next pass: require corroboration across unrelated creators, widen the query families, label adjacency findings as hypotheses, and weight unedited working footage above polished explainer content. The query discipline that implements the second one is in searching YouTube like a researcher.
- ✗Reviews execution, not evidence
- ✗Conclusions never restated as predictions
- ✗Wrong calls attributed to bad luck
- ✗Same sampling habits carried into the next pass
- ✓Every load-bearing claim scored
- ✓Compound claims split before scoring
- ✓Failures diagnosed by source pattern
- ✓One concrete change to the next pass
Was it a bad sample or a moving market?
Not every wrong prediction is a research failure. Markets move: a platform changes its pricing, a default tool ships the feature you were building around, a regulation lands. Distinguishing a sampling error from a genuine change matters, because the remedies are opposite — one calls for better method, the other for shorter research half-life.
The test is whether the evidence available at the time supported the conclusion. If it did and the world changed afterwards, the fix is cadence: re-run the corpus more often, which is exactly the case for monitoring a niche with recurring research. If the evidence was thin even then, the fix is sampling, and no amount of frequency will help.
Study what you got right, too
Correct predictions are treated as unremarkable and they carry as much method information as the failures. Look at how the true ones were evidenced: how many independent sources, what kind of content, how specific the original phrasing was.
In most reviews the winners share a profile — a specific behaviour observed in unedited footage, corroborated across three or more unrelated practitioners, phrased narrowly enough to be checkable. That profile is a reusable bar for the next synthesis, and applying it is what upgrades prioritising features with research evidence from a slogan into a rule.
The review should produce four artefacts
A scorecard of predictions with verdicts. A short diagnosis naming the dominant failure pattern. One concrete change to the next research pass — a query family added, a source type down-weighted, a claim class that now requires an interview. And an updated context file, because everything you just retired is still sitting in it telling a coding agent what to believe.
That last one is the step teams skip and the one with the fastest consequences, since a retired conclusion left in place keeps shaping work — the maintenance argument in keeping a CLAUDE.md current as research evolves. If the review also changed who you think the customer is, the ICP document needs the same treatment, per defining your ICP from video research.
What the pass costs
A review pass is mostly reading your own artefacts, plus five to twelve new sources to re-check the claims that failed. As of September 2026 that sits inside the Hobby plan at $19 a month with 25 videos, 2 projects and 3 syntheses; Pro at $59 covers 80 videos, 8 projects and 6 syntheses; and Studio at $199 covers 250 videos, 20 projects, 25 syntheses and 3 seats. Every plan starts with a 7-day free trial — see the pricing page.
Keep the original corpus, add the sources that test the failed predictions, and regenerate a synthesis that reflects what you now know. 7-day free trial.
Closing thought
Teams that run this review twice stop arguing about whether research works and start arguing about which sources to trust, which is a far more productive argument to be having.
Frequently asked
When should I run a post-launch research review?
Once you have roughly thirty real conversations or ninety days of usage, whichever comes first. Earlier than that and you are grading yourself on noise; much later and the wrong assumptions have already shaped a quarter of roadmap.
What exactly am I reviewing?
The specific predictions your research made, not the product's performance. Each conclusion that reached the roadmap was a claim about the world — that a step was painful, that a segment would pay, that an objection would dominate. The review asks which of those turned out to be true.
How do I score a prediction that was partly right?
Split it. A claim that a workflow step was painful but not painful enough to pay for is two predictions with different verdicts, and separating them is usually where the useful learning is. Compound claims are the ones that survive review by being too vague to fail.
What does a wrong prediction tell me about my method?
Look at the source pattern rather than the topic. Predictions that fail because they rested on one voice, on a single channel, or on adjacent-industry inference point at a fixable sampling habit; predictions that fail despite broad corroboration usually point at a real change in the market.
Should the review change the research process or the product?
Both, but the process change is the more valuable output. A wrong feature costs one quarter; a sampling bias that keeps producing wrong features costs every quarter until someone names it.
Who should be in the review?
Whoever wrote the research and whoever built against it, at minimum. The failure mode of a solo review is charitable scoring — the person who made a prediction is rarely the harshest reader of it.
What does the re-research pass cost?
As of September 2026, Hobby is $19 a month for 25 videos, 2 projects and 3 syntheses, Pro is $59 for 80 videos, 8 projects and 6 syntheses, and Studio is $199 for 250 videos, 20 projects, 25 syntheses and 3 seats. A review pass normally adds five to twelve sources to the original project.