Prioritizing a backlog with evidence instead of opinion
Reach and impact are usually guesses wearing numbers. Replace them with counted independent sources and the ranking changes — often dramatically.
The short answer: give every backlog item an evidence record — how many independent sources described the problem, how directly they experienced it, and when — then rank on that before effort enters the conversation. Items with no evidence record move to a separate bets lane rather than to the bottom of the list. The change is mundane and it reorders most backlogs immediately.
Scoring frameworks are not the problem. The problem is that reach and impact get filled in from recollection, and recollection is dominated by whoever spoke most recently and most persuasively. The framework then launders that into a number, and the number ends the argument.
What an evidence record contains
Four fields do the work. Anything more is bureaucracy; anything less leaves the ranking to memory.
| Field | What it records | Why the ranking needs it |
|---|---|---|
| Independent sources | Count of unrelated origins describing the problem | Distinguishes widespread from loud |
| Proximity | First-hand, observed, or second-hand | A report of a problem is weaker than the problem |
| Recency | When the most recent source described it | Evidence expires when competitors ship |
| Cost described | What the problem cost the person, in their words | Separates irritation from expense |
The independence field is the one that changes rankings. A backlog item supported by seven mentions that all trace to one enthusiastic customer is a one-source item, and labelling it correctly usually moves it down several places without anyone needing to argue.
Loud is not the same as widespread
Internal signal is generated by whoever is willing to generate it. One articulate customer who emails weekly can produce more apparent demand than thirty quiet accounts with the same problem, and the quiet thirty are the ones who churn without explaining why.
Public research is a useful counterweight precisely because it is not filtered through your support inbox. If a problem your loudest customer raises constantly appears nowhere in a corpus of practitioners discussing the same workflow, that is informative — not decisive, but informative.
- ✗Reach and impact scored from memory
- ✗Recent conversations dominate the ranking
- ✗One loud account counted many times
- ✗Nobody can reconstruct why the order is this
- ✓Each item carries an independent-source count
- ✓Recency visible, so stale evidence is obvious
- ✓One account counts once, regardless of volume
- ✓The order can be defended six months later
Starting without a year of research behind you
Most teams arrive at this with an existing backlog and no evidence base, and the reasonable objection is that retro-fitting records to two hundred items is a quarter of work nobody will fund.
It is also unnecessary. The tail of a backlog does not need records because it is not going to be built; only the contested top matters. Taking the top fifteen or twenty items and giving each an evidence record is an afternoon, and it resolves the only argument anyone was actually having.
The second shortcut is that the evidence usually already exists in unusable form — support tickets, sales-call notes, churn interviews. Converting those into counted independent origins is mostly clerical, and the gap it leaves is coverage of people who are not yet your customers. That gap is what a public corpus fills, and it is the half most likely to change the ranking, since existing customers can only report problems that did not stop them from buying — the discovery angle covered in reading demand signals from public video.
Weighting by proximity, not by audience
Not all evidence of equal count is of equal quality. The axis that matters is how directly the speaker experienced what they describe.
Someone walking through a job they ran last week is first-hand. Someone summarising what their team reports is observed. Someone repeating an industry claim is second-hand, and second-hand material should count toward awareness of a problem rather than toward its existence. Audience size correlates with none of this — the largest channel in a category is often the furthest from doing the work, which is one of the reliable distortions catalogued in separating hype from signal in creator content.
“Users struggle with import” cannot be compared with anything. “Import failure described first-hand by five independent sources, most recent in July” can be compared, re-checked, and watched for decay. Prose is not rankable; records are.
Keep the bets lane open
An evidence-only backlog has a real failure mode: it can only build things people already know they want. Nobody describes a feature that does not exist, so a strict evidence rule quietly filters out anything genuinely new.
The fix is a second lane rather than a loosened rule. Bets are labelled as bets, given a defined check — a landing page test, a manual version run for a handful of customers, a limited release with a named success threshold — and allocated a fixed share of capacity. What makes this work is that a bet cannot be justified by evidence it does not have, so it never competes on the ranked list and never gets quietly promoted by someone senior liking it.
Presenting a ranking someone disagrees with
The moment an evidence-ranked backlog earns its cost is when a senior stakeholder wants their item moved up. In an opinion-ranked backlog that conversation is unwinnable, because there is nothing to point at — one person's confidence against another's, settled by seniority.
With records attached, the conversation changes shape entirely. The item has three independent sources and the one above it has eight; the disagreement is now about whether those counts are right, which is checkable, or about whether evidence should govern this decision at all, which is a legitimate argument worth having explicitly. Both are better than the alternative.
What makes this work in practice is showing the sources rather than the score. A ranking presented as numbers invites haggling over the numbers; the same ranking presented as “here are the five people describing this problem, and here are the two describing that one” tends to settle itself. Evidence assembled for external strategy work gets presented the same way, for the same reason — see running client market research at agency scale. The same discipline settles integration requests, which arrive with urgency attached and no frequency data behind them: picking your first integrations from research.
Evidence expires
The most common failure of evidence-based prioritization is not bad evidence, it is old evidence treated as current. A problem that ranked first eighteen months ago may have been solved by a competitor, and a backlog does not notice that on its own.
A quarterly refresh against fresh sources catches it, and the finding to watch for is the complaint that has stopped appearing — a disappearance is a signal that something got fixed, and it is the single most under-read signal in ongoing research. The mechanics of keeping a baseline and reading the diff are in monitoring a niche with recurring research.
What this looks like in practice
A quarterly cycle that works: refresh the corpus for the product's core workflow, update the independent-source counts on the top twenty backlog items, mark anything whose most recent evidence is older than a year as stale, then rank. Effort estimates come in last, as a filter on an already-ordered list rather than as an input to the order.
The discovery side of this — finding candidate items rather than ranking existing ones — runs on the same corpus and is covered in YouTube research for product managers. Ranking and discovery working from one evidence base is what keeps them from disagreeing.
An evidence record is an input to a judgement, not a replacement for one. Its real value is that disagreement becomes specific — someone can challenge a source count or a proximity rating, which is a conversation that can be resolved, unlike a disagreement about impact scores.
What maintaining an evidence base costs
The cost is a standing project rather than a burst of volume, since the evidence base needs to persist between quarters for the comparison to work.
As of August 2026, that fits the $59 Pro tier — 8 active projects and 80 videos a month — which supports a quarterly refresh alongside other research lines. A single product with one evidence base fits the $19 Hobby plan's 2 projects and 25 videos, and Studio at $199 carries 20 projects, 250 videos and 3 team seats for multi-product teams. Detail is on the pricing page.
Keep a standing corpus for your product's core workflow, refresh it quarterly, and get source counts you can attach to backlog items. 7-day free trial.
Closing thought
The point of counting sources is not precision — the counts are rough and everybody knows it. The point is that a rough count is checkable and a confident opinion is not, and a backlog you can argue with specifically is worth considerably more than one that was settled by whoever spoke last.
Frequently asked
How do you prioritize a backlog using research evidence?
Attach an evidence record to each item — how many independent sources described the problem, how directly they experienced it, and when they said it — then sort by evidence strength before applying effort. Items with no evidence record are not deprioritized, they are separated out as bets, which is a different conversation.
What is wrong with RICE and similar scoring frameworks?
Nothing structurally, but reach and impact are usually filled in from memory, so the score inherits whatever the team already believed with a number attached. Replacing those two inputs with counted evidence is the change that makes the framework mean something.
Does the loudest customer problem deserve the highest priority?
Only if loud correlates with widespread, and it frequently does not. One articulate customer who emails often can generate more internal signal than thirty quiet ones, which is exactly the distortion an independent-source count is designed to correct.
How do you weight evidence from different sources?
By independence and proximity — how many unrelated origins described the problem, and how directly each speaker experienced it. A first-hand account outweighs a summary of other people's accounts, regardless of the audience either one reached.
Should features with no supporting evidence ever be built?
Yes, deliberately and in a separate lane. Some of the best features have no prior evidence because nobody can describe a thing that does not exist. The discipline is labelling those as bets with a defined check, not smuggling them into the evidence-ranked list.
How often should the evidence behind a backlog be refreshed?
Quarterly for most products. Evidence has a shelf life — a problem that ranked highly eighteen months ago may have been solved by a competitor since, and nothing in a backlog notices that on its own.
What does maintaining an evidence base cost?
As of August 2026, a standing project holding a product's evidence base fits within the $59 Pro tier's 8 active projects and 80 videos a month, which supports a quarterly refresh alongside other research. The $19 entry plan carries 2 projects and 25 videos.