Adding a pain phrase or excluding a noisy term can change which conversations you see. More results do not establish that the change helped, and a quieter feed can hide useful losses. Evaluate the current and proposed queries against the same review question: which unique conversations deserve attention for this project? This guide proposes a manual comparison worksheet. The numerical examples are illustrative calculations, not ReplyRadar performance results.
Write the hypothesis before the query
State the specific gap: for example, the current query misses people describing a manual workaround without naming the category. Predict which conversations the new wording should add. A hypothesis makes it possible to reject a change even when total alerts rise.
Hold the surrounding conditions steady
Keep source, language, time window, project profile, and review criteria the same where possible. Record unavoidable differences such as ranking or result caps. Comparing yesterday's narrow search with today's broad search mixes query effects with changes in available conversations.
Count conversations once
Build a union of both result sets using source identity. Mark each item current-only, candidate-only, or shared. Repeated matches from several terms are useful detection history, but they should not inflate the count of new opportunities.
Useful is an explicit review label
Define usefulness for the experiment: a research question, a qualified reply opportunity, or an active replacement decision. Keep those outcomes separate. A query that improves research coverage may still be unsuitable for a reply queue.
A query replacement decision sheet
Save the old query and decide the acceptance conditions before reading the candidate results. This is a paired review, not a randomized causal experiment.
Freeze the comparison
Record query versions, source scope, collection window, result limits, and the exact useful/not-useful/uncertain definitions. If supported, run both searches over the same archived interval; otherwise record when each search ran and the comparability limit.
Review the union
Review all unique results when feasible. If a sample is necessary, sample within shared and exclusive groups and report the selection rule. Hide the query label during qualification when practical so expectations do not decide the result.
Account for gains and losses
Record useful candidate-only items, useful current-only items, uncertain items, and review minutes. Within the evaluated union, net useful gain equals useful candidate-only minus useful current-only. This does not measure every relevant conversation on the platform.
Check known relevant examples
Try both queries against a small, separately selected set of relevant source examples. A query can perform well on its own returned results while missing an entire vocabulary group. Document whether a miss came from wording, source access, or search limitations.
Replace, combine, revise, or retain
Replace only if the gains meet the prewritten usefulness and review-capacity conditions. Combine queries if each contributes distinct useful items within capacity. Revise the candidate if a correctable exclusion causes losses; retain the current version if the result remains unclear.
Representative: more alerts, limited incremental value
The current query returns 40 conversations, the candidate 60, with 30 shared. Review finds 6 useful candidate-only items and 4 useful current-only items.
Why it matters: The evaluated union shows a net gain of 2 useful conversations, not 20. Decide whether that gain justifies the additional review time and whether the 4 losses include a critical buyer segment.
Representative: a harmful exclusion
Excluding free removes giveaway threads but also removes relevant requests for a free trial before a paid purchase.
Why it matters: Retain the relevant examples as regression checks and narrow the exclusion. Do not equate every occurrence of a noisy word with an irrelevant conversation.
Representative: incomparable search results
One query hits a source's result cap while the other does not, and their returned date ranges differ.
Why it matters: Report the comparison as incomplete. Narrow the common interval or evaluate an accessible common corpus before attributing the difference to query quality.
Keep a reversible change log
Store the old and new wording, hypothesis, collection conditions, decision, owner, and a review date. Change one meaningful condition at a time so an unexpected loss has an identifiable cause.
Keep discovery and evaluation separate
Examples used to invent a phrase show why it might help; fresh examples show whether it generalizes. Reserve some independently found relevant conversations for the final check.
Apply the method manually
Use a worksheet and the source searches available to you. ReplyRadar's relevance filtering supports review, but this guide does not describe a built-in query experiment runner or guarantee complete platform coverage.
Make query changes answerable to the conversations they surface.
Use the worksheet to evaluate the input, then inspect how relevance filtering affects the resulting review work.
How long should a query test run?
Choose a common window that includes the workflow situations you need to evaluate. Sparse results may require a longer observation period. Do not claim success from a fixed number of days when the useful examples are still too few or the sources are incomparable.