If two answer engines receive different question sets, a difference in brand visibility may reflect the questions rather than the engines. Use paired observations when the goal is to compare responses to the same panel.
Keep the exact prompt, context, collection date, interface, and available settings for each side. Record differences that cannot be aligned. This produces a documented comparison, not a claim that the two systems are identical experiments.
The example below uses fictional Engine A and Engine B and invented counts. It demonstrates denominator handling without asserting any real provider's performance.
Keep missing responses visible
Suppose a twelve-question panel produces ten valid responses from Engine A and eleven from Engine B. Nine questions have valid responses from both. One is valid only for A, and two are valid only for B.
Report the collection coverage first: A has 10/12 valid responses, B has 11/12, and the matched comparison contains nine pairs. Do not silently substitute a retry for one engine while keeping the first attempt for the other unless that is the stated method.
| Set | Count | Use |
|---|---|---|
| Valid for both | 9 questions | Paired comparison |
| Valid only for A | 1 question | Separate observation; missing counterpart |
| Valid only for B | 2 questions | Separate observations; missing counterparts |
Describe why responses are missing when known. A collection error, unavailable tool, and valid answer with no brand mention are different states. Only the first two concern missing observation coverage.
Compare outcomes within the nine pairs
In the invented matched set, both engines mention the brand on two questions. A alone mentions it on one, B alone on three, and neither on three. That sums to nine paired questions.
A therefore has three mentions in nine matched observations; B has five in nine. The paired table also shows where the difference occurred. A single aggregate percentage would hide whether the engines agree on the same buyer questions.
Keep recommendations and citations separate from mentions. A product might be named on both sides while only one answer proposes it as suitable. Review accuracy and conditions if the business decision depends on the quality of the explanation.
Do not stretch the comparison beyond the panel
The result describes the selected questions and collection conditions. It is not an estimate of all engine users, total market share, or future behavior. A hand-selected panel remains hand-selected even when the arithmetic is correct.
If repeated observations are part of the design, specify their timing and treatment before collection. Repetition can reveal variability, but repeated answers to the same question should not be presented as unrelated independent buyers.
Avoid introducing a significance claim unless the study design and analysis support it. A small operational review can still identify useful content gaps without pretending to be a formal population study.
Turn differences into inspectable questions
For each disagreement, inspect the answer's cited evidence, product conditions, and interpretation of the task. One engine may discuss a different requirement or rely on a source that omits your supported use case.
Write a follow-up hypothesis tied to a specific resource or explanation. Do not assume that copying a competitor's wording will cause convergence. The observation suggests where to investigate; it does not prove the mechanism behind the answer.
Use the visibility measurement guide and preserve the paired records. A useful comparison makes shared outcomes, disagreements, missing evidence, and practical next questions easy to inspect.
