Two reviewers can read the same AI answer and disagree about whether it recommends a product. If that disagreement remains hidden, the report may look precise while depending heavily on who coded the answers.
Run a short calibration exercise before the full review. Give both reviewers the same evidence and definitions, have them code independently, then discuss disagreements. The purpose is to improve the rule and expose ambiguous cases. It is not to pressure the reviewers into a flattering result.
This exercise uses invented answer excerpts and simple agreement arithmetic. It does not estimate the reliability of a large production study or use a named statistical agreement coefficient.
Define recommendation narrowly enough to apply
For this example, code a recommendation when the answer affirmatively proposes the named product as a suitable option for the stated task. A product can be recommended conditionally, such as when a particular integration is required. Record that condition separately.
Do not count a neutral list of company names as a recommendation under this rule. Do not count a warning against a product. A comparison can contain both a recommendation and a limitation, so retain the surrounding passage rather than extracting a favorable sentence alone.
Keep mention, recommendation, citation, and claim accuracy as distinct fields. An answer can recommend a product while misstating what it does. That is a recommendation observation with an accuracy problem, not a successful product explanation.
Resolve the rules behind the disagreement
Case 4 exposes the boundary between a research suggestion and an affirmative suitability judgment. Under the definition above, merely naming examples to investigate is not enough. Record that clarification and recode both reviewers' answers using the clarified rule.
Case 7 exposes conditional recommendations. The answer proposes suitability subject to a stated requirement, so this method counts it as a conditional recommendation. Add a condition field instead of forcing all nuance into a yes-or-no cell.
Preserve the original independent codes, the adjudicated result, and the reason. Do not overwrite the evidence of disagreement. The record shows which parts of the rubric needed clarification and makes later changes easier to audit.
Add an unresolved state when the evidence is incomplete
Sometimes the answer capture is truncated, the product name is ambiguous, or a citation opens a different destination from the one recorded. In those cases, “uncertain” or “insufficient evidence” can be more honest than forcing a binary code.
Define how those cases enter the denominator before reporting. For example, report completed codes separately from unresolved observations and show the coverage. Do not quietly classify every missing capture as no recommendation or omit it without explanation.
If a disagreement reflects a genuine limitation in the evidence, retrieve the missing context when permitted. If it reflects a policy choice, document the choice. These are different problems and require different follow-up work.
Use calibration to improve the production review
After resolving the initial set, code a fresh small set independently. Look for the same ambiguity recurring or a new boundary case. Stop when the rubric is usable for the intended review; do not keep selecting easy examples to inflate agreement.
For a larger review, keep a defined overlap sample that both reviewers code. Record any rubric change and revisit affected earlier observations. A midstream definition change can alter a trend even if the underlying engine answers did not change.
Use the reporting distinctions as the shared vocabulary. The final report should include the definition, panel context, unresolved count, and meaningful disagreements. That gives readers a way to assess the measurement instead of relying on a score with invisible judgment behind it.
For the next implementation step, use Record the conditions attached to a competitor recommendation.
