TREC 2016 Total Recall Track Overview
The Twenty-Fifth Text REtrieval Conference (TREC 2016), NIST Special Publication 500-321 (2017)
Introduces a statistically sound framework for comparing high-recall review methods in the absence of an infallible gold standard.
Overview
The TREC Total Recall Track created a common experimental framework for evaluating methods intended to find nearly all relevant documents with reasonable effort. Participating systems selected documents for review, while effectiveness was measured through independent assessment of a probability sample with unequal inclusion probabilities.
This design permitted sound quantitative comparison even though relevance assessments were necessarily incomplete and imperfect. It separated the review method from the evidence used to evaluate it and produced estimates that remained valid in spite of uncertainty in human judgment.
Key contributions
- Establishes independent assessment as the basis for comparing competing review methods.
- Uses unequal-inclusion-probability sampling to concentrate assessment effort where it is most informative.
- Supports unbiased estimation of effectiveness from a manageable sample.
- Provides quantitative measures of uncertainty rather than treating one assessor’s judgments as perfect truth.
- Creates a reproducible common task for comparing different approaches.
Citation
Maura R. Grossman, Gordon V. Cormack & Adam Roegiest, TREC 2016 Total Recall Track Overview, in The Twenty-Fifth Text REtrieval Conference Proceedings (TREC 2016), NIST Special Publication 500-321 (2017).