Automatic and Semi-Automatic Document Selection for Technology-Assisted Review
Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 905–908 (2017)
Tests whether observed differences among competing review methods survive sampling variability, assessment disagreement, selection bias, and other possible confounds.
Overview
The TREC 2016 Total Recall Track reported that a fully automatic baseline outperformed two semi-automatic efforts. This paper asks whether that ranking might instead have resulted from chance, inconsistent adherence to the Track guidelines, selection bias in the evaluation method, or discordant relevance assessments.
The analysis found no evidence that any of those factors could produce relative effectiveness scores inconsistent with the official ranking. The work therefore demonstrates how a comparative result can be stress-tested rather than accepted at face value.
Key contributions
- Examines chance variation as a possible explanation for observed differences.
- Investigates inconsistent adherence to the evaluation protocol.
- Tests for selection bias in the sampling and evaluation procedure.
- Measures the effect of discordant human relevance assessments.
- Confirms that the reported ranking remains robust across these challenges.
Citation
Maura R. Grossman, Gordon V. Cormack & Adam Roegiest, Automatic and Semi-Automatic Document Selection for Technology-Assisted Review, in Proceedings of the 40th International ACM SIGIR Conference 905–908 (2017), doi:10.1145/3077136.3080675.