Maura R. GrossmanSelected publications / 2016
Continuous active learning · Scalability · CIKM 2016

Scalability of Continuous Active Learning for Reliable High-Recall Text Classification

Gordon V. Cormack and Maura R. Grossman

Proceedings of the 25th ACM International Conference on Information and Knowledge Management, pp. 1039–1048 (2016)

Shows that Continuous Active Learning can maintain reliable high recall with a fixed amount of training effort, regardless of corpus size.

Read PDFPublication recordDOIBibTeXPlain citation
Authors
Gordon V. Cormack and Maura R. Grossman
Published
Proceedings of the 25th ACM International Conference on Information and Knowledge Management, pp. 1039–1048 (2016)
Format
Article / scholarly paper

Overview

Continuous Active Learning repeatedly presents the highest-ranked unreviewed documents for assessment and retrains after each judgment. This paper investigates whether the amount of training needed to achieve reliable high recall must grow with the size of the collection.

Across collections ranging from thousands to millions of documents, the study found that a fixed amount of training could provide effective review even as the corpus grew. The result made CAL practical for very large matters and supplied the review method later examined through unbiased validation.

Why this paper matters. The contribution is not merely that CAL works well, but that its training burden does not have to increase with corpus size. That makes it possible to compare review methods at realistic scale without allowing collection size alone to dictate reviewer effort.

Key contributions

Citation

Gordon V. Cormack & Maura R. Grossman, Scalability of Continuous Active Learning for Reliable High-Recall Text Classification, in Proceedings of the 25th ACM International Conference on Information and Knowledge Management 1039–1048 (2016), doi:10.1145/2983323.2983776.

Related publications