Scalability of Continuous Active Learning for Reliable High-Recall Text Classification
Proceedings of the 25th ACM International Conference on Information and Knowledge Management, pp. 1039–1048 (2016)
Shows that Continuous Active Learning can maintain reliable high recall with a fixed amount of training effort, regardless of corpus size.
Overview
Continuous Active Learning repeatedly presents the highest-ranked unreviewed documents for assessment and retrains after each judgment. This paper investigates whether the amount of training needed to achieve reliable high recall must grow with the size of the collection.
Across collections ranging from thousands to millions of documents, the study found that a fixed amount of training could provide effective review even as the corpus grew. The result made CAL practical for very large matters and supplied the review method later examined through unbiased validation.
Key contributions
- Demonstrates reliable high-recall classification on collections spanning several orders of magnitude.
- Shows that effective CAL training effort can remain fixed as corpus size grows.
- Establishes a practical basis for applying CAL to very large legal and regulatory collections.
- Provides the CAL methodology evaluated in later work on unbiased validation.
Citation
Gordon V. Cormack & Maura R. Grossman, Scalability of Continuous Active Learning for Reliable High-Recall Text Classification, in Proceedings of the 25th ACM International Conference on Information and Knowledge Management 1039–1048 (2016), doi:10.1145/2983323.2983776.