Abstract:Active learning (AL) attempts to maximize a model’s performance gain while annotating the fewest samples possible. Deep learning (DL) is greedy for data and requires a large amount of data supply to optimize a massive number of parameters if the model is to learn how to extract high-quality features. In recent years, due to the rapid development of internet technology, we have entered an era of information abundance characterized by massive amounts of available data. As a result, DL has attracted significant attention from researchers and has been rapidly developed. Compared with DL, however, researchers have a relatively low interest in AL. This is mainly because before the rise of DL, traditional machine learning requires relatively few labeled samples, meaning that early AL is rarely according the value it deserves. Although DL has made breakthroughs in various fields, most of this success is due to a large number of publicly available annotated datasets. However, the acquisition of a large number of high-quality annotated datasets consumes a lot of manpower, making it unfeasible in fields that require high levels of expertise (such as speech recognition, information extraction, medical images, etc.). Therefore, AL is gradually coming to receive the attention it is due. It is therefore natural to investigate whether AL can be used to reduce the cost of sample annotation while retaining the powerful learning capabilities of DL. As a result of such investigations, deep active learning (DeepAL) has emerged. Although research on this topic is quite abundant, there has not yet been a comprehensive survey of DeepAL-related works; accordingly, this article aims to fill this gap. We provide a formal classification method for the existing work, along with a comprehensive and systematic overview. In addition, we also analyze and summarize the development of DeepAL from an application perspective. Finally, we discuss the confusion and problems associated with DeepAL and provide some possible development directions.

LEAF: A Less Expert Annotation Framework with Active Learning

ALE: A Simulation-Based Active Learning Evaluation Framework for the Parameter-Driven Comparison of Query Strategies for NLP

FreeAL: Towards Human-Free Active Learning in the Era of Large Language Models

Optimizing Active Learning for Low Annotation Budgets

Reducing Workload of Manual Annotation for User Requests Via a Novel Active Learning Framework

AutoAL: Automated Active Learning with Differentiable Query Strategy Search

A meta-framework for multi-label active learning based on deep reinforcement learning

Active Learning for NLP with Large Language Models

LLMaAA: Making Large Language Models as Active Annotators

A Survey on Cost Types, Interaction Schemes, and Annotator Performance Models in Selection Algorithms for Active Learning in Classification

Context Aware Image Annotation in Active Learning

A Survey of Deep Active Learning

Beyond Active Learning: Leveraging the Full Potential of Human Interaction via Auto-Labeling, Human Correction, and Human Verification

LegalATLE: an Active Transfer Learning Framework for Legal Triple Extraction

Federated Active Learning (F-AL): An Efficient Annotation Strategy for Federated Learning

Efficient Human-in-the-Loop Active Learning: A Novel Framework for Data Labeling in AI Systems

Improving Active Learning by Data Balance to Reduce Annotation Efforts

Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning

Learning from conflicting data with hidden contexts

Practical Obstacles to Deploying Active Learning

The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators