Abstract:Classical supervised machine learning techniques have been explored for semantically annotating unstructured textual data such as consumers' comments archived at social media websites to extract business intelligence. However, these techniques often require a large number of manually labeled training examples to produce accurate annotations. Several active learning approaches that are designed based on probabilistic sequence models have been explored to minimize the number of labeled training examples for semantic annotation tasks. Recent research has shown that large-margin classifiers are viable alternatives to automated semantic annotation, given their strong generalization capabilities and the ability to process high-dimensional data. However, the existing active learning methods that are designed for probabilistic sequence models cannot be easily adapted and applied to large-margin classifiers. The main contribution of this paper is the development of novel active learning methods for large-margin classifiers to fill the aforementioned research gap. In particular, we propose an innovative perspective of taking active learning as a search of optimal parameters for large-margin classifiers. A rigorous evaluation involving two benchmark tests and an empirical test based on real-world data extracted from Amazon.com reveals that the proposed active learning methods can train effective classifiers with significantly fewer training examples while achieving similar annotation performance, compared to a typical state-of-the-art classifier that only uses several labeled training examples. More specifically, one of our proposed active learning methods can reduce the number of training examples by 19.74% at the 68% level of F 1 when compared to the best baseline method, as evaluated based on the Amazon data set. Our research opens the door to the application of intelligent semantic annotation techniques to support real-world applications such as automatically analyzing consumer comments for customer relationship management.

Pre-trained Language Model Based Active Learning for Sentence Matching

ActiveMatch: End-to-end Semi-supervised Active Representation Learning

Active Learning for NLP with Large Language Models

Learning to Label with Active Learning and Reinforcement Learning.

Active Sentence Learning by Adversarial Uncertainty Sampling in Discrete Space

Multi-domain active learning for text classification.

Investigating the Effectiveness of Representations Based on Pretrained Transformer-based Language Models in Active Learning for Labelling Text Datasets

ActiveLLM: Large Language Model-based Active Learning for Textual Few-Shot Scenarios

The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active Learning

Effective Active Learning Strategies for the Use of Large-Margin Classifiers in Semantic Annotation: an Optimal Parameter Discovery Perspective.

Active Learning for Mention Detection: A Comparison of Sentence Selection Strategies

Phrase-level Active Learning for Neural Machine Translation

FreeAL: Towards Human-Free Active Learning in the Era of Large Language Models

Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation

An Active Learning Approach to Task Adaptation.

LLMaAA: Making Large Language Models as Active Annotators

ActiveLab: Active Learning with Re-Labeling by Multiple Annotators

Active Learning for Vision-Language Models

Active Discriminative Text Representation Learning

Language Model-Driven Data Pruning Enables Efficient Active Learning

Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study