Abstract:Representation learning has been proven to play an important role in the unprecedented success of machine learning models in numerous tasks, such as machine translation, face recognition and recommendation. The majority of existing representation learning approaches often require a large number of consistent and noise-free labels. However, due to various reasons such as budget constraints and privacy concerns, labels are very limited in many real-world scenarios. Directly applying standard representation learning approaches on small labeled data sets will easily run into over-fitting problems and lead to sub-optimal solutions. Even worse, in some domains such as education, the limited labels are usually annotated by multiple workers with diverse expertise, which yields noises and inconsistency in such crowdsourcing settings. In this paper, we propose a novel framework which aims to learn effective representations from limited data with crowdsourced labels. Specifically, we design a grouping based deep neural network to learn embeddings from a limited number of training samples and present a Bayesian confidence estimator to capture the inconsistency among crowdsourced labels. Furthermore, to expedite the training process, we develop a hard example selection procedure to adaptively pick up training examples that are misclassified by the model. Extensive experiments conducted on three real-world data sets demonstrate the superiority of our framework on learning representations from limited data with crowdsourced labels, comparing with various state-of-the-art baselines. In addition, we provide a comprehensive analysis on each of the main components of our proposed framework and also introduce the promising results it achieved in our real production to fully understand the proposed framework. To encourage reproducible results, we make our code available online at https://github.com/tal-ai/RECLE .

Deep Learning from Crowds

Learning from Crowds under Experts' Supervision

A Formalized Framework for Incorporating Expert Labels in Crowdsourcing Environment

Deep Learning From Crowdsourced Labels: Coupled Cross-entropy Minimization, Identifiability, and Regularization

Training deep networks for facial expression recognition with crowd-sourced label distribution

Semi-Supervised Learning From Crowds Using Deep Generative Models

NeuCrowd: Neural Sampling Network for Representation Learning with Crowdsourced Labels

Learning From Crowdsourced Noisy Labels: A Signal Processing Perspective

Label Selection Approach to Learning from Crowds

Representation Learning from Limited Educational Data with Crowdsourced Labels

Leveraging Crowdsourcing Data For Deep Active Learning - An Application: Learning Intents in Alexa

EPDL: an Efficient and Privacy-Preserving Deep Learning for Crowdsensing

Deep Learning for Mobile Crowdsourcing Techniques, Methods, and Challenges: A Survey

Learning from Crowds with Annotation Reliability

Learning from Crowds Using Graph Neural Networks with Attention Mechanism

GREIL-Crowds: Crowd Simulation with Deep Reinforcement Learning and Examples

Disambiguating Named Entities with Deep Supervised Learning Via Crowd Labels.

Hierarchical Crowdsourcing for Data Labeling with Heterogeneous Crowd.

Curriculum Learning for Speech Emotion Recognition from Crowdsourced Labels

Crowdlearning: Crowded Deep Learning with Data Privacy

Deep Learning Approaches for Crowd Density Estimation: A Review