Abstract:Concept Bottleneck Models (CBM) map images to human-interpretable concepts before making class predictions. Recent approaches automate CBM construction by prompting Large Language Models (LLMs) to generate text concepts and employing Vision Language Models (VLMs) to score these concepts for CBM training. However, it is desired to build CBMs with concepts defined by human experts rather than LLM-generated ones to make them more trustworthy. In this work, we closely examine the faithfulness of VLM concept scores for such expert-defined concepts in domains like fine-grained bird species and animal classification. Our investigations reveal that VLMs like CLIP often struggle to correctly associate a concept with the corresponding visual input, despite achieving a high classification performance. This misalignment renders the resulting models difficult to interpret and less reliable. To address this issue, we propose a novel Contrastive Semi-Supervised (CSS) learning method that leverages a few labeled concept samples to activate truthful visual concepts and improve concept alignment in the CLIP model. Extensive experiments on three benchmark datasets demonstrate that our method significantly enhances both concept (+29.95) and classification (+3.84) accuracies yet requires only a fraction of human-annotated concept labels. To further improve the classification performance, we introduce a class-level intervention procedure for fine-grained classification problems that identifies the confounding classes and intervenes in their concept space to reduce errors.

Visual Concept-Metaconcept Learning

Visual Superordinate Abstraction for Robust Concept Learning

Visual Concept Learning: Combining Machine Vision and Bayesian Generalization on Concept Hierarchies

Pre-trained Vision-Language Models Learn Discoverable Visual Concepts

Language-Informed Visual Concept Learning

A Concept-Centric Approach to Multi-Modality Learning

MetaReVision: Meta-Learning with Retrieval for Visually Grounded Compositional Concept Acquisition

Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?

A Hippocampal–Entorhinal System Inspired Model for Visual Concept Representation

Fundamental Visual Concept Learning from Correlated Images and Text.

ConceptLearner: Discovering Visual Concepts from Weakly Labeled Image Collections

Conceptual Codebook Learning for Vision-Language Models

Context-Aware Meta-Learning

Flexible Compositional Learning of Structured Visual Concepts

3D Concept Learning and Reasoning from Multi-View Images

VL-Meta: Vision-Language Models for Multimodal Meta-Learning

Understanding Visual Concepts Across Models

Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency

Improving Concept Alignment in Vision-Language Concept Bottleneck Models

Learning Unseen Concepts Via Hierarchical Decomposition and Composition

Automated Construction of Visual-Linguistic Knowledge via Concept Learning from Cartoon Videos