Data Alignment for Zero-Shot Concept Generation in Dermatology AI

Soham Gadgil,Mahtab Bigverdi

2024-04-20

Abstract:AI in dermatology is evolving at a rapid pace but the major limitation to training trustworthy classifiers is the scarcity of data with ground-truth concept level labels, which are meta-labels semantically meaningful to humans. Foundation models like CLIP providing zero-shot capabilities can help alleviate this challenge by leveraging vast amounts of image-caption pairs available on the internet. CLIP can be fine-tuned using domain specific image-caption pairs to improve classification performance. However, CLIP's pre-training data is not well-aligned with the medical jargon that clinicians use to perform diagnoses. The development of large language models (LLMs) in recent years has led to the possibility of leveraging the expressive nature of these models to generate rich text. Our goal is to use these models to generate caption text that aligns well with both the clinical lexicon and with the natural human language used in CLIP's pre-training data. Starting with captions used for images in PubMed articles, we extend them by passing the raw captions through an LLM fine-tuned on the field's several textbooks. We find that using captions generated by an expressive fine-tuned LLM like GPT-3.5 improves downstream zero-shot concept classification performance.

Computer Vision and Pattern Recognition,Computation and Language,Machine Learning

What problem does this paper attempt to address?

The paper aims to address a key issue in the field of dermatology artificial intelligence: the difficulty in training reliable classifiers due to the lack of data with real concept labels. Specifically, the paper attempts to improve zero-shot concept generation through the following methods: 1. **Utilizing Large Language Models (LLM) to Generate Descriptions**: The paper proposes using large language models such as GPT-2 and GPT-3.5 to generate image descriptions aligned with dermatological clinical terms. These descriptions need to conform to medical professional vocabulary and align with the pre-training data of the CLIP model. 2. **Improving CLIP Model Performance**: By using the generated descriptions to fine-tune the CLIP model, the paper aims to enhance its zero-shot concept classification performance on dermatological images. Various dermatology textbooks were used as fine-tuning data sources to ensure the generated descriptions are more professional and accurate. 3. **Evaluating the Method's Effectiveness**: The effectiveness of this method was validated through zero-shot concept classification experiments on the SKINCON dataset. The results showed that descriptions generated by the fine-tuned GPT-3.5 model significantly improved the CLIP model's classification performance. In summary, the main goal of the paper is to improve the performance of the CLIP model in dermatological image analysis by combining the power of large language models with domain-specific knowledge to generate high-quality image descriptions.

Data Alignment for Zero-Shot Concept Generation in Dermatology AI

Domain Adaptation Meets Zero-Shot Learning: an Annotation-Efficient Approach to Multi-Modality Medical Image Segmentation

VPL: Visual Proxy Learning Framework for Zero-Shot Medical Image Diagnosis

A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis

VisionCLIP: An Med-AIGC based Ethical Language-Image Foundation Model for Generalizable Retina Image Analysis

Exploring the Versatility of Zero-Shot CLIP for Interstitial Lung Disease Classification

Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training

Grounding Descriptions in Images informs Zero-Shot Visual Recognition

Pushing Boundaries: Exploring Zero Shot Object Classification with Large Multimodal Models

SkinCAP: A Multi-modal Dermatology Dataset Annotated with Rich Medical Captions

Transparent medical image AI via an image–text foundation model grounded in medical literature

Towards Cognition-Aligned Visual Language Models via Zero-Shot Instance Retrieval

XAI4LLM. Let Machine Learning Models and LLMs Collaborate for Enhanced In-Context Learning in Healthcare

Zero-Shot Clinical Acronym Expansion via Latent Meaning Cells

Advancing High Resolution Vision-Language Models in Biomedicine

Exploring Zero-Shot Anomaly Detection with CLIP in Medical Imaging: Are We There Yet?

Fostering transparent medical image AI via an image-text foundation model grounded in medical literature

DS@BioMed at ImageCLEFmedical Caption 2024: Enhanced Attention Mechanisms in Medical Caption Generation through Concept Detection Integration

Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment

Zero-shot Medical Entity Retrieval without Annotation: Learning From Rich Knowledge Graph Semantics

Revamping AI Models in Dermatology: Overcoming Critical Challenges for Enhanced Skin Lesion Diagnosis