Abstract:Most advances in medical image recognition supporting clinical auxiliary diagnosis meet challenges due to the low-resource situation in the medical field, where annotations are highly expensive and professional. This low-resource problem can be alleviated by leveraging the transferable representations of large-scale pre-trained vision-language models like CLIP. After being pre-trained using large-scale unlabeled medical images and texts (such as medical reports), the vision-language models can learn transferable representations and support flexible downstream clinical tasks such as medical image classification via relevant medical text prompts. However, existing pre-trained vision-language models require domain experts (clinicians) to carefully design the medical text prompts based on different datasets when applied to specific medical image tasks, which is extremely time-consuming and greatly increases the burden on clinicians. To address this problem, we propose a weakly supervised prompt learning method MedPrompt for automatically generating medical prompts, which includes an unsupervised pre-trained vision-language model and a weakly supervised prompt learning model. The unsupervised pre-trained vision-language model adopts large-scale medical images and texts for pre-training, utilizing the natural correlation between medical images and corresponding medical texts without manual annotations. The weakly supervised prompt learning model only utilizes the classes of images in the dataset to guide the learning of the specific class vector in the prompt, while the learning of other context vectors in the prompt does not require any manual annotations for guidance. To the best of our knowledge, this is the first model to automatically generate medical prompts. With the assistance of these prompts, the pre-trained vision-language model can be freed from the strong expert dependency of manual annotation and manual prompt design, thus achieving end-to-end, low-cost medical image classification. Experimental results show that the model using our automatically generated prompts outperforms all its hand-crafted prompts counterparts in full-shot learning on all four datasets, and achieves superior accuracy on zero-shot image classification and few-shot learning in three of the four medical benchmark datasets and comparable accuracy in the remaining one. In addition, the proposed prompt generator is lightweight and therefore has the potential to be embedded into any network architecture.

XCoOp: Explainable Prompt Learning for Computer-Aided Diagnosis via Concept-guided Context Optimization

BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models

Visual Attention Prompted Prediction and Learning

Exploring low-resource medical image classification with weakly supervised prompt learning

Aligning Medical Images with General Knowledge from Large Language Models

MCPL: Multi-modal Collaborative Prompt Learning for Medical Vision-Language Model

Report-Concept Textual-Prompt Learning for Enhancing X-ray Diagnosis

Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification

Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification

Visual Prompt Engineering for Medical Vision Language Models in Radiology

Learning to Prompt for Vision-Language Models

Candidate-Heuristic In-Context Learning: A new framework for enhancing medical visual question answering with LLMs

Domain-Controlled Prompt Learning

ECNU-LLM@CHIP-PromptCBLUE: Prompt Optimization and In-Context Learning for Chinese Medical Tasks

CMed-Baichuan: Task Explanation-Enhanced Prompt Method on PromptCBLUE Benchmark

MedPromptX: Grounded Multimodal Prompting for Chest X-ray Diagnosis

Concept-Guided Prompt Learning for Generalization in Vision-Language Models

MMGPL: Multimodal Medical Data Analysis with Graph Prompt Learning

Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models

PromptExp: Multi-granularity Prompt Explanation of Large Language Models