Towards Specific Domain Prompt Learning Via Improved Text Label Optimization

Liangchen Liu,Nannan Wang,Decheng Liu,Xi Yang,Xinbo Gao,Tongliang Liu
DOI: https://doi.org/10.1109/tmm.2024.3413318
IF: 7.3
2024-01-01
IEEE Transactions on Multimedia
Abstract:Prompt learning has emerged as a thriving parameter-efficient fine-tuning technique for adapting pre-trained vision-language models (VLMs) to various downstream tasks. However, existing prompt learning approaches still exhibit limited capability for adapting foundational VLMs to specific domains that require specialized and expert-level knowledge. Since this kind of specific knowledge is primarily embedded in the pre-defined text labels, we infer that foundational VLMs cannot directly interpret semantic meaningful information from these specific text labels, which causes the above limitation. From this perspective, this paper additionally models text labels with learnable tokens and casts this operation into traditional prompt learning framework. By optimizing label tokens, semantic meaningful text labels are automatically learned for each class. Nevertheless, directly optimizing text label still remains two critical problems, i.e., insufficient optimization and biased optimization. We further address these problems by proposing Modality Interaction Text Label Optimization (MITLOp) and Color-based Consistency Augmentation (CCAug) respectively, thereby effectively improving the quality of the optimized text labels. Extensive experiments indicate that our proposed method achieves significant improvements in VLM adaptation on specific domains. Code is available at https://github.com/llcllc1997/MITLOp-CCAug .
What problem does this paper attempt to address?