Abstract:Image-to-recipe retrieval is a challenging vision-to-language task of significant practical value. The main challenge of the task lies in the ultra-high redundancy in the long recipe and the large variation reflected in both food item combination and food item appearance. A de-facto idea to address this task is to learn a shared feature embedding space in which a food image is aligned better to its paired recipe than other recipes. However, such supervised global matching is prone to supervision collapse, i.e., only partial information that is necessary for distinguishing training pairs can be identified, while other information that is potentially useful in generalization could be lost. To mitigate such a problem, we propose a mask-augmentation-based local matching network (MALM), where an image-text matching module and a masked self-distillation module benefit each other mutually to learn generalizable cross-modality representations. On one hand, we perform local matching between the tokenized representations of image and text to locate fine-grained cross-modality correspondence explicitly. We involve representations of masked image patches in this process to alleviate overfitting resulting from local matching especially when some food items are underrepresented. On the other hand, predicting the hidden representations of the masked patches through self-distillation helps to learn general-purpose image representations that are expected to generalize better. And the multi-task nature of the model enables the representations of masked patches to be text-aware and thus facilitates the lost information reconstruction. Experimental results on Recipe1M dataset show our method can clearly outperform state-of-the-art (SOTA) methods. Our code will be available at <a class="link-external link-https" href="https://github.com/MyFoodChoice/MALM_Mask_Augmentation_based_Local_Matching-_for-_Food_Recipe_Retrieval" rel="external noopener nofollow">this https URL</a>

Efficient low-rank multi-component fusion with component-specific factors in image-recipe retrieval

Exploring latent weight factors and global information for food-oriented cross-modal retrieval

MCEN: Bridging Cross-Modal Gap Between Cooking Recipes and Dish Images with Latent Variable Model

MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable Model.

Revamping Image-Recipe Cross-Modal Retrieval with Dual Cross Attention Encoders

Cross-modal recipe retrieval based on unified text encoder with fine-grained contrastive learning

Enhancing Recipe Retrieval with Foundation Models: A Data Augmentation Perspective

ChefFusion: Multimodal Foundation Model Integrating Recipe and Food Image Generation

CREAMY: Cross-Modal Recipe Retrieval By Avoiding Matching Imperfectly

Transformer Decoders with MultiModal Regularization for Cross-Modal Food Retrieval

Cross-Modal Recipe Retrieval: How to Cook This Dish?

Recognize after early fusion: the Chinese food recognition based on the alignment of image and ingredients

Deep-based Ingredient Recognition for Cooking Recipe Retrieval

Retrieval Augmented Recipe Generation

Cross-modal Recipe Retrieval with Rich Food Attributes

Deep Understanding Of Cooking Procedure For Cross-Modal Recipe Retrieval

Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models

Cross-domain Cross-modal Food Transfer.

Cross-Modal Food Retrieval: Learning a Joint Embedding of Food Images and Recipes with Semantic Consistency and Attention Mechanism

MALM: Mask Augmentation based Local Matching for Food-Recipe Retrieval

Video-based Recipe Retrieval