Efficient Visual Metaphor Image Generation Based on Metaphor Understanding

Chang Su,Xingyue Wang,Shupin Liu,Yijiang Chen
DOI: https://doi.org/10.1007/s11063-024-11609-w
IF: 2.565
2024-04-17
Neural Processing Letters
Abstract:Metaphor has significant implications for revealing cognitive and thinking mechanisms. Visual metaphor image generation not only presents metaphorical connotations intuitively but also reflects AI's understanding of metaphor through the generated images. This paper investigates the task of generating images based on text with visual metaphors. We explore metaphor image generation and create a dataset containing sentences with visual metaphors. Then, we propose a visual metaphor generation image framework based on metaphor understanding, which is more tailored to the essence of metaphor, better utilizes visual features, and has stronger interpretability. Specifically, the framework extracts the source domain, target domain, and metaphor interpretation from metaphorical sentences, separating the elements of the metaphor to deepen the understanding of its themes and intentions. Additionally, the framework introduces image data from the source domain to capture visual similarities and generate visual enhancement prompts specific to the domain. Finally, these prompts are combined with metaphorical interpretation sentences to form the final prompt text. Experimental results demonstrate that this approach effectively captures the essence of metaphor and generates metaphorical images consistent with the textual meaning.
computer science, artificial intelligence
What problem does this paper attempt to address?
### Problems the Paper Aims to Solve This paper primarily investigates the task of generating visual metaphor images based on text and proposes a new framework to improve the quality and interpretability of the generated images. Specifically: 1. **Generating Visual Metaphor Images**: - The study explores how to generate corresponding images from text containing visual metaphors, which not only intuitively showcases the content of the metaphor but also reflects the AI's understanding of the metaphor. 2. **Understanding Metaphors**: - By combining metaphor understanding with image generation, the paper aims to deepen the understanding of metaphor themes and intentions through the extraction of source domains, target domains, and metaphor explanations. 3. **Utilizing Visual Features**: - The paper introduces image data from the source domain to capture visual similarities and generate domain-specific visual enhancement prompts. These prompts, combined with metaphor explanations, form the final prompt text. 4. **Optimizing Prompts**: - The paper employs prompt optimization methods to improve the performance of existing models in handling metaphor-related image generation tasks, enhancing the consistency and effectiveness of the generated images. Through these methods, the paper aims to address issues present in existing models when generating visual metaphor images, such as the inability to accurately capture metaphorical mappings and generating images that do not match expectations. Experimental results indicate that this method effectively captures the essence of metaphors and generates metaphor images consistent with the textual meaning.