Abstract:We can visually discriminate and recognize a wide range of materials. Meanwhile, we use language to express our subjective understanding of visual input and communicate relevant information about the materials. Here, we investigate the relationship between visual judgment and language expression in material perception to understand how visual features relate to semantic representations. We use deep generative networks to construct an expandable image space to systematically create materials of well-defined and ambiguous categories. From such a space, we sampled diverse stimuli and compared the representations of materials from two behavioral tasks: visual material similarity judgments and free-form verbal descriptions. Our findings reveal a moderate but significant correlation between vision and language on a categorical level. However, analyzing the representations with an unsupervised alignment method, we discover structural differences that arise at the image-to-image level, especially among materials morphed between known categories. Moreover, visual judgments exhibit more individual differences compared to verbal descriptions. Our results show that while verbal descriptions capture material qualities on the coarse level, they may not fully convey the visual features that characterize the material's optical properties. Analyzing the image representation of materials obtained from various pre-trained data-rich deep neural networks, we find that human visual judgments' similarity structures align more closely with those of the text-guided visual-semantic model than purely vision-based models. Our findings suggest that while semantic representations facilitate material categorization, non-semantic visual features also play a significant role in discriminating materials at a finer level. This work illustrates the need to consider the vision-language relationship in building a comprehensive model for material perception. Moreover, we propose a novel framework for quantitatively evaluating the alignment and misalignment between representations from different modalities, leveraging information from human behaviors and computational models.

Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

Probing Multimodal Large Language Models for Global and Local Semantic Representations

Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Probing the Link Between Vision and Language in Material Perception Using Psychophysics and Unsupervised Learning

Perceptual Grouping in Contrastive Vision-Language Models

Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?

Explaining Multi-modal Large Language Models by Analyzing their Vision Perception

Refining Skewed Perceptions in Vision-Language Models through Visual Representations

Visually-Augmented Language Modeling

Universal Multimodal Representation for Language Understanding

Enhancing Sentence Representation with Visually-supervised Multimodal Pre-training

Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning

Probing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case

Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach

Identifying and interpreting non-aligned human conceptual representations using language modeling

Superpixel Semantics Representation and Pre-training for Vision-Language Task

Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels

Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations

Pixel Aligned Language Models

Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models