Abstract:We can visually discriminate and recognize a wide range of materials. Meanwhile, we use language to express our subjective understanding of visual input and communicate relevant information about the materials. Here, we investigate the relationship between visual judgment and language expression in material perception to understand how visual features relate to semantic representations. We use deep generative networks to construct an expandable image space to systematically create materials of well-defined and ambiguous categories. From such a space, we sampled diverse stimuli and compared the representations of materials from two behavioral tasks: visual material similarity judgments and free-form verbal descriptions. Our findings reveal a moderate but significant correlation between vision and language on a categorical level. However, analyzing the representations with an unsupervised alignment method, we discover structural differences that arise at the image-to-image level, especially among materials morphed between known categories. Moreover, visual judgments exhibit more individual differences compared to verbal descriptions. Our results show that while verbal descriptions capture material qualities on the coarse level, they may not fully convey the visual features that characterize the material's optical properties. Analyzing the image representation of materials obtained from various pre-trained data-rich deep neural networks, we find that human visual judgments' similarity structures align more closely with those of the text-guided visual-semantic model than purely vision-based models. Our findings suggest that while semantic representations facilitate material categorization, non-semantic visual features also play a significant role in discriminating materials at a finer level. This work illustrates the need to consider the vision-language relationship in building a comprehensive model for material perception. Moreover, we propose a novel framework for quantitatively evaluating the alignment and misalignment between representations from different modalities, leveraging information from human behaviors and computational models.

Towards Visual Semantics

Aligning Visual and Lexical Semantics

Object Recognition as Classification via Visual Properties

Visual information in semantic representation

Visual Analytics for Fine‐grained Text Classification Models and Datasets

Learning Semantics for Image Annotation

Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories

Towards Semantic Embedding In Visual Vocabulary

Visual Superordinate Abstraction for Robust Concept Learning

Probing the Link Between Vision and Language in Material Perception Using Psychophysics and Unsupervised Learning

Multimodal Distributional Semantics

Visual Semantic Information Pursuit: A Survey

Understanding Visualization: A Formal Approach using Category Theory and Semiotics

SeMap: A Concept for the Visualization of Semantics as Maps

Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models

There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks

Visual Analytics for Fine-grained Text Classification Models and Datasets

Attributes as Semantic Units between Natural Language and Visual Recognition

Semantic-Based Active Perception for Humanoid Visual Tasks with Foveal Sensors

Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories