BAT: Learning to Reason about Spatial Sounds with Large Language Models

Zhisheng Zheng,Puyuan Peng,Ziyang Ma,Xie Chen,Eunsol Choi,David Harwath

2024-05-26

Abstract:Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.

Audio and Speech Processing,Artificial Intelligence,Computation and Language,Sound

What problem does this paper attempt to address?

The paper aims to address the current shortcomings of large language models (LLMs) in handling spatial audio inputs in real-world environments. Specifically, although existing multimodal language models have excelled in image understanding and audio tasks, they are still unable to handle complex, real-world 3D spatial audio perception and reasoning tasks. The main goal of the paper is to bridge this gap by introducing a new model—BAT (which combines the spatial awareness capabilities of binaural scene analysis models with the natural language reasoning abilities of large language models). Additionally, due to the lack of real-world spatial audio datasets, the authors synthesized a large-scale binaural audio dataset and developed a spatial audio question-answering dataset called SPATIAL SOUND QA, used to train and evaluate the model's spatial audio understanding capabilities under different complexities. Through these efforts, BAT not only excels in spatial audio perception but also demonstrates significant potential in spatial reasoning within mixed-source scenarios, showcasing the immense potential of LLMs in handling complex spatial audio environments.

BAT: Learning to Reason about Spatial Sounds with Large Language Models

Can Large Language Models Understand Spatial Audio?

Learning Spatially-Aware Language and Audio Embeddings

Semantic Object Prediction and Spatial Sound Super-Resolution with Binaural Sounds

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning

Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models

SpatialBot: Precise Spatial Understanding with Vision Language Models

BatVision: Learning to See 3D Spatial Layout with Two Ears

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models

Exploring and Improving the Spatial Reasoning Abilities of Large Language Models

What Are They Doing? Joint Audio-Speech Co-Reasoning

STBench: Assessing the Ability of Large Language Models in Spatio-Temporal Analysis

Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning

Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities

State-Space Large Audio Language Models

LAVSS: Location-Guided Audio-Visual Spatial Audio Separation

Soundscape Captioning using Sound Affective Quality Network and Large Language Model

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding