Abstract:Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we focus on the problem of compositional representation learning for music data, specifically targeting the fully-unsupervised setting. We propose a simple and extensible framework that leverages an explicit compositional inductive bias, defined by a flexible auto-encoding objective that can leverage any of the current state-of-art generative models. We demonstrate that our framework, used with diffusion models, naturally addresses the task of unsupervised audio source separation, showing that our model is able to perform high-quality separation. Our findings reveal that our proposal achieves comparable or superior performance with respect to other blind source separation methods and, furthermore, it even surpasses current state-of-art supervised baselines on signal-to-interference ratio metrics. Additionally, by learning an a-posteriori masking diffusion model in the space of composable representations, we achieve a system capable of seamlessly performing unsupervised source separation, unconditional generation, and variation generation. Finally, as our proposal works in the latent space of pre-trained neural audio codecs, it also provides a lower computational cost with respect to other neural baselines.

Problems using deep generative models for probabilistic audio source separation

A Variational Bayesian Approximation Approach Via A Sparsity Enforcing Prior In Acoustic Imaging

A Hierarchical Variational Bayesian Approximation Approach in Acoustic Imaging

An Efficient Variational Bayesian Inference Approach Via Studient's-t Priors for Acoustic Imaging in Colored Noises

Deep Neural Network Based Audio Source Separation

Music Source Separation With Generative Flow

Unsupervised Composable Representations for Audio

Music Separation Enhancement with Generative Modeling

Unsupervised Source Separation By Steering Pretrained Music Models

Music Source Separation Via Hybrid Waveform and Spectrogram Based Generative Adversarial Network

Deep generative models for musical audio synthesis

Maximum Discrepancy Generative Regularization and Non-Negative Matrix Factorization for Single Channel Source Separation

Generative adversarial networks with physical sound field priors

Unified Gradient Reweighting for Model Biasing with Applications to Source Separation

Class-conditional Embeddings for Music Source Separation

Multi-Source Diffusion Models for Simultaneous Music Generation and Separation

Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation

Empirical Bayesian Independent Deeply Learned Matrix Analysis For Multichannel Audio Source Separation

Music source separation conditioned on 3D point clouds

Deep Audio Waveform Prior

High-Quality Visually-Guided Sound Separation from Diverse Categories