Abstract:Object Based Audio (OBA) provides a new kind of audio experience, delivered to the audience to personalize and customize their experience of listening and to give them choice of what and how to hear their audio content. OBA can be applied to different platforms such as broadcasting, streaming and cinema sound. This paper presents a novel approach for creating object-based audio on the production side. The approach here presents Sample-by-Sample Object Based Audio (SSOBA) embedding. SSOBA places audio object samples in such a way that allows audiences to easily individualize their chosen audio sources according to their interests and needs. SSOBA is an extra service and not an alternative, so it is also compliant with legacy audio players. The biggest advantage of SSOBA is that it does not require any special additional hardware in the broadcasting chain and it is therefore easy to implement and equip legacy players and decoders with enhanced ability. Input audio objects, number of output channels and sampling rates are three important factors affecting SSOBA performance and specifying it to be lossless or lossy. SSOBA adopts interpolation at the decoder side to compensate for eliminated samples. Both subjective and objective experiments are carried out to evaluate the output results at each step. MUSHRA subjective experiments conducted after the encoding step shows good-quality performance of SSOBA with up to five objects. SNR measurements and objective experiments, performed after decoding and interpolation, show significant successful recovery and separation of audio objects. Experimental results show that a minimum sampling rate of 96 kHz is indicated to encode up to five objects in a Stereo-mode channel to acquire good subjective and objective results simultaneously.

Stacked Sparse Autoencoder for Audio Object Coding.

Sparse Autoencoder Based Multiple Audio Objects Coding Method

Low Bitrates Audio Object Coding Using Convolutional Auto-Encoder and Densenet Mixture Model.

Adaptive subband partition encoding scheme for multiple audio objects using CNN and residual dense blocks mixture network

SpatialCodec: Neural Spatial Speech Coding

Frequency Domain Singular Value Decomposition for Efficient Spatial Audio Coding

Layered Image Compression Using Scalable Auto-Encoder

A High Fidelity and Low Complexity Neural Audio Coding

APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm

Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder

A Novel Approach for Object Based Audio Broadcasting

SuperCodec: A Neural Speech Codec with Selective Back-Projection Network

Source-Aware Neural Speech Coding for Noisy Speech Compression

AHCM: Adaptive Huffman Code Mapping for Audio Steganography Based on Psychoacoustic Model

Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Multimodal Variational Auto-encoder based Audio-Visual Segmentation

Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec