Abstract:Subtitles are a crucial component of Digital Entertainment Content (DEC such as movies and TV shows) localization. With ever increasing catalog (≈ 2M titles) and localization expansion (30+ languages), automated subtitle quality checks becomes paramount. Being a manual creation process, subtitles can have errors such as missing transcriptions, out-of-sync subtitle blocks with the audio and incorrect translations. Such erroneous subtitles result in an unpleasant viewing experience and impact the viewership. Moreover, manual correction is laborious, highly costly and requires expertise of audio and subtitle languages. A typical subtitle correction process consists of (1) linear watch of the movie, (2) identification of time stamps associated with erroneous subtitle blocks, and (3) correcting procedure. Among the three, time taken to watch the entire movie by a human expert is the most time consuming step. This paper discusses the problem of missing transcription, where the subtitle blocks corresponding to some speech segments in the DEC are non-existent. We present a solution to augment human correction process by automatically identifying the timings associated with the non-transcribed dialogues in a language agnostic manner. The correction step can then be performed by either human-in-the-loop mechanism or automatically using neural transcription (speech-to-text in same language) and translation (text-to-text in different languages) engines. Our method uses a language agnostic neural voice activity detector (VAD) and an audio classifier (AC) trained explicitly on DEC corpora for better generalization. The method consists of three steps: first, we use VAD to identify the timings associated with dialogues (predicted speech blocks). Second, we refine those timings using the AC module by removing the timings associated with the leading and trailing non-speech segments identified as speech by VAD. Finally, we compare the predicted dialogue timings to the dialogue timings present in the subtitle file (subtitle speech blocks) and flag the missing transcriptions. We empirically demonstrate that the proposed method (a) reduces incorrect predicted missing subtitle timings by 10%, (b) improves the predicted missing subtitle timings by 2.5%, (c) reduces false positive rate (FPR) of overextending the predicted timings by 77%, and (d) improves the predicted speech block-level precision by a 119% over VAD baseline on a human-annotated dataset of missing subtitle speech blocks.

Speaker Diarization of Scripted Audiovisual Content

Audiovisual speaker diarization of TV series

D{é}tection de locuteurs dans les s{é}ries TV

Character-aware audio-visual subtitling in context

Audio-Visual Speaker Diarization Based on Spatiotemporal Bayesian Fusion

An Integrated Top-Down/Bottom-Up Approach To Speaker Diarization

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR

Audio-Visual Talker Localization in Video for Spatial Sound Reproduction

Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing

One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition

Creating Speech-to-Speech Corpus from Dubbed Series

Diarization of Legal Proceedings. Identifying and Transcribing Judicial Speech from Recorded Court Audio

Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments

From Modular to End-to-End Speaker Diarization

Language agnostic missing subtitle detection

A lightweight approach to real-time speaker diarization: from audio toward audio-visual data streams

Speaker Diarisation and Longitudinal Linking in Multi-Genre Broadcast Data.

Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models

A Review of Common Online Speaker Diarization Methods

Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network

Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings