Abstract:This paper presents a maximum-likelihood approach to multiple fundamental frequency (F0) estimation for a mixture of harmonic sound sources, where the power spectrum of a time frame is the observation and the F0s are the parameters to be estimated. When defining the likelihood model, the proposed method models both spectral peaks and non-peak regions (frequencies further than a musical quarter tone from all observed peaks). It is shown that the peak likelihood and the non-peak region likelihood act as a complementary pair. The former helps find F0s that have harmonics that explain peaks, while the latter helps avoid F0s that have harmonics in non-peak regions. Parameters of these models are learned from monophonic and polyphonic training data. This paper proposes an iterative greedy search strategy to estimate F0s one by one, to avoid the combinatorial problem of concurrent F0 estimation. It also proposes a polyphony estimation method to terminate the iterative process. Finally, this paper proposes a postprocessing method to refine polyphony and F0 estimates using neighboring frames. This paper also analyzes the relative contributions of different components of the proposed method. It is shown that the refinement component eliminates many inconsistent estimation errors. Evaluations are done on ten recorded four-part J. S. Bach chorales. Results show that the proposed method shows superior F0 estimation and polyphony estimation compared to two state-of-the-art algorithms.

Statistical Models For Dealing With Discontinuity Of Fundamental Frequency

Cross-stream Dependency Modeling Using Continuous F0 Model for HMM-based Speech Synthesis

Asynchronous F0 and Spectrum Modeling for HMM-based Speech Synthesis

Statistical modeling of syllable-level F0 features for HMM-based unit selection speech synthesis

Cross-Stream Dependency Modeling for HMM-Based Speech Synthesis

A Novel HTS System Using both Continuous HMMs and Discrete HMMs

A Hierarchical F0 Modeling Method for HMM-based Speech Synthesis

A Novel Hmm-Based Tts System Using Both Continuous Hmms And Discrete Hmms

Robust F0 Modeling for Mandarin Speech Recognition in Noise.

Modeling F0 Trajectories in Hierarchically Structured Deep Neural Networks.

Formant-Controlled HMM-Based Speech Synthesis.

Multi-Layer F0 Modeling for HMM-Based Speech Synthesis

Modeling DCT Parameterized F0 Trajectory at Intonation Phrase Level with DNN or Decision Tree

Improving F0 prediction using bidirectional associative memories and syllable-level F0 features for HMM-based Mandarin speech synthesis

Voiced/unvoiced Decision Algorithm for HMM-based Speech Synthesis

A Full Training Framework of Cross-Stream Dependence Modelling for HMM-based Singing Voice Synthesis

Global Variance Modeling on Frequency Domain Delta LSP for HMM-based Speech Synthesis

Investigation of Prosodie FO Layers in Hierarchical FO Modeling for HMM-based Speech Synthesis

F0 Modeling In Hmm-Based Speech Synthesis System Using Deep Belief Network

Multiple Fundamental Frequency Estimation by Modeling Spectral Peaks and Non-Peak Regions

Modeling the cross-linguistic variations of tonal systems