Publications
Automatic speaker age and gender recognition using acoustic and prosodic level information fusion
Abstract
The paper presents a novel automatic speaker age and gender identification approach which combines seven different methods at both acoustic and prosodic levels to improve the baseline performance. The three baseline subsystems are (1) Gaussian mixture model (GMM) based on mel-frequency cepstral coefficient (MFCC) features, (2) Support vector machine (SVM) based on GMM mean supervectors and (3) SVM based on 450-dimensional utterance level features including acoustic, prosodic and voice quality information. In addition, we propose four subsystems: (1) SVM based on UBM weight posterior probability supervectors using the Bhattacharyya probability product kernel, (2) Sparse representation based on UBM weight posterior probability supervectors, (3) SVM based on GMM maximum likelihood linear regression (MLLR) matrix supervectors and (4) SVM based on the polynomial expansion …
- Date
- 2013
- Authors
- Ming Li, Kyu J Han, Shrikanth Narayanan
- Journal
- Computer Speech & Language
- Volume
- 27
- Issue
- 1
- Pages
- 151-167
- Publisher
- Academic Press