Publications

Analyzing quality of crowd-sourced speech transcriptions of noisy audio for acoustic model adaptation

Abstract

The accuracy of crowd-sourced speech transcriptions varies depending on a variety of factors. This paper studies the impact of one such factor, namely, the quality of audio. We employed a speech database with babble noise at three SNR levels (clean, 2 dB and -2 dB) and asked workers on Amazon Mechanical Turk to transcribe it. Two interesting observations emerge. First, as expected, the quality of transcripts combined by word frequency based ROVER decreases with decreasing SNR. Further, we demonstrate that the use of some unsupervised reliability scores can improve the transcription quality, with increasing benefits at lower SNR. Second, we do not observe a significant drop in the performance of acoustic models adapted with increasing transcription noise. This highlights the surprising robustness of crowd-sourced transcripts for acoustic model adaptation.

Date
2012
Authors
Kartik Audhkhasi, Panayiotis G Georgiou, Shrikanth S Narayanan
Conference
2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Pages
4137-4140
Publisher
IEEE