Publications
Supplement for “Rate my therapist”: Automated detection of empathy in drug and alcohol counseling via speech and language processing
Abstract
Voice Activity Detection (VAD) aims to separate regions of the audio signal that contain speech from those that do not contain speech (eg, silences and environmental noises). The present research employed the VAD module developed by Van Segbroeck et al.[1]. The module uses a variety of spectral and acoustic features extracted from the audio signal, including:(I) spectral shape,(II) spectro-temporal modulations,(III) periodicity structure due to the presence of pitch harmonics, and (IV) the long-term spectral variability profile. After extracting the raw features, we normalize each feature dimension by the variance.
Using training data that are already manually annotated into speech and non-speech regions, we train a neural network model for the VAD task. The parameters of the VAD system are optimized on a separate set of audio signals referred to as the development set. Both the training and development data are independent of the data used for the evaluation of our system (ie, predicting empathy codes) described in the body of this work. The two sets employ a sample of 67 sessions drawn from the MI randomized trials, totaling 5.2 hours and 2.6 hours in length, respectively.
- Date
- 2015
- Authors
- Bo Xiao, Zac E Imel, Panayiotis G Georgiou, David C Atkins, Shrikanth S Narayanan