Publications

Lightly-supervised utterance-level emotion identification using latent topic modeling of multimodal words

Abstract

Research on multimodal emotion recognition has drawn much attention recently in diverse disciplines. With the increasing amount of multimodal data, unsupervised or semi-supervised learning has become highly desirable to automatically discover expression of emotion patterns in behavioral data. We present a novel approach for multimodal emotion learning using only a small amount of labels. Our approach is hinging on probabilistic latent semantic analysis (pLSA) that defines the latent variable as the emotion class, motivated by the conceptualization that human emotion acts as a latent control variable that regulates the external behavior manifestations, such as through speech and body gesture. In our approach, we represent the audio-visual information in an utterance as a bag of multimodal words. To exploit the interrelation between speech and gesture modalities, we propose a canonical correlation …

Date
2016
Authors
Zhaojun Yang, Shrikanth Narayanan
Conference
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Pages
2767-2771
Publisher
IEEE