Publications

Role annotated speech recognition for conversational interactions

Abstract

Speaker Role Recognition (SRR) assigns a specific speaker role to each speaker-homogeneous speech segment in a conversation. Typically, those segments have to be identified first through a diarization step. Additionally, since SRR is usually based on the different linguistic patterns observed between the roles to be recognized, an Automatic Speech Recognition (ASR) system is also indispensable for the task in hand to convert speech to text. In this work we introduce a Role Annotated Speech Recognition (RASR) system which, given a speech signal, outputs a sequence of words annotated with the corresponding speaker roles. Thus, the need of different component modules which are connected in a way that may lead to error propagation is eliminated. We present, analyze, and test our system for the case of two speaker roles to show-case an end-to-end approach for automatic rich transcription with …

Date
2018
Authors
Nikolaos Flemotomos, Zhuohao Chen, David C Atkins, Shrikanth Narayanan
Conference
2018 IEEE Spoken Language Technology Workshop (SLT)
Pages
1036-1043
Publisher
IEEE