Publications
Role annotated speech recognition for conversational interactions
Abstract
Speaker Role Recognition (SRR) assigns a specific speaker role to each speaker-homogeneous speech segment in a conversation. Typically, those segments have to be identified first through a diarization step. Additionally, since SRR is usually based on the different linguistic patterns observed between the roles to be recognized, an Automatic Speech Recognition (ASR) system is also indispensable for the task in hand to convert speech to text. In this work we introduce a Role Annotated Speech Recognition (RASR) system which, given a speech signal, outputs a sequence of words annotated with the corresponding speaker roles. Thus, the need of different component modules which are connected in a way that may lead to error propagation is eliminated. We present, analyze, and test our system for the case of two speaker roles to show-case an end-to-end approach for automatic rich transcription with …
- Date
- 2018
- Authors
- Nikolaos Flemotomos, Zhuohao Chen, David C Atkins, Shrikanth Narayanan
- Conference
- 2018 IEEE Spoken Language Technology Workshop (SLT)
- Pages
- 1036-1043
- Publisher
- IEEE