Publications
4.4 Unsupervised Learning for Audio
Abstract
After a more fundamental discussion on unsupervised learning for audio, our group decided to focus on the use of unsupervised learning in a concrete application scenario. Even though unsupervised learning could be used in many processing stages of a computational audio analysis system (eg, to develop the feature extraction front-end), in practical scenarios one often takes advantage of some prior information. In particular, defining a specific application and ways to evaluate the performance of the audio processing system already imposes some prior information.
We considered an application where the goal is to detect car crash sounds from continuous audio recordings. In the unsupervised learning scenario, one has audio recordings that can be used as a training data, but no reference times of crashes are actually annotated. There are several kinds of prior information that can be used in the given application. First, one may assume that the events one is looking for can be characterizing by a specific set of audio features such as MFCCs. Second, one may assume to have a metric that is appropriate for describing distances between acoustic features. Third, we know that events are rare, ie, only a small number of target events are present in the training data. Finally, we assume that each event is localized in time and the duration of events is approximately known (eg one or two seconds). The above prior information can be used for novelty-based audio segmentation using the calculated features and the distance metric. Alternatively, unsupervised learning can be used to learn features and a distance metric. The resulting segments can be …
- Date
- 2014
- Authors
- Tuomas Virtanen, Jon Barker, Shrikanth Narayanan, Alexandros Potamianos, Bhiksha Raj, Gaël Richard, Rita Singh, Paris Smaragdis, Stefano Squartini, Shiva Sundaram
- Journal
- Computational Audio Analysis
- Pages
- 14