Publications
Multimodal representation of advertisements using segment-level autoencoders
Abstract
Automatic analysis of advertisements (ads) poses an interesting problem for learning multimodal representations. A promising direction of research is the development of deep neural network autoencoders to obtain inter-modal and intra-modal representations. In this work, we propose a system to obtain segment-level unimodal and joint representations. These features are concatenated, and then averaged across the duration of an ad to obtain a single multimodal representation. The autoencoders are trained using segments generated by time-aligning frames between the audio and video modalities with forward and backward context. In order to assess the multimodal representations, we consider the tasks of classifying an ad as funny or exciting in a publicly available dataset of 2,720 ads. For this purpose we train the segment-level autoencoders on a larger, unlabeled dataset of 9,740 ads, agnostic of the test set …
- Date
- 2018
- Authors
- Krishna Somandepalli, Victor Martinez, Naveen Kumar, Shrikanth Narayanan
- Book
- Proceedings of the 20th ACM International Conference on Multimodal Interaction
- Pages
- 418-422