Audio-visual speech synchronization detection using a bimodal linear prediction model

Kshitiz Kumar; Jiri Navratil; Etienne Marcheret; Vit Libal; Ganesh Ramaswamy; Gerasimos Potamianos

doi:10.1109/CVPR.2009.5204303

CVPRW 2009

Conference paper

20 Jun 2009

Audio-visual speech synchronization detection using a bimodal linear prediction model

View publication

Abstract

In this work, we study the problem of detecting audiovisual (AV) synchronization in video segments containing a speaker in frontal head pose. The problem holds important applications in biometrics, for example spoofing detection, and it constitutes an important step in AV segmentation necessary for deriving AV fingerprints in multimodal speaker recognition. To attack the problem, we propose a timeevolution model for AV features and derive an analytical approach to capture the notion of synchronization between them. We report results on an appropriate AV database, using two types of visual features extracted from the speaker's facial area: geometric ones and features based on the discrete cosine image transform. Our results demonstrate that the proposed approach provides substantially better AV synchrony detection over a baseline method that employs mutual information, with the geometric visual features outperforming the image transform ones. Audio-Visual Synchronization, Mutual Information, Linear Prediction, Visual Features © 2009 IEEE.

Conference paper