Automatic generation and selection of multiple pronunciations for dynamic vocabularies

S. Deligne; B. Maison; R.A. Gopinath

ICASSP 2001

Conference paper

26 Sep 2001

Automatic generation and selection of multiple pronunciations for dynamic vocabularies

Abstract

In this paper, we present a new scheme for the acoustic modeling of speech recognition applications requiring dynamic vocabularies. It applies especially to the acoustic modeling of out-of-vocabulary words which need to be added to a recognition lexicon based on the observation of a few (say one or two) speech utterances of these words. Standard approaches to this problem derive a single pronunciation from each speech utterance by combining acoustic and phone transition scores. In our scheme, multiple pronunciations are generated from each speech utterance of a word to enroll by varying the relative weights assigned to the acoustic and phone transition models. In our experiments, the use of these multiple baseforms dramatically outperforms the standard approach with a relative decrease of the word error rate ranging from 20% to 40% on all our test sets.

Conference paper