IBM MASTOR: Multilingual automatic speech-to-speech translator
Yuqing Gao, Bowen Zhou, et al.
ICASSP 2006
This paper presents the theoretical framework of a new statistical model for phoneme recognition. In contrast with traditional HMMs, the posterior probability of a state sequence given an observation sequence is computed directly with the new model. The development of this paper is based on Maximum Entropy Markov Models (MEMMs[5]), appearing as a result of the application of Maximum Entropy principle to sequential processes. The main contributions of our work include modifying the MEMM to large-scale speech recognition problem and introduction of another direct model (NDM), which overcomes the shortcome of the MEMM of poor representation of contextual information. Direct comparison of direct model phoneme recognizers with HMM-based recognizers demonstrates the superiority of the new models, particularly on smaller training sets.
Yuqing Gao, Bowen Zhou, et al.
ICASSP 2006
Hakan Erdogan, Ruhi Sarikaya, et al.
ICSLP 2002
Liang Gu, Yonggang Deng, et al.
SLT 2006
Ruhi Sarikaya, Yonggang Deng, et al.
INTERSPEECH 2008