A simple sentence-level extraction algorithm for comparable data

Christoph Tillmann; Jian-Ming Xu

NAACL-HLT 2009

Conference paper

31 May 2009

A simple sentence-level extraction algorithm for comparable data

Abstract

The paper presents a novel sentence pair extraction algorithm for comparable data, where a large set of candidate sentence pairs is scored directly at the sentence-level. The sentence-level extraction relies on a very efficient implementation of a simple symmetric scoring function: a computation speed-up by a factor of 30 is reported. On Spanish-English data, the extraction algorithm finds the highest scoring sentence pairs from close to 1 trillion candidate pairs without search errors. Significant improvements in BLEU are reported by including the extracted sentence pairs into the training of a phrase-based SMT (Statistical Machine Translation) system.

Conference paper