Arabic Named Entity Recognition: Using features extracted from noisy data

Yassine Benajiba; Imed Zitouni; Mona Diab; Paolo Rosso

ACL 2010

Conference paper

01 Dec 2010

Arabic Named Entity Recognition: Using features extracted from noisy data

Abstract

Building an accurate Named Entity Recognition (NER) system for languages with complex morphology is a challenging task. In this paper, we present research that explores the feature space using both gold and bootstrapped noisy features to build an improved highly accurate Arabic NER system. We bootstrap noisy features by projection from an Arabic-English parallel corpus that is automatically tagged with a baseline NER system. The feature space covers lexical, morphological, and syntactic features. The proposed approach yields an improvement of up to 1.64 F-measure (absolute). © 2010 Association for Computational Linguistics.

Conference paper