TR9856: A multi-word term relatedness benchmark

Ran Levy; Liat Ein-Dor; Shay Hummel; Ruty Rinott; Noam Slonim

doi:10.3115/v1/p15-2069

ACL-IJCNLP 2015

Conference paper

26 Jul 2015

TR9856: A multi-word term relatedness benchmark

View publication

Abstract

Measuring word relatedness is an impor-tant ingredient of many NLP applications. Several datasets have been developed in order to evaluate such measures. The main drawback of existing datasets is the fo-cus on single words, although natural lan-guage contains a large proportion of multi-word terms. We propose the new TR9856 dataset which focuses on multi-word terms and is significantly larger than existing datasets. The new dataset includes many real world terms such as acronyms and named entities, and further handles term ambiguity by providing topical context for all term pairs. We report baseline results for common relatedness methods over the new data, and exploit its magni-tude to demonstrate that a combination of these methods outperforms each individ-ual method.

Paper