Interpreting Data from Scanned Tables

Waleed Farrukh; Antonio Foncubierta-Rodriguez; Anca Nicoleta Ciubotaru; Guillaume Jaume; Costas Bejas; Orcun Goksel; Maria Gabrani

doi:10.1109/ICDAR.2017.250

GREC 2017

Conference paper

25 Jan 2018

Interpreting Data from Scanned Tables

View publication

Abstract

Densely-packed but structured scientific data are typically presented in the form of tables, which often appear in raster image form. To interpret data from scanned tables, understanding their hierarchical structure is vital. To further address the vast variability of table representations, we propose a fully automatic methodology that uses a bottom-up reasoning that is independent on the existence of representation features, such as lines. We evaluate our approach on the ICDAR 2013 dataset and demonstrate its effectiveness on detecting tables cells and their content and classifying header and data cells. For detecting the cell hierarchy, we demonstrate results on synthetic data due to lack of ground truth.

Conference paper