Managing data quality by identifying the noisiest data samples

K. Hima Prasad; Snigdha Chaturvedi; Tanveer A. Faruquie; L. Venkata Subramaniam; Mukesh K. Mohania

doi:10.1109/SOLI.2012.6273510

SOLI 2012

Conference paper

12 Oct 2012

Managing data quality by identifying the noisiest data samples

View publication

Abstract

Enterprise datasets are often noisy. Several columns can have non-standard, erroneous or missing information. Poor quality data can lead to incorrect reporting and wrong conclusions being drawn. Data cleansing involves standardizing such data to improve its quality. Often data cleansing tasks involve writing rules manually. The step involves understanding the data quality issues and then writing data transformation rules to correct these issues. This is a human intensive task. In this study we propose a method to identify noisy subsets of huge unlabelled textual datasets. This is a two step process where in the first step we develop an estimation tool to predict the data quality on an unlabelled text dataset as produced by a segmentation model. The accuarcy of the proposed method is shown on a real life dataset. © 2012 IEEE.

Conference paper