Removing data with noisy responses in regression analysis

Alan Wisler; Visar Berisha; Karthikeyan Natesan Ramamurthy; Andreas Spanias; Julie Liss

doi:10.1109/ICASSP.2015.7178334

ICASSP 2015

Conference paper

04 Aug 2015

Removing data with noisy responses in regression analysis

View publication

Abstract

In regression analysis, outliers in the data can induce a bias in the learned function, resulting in larger errors. In this paper we derive an empirically estimable bound on the regression error based on a Euclidean minimum spanning tree generated from the data. Using this bound as motivation, we propose an iterative approach to remove data with noisy responses from the training set. We evaluate the performance of the algorithm on experiments with real-world pathological speech (speech from individuals with neurogenic disorders). Comparative results show that removing noisy examples during training using the proposed approach yields better predictive performance on out-of-sample data.

Conference paper