Document clustering via adaptive subspace iteration

Tao Li; Sheng Ma; Mitsunori Ogihara

doi:10.1145/1008992.1009031

SIGIR 2004

Conference paper

25 Jul 2004

Document clustering via adaptive subspace iteration

View publication

Abstract

Document clustering has long been an important problem in information retrieval. In this paper, we present a new clustering algorithm ASI1, which uses explicitly modeling of the subspace structure associated with each cluster. ASI simultaneously performs data reduction and subspace identification via an iterative alternating optimization procedure. Motivated from the optimization procedure, we then provide a novel method to determine the number of clusters. We also discuss the connections of ASI with various existential clustering approaches. Finally, extensive experimental results on real data sets show the effectiveness of ASI algorithm.

Conference paper