Sequential sampling procedures for query size estimation

Peter J. Haas; Arun N. Swami

doi:10.1145/130283.130321

SIGMOD 1992

Conference paper

02 Jun 1992

Sequential sampling procedures for query size estimation

Download paper

Abstract

We provide a procedure, based on random sampling for estimation of the size of a query result. The proceterminates is sequential in that sampling terminates after a random number of steps according to a stopping rule that depends upon the observations obtained so far. Enough observations are obtained so that, with a prespecified probability, the estimate differs from the true size of the query result by no more than a prespecified amount. Unlike previous sequential estimation procedures for queries, our procedure is asymptotically efficient and requires no ad hoc pilot sample or a priori assumptions about data characteristics. In addition to establishing the asymptotic properties of the estimation procedure, we provide techniques for reducing undercoverage at small sample sizes and show that the sampling cost of the procedure can be reduced through stratified sampling techniques.

Conference paper