{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T01:16:51Z","timestamp":1778807811611,"version":"3.51.4"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2008,8]]},"abstract":"<jats:p>\n            Summaries of massive data sets support approximate query processing over the original data. A basic aggregate over a set of records is the weight of subpopulations specified as a predicate over records' attributes.\n            <jats:italic>Bottom-k<\/jats:italic>\n            sketches are a powerful summarization format of weighted items that includes priority sampling [22], and the classic weighted sampling without replacement. They can be computed efficiently for many representations of the data including distributed databases and data streams and support coordinated and all-distances sketches.\n          <\/jats:p>\n          <jats:p>\n            We derive novel unbiased estimators and confidence bounds for subpopulation weight. Our\n            <jats:italic>rank conditioning<\/jats:italic>\n            (RC) estimator is applicable when the total weight of the sketched set cannot be computed by the summarization algorithm without a significant use of additional resources (such as for sketches of network neighborhoods) and the tighter\n            <jats:italic>subset conditioning<\/jats:italic>\n            (SC) estimator that is applicable when the total weight is available (sketches of data streams).\n          <\/jats:p>\n          <jats:p>\n            Our estimators are derived using clever applications of the Horvitz-Thompson estimator (that is not directly applicable to bottom-\n            <jats:italic>k<\/jats:italic>\n            sketches). We develop efficient computational methods and conduct performance evaluation using a range of synthetic and real data sets. We demonstrate considerable benefits of the SC estimator on larger subpopulations (over all other estimators); of the RC estimator (over existing estimators for weighted sampling without replacement); and of our confidence bounds (over all previous approaches).\n          <\/jats:p>","DOI":"10.14778\/1453856.1453884","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"213-224","source":"Crossref","is-referenced-by-count":46,"title":["Tighter estimation using bottom k sketches"],"prefix":"10.14778","volume":"1","author":[{"given":"Edith","family":"Cohen","sequence":"first","affiliation":[{"name":"AT&amp;T Labs-Research, Florham Park, NJ"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haim","family":"Kaplan","sequence":"additional","affiliation":[{"name":"Tel Aviv University, Tel Aviv, Israel"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2008,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065167.1065209"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1247480.1247504"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1016\/S1389-1286(99)00021-3"},{"key":"e_1_2_1_4_1","first-page":"21","volume-title":"Proceedings of the Compression and Complexity of Sequences","author":"Broder A. Z.","year":"1997","unstructured":"A. Z. Broder . On the resemblance and containment of documents . In Proceedings of the Compression and Complexity of Sequences , pages 21 -- 29 . ACM, 1997 . A. Z. Broder. On the resemblance and containment of documents. In Proceedings of the Compression and Complexity of Sequences, pages 21--29. ACM, 1997."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/647819.736184"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1997.1534"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1298306.1298344"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1265530.1265566"},{"key":"e_1_2_1_10_1","volume-title":"Summarization framework for unaggregated data. Submitted","author":"Cohen E.","year":"2008","unstructured":"E. Cohen , N. Duffield , C. Lund , M. Thorup , and H. Kaplan . Summarization framework for unaggregated data. Submitted , 2008 . E. Cohen, N. Duffield, C. Lund, M. Thorup, and H. Kaplan. Summarization framework for unaggregated data. Submitted, 2008."},{"key":"e_1_2_1_11_1","volume-title":"Proc. 15th ACM-SIAM Symposium on Discrete Algorithms. ACM-SIAM","author":"Cohen E.","year":"2004","unstructured":"E. Cohen and H. Kaplan . Efficient estimation algorithms for neighborhood variance and other moments . In Proc. 15th ACM-SIAM Symposium on Discrete Algorithms. ACM-SIAM , 2004 . E. Cohen and H. Kaplan. Efficient estimation algorithms for neighborhood variance and other moments. In Proc. 15th ACM-SIAM Symposium on Discrete Algorithms. ACM-SIAM, 2004."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007647"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1254882.1254926"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcss.2006.10.016"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281100.1281133"},{"key":"e_1_2_1_16_1","volume-title":"Manuscript","author":"Cohen E.","year":"2008","unstructured":"E. Cohen and H. Kaplan . Estimating aggregates over multiple subsets . Manuscript , 2008 . E. Cohen and H. Kaplan. Estimating aggregates over multiple subsets. Manuscript, 2008."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/773153.773175"},{"key":"e_1_2_1_19_1","first-page":"66","volume-title":"Proc. Pacific Rim International Symposium on Fault-Tolerant Systems","author":"Cohen E.","year":"1995","unstructured":"E. Cohen , Y.-M. Wang , and G. Suri . When piecewise determinism is almost true . In Proc. Pacific Rim International Symposium on Fault-Tolerant Systems , pages 66 -- 71 , Dec. 1995 . E. Cohen, Y.-M. Wang, and G. Suri. When piecewise determinism is almost true. In Proc. Pacific Rim International Symposium on Fault-Tolerant Systems, pages 66--71, Dec. 1995."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/564691.564719"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2005.846400"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1314690.1314696"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipl.2005.11.003"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/276304.276334"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1952.10483446"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 33rd VLDB Conference","author":"Hua M.","year":"2007","unstructured":"M. Hua , J. Pei , A. W. C. Fu , X. Lin , and H.-F. Leung . Efficiently answering top-k typicality queries on large databases . In Proceedings of the 33rd VLDB Conference , 2007 . M. Hua, J. Pei, A. W. C. Fu, X. Lin, and H.-F. Leung. Efficiently answering top-k typicality queries on large databases. In Proceedings of the 33rd VLDB Conference, 2007."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/1109557.1109611"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1146381.1146401"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.908981"},{"key":"e_1_2_1_30_1","unstructured":"The Netflix Prize http:\/\/www.netflixprize.com\/.  The Netflix Prize http:\/\/www.netflixprize.com\/."},{"key":"e_1_2_1_31_1","unstructured":"Cisco NetFlow. http:\/\/www.cisco.com\/warp\/public\/732\/Tech\/netflow.  Cisco NetFlow. http:\/\/www.cisco.com\/warp\/public\/732\/Tech\/netflow."},{"key":"e_1_2_1_32_1","volume-title":"Sampling Theory and Methods","author":"Sampath S.","year":"2000","unstructured":"S. Sampath . Sampling Theory and Methods . CRC press , 2000 . S. Sampath. Sampling Theory and Methods. CRC press, 2000."},{"key":"e_1_2_1_33_1","volume-title":"Theory, Practice and Visualization","author":"Scott D. W.","year":"1992","unstructured":"D. W. Scott . Multivariate Density Estimation : Theory, Practice and Visualization . John Wiley & Sons , New York , 1992 . D. W. Scott. Multivariate Density Estimation: Theory, Practice and Visualization. John Wiley & Sons, New York, 1992."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-017-1404-4"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/347059.347408"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1132516.1132539"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1140103.1140307"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/1453856.1453884","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:02:07Z","timestamp":1672225327000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/1453856.1453884"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,8]]},"references-count":35,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2008,8]]}},"alternative-id":["10.14778\/1453856.1453884"],"URL":"https:\/\/doi.org\/10.14778\/1453856.1453884","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2008,8]]}}}