{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T09:56:59Z","timestamp":1773482219211,"version":"3.50.1"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"14","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2013,9]]},"abstract":"<jats:p>As of 2005, sampling has been incorporated in all major database systems. While efficient sampling techniques are realizable, determining the accuracy of an estimate obtained from the sample is still an unresolved problem. In this paper, we present a theoretical framework that allows an elegant treatment of the problem. We base our work on generalized uniform sampling (GUS), a class of sampling methods that subsumes a wide variety of sampling techniques. We introduce a key notion of equivalence that allows GUS sampling operators to commute with selection and join, and derivation of confidence intervals. We illustrate the theory through extensive examples and give indications on how to use it to provide meaningful estimates in database systems.<\/jats:p>","DOI":"10.14778\/2556549.2556563","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"1798-1809","source":"Crossref","is-referenced-by-count":17,"title":["A sampling algebra for aggregate estimation"],"prefix":"10.14778","volume":"6","author":[{"given":"Supriya","family":"Nirkhiwale","sequence":"first","affiliation":[{"name":"University of Florida"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alin","family":"Dobra","sequence":"additional","affiliation":[{"name":"University of Florida"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher","family":"Jermaine","sequence":"additional","affiliation":[{"name":"Rice University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Sql-2003 standard 2003."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/645925.671347"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304581"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304207"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807224"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-47534-9_7"},{"key":"e_1_2_1_8_1","unstructured":"T.-H. Benchmark. http:\/\/www.tpc.org\/tpch\/."},{"issue":"4","key":"e_1_2_1_9_1","first-page":"41","article-title":"On sampling and relational operators","volume":"22","author":"Chaudhuri S.","year":"1999","unstructured":"S. Chaudhuri and R. Motwani. On sampling and relational operators. IEEE Data Eng. Bull., 22(4):41-46, 1999.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304206"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000006"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687675"},{"key":"e_1_2_1_13_1","volume-title":"Aqua: System and techniques for approximate query answering. Technical report","author":"Gibbons P. B.","year":"1998","unstructured":"P. B. Gibbons, V. Poosala, S. Acharya, Y. Bartal, Y. Matias, S. Muthukrishnan, S. Ramaswamy, and T. Suel. Aqua: System and techniques for approximate query answering. Technical report, 1998."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007664"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/646496.695465"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/304181.304208"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/375663.375800"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1996.0041"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1996.0041"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/253262.253291"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1412331.1412335"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1412331.1412335"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1189769.1189775"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/0304-3975(93)90224-H"},{"key":"e_1_2_1_25_1","volume-title":"Random sampling from databases","author":"Olken F.","year":"1993","unstructured":"F. Olken. Random sampling from databases, 1993."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/602259.602294"},{"key":"e_1_2_1_27_1","unstructured":"SWI-Prolog. http:\/\/www.swi-prolog.org."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2556549.2556563","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,23]],"date-time":"2024-10-23T22:34:34Z","timestamp":1729722874000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2556549.2556563"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,9]]},"references-count":26,"journal-issue":{"issue":"14","published-print":{"date-parts":[[2013,9]]}},"alternative-id":["10.14778\/2556549.2556563"],"URL":"https:\/\/doi.org\/10.14778\/2556549.2556563","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2013,9]]}}}