{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:08:25Z","timestamp":1750306105954,"version":"3.41.0"},"reference-count":13,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2017,5,11]],"date-time":"2017-05-11T00:00:00Z","timestamp":1494460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMOD Rec."],"published-print":{"date-parts":[[2017,5,11]]},"abstract":"<jats:p>Variance is a popular and often necessary component of aggregation queries. It is typically used as a secondary measure to ascertain statistical properties of the result such as its error. Yet, it is more expensive to compute than primary measures such as SUM, MEAN, and COUNT.<\/jats:p>\n          <jats:p>There exist numerous techniques to compute variance. While the definition of variance implies two passes overthe data, other mathematical formulations lead to a singlepass computation. Some single-pass formulations, however, can suffer from severe precision loss, especially for large datasets.<\/jats:p>\n          <jats:p>In this paper, we study variance implementations in various real-world systems and find that major database systems such as PostgreSQL and most likely System X, a major commercial closed-source database, use a representation that is efficient, but suffers from floating point precision loss resulting from catastrophic cancellation. We review literature over the past five decades on variance calculation in both the statistics and database communities, and summarize recommendations on implementing variance functions in various settings, such as approximate query processing and large-scale distributed aggregation. Interestingly, we recommend using the mathematical formula for computing variance if two passes over the data are acceptable due to its precision, parallelizability, and surprisingly computation speed.<\/jats:p>","DOI":"10.1145\/3092931.3092936","type":"journal-article","created":{"date-parts":[[2017,5,12]],"date-time":"2017-05-12T12:16:12Z","timestamp":1494591372000},"page":"28-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["A Closer Look at Variance Implementations In Modern Database Systems"],"prefix":"10.1145","volume":"45","author":[{"given":"Niranjan","family":"Kamat","sequence":"first","affiliation":[{"name":"The Ohio State University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arnab","family":"Nandi","sequence":"additional","affiliation":[{"name":"The Ohio State University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,5,11]]},"reference":[{"volume-title":"Am. Stat.","year":"1983","author":"T. Chan","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/360569.360662"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367510"},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","DOI":"10.1137\/1.9780898718027","volume-title":"Accuracy and Stability of Numerical Algorithms","author":"Higham N. J.","year":"2002"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/363707.363723"},{"volume-title":"Arxiv TR","year":"2016","author":"Kamat N.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254556.2254659"},{"volume-title":"Volume 2: Seminumerical Algorithms","year":"2014","author":"Knuth D. E.","key":"e_1_2_1_8_1"},{"volume-title":"StRD: Statistical Reference Datasets for Testing the Numerical Accuracy of Statistical Software","year":"1998","author":"Rogers J.","key":"e_1_2_1_9_1"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.12"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1962.10490022"},{"volume-title":"An Improved Method","year":"1979","author":"West D.","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1971.10488826"}],"container-title":["ACM SIGMOD Record"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3092931.3092936","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3092931.3092936","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:39Z","timestamp":1750217439000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3092931.3092936"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,5,11]]},"references-count":13,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,5,11]]}},"alternative-id":["10.1145\/3092931.3092936"],"URL":"https:\/\/doi.org\/10.1145\/3092931.3092936","relation":{},"ISSN":["0163-5808"],"issn-type":[{"type":"print","value":"0163-5808"}],"subject":[],"published":{"date-parts":[[2017,5,11]]},"assertion":[{"value":"2017-05-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}