{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:37:58Z","timestamp":1777696678690,"version":"3.51.4"},"reference-count":32,"publisher":"SAGE Publications","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IDA"],"published-print":{"date-parts":[[2024,9,19]]},"abstract":"<jats:p>Histograms are among the most popular methods used in exploratory analysis to summarize univariate distributions. In particular, irregular histograms are good non-parametric density estimators that require very few parameters: the number of bins with their lengths and frequencies. Although many approaches have been proposed in the literature to infer these parameters, most existing histogram methods are difficult to exploit for exploratory analysis in the case of real-world data sets, with scalability issues, truncated data, outliers or heavy-tailed distributions. In this paper, we focus on the G-Enum histogram method, which exploits the Minimum Description Length (MDL) principle to build histograms without any user parameter. We then propose to extend this method by exploiting a new modeling space based on floating-point representation, with the objective of building histograms resistant to outliers or heavy-tailed distributions. We also suggest several heuristics and a methodology suitable for the exploratory analysis of large scale real-world data sets, whose underlying patterns are difficult to recover for digitization reasons. Extensive experiments show the benefits of the approach, evaluated with a dual objective: the accuracy of density estimation in the case of outliers or heavy-tailed distributions, and the effectiveness of the approach for exploratory data analysis.<\/jats:p>","DOI":"10.3233\/ida-230638","type":"journal-article","created":{"date-parts":[[2024,2,2]],"date-time":"2024-02-02T10:46:24Z","timestamp":1706870784000},"page":"1347-1394","source":"Crossref","is-referenced-by-count":0,"title":["Floating-point histograms for exploratory analysis of large scale real-world data sets"],"prefix":"10.1177","volume":"28","author":[{"given":"Marc","family":"Boull\u00e9","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"issue":"2","key":"10.3233\/IDA-230638_ref1","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1109\/18.119689","article-title":"Density estimation by stochastic complexity","volume":"38","author":"Rissanen","year":"1992","journal-title":"IEEE Transactions on Information Theory"},{"key":"10.3233\/IDA-230638_ref2","unstructured":"P. Kontkanen and P. Myllym\u00c3\u00a4ki, MDL Histogram Density Estimation, in: Proceedings of the Eleventh International Conference on Artificial Intelligence and Statistics, M. Meila and X. Shen, eds, Proceedings of Machine Learning Research, Vol. 2, PMLR, 2007, pp.\u00a0219\u2013226."},{"issue":"3","key":"10.3233\/IDA-230638_ref3","doi-asserted-by":"publisher","first-page":"1093","DOI":"10.1214\/009053604000000364","article-title":"Densities, spectral densities and modality","volume":"32","author":"Davies","year":"2004","journal-title":"Ann. Statist."},{"issue":"12","key":"10.3233\/IDA-230638_ref4","doi-asserted-by":"publisher","first-page":"3313","DOI":"10.1016\/j.csda.2010.04.021","article-title":"Combining regular and irregular histograms by penalized likelihood","volume":"54","author":"Rozenholc","year":"2010","journal-title":"Computational Statistics and Data Analysis"},{"issue":"2","key":"10.3233\/IDA-230638_ref5","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1088\/0004-637x\/764\/2\/167","article-title":"Studies in astronomical time series analysis. vi. bayesian block representations","volume":"764","author":"Scargle","year":"2013","journal-title":"The Astrophysical Journal"},{"key":"10.3233\/IDA-230638_ref6","doi-asserted-by":"publisher","first-page":"107668","DOI":"10.1016\/j.csda.2022.107668","article-title":"Fast and fully-automated histograms for large-scale data sets","volume":"180","author":"Zelaya Mendiz\u00e1bal","year":"2023","journal-title":"Computational Statistics & Data Analysis"},{"issue":"2","key":"10.3233\/IDA-230638_ref7","doi-asserted-by":"publisher","first-page":"416","DOI":"10.1214\/aos\/1176346150","article-title":"A universal prior for integers and estimation by minimum description length","volume":"11","author":"Rissanen","year":"1983","journal-title":"Ann. Statist."},{"key":"10.3233\/IDA-230638_ref9","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1051\/ps:2008005","article-title":"A comparison of automatic histogram constructions","volume":"13","author":"Davies","year":"2009","journal-title":"ESAIM: Probability and Statistics"},{"issue":"4","key":"10.3233\/IDA-230638_ref10","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1007\/BF01025868","article-title":"On the histogram as a density estimator: L2 theory","volume":"57","author":"Freedman","year":"1981","journal-title":"Zeitschrift f\u00fcr Wahrscheinlichkeitstheorie und Verwandte Gebiete"},{"key":"10.3233\/IDA-230638_ref11","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1023\/B:AIRE.0000045502.10941.a9","article-title":"A survey of outlier detection methodologies","volume":"22","author":"Hodge","year":"2004","journal-title":"Artificial Intelligence Review"},{"issue":"1","key":"10.3233\/IDA-230638_ref12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4108\/trans.sis.2013.01-03.e2","article-title":"Advancements of outlier detection: A survey","volume":"13","author":"Zhang","year":"2013","journal-title":"ICST Transactions on Scalable Information Systems"},{"key":"10.3233\/IDA-230638_ref13","unstructured":"M. Gebski and R.K. Wong, An efficient histogram method for outlier detection, in: Advances in Databases: Concepts, Systems and Applications: 12th International Conference on Database Systems for Advanced Applications, DASFAA 2007, Bangkok, Thailand, April 9\u201312, 2007. Proceedings 12, Springer, 2007, pp.\u00a0176\u2013187."},{"key":"10.3233\/IDA-230638_ref15","doi-asserted-by":"publisher","DOI":"10.1109\/IEEESTD.1985.82928"},{"key":"10.3233\/IDA-230638_ref17","doi-asserted-by":"crossref","unstructured":"K.H. Knuth, J.P. Castle and K.R. Wheeler, Identifying excessively rounded or truncated data, in: Compstat 2006 \u2013 Proceedings in Computational Statistics, A. Rizzi and M. Vichi, eds, Springer, 2006, pp.\u00a0313\u2013323.","DOI":"10.1007\/978-3-7908-1709-6_24"},{"issue":"3","key":"10.3233\/IDA-230638_ref18","first-page":"357","article-title":"A look at some data on the old faithful geyser","volume":"39","author":"Azzalini","year":"1990","journal-title":"Journal of the Royal Statistical Society. Series C (Applied Statistics)"},{"key":"10.3233\/IDA-230638_ref19","doi-asserted-by":"crossref","unstructured":"A. Marx, L. Yang and M. van Leeuwen, Estimating Conditional Mutual Information for Discrete-Continuous Mixtures using Multi-Dimensional Adaptive Histograms, in: Proceedings of the 2021 SIAM International Conference on Data Mining, SDM 2021, Virtual Event, April 29\u2013May 1, 2021, C. Demeniconi and I. Davidson, eds, SIAM, 2021, pp.\u00a0387\u2013395.","DOI":"10.1137\/1.9781611976700.44"},{"issue":"1","key":"10.3233\/IDA-230638_ref21","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1016\/j.pss.2011.09.003","article-title":"LU60645GT and MA132843GT catalogues of Lunar and Martian impact craters developed using a Crater Shape-based interpolation crater detection algorithm for topography data","volume":"60","author":"Salamuni\u0107car","year":"2012","journal-title":"Planetary and Space Science"},{"issue":"4","key":"10.3233\/IDA-230638_ref22","doi-asserted-by":"publisher","first-page":"871","DOI":"10.1029\/2018JE005592","article-title":"A New Global Database of Lunar Impact Craters > 1\u20132 km: 1. Crater Locations and Sizes, Comparisons With Published Databases, and Global Analysis","volume":"124","author":"Robbins","year":"2019","journal-title":"Journal of Geophysical Research (Planets)"},{"issue":"12","key":"10.3233\/IDA-230638_ref23","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1088\/1674-4527\/16\/12\/185","article-title":"Determining proportions of lunar crater populations by fitting crater size distribution","volume":"16","author":"Wang","year":"2016","journal-title":"Research in Astronomy and Astrophysics"},{"key":"10.3233\/IDA-230638_ref24","doi-asserted-by":"crossref","unstructured":"P. Boldi and S. Vigna, The WebGraph Framework I: Compression Techniques, in: Proc. of the Thirteenth International World Wide Web Conference (WWW 2004), ACM Press, 2004, pp.\u00a0595\u2013601.","DOI":"10.1145\/988672.988752"},{"key":"10.3233\/IDA-230638_ref25","doi-asserted-by":"crossref","unstructured":"P. Boldi, M. Rosa, M. Santini and S. Vigna, Layered Label Propagation: A MultiResolution Coordinate-Free Ordering for Compressing Social Networks, in: Proceedings of the 20th international conference on World Wide Web, S. Srinivasan, K. Ramamritham, A. Kumar, M.P. Ravindra, E. Bertino and R. Kumar, eds, ACM Press, 2011, pp.\u00a0587\u2013596.","DOI":"10.1145\/1963405.1963488"},{"key":"10.3233\/IDA-230638_ref27","unstructured":"B. Mandelbrot, An Information Theory of the Statistical Structure of Language, in: Communication Theory, Academic Press, 1953, pp.\u00a0486\u2013502."},{"issue":"5","key":"10.3233\/IDA-230638_ref28","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1080\/00107510500052444","article-title":"Power laws, Pareto distributions and Zipf\u2019s law","volume":"46","author":"Newman","year":"2005","journal-title":"Contemporary Physics"},{"key":"10.3233\/IDA-230638_ref32","unstructured":"T. Mononen and P. Myllym\u00e4ki, Computing the multinomial stochastic complexity in sub-linear time, in: Proceedings of the 4th European Workshop on Probabilistic Graphical Models (PGM-08), September 17\u201319, 2008, Hirtshals, Denmark, 2008, pp.\u00a0209\u2013216, Volume: Proceeding volume."},{"issue":"1","key":"10.3233\/IDA-230638_ref33","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/18.481776","article-title":"Fisher information and stochastic complexity","volume":"42","author":"Rissanen","year":"1996","journal-title":"IEEE Transactions on Information Theory"},{"issue":"2","key":"10.3233\/IDA-230638_ref34","first-page":"142","article-title":"On asymptotics of certain recurrences arising in universal coding","volume":"34","author":"Szpankowski","year":"1998","journal-title":"Problems of Information Transmission"},{"issue":"1","key":"10.3233\/IDA-230638_ref37","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1007\/s10994-006-8364-x","article-title":"MODL: a Bayes optimal discretization method for continuous attributes","volume":"65","author":"Boull\u00e9","year":"2006","journal-title":"Machine Learning"},{"issue":"1","key":"10.3233\/IDA-230638_ref38","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1109\/18.481776","article-title":"Fisher information and stochastic complexity","volume":"42","author":"Rissanen","year":"1996","journal-title":"IEEE Transactions on Information Theory"},{"key":"10.3233\/IDA-230638_ref40","unstructured":"P. Kontkanen, W.L. Buntine, P. Myllym\u00e4ki, J. Rissanen and H. Tirri, Efficient Computing of Stochastic Complexity, in: Proceedings of the Ninth International Workshop on Artificial Intelligence and Statistics, AISTATS 2003, Key West, Florida, USA, January 3\u20136, 2003, 2003. http:\/\/research.microsoft.com\/en-us\/um\/cambridge\/events\/aistats2003\/proceedings\/172.pdf."},{"key":"10.3233\/IDA-230638_ref41","doi-asserted-by":"crossref","first-page":"465","DOI":"10.1016\/0005-1098(78)90005-5","article-title":"Modeling by shortest data description","volume":"14","author":"Rissanen","year":"1978","journal-title":"Automatica"},{"issue":"7825","key":"10.3233\/IDA-230638_ref44","doi-asserted-by":"publisher","first-page":"357","DOI":"10.1038\/s41586-020-2649-2","article-title":"Array programming with NumPy","volume":"585","author":"Harris","year":"2020","journal-title":"Nature"},{"issue":"2","key":"10.3233\/IDA-230638_ref45","doi-asserted-by":"publisher","first-page":"167","DOI":"10.3847\/1538-4357\/ac7c74","article-title":"The Astropy Project: Sustaining and Growing a Community-oriented Open-source Project and the Latest Major Release (v5.0) of the Core Package","volume":"935","author":"Price-Whelan","year":"2022","journal-title":"apj"}],"container-title":["Intelligent Data Analysis"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/IDA-230638","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:20:41Z","timestamp":1777454441000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.medra.org\/servlet\/aliasResolver?alias=iospress&doi=10.3233\/IDA-230638"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,19]]},"references-count":32,"journal-issue":{"issue":"5"},"URL":"https:\/\/doi.org\/10.3233\/ida-230638","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,19]]}}}