{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:33:41Z","timestamp":1750307621146,"version":"3.41.0"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2009,1,1]],"date-time":"2009-01-01T00:00:00Z","timestamp":1230768000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001868","name":"National Science Council Taiwan","doi-asserted-by":"publisher","award":["NSC95-2752-E-002-006-PAE"],"award-info":[{"award-number":["NSC95-2752-E-002-006-PAE"]}],"id":[{"id":"10.13039\/501100001868","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2009,1]]},"abstract":"<jats:p>In this article, we explore a novel sampling model, called<jats:italic>feature preserved sampling<\/jats:italic>(<jats:italic>FPS<\/jats:italic>) that sequentially generates a high-quality sample over sliding windows. The sampling quality we consider refers to the degree of consistency between the sample proportion and the population proportion of each attribute value in a window. Due to the time-variant nature of real-world datasets, users are more likely to be interested in the most recent data. However, previous works have not been able to generate a high-quality sample over sliding windows that precisely preserves up-to-date population characteristics. Motivated by this shortcoming, we have developed the<jats:italic>FPS<\/jats:italic>algorithm, which has several advantages: (1) it sequentially generates a sample from a time-variant data source over sliding windows; (2) the execution time of<jats:italic>FPS<\/jats:italic>is linear with respect to the database size; (3) the<jats:italic>relative<\/jats:italic>proportional differences between the sample proportions and population proportions of most distinct attribute values are guaranteed to be below a specified error threshold, \u03b5, while the<jats:italic>relative<\/jats:italic>proportion differences of the remaining attribute values are as close to \u03b5 as possible, which ensures that the generated sample is of high quality; (4) the sample rate is close to the user specified rate so that a high quality sampling result can be obtained without increasing the sample size; (5) by a thorough analytical and empirical study, we prove that<jats:italic>FPS<\/jats:italic>has acceptable space overheads, especially when the attribute values have Zipfian distributions, and<jats:italic>FPS<\/jats:italic>can also excellently preserve the population proportion of multivariate features in the sample; and (6)<jats:italic>FPS<\/jats:italic>can be applied to infinite streams and finite datasets equally, and the generated samples can be used for various applications. Our experiments on both real and synthetic data validate that<jats:italic>FPS<\/jats:italic>can effectively obtain a high quality sample of the desired size. In addition, while using the sample generated by<jats:italic>FPS<\/jats:italic>in various mining applications, a significant improvement in efficiency can be achieved without compromising the model's precision.<\/jats:p>","DOI":"10.1145\/1460797.1460798","type":"journal-article","created":{"date-parts":[[2009,1,13]],"date-time":"2009-01-13T13:15:48Z","timestamp":1231852548000},"page":"1-45","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Feature-preserved sampling over streaming data"],"prefix":"10.1145","volume":"2","author":[{"given":"Kun-Ta","family":"Chuang","sequence":"first","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hung-Leng","family":"Chen","sequence":"additional","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ming-Syan","family":"Chen","sequence":"additional","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,1,16]]},"reference":[{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Agrawal R.","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","unstructured":"Aho A. V. Sethi R. and Ullman J. D. 1986. Compliers. Principles Techniques and Tools. Addison-Wesley. Aho A. V. Sethi R. and Ullman J. D. 1986. Compliers. Principles Techniques and Tools. Addison-Wesley."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055558.1055598"},{"key":"e_1_2_1_4_1","unstructured":"Asuncion A. and Newman D. 2007. UCI Machine Learning Repository. http:www.ics.uv.edu\/-mlearn\/MLRepository.html. Asuncion A. and Newman D. 2007. UCI Machine Learning Repository. http:www.ics.uv.edu\/-mlearn\/MLRepository.html."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/543613.543615"},{"volume-title":"Proceedings of ACM-SIAM Symposium on Discrete Algorithms.","author":"Babcock B.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/773153.773176"},{"key":"e_1_2_1_8_1","unstructured":"Baeza-Yates R. and Ribeiro-Neto B. 1999. Modern Information Retrieval. Addison-Wesley. Baeza-Yates R. and Ribeiro-Neto B. 1999. Modern Information Retrieval. Addison-Wesley."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0377-0427(00)00336-8"},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","unstructured":"Br\u00f6nnimann H. Chen B. Dash M. Haas P. Qiao Y. and Scheuermann P. 2004. Efficient data reduction methods for on-line association rule discovery. In Data Mining: Next Generation Challenges and Future Directions. AAAT Press 190--208. Br\u00f6nnimann H. Chen B. Dash M. Haas P. Qiao Y. and Scheuermann P. 2004. Efficient data reduction methods for on-line association rule discovery. In Data Mining: Next Generation Challenges and Future Directions. AAAT Press 190--208.","DOI":"10.1145\/956750.956761"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956761"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775114"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/11430919_59"},{"key":"e_1_2_1_14_1","unstructured":"Cochran W. G. 1977. Sampling Techniques. John Wiley and Sons. Cochran W. G. 1977. Sampling Techniques. John Wiley and Sons."},{"volume-title":"Proceedings of the SIAM International Conference on Data Mining.","author":"Cormode G.","key":"e_1_2_1_15_1"},{"volume-title":"Proceedings of the Annual ACM\/SIAM Symposium on Discrete Algorithms.","author":"Datar M.","key":"e_1_2_1_16_1"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1014091514039"},{"volume-title":"Proceedings of the of International Conference on Machine Learning.","author":"Dougherty J.","key":"e_1_2_1_18_1"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Ganti V.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/564691.564746"},{"key":"e_1_2_1_21_1","unstructured":"Ghahramani S. 1999. Fundamentals of Probability 2nd Edu. Prentice Hall. Ghahramani S. 1999. Fundamentals of Probability 2nd Edu. Prentice Hall."},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","year":"2001","author":"Gibbons P.","key":"e_1_2_1_22_1"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/276304.276334"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Gibbons P. B.","key":"e_1_2_1_24_1"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/276304.276312"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Guha S.","key":"e_1_2_1_26_1"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/130283.130335"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/115790.115837"},{"volume-title":"Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.","author":"John G. H.","key":"e_1_2_1_29_1"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066159"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2003.1232271"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/502585.502630"},{"key":"e_1_2_1_33_1","unstructured":"Levy P. S. and Lemeshow S. 1991. Sampling of Populations: Methods and Applications. John Wiley and Sons. Levy P. S. and Lemeshow S. 1991. Sampling of Populations: Methods and Applications. John Wiley and Sons."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1995.1050"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Manku G. S.","key":"e_1_2_1_35_1"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Matias Y.","key":"e_1_2_1_36_1"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30570-5_27"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/584792.584858"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/342009.335384"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/223784.223813"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/844380.844755"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/312129.312188"},{"key":"e_1_2_1_43_1","unstructured":"Rice J. A. 1995. Mathematical Statistics and Data Analysis. Duxbury Press. Rice J. A. 1995. Mathematical Statistics and Data Analysis. Duxbury Press."},{"key":"e_1_2_1_44_1","unstructured":"Scheaffer R. L. Mendenhall W. and Ott R. L. 1995. Elementary Survey Sampling. Duxbury Press. Scheaffer R. L. Mendenhall W. and Ott R. L. 1995. Elementary Survey Sampling. Duxbury Press."},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Teng W.-G.","key":"e_1_2_1_45_1"},{"key":"e_1_2_1_46_1","unstructured":"Thompson S. K. and Seber G. A. F. 1996. Adaptive Sampling. WILEY Series in Probability and Statistics. Thompson S. K. and Seber G. A. F. 1996. Adaptive Sampling. WILEY Series in Probability and Statistics."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3147.3165"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/23002.23003"},{"key":"e_1_2_1_49_1","unstructured":"Wald A. 1947. Sequential Analysis. Wiley. Wald A. 1947. Sequential Analysis. Wiley."},{"volume-title":"Data Mining: Practical Machine Learning Tools with Java Implementations. Morgan Kaufmann.","year":"1999","author":"Witten I. H.","key":"e_1_2_1_50_1"},{"volume-title":"Proceedings of the International Conference on Very Large Data Bases.","author":"Yu J. X.","key":"e_1_2_1_51_1"},{"volume-title":"Proceedings of the International Workshop on Research Issues in Data Engineering.","author":"Zaki M.","key":"e_1_2_1_52_1"},{"key":"e_1_2_1_53_1","unstructured":"Zipf G. 1949. Human Behavior and the Principle of Least Effort. Addison-Wesley Press. Zipf G. 1949. Human Behavior and the Principle of Least Effort. Addison-Wesley Press."}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1460797.1460798","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1460797.1460798","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T12:45:49Z","timestamp":1750250749000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1460797.1460798"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,1]]},"references-count":53,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2009,1]]}},"alternative-id":["10.1145\/1460797.1460798"],"URL":"https:\/\/doi.org\/10.1145\/1460797.1460798","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2009,1]]},"assertion":[{"value":"2006-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-01-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}