{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T07:56:26Z","timestamp":1774943786708,"version":"3.50.1"},"reference-count":32,"publisher":"World Scientific Pub Co Pte Ltd","issue":"06","funder":[{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2021MF085"],"award-info":[{"award-number":["ZR2021MF085"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2023QC116"],"award-info":[{"award-number":["ZR2023QC116"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100012906","name":"Department of Education of Shandong Province","doi-asserted-by":"publisher","award":["J17KB183"],"award-info":[{"award-number":["J17KB183"]}],"id":[{"id":"10.13039\/100012906","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100012906","name":"Department of Education of Shandong Province","doi-asserted-by":"publisher","award":["M2024314"],"award-info":[{"award-number":["M2024314"]}],"id":[{"id":"10.13039\/100012906","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:p>Unsupervised learning is one of the fundamental machine learning methods. Clustering is a vital unsupervised learning task and can significantly contribute to the detection of hidden structures in unknown datasets. Clusterability is an important concept due to the fact that it can theoretically portray the extent to which a clustering algorithm can recover a benchmark clustering, with the absence of excessive experimental validations. Moreover, conventional batch-mode-clustering-oriented clusterability analysis should be extended to the incremental setting when the clustering algorithm is required to handle stream data. However, such clusterability analysis is facing two barriers. First, the incremental clustering algorithm proceeds in a step-wise manner and can merely access the newly arrived data of the current step. This extremely fragmentary view of the entire input data stream inevitably results in a biased perception of the underlying benchmark clustering. Second, incremental clustering is conventionally applied to real-time or massive-data scenarios. Such application scenarios typically require the computational power of mainstream SIMD (Single Instruction Multiple Data) hardware accelerators. However, strong data dependency inherently exists between two successive steps of an incremental clustering algorithm, which dramatically impairs data parallelism. In view of these constraints, we propose our roadmap to theoretically analyze and ensure the clusterability under an incremental setting in terms of a general clusterability metric: niceness (higher intra-cluster similarity than inter-cluster similarity). In our work, a nice-k clustering (a clustering that has k clusters and satisfies the niceness metric) is supposed to exist in the input data stream. Meanwhile, the input data stream is supposed to be divided into a series of micro-clusters, and the micro-clusters are incrementally clustered into clusters. In addition, we rely on an assumption (homogeneity assumption) that every micro-cluster merely contains homogenous data. First, we point out that a vital reason for the induction of heterogeneous clusters is the lack of representative micro-clusters. We propose Theorem 1 to iteratively identify a set of 2[Formula: see text] representative micro-clusters that can cover all k benchmark clusters. Therefore, we can trade the number of clusters for homogeneity and thus assure clusterability. Second, we demonstrate that evolution in the granularity of a micro-cluster can prompt SIMD-parallelism more than in the granularity of a single data point. Consequently, the clusterability-assured method of Theorem 1 is furthermore parallel-friendly. In all, we depict a roadmap to assure clusterability under both incremental and SIMD-friendly constraints.<\/jats:p>","DOI":"10.1142\/s021800142651002x","type":"journal-article","created":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T04:01:51Z","timestamp":1768363311000},"source":"Crossref","is-referenced-by-count":0,"title":["Unsupervised Learning on Stream Data: Clusterability Analysis in a Joint Perspective Under Incremental and Parallel Constraints"],"prefix":"10.1142","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0883-0159","authenticated-orcid":false,"given":"Chunlei","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6548-3299","authenticated-orcid":false,"given":"Jinkui","family":"Hou","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1695-4629","authenticated-orcid":false,"given":"Jiangyan","family":"Dai","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1012-8089","authenticated-orcid":false,"given":"Huihui","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3668-4569","authenticated-orcid":false,"given":"Guoxu","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8527-9041","authenticated-orcid":false,"given":"Yujie","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5155-5762","authenticated-orcid":false,"given":"Lu","family":"Hong","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Weifang University, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8942-5392","authenticated-orcid":false,"given":"Jia","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Big Data, Weifang Institute of Technology, Weifang, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2026,2,21]]},"reference":[{"key":"S021800142651002XBIB001","first-page":"307","volume-title":"Proc. 27th Int. Conf. Neural Information Processing Systems","volume":"1","author":"Ackerman M.","year":"2014"},{"key":"S021800142651002XBIB002","doi-asserted-by":"publisher","DOI":"10.1145\/3543507.3583213"},{"issue":"6","key":"S021800142651002XBIB003","first-page":"1452","volume":"12","author":"Al-Sudany S. M.","year":"2020","journal-title":"J. Xi\u2019an Univ. Archit. Technol."},{"key":"S021800142651002XBIB004","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2019.09.012"},{"key":"S021800142651002XBIB005","doi-asserted-by":"publisher","DOI":"10.1145\/3565557"},{"key":"S021800142651002XBIB006","doi-asserted-by":"publisher","DOI":"10.1145\/1374376.1374474"},{"key":"S021800142651002XBIB007","doi-asserted-by":"publisher","DOI":"10.23919\/ACC55779.2023.10155791"},{"key":"S021800142651002XBIB008","doi-asserted-by":"crossref","unstructured":"D. Cole, S. Shin, F. Pacaud, V. M. Zavala and M. Anitescu, How to make the most out of SIMD on AArch64? in\n                      Proc. ISC High Performance 2025 Research Paper (40th Int. Conf.)\n                      (IEEE 2025), pp. 1\u201311.","DOI":"10.23919\/ISC.2025.11018308"},{"key":"S021800142651002XBIB009","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3306024"},{"key":"S021800142651002XBIB010","doi-asserted-by":"publisher","DOI":"10.3390\/stats6030048"},{"key":"S021800142651002XBIB011","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2022.102162"},{"key":"S021800142651002XBIB012","unstructured":"M. Krishnamoorthy, and M. Zaki. Clusterability detection and initial seed selection in large datasets. In The International Conference on Knowledge Discovery in Databases, volume 7, paper NO. 368, ACM, 1999."},{"key":"S021800142651002XBIB013","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2022.104743"},{"key":"S021800142651002XBIB014","doi-asserted-by":"publisher","DOI":"10.1016\/j.softx.2022.101270"},{"key":"S021800142651002XBIB015","doi-asserted-by":"publisher","DOI":"10.1109\/ARITH58626.2023.00010"},{"key":"S021800142651002XBIB016","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2018.2794998"},{"key":"S021800142651002XBIB017","author":"Hu L.","year":"2023","journal-title":"CoRR"},{"key":"S021800142651002XBIB018","doi-asserted-by":"publisher","DOI":"10.1145\/3597635.3598021"},{"key":"S021800142651002XBIB019","doi-asserted-by":"publisher","DOI":"10.61822\/amcs-2024-0010"},{"key":"S021800142651002XBIB020","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-023-06308-x"},{"issue":"2","key":"S021800142651002XBIB021","first-page":"288","volume":"17","author":"Krishnaprasad S.","year":"2001","journal-title":"J. Comput. Sci. Coll."},{"key":"S021800142651002XBIB022","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-023-05210-6"},{"key":"S021800142651002XBIB023","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2022.11.033"},{"key":"S021800142651002XBIB024","first-page":"19","volume":"660","author":"Ma L.","year":"2024","journal-title":"Inf. Sci."},{"key":"S021800142651002XBIB025","doi-asserted-by":"publisher","DOI":"10.1007\/s41060-023-00389-6"},{"key":"S021800142651002XBIB026","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-025-11240-8"},{"key":"S021800142651002XBIB027","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.2021.3117537"},{"key":"S021800142651002XBIB028","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-023-05193-w"},{"key":"S021800142651002XBIB029","doi-asserted-by":"publisher","DOI":"10.1016\/0196-6774(82)90013-X"},{"key":"S021800142651002XBIB030","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2024.102799"},{"key":"S021800142651002XBIB031","doi-asserted-by":"publisher","DOI":"10.1109\/I3CEET61722.2024.10993534"},{"key":"S021800142651002XBIB032","doi-asserted-by":"publisher","DOI":"10.1145\/3689036"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S021800142651002X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T06:22:28Z","timestamp":1774938148000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S021800142651002X"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,21]]},"references-count":32,"journal-issue":{"issue":"06","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["10.1142\/S021800142651002X"],"URL":"https:\/\/doi.org\/10.1142\/s021800142651002x","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"value":"0218-0014","type":"print"},{"value":"1793-6381","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,21]]},"article-number":"2651002"}}