{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:40:22Z","timestamp":1760240422037,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2019,6,10]],"date-time":"2019-06-10T00:00:00Z","timestamp":1560124800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R\\&amp;D Program of China","award":["2017YFC0822604-2"],"award-info":[{"award-number":["2017YFC0822604-2"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>In this paper, we propose a latent feature group learning (LFGL) algorithm to discover the feature grouping structures and subspace clusters for high-dimensional data. The feature grouping structures, which are learned in an analytical way, can enhance the accuracy and efficiency of high-dimensional data clustering. In LFGL algorithm, the Darwinian evolutionary process is used to explore the optimal feature grouping structures, which are coded as chromosomes in the genetic algorithm. The feature grouping weighting k-means algorithm is used as the fitness function to evaluate the chromosomes or feature grouping structures in each generation of evolution. To better handle the diverse densities of clusters in high-dimensional data, the original feature grouping weighting k-means is revised with the mass-based dissimilarity measure rather than the Euclidean distance measure and the feature weights are optimized as a nonnegative matrix factorization problem under the orthogonal constraint of feature weight matrix. The genetic operations of mutation and crossover are used to generate the new chromosomes for next generation. In comparison with the well-known clustering algorithms, LFGL algorithm produced encouraging experimental results on real world datasets, which demonstrated the better performance of LFGL when clustering high-dimensional data.<\/jats:p>","DOI":"10.3390\/info10060208","type":"journal-article","created":{"date-parts":[[2019,6,10]],"date-time":"2019-06-10T11:39:47Z","timestamp":1560166787000},"page":"208","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Latent Feature Group Learning for High-Dimensional Data Clustering"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5041-7988","authenticated-orcid":false,"given":"Wenting","family":"Wang","sequence":"first","affiliation":[{"name":"Big Data Institute, College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"},{"name":"National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yulin","family":"He","sequence":"additional","affiliation":[{"name":"Big Data Institute, College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"},{"name":"National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liheng","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer Science, McGill University, Montreal, QC H3A OG4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joshua Zhexue","family":"Huang","sequence":"additional","affiliation":[{"name":"Big Data Institute, College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"},{"name":"National Engineering Laboratory for Big Data System Computing Technology, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,6,10]]},"reference":[{"key":"ref_1","first-page":"375","article-title":"High-dimensional data analysis: The curses and blessings of dimensionality","volume":"1","author":"Donoho","year":"2000","journal-title":"AMS Math Chall. Lect."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1145\/1007730.1007731","article-title":"Subspace clustering for high dimensional data: A review","volume":"6","author":"Parsons","year":"2004","journal-title":"ACM Sigkdd Explor. Newsl."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1497577.1497578","article-title":"Clustering high-dimensional data: A survey on subspace clustering, pattern-based clustering, and correlation clustering","volume":"3","author":"Kriegel","year":"2009","journal-title":"ACM Trans. Knowl. Discov. Data"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1026","DOI":"10.1109\/TKDE.2007.1048","article-title":"An entropy weighting k-means algorithm for subspace clustering of high-dimensional sparse data","volume":"8","author":"Jing","year":"2007","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1109\/TPAMI.2005.95","article-title":"Automated variable weighting in k-means type clustering","volume":"5","author":"Huang","year":"2005","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"932","DOI":"10.1109\/TKDE.2011.262","article-title":"TW-k-means: Automated two-level variable weighting clustering algorithm for multiview data","volume":"25","author":"Chen","year":"2013","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"434","DOI":"10.1016\/j.patcog.2011.06.004","article-title":"A feature group weighting method for subspace clustering of high-dimensional data","volume":"45","author":"Chen","year":"2012","journal-title":"Pattern Recognit."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"113","DOI":"10.1007\/BF01202271","article-title":"Weighting and selection of variables for cluster analysis","volume":"12","author":"Gnanadesikan","year":"1995","journal-title":"J. Classif."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1007\/BF00227423","article-title":"Optimal variable weighting for ultrametric and additive tree clustering","volume":"20","year":"1986","journal-title":"Qual. Quant."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1007\/BF01901677","article-title":"OVWTRE: A program for optimal variable weighting for ultrametric and additive tree fitting","volume":"5","year":"1988","journal-title":"J. Classif."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1007\/BF01897164","article-title":"Variable selection in clustering","volume":"5","author":"Fowlkes","year":"1988","journal-title":"J. Classif."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/s003579900040","article-title":"An algorithm for the fitting of a tree metric according to a weighted least-squares criterion","volume":"16","author":"Makarenkov","year":"1999","journal-title":"J. Classif."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1007\/s00357-001-0018-x","article-title":"Optimal variable weighting for ultrametric and additive trees and K-means partitioning: Methods and software","volume":"18","author":"Makarenkov","year":"2001","journal-title":"J. Classif."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"217","DOI":"10.1023\/A:1024016609528","article-title":"Feature weighting in k-means clustering","volume":"52","author":"Modha","year":"2003","journal-title":"Mach. Learn."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"815","DOI":"10.1111\/j.1467-9868.2004.02059.x","article-title":"Clustering objects on subsets of attributes (with discussion)","volume":"66","author":"Friedman","year":"2004","journal-title":"J. R. Stat. Soc. B"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1007\/s10618-006-0060-8","article-title":"Locally adaptive metrics for clustering high dimensional data","volume":"14","author":"Domeniconi","year":"2007","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1214\/06-BA111","article-title":"Model-based subspace clustering","volume":"1","author":"Hoff","year":"2006","journal-title":"Bayesian Anal."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"502","DOI":"10.1016\/j.csda.2007.02.009","article-title":"High-dimensional data clustering","volume":"52","author":"Bouveyron","year":"2007","journal-title":"Comput. Stat. Data Anal."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"4658","DOI":"10.1016\/j.csda.2008.03.002","article-title":"Developing a feature weight self-adjustment mechanism for a K-means clustering algorithm","volume":"52","author":"Tsai","year":"2008","journal-title":"Comput. Stat. Data Anal."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"767","DOI":"10.1016\/j.patcog.2009.09.010","article-title":"Enhanced soft subspace clustering integrating within-cluster and between-cluster information","volume":"43","author":"Deng","year":"2010","journal-title":"Pattern Recognit."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"90","DOI":"10.14778\/1453856.1453871","article-title":"Constrained locally weighted clustering","volume":"1","author":"Cheng","year":"2008","journal-title":"Proc. VLDB Endow."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1016\/j.patcog.2011.08.012","article-title":"Minkowski metric, feature weighting and anomalous cluster initializing in K-Means clustering","volume":"45","author":"Mirkin","year":"2012","journal-title":"Pattern Recognit."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1670","DOI":"10.1109\/TKDE.2012.101","article-title":"Unsupervised hybrid feature extraction selection for high-dimensional non-Gaussian data clustering with variational inference","volume":"25","author":"Fan","year":"2013","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2562","DOI":"10.1016\/j.patcog.2013.02.005","article-title":"Novel soft subspace clustering with multi-objective evolutionary approach for high-dimensional data","volume":"46","author":"Xia","year":"2013","journal-title":"Pattern Recognit."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cai, Y., Chen, X., Peng, P.X., and Huang, J.Z. (2014, January 13). A LDA feature grouping method for subspace clustering of text data. Proceedings of the Pacific-Asia Workshop on Intelligence and Security Informatics, Tainan, Taiwan.","DOI":"10.1007\/978-3-319-06677-6_7"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"3703","DOI":"10.1016\/j.patcog.2015.05.016","article-title":"Subspace clustering with automatic feature grouping","volume":"48","author":"Gan","year":"2015","journal-title":"Pattern Recognit."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ting, K.M., Zhu, Y., Carman, M., Zhu, Y., and Zhou, Z.H. (2016, January 13\u201317). Overcoming key weaknesses of distance-based neighbourhood methods using a data dependent dissimilarity measure. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939779"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"617","DOI":"10.1109\/TEVC.2008.920670","article-title":"Darwinian, Lamarckian, and Baldwinian (Co) Evolutionary Approaches for Feature Weighting in K-means-Based Algorithms","volume":"12","year":"2008","journal-title":"IEEE Trans. Evol. Comput."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1007\/s41019-015-0003-8","article-title":"Clustering embedded approaches for efficient information network inference","volume":"1","author":"Hu","year":"2016","journal-title":"Data Sci. Eng."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1007\/s41019-018-0080-6","article-title":"Evolutionary active constrained clustering for obstructive sleep apnea analysis","volume":"3","author":"Mai","year":"2018","journal-title":"Data Sci. Eng."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TKDE.2011.181","article-title":"A fast clustering-based feature subset selection algorithm for high-dimensional data","volume":"25","author":"Song","year":"2013","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1016\/j.ins.2015.07.041","article-title":"High-dimensional feature selection via feature grouping: A Variable Neighborhood Search approach","volume":"326","year":"2016","journal-title":"Inf. Sci."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1007\/s41019-016-0022-0","article-title":"Big data reduction methods: a survey","volume":"1","author":"Liew","year":"2016","journal-title":"Data Sci. Eng."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1007\/s41019-017-0043-3","article-title":"Big Data Management: What to Keep from the Past to Face Future Challenges?","volume":"2","year":"2017","journal-title":"Data Sci. Eng."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/10\/6\/208\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:57:17Z","timestamp":1760187437000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/10\/6\/208"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,10]]},"references-count":34,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2019,6]]}},"alternative-id":["info10060208"],"URL":"https:\/\/doi.org\/10.3390\/info10060208","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2019,6,10]]}}}