{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T18:33:15Z","timestamp":1777573995283,"version":"3.51.4"},"reference-count":26,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,3,29]],"date-time":"2020-03-29T00:00:00Z","timestamp":1585440000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61902405"],"award-info":[{"award-number":["61902405"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Categorical data are ubiquitous in machine learning tasks, and the representation of categorical data plays an important role in the learning performance. The heterogeneous coupling relationships between features and feature values reflect the characteristics of the real-world categorical data which need to be captured in the representations. The paper proposes an enhanced categorical data embedding method, i.e., CDE++, which captures the heterogeneous feature value coupling relationships into the representations. Based on information theory and the hierarchical couplings defined in our previous work CDE (Categorical Data Embedding by learning hierarchical value coupling), CDE++ adopts mutual information and margin entropy to capture feature couplings and designs a hybrid clustering strategy to capture multiple types of feature value clusters. Moreover, Autoencoder is used to learn non-linear couplings between features and value clusters. The categorical data embeddings generated by CDE++ are low-dimensional numerical vectors which are directly applied to clustering and classification and achieve the best performance comparing with other categorical representation learning methods. Parameter sensitivity and scalability tests are also conducted to demonstrate the superiority of CDE++.<\/jats:p>","DOI":"10.3390\/e22040391","type":"journal-article","created":{"date-parts":[[2020,3,31]],"date-time":"2020-03-31T13:27:19Z","timestamp":1585661239000},"page":"391","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["CDE++: Learning Categorical Data Embedding by Enhancing Heterogeneous Feature Value Coupling Relationships"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1869-3784","authenticated-orcid":false,"given":"Bin","family":"Dong","sequence":"first","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Songlei","family":"Jian","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ke","family":"Zuo","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,29]]},"reference":[{"key":"ref_1","first-page":"413","article-title":"A link-based cluster ensemble approach for categorical data clustering","volume":"24","author":"Boongeon","year":"2010","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1109\/TKDE.2018.2848902","article-title":"Cure: Flexible categorical data representation by hierarchical coupling learning","volume":"31","author":"Jian","year":"2018","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3142","DOI":"10.1016\/j.eswa.2014.12.002","article-title":"Nearest neighbor classification of categorical data by attributes weighting","volume":"42","author":"Chen","year":"2015","journal-title":"Expert Syst. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Alamuri, M., Surampudi, B.R., and Negi, A. (2014, January 6\u201311). A survey of distance\/similarity measures for categorical data. Proceedings of the 2014 International Joint Conference on Neural Networks (IJCNN), Beijing, China.","DOI":"10.1109\/IJCNN.2014.6889941"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Jian, S., Cao, L., Pang, G., Lu, K., and Gao, H. (2017, January 19\u201325). Embedding-based Representation of Categorical Data by Hierarchical Value Coupling Learning. Proceedings of the International Joint Conference on Artificial Intelligence, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/269"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"593","DOI":"10.1613\/jair.5228","article-title":"ZERO++: Harnessing the Power of Zero Appearances to Detect Anomalies in Large-Scale Data Sets","volume":"57","author":"Pang","year":"2016","journal-title":"J. Artif. Intell. Res."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1016\/S0306-4573(02)00021-3","article-title":"An information-theoretic perspective of tf\u2013idf measures","volume":"39","author":"Aizawa","year":"2003","journal-title":"Inf. Process. Manag."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1016\/j.datak.2007.03.016","article-title":"A k-mean clustering algorithm for mixed numeric and categorical data","volume":"63","author":"Ahmad","year":"2007","journal-title":"Data Knowl. Eng."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2133360.2133361","article-title":"From context to distance: Learning dissimilarity for categorical data clustering","volume":"6","author":"Ienco","year":"2012","journal-title":"ACM Trans. Knowl. Discov. Data (Tkdd)"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1065","DOI":"10.1109\/TNNLS.2015.2436432","article-title":"A new distance metric for unsupervised learning of categorical data","volume":"27","author":"Jia","year":"2015","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"781","DOI":"10.1109\/TNNLS.2014.2325872","article-title":"Coupled attribute similarity learning on categorical data","volume":"26","author":"Wang","year":"2014","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Jolliffe, I. (2011). Principal Component Analysis, Springer.","DOI":"10.1007\/978-3-642-04898-2_455"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, K., Wang, Q., Chen, Z., Marsic, I., Kumar, V., Jiang, G., and Zhang, J. (May, January 30). From categorical to numerical: Multiple transitive distance learning and embedding. Proceedings of the 2015 SIAM International Conference on Data Mining, Vancouver, BC, Canada.","DOI":"10.1137\/1.9781611974010.6"},{"key":"ref_14","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9","article-title":"Indexing by latent semantic analysis","volume":"41","author":"Deerwester","year":"1990","journal-title":"J. Am. Soc. Inf. Sci."},{"key":"ref_16","first-page":"993","article-title":"Latent dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_17","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hofmann, T. (1999, January 15\u201319). Probabilistic latent semantic indexing. Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Berkeley, CA, USA.","DOI":"10.1145\/312624.312649"},{"key":"ref_19","unstructured":"Wilson, A.T., and Chew, P.A. (2010, January 1\u20136). Term weighting schemes for latent dirichlet allocation. Proceedings of the Human language technologies: The 2010 annual conference of the North American Chapter of the Association for Computational Linguistics, Los Angeles, CA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Martino, A., Giuliani, A., and Rizzi, A. (2018). Granular computing techniques for bioinformatics pattern recognition problems in non-metric spaces. Computational Intelligence for Pattern Recognition, Springer.","DOI":"10.1007\/978-3-319-89629-8_3"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"107187","DOI":"10.1016\/j.compbiolchem.2019.107187","article-title":"Metabolic networks classification and knowledge discovery by information granulation","volume":"84","author":"Martino","year":"2020","journal-title":"Comput. Biol. Chem."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Martino, A., Giuliani, A., and Rizzi, A. (2019). (Hyper) Graph Embedding and Classification via Simplicial Complexes. Algorithms, 12.","DOI":"10.3390\/a12110223"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cox, T.F., and Cox, M.A. (2000). Multidimensional Scaling, Chapman and Hall\/CRC.","DOI":"10.1201\/9781420036121"},{"key":"ref_24","unstructured":"Hinton, G.E., and Roweis, S.T. (2003). Stochastic neighbor embedding. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1109\/TNN.2008.2005601","article-title":"Normalized mutual information feature selection","volume":"20","author":"Tesmer","year":"2009","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_26","first-page":"226","article-title":"A density-based algorithm for discovering clusters in large spatial databases with noise","volume":"96","author":"Ester","year":"1996","journal-title":"Kdd"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/4\/391\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:13:04Z","timestamp":1760173984000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/4\/391"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,29]]},"references-count":26,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,4]]}},"alternative-id":["e22040391"],"URL":"https:\/\/doi.org\/10.3390\/e22040391","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,29]]}}}