{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T04:14:56Z","timestamp":1784261696753,"version":"3.55.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,2,15]],"date-time":"2025-02-15T00:00:00Z","timestamp":1739577600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["2426340, 2416727, 2421864, 2421865, 2421803"],"award-info":[{"award-number":["2426340, 2416727, 2421864, 2421865, 2421803"]}]},{"name":"National Academy of Engineering Grainger Foundation Frontiers of Engineering"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>\n            Feature selection aims to identify the optimal feature subset for enhancing downstream models. Effective feature selection can remove redundant features, save computational resources, accelerate the model learning process, and improve the model overall performance. However, existing works are often time-intensive to identify the effective feature subset within high-dimensional feature spaces. Meanwhile, these methods mainly utilize a single downstream task performance as the selection criterion, leading to the selected subsets that are not only redundant but also lack generalizability. To bridge these gaps, we reformulate feature selection through a neuro-symbolic lens and introduce a novel generative framework aimed at identifying short and effective feature subsets. More specifically, we found that feature ID tokens of the selected subset can be formulated as symbols to reflect the intricate correlations among features. Thus, in this framework, we first create a data collector to automatically collect numerous feature selection samples consisting of feature ID tokens, model performance, and the measurement of feature subset redundancy. Building on the collected data, an encoder-decoder-evaluator learning paradigm is developed to preserve the intelligence of feature selection into a continuous embedding space for efficient search. Within the learned embedding space, we leverage a multi-gradient search algorithm to find more robust and generalized embeddings with the objective of improving model performance and reducing feature subset redundancy. These embeddings are then utilized to reconstruct the feature ID tokens for executing the final feature selection. Ultimately, comprehensive experiments and case studies are conducted to validate the effectiveness of the proposed framework. The associated data and code are publicly available (\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/NanxuGong\/feature-selection-via-autoregreesive-generation\">https:\/\/github.com\/NanxuGong\/feature-selection-via-autoregreesive-generation<\/jats:ext-link>\n            ).\n          <\/jats:p>","DOI":"10.1145\/3709011","type":"journal-article","created":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T15:56:41Z","timestamp":1734710201000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Neuro-Symbolic Embedding for Short and Effective Feature Selection via Autoregressive Generation"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-4534-8395","authenticated-orcid":false,"given":"Nanxu","family":"Gong","sequence":"first","affiliation":[{"name":"School of Computing and Augmented Intelligence, Arizona State University, Tempe, Arizona, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6196-0287","authenticated-orcid":false,"given":"Wangyang","family":"Ying","sequence":"additional","affiliation":[{"name":"School of Computing and Augmented Intelligence, Arizona State University, Tempe, Arizona, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3948-0059","authenticated-orcid":false,"given":"Dongjie","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Kansas, Lawrence, Kansas, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1767-8024","authenticated-orcid":false,"given":"Yanjie","family":"Fu","sequence":"additional","affiliation":[{"name":"School of Computing and Augmented Intelligence, Arizona State University, Tempe, Arizona, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,2,15]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISIEA49364.2020.9188198"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2906757"},{"key":"e_1_3_1_4_2","first-page":"214","volume-title":"International Conference on Machine Learning","author":"Arjovsky Martin","year":"2017","unstructured":"Martin Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In International Conference on Machine Learning. PMLR, 214\u2013223."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-020-05399-0"},{"key":"e_1_3_1_6_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCIII.2007.367361"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611976700.39"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944974"},{"key":"e_1_3_1_10_2","unstructured":"Nanxu Gong Chandan K. Reddy Wangyang Ying and Yanjie Fu. 2024. Evolutionary large language model for automated feature transformation. arXiv:2405.16203. Retrieved from https:\/\/arxiv.org\/abs\/2405.16203"},{"key":"e_1_3_1_11_2","first-page":"27","article-title":"Generative adversarial nets","volume":"27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in Neural Information Processing Systems 27 (2014), 27.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chemolab.2006.01.007"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-021-01347-z"},{"key":"e_1_3_1_14_2","first-page":"18","article-title":"Laplacian score for feature selection","author":"He Xiaofei","year":"2005","unstructured":"Xiaofei He, Deng Cai, and Partha Niyogi. 2005. Laplacian score for feature selection. Advances in Neural Information Processing Systems 18 (2005).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627673.3680105"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3083165.3083171"},{"key":"e_1_3_1_18_2","first-page":"10681","volume-title":"IEEE Transactions on Knowledge and Data Engineering","volume":"35","author":"Huang Yanyong","year":"2023","unstructured":"Yanyong Huang, Zongxin Shen, Yuxin Cai, Xiuwen Yi, Dongjie Wang, Fengmao Lv, and Tianrui Li. 2023. IMUFS: Complementary and consensus learning-based incomplete multi-view unsupervised feature selection. IEEE Transactions on Knowledge and Data Engineering 35, 10 (2023), 10681\u201310694."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/347090.347169"},{"key":"e_1_3_1_20_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(97)00043-X"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-15-4409-5_71"},{"key":"e_1_3_1_23_2","first-page":"10","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Lemhadri Ismael","year":"2021","unstructured":"Ismael Lemhadri, Feng Ruan, and Rob Tibshirani. 2021. Lassonet: Neural networks with feature sparsity. In International Conference on Artificial Intelligence and Statistics. PMLR, 10\u201318."},{"key":"e_1_3_1_24_2","unstructured":"Chunyuan Li Xiang Gao Yuan Li Baolin Peng Xiujun Li Yizhe Zhang and Jianfeng Gao. 2020. Optimus: Organizing sentences via pre-trained modeling of a latent space. arXiv:2004.04092. Retrieved from https:\/\/arxiv.org\/abs\/2004.04092"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1089\/cmb.2015.0189"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330868"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-022-01812-3"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM51629.2021.00051"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1037\/0033-2909.111.1.172"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1977.1674939"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/BIBM.2016.7822569"},{"key":"e_1_3_1_33_2","unstructured":"Milad Zafar Nezhad Dongxiao Zhu Najibesadat Sadati and Kai Yang. 2018. A predictive approach using deep feature learning for electronic medical records: A comparative study. arXiv:1801.02961. Retrieved from https:\/\/arxiv.org\/abs\/1801.02961"},{"key":"e_1_3_1_34_2","first-page":"11","article-title":"Hybrid genetic algorithms for feature selection","volume":"26","author":"Oh Il-Seok","year":"2004","unstructured":"Il-Seok Oh, Jin-Seon Lee, and Byung-Ro Moon. 2004. Hybrid genetic algorithms for feature selection. IEEE Transactions on Pattern Analysis and Machine Intelligence 26, 11 (2004), 1424\u20131437.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_35_2","first-page":"309","article-title":"The power of student\u2019s t-test","volume":"60","author":"Owen Donald B.","year":"1965","unstructured":"Donald B. Owen. 1965. The power of student\u2019s t-test. Journal of American Statististical Association. 60, 309 (1965), 320\u2013333.","journal-title":"Journal of American Statististical Association"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Seongmin Park and Jihwa Lee. 2021. Finetuning pretrained transformers Into variational autoencoders. arXiv:2108.02446. Retrieved from https:\/\/arxiv.org\/abs\/2108.02446","DOI":"10.18653\/v1\/2021.insights-1.5"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2005.159"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-017-2988-6"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2016.11.017"},{"key":"e_1_3_1_41_2","first-page":"3145","volume-title":"International Conference on Machine Learning","author":"Shrikumar Avanti","year":"2017","unstructured":"Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017. Learning important features through propagating activation differences. In International Conference on Machine Learning. PMLR, 3145\u20133153."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/0169-7439(89)80095-4"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2006.05.004"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chemolab.2017.02.004"},{"issue":"2","key":"e_1_3_1_45_2","first-page":"229","article-title":"Reinforcement learning: An introduction","volume":"17","author":"Sutton Richard S.","year":"1999","unstructured":"Richard S. Sutton and Andrew G. Barto. 1999. Reinforcement learning: An introduction. Robotica 17, 2 (1999), 229\u2013235.","journal-title":"Robotica"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1996.tb02080.x"},{"key":"e_1_3_1_47_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539278"},{"key":"e_1_3_1_49_2","article-title":"Reinforcement-enhanced autoregressive feature transformation: Gradient-steered search in continuous space for postfix expressions","volume":"36","author":"Wang Dongjie","year":"2024","unstructured":"Dongjie Wang, Meng Xiao, Min Wu, Yuanchun Zhou, Yanjie Fu. 2024. Reinforcement-enhanced autoregressive feature transformation: Gradient-steered search in continuous space for postfix expressions. Advances in Neural Information Processing Systems 36 (2024).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_50_2","unstructured":"Xinyuan Wang Dongjie Wang Wangyang Ying Rui Xie Haifeng Chen and Yanjie Fu. 2024. Knockoff-guided feature selection via a single pre-trained reinforced agent. arXiv:2403.04015. Retrieved from https:\/\/arxiv.org\/abs\/2403.04015"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638780"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611977653.ch87"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM58522.2023.00078"},{"key":"e_1_3_1_54_2","first-page":"10648","volume-title":"International Conference on Machine Learning","author":"Yamada Yutaro","year":"2020","unstructured":"Yutaro Yamada, Ofir Lindenbaum, Sahand Negahban, and Yuval Kluger. 2020. Feature selection using stochastic gates. In International Conference on Machine Learning. PMLR, 10648\u201310659."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/5254.671091"},{"key":"e_1_3_1_56_2","first-page":"412","volume-title":"International Conference on Machine Learning (ICML \u201997)","author":"Yang Yiming","year":"1997","unstructured":"Yiming Yang and Jan O. Pedersen. 1997. A comparative study on feature selection in text categorization. In International Conference on Machine Learning (ICML \u201997), 412\u2013420."},{"key":"e_1_3_1_57_2","unstructured":"Wangyang Ying Haoyue Bai Kunpeng Liu and Yanjie Fu. 2024. Topology-aware reinforcement feature space Reconstruction for graph data. arXiv:2411.05742. Retrieved from https:\/\/arxiv.org\/abs\/2411.05742"},{"key":"e_1_3_1_58_2","unstructured":"Wangyang Ying Dongjie Wang Haifeng Chen and Yanjie Fu. 2024. Feature selection as deep sequential generative learning. arXiv:2403.03838. Retrieved from https:\/\/arxiv.org\/abs\/2403.03838"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3637528.3672015"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM58522.2023.00084"},{"key":"e_1_3_1_61_2","first-page":"856","volume-title":"Proceedings of the 20th International Conference on Machine Learning (ICML \u201903)","author":"Yu Lei","year":"2003","unstructured":"Lei Yu and Huan Liu. 2003. Feature selection for high-dimensional data: A fast correlation-based filter solution. In Proceedings of the 20th International Conference on Machine Learning (ICML \u201903). 856\u2013863."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2016.06.004"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709011","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709011","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:55Z","timestamp":1750295875000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709011"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,15]]},"references-count":61,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3709011"],"URL":"https:\/\/doi.org\/10.1145\/3709011","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,15]]},"assertion":[{"value":"2024-04-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}