{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T17:39:08Z","timestamp":1767980348195,"version":"3.49.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,3,25]],"date-time":"2023-03-25T00:00:00Z","timestamp":1679702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Science Foundation of China","doi-asserted-by":"crossref","award":["61732005 and 61876035"],"award-info":[{"award-number":["61732005 and 61876035"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key R&D Project of China","award":["2019QY1801"],"award-info":[{"award-number":["2019QY1801"]}]},{"name":"China HTRD Center","award":["2020AAA0107904"],"award-info":[{"award-number":["2020AAA0107904"]}]},{"name":"Yunnan Provincial Major Science and Technology Special Plan Projects","award":["201902D08001905 and 202103AA080015"],"award-info":[{"award-number":["201902D08001905 and 202103AA080015"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>Neural architecture search (NAS) has shown the strong performance of learning neural models automatically in recent years. But most NAS systems are unreliable due to the architecture gap brought by discrete representations of atomic architectures. In this article, we improve the performance and robustness of NAS via narrowing the gap between architecture representations. More specifically, we apply a general contraction mapping to model neural networks with distributed representations (Neural Architecture Search with Distributed Architecture Representations (ArchDAR)). Moreover, for a better search result, we present a joint learning approach to integrating distributed representations with advanced architecture search methods. We implement our ArchDAR in a differentiable architecture search model and test learned architectures on the language modeling task. On the Penn Treebank data, it outperforms a strong baseline significantly by 1.8 perplexity scores. Also, the search process with distributed representations is more stable, which yields a faster structural convergence when it works with the differentiable architecture search model.<\/jats:p>","DOI":"10.1145\/3578709","type":"journal-article","created":{"date-parts":[[2023,1,4]],"date-time":"2023-01-04T12:55:31Z","timestamp":1672836931000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Learning Reliable Neural Networks with Distributed Architecture Representations"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7051-693X","authenticated-orcid":false,"given":"Yinqiao","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Northeastern University, Shenyang, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9870-4131","authenticated-orcid":false,"given":"Runzhe","family":"Cao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Northeastern University, Shenyang, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5927-0502","authenticated-orcid":false,"given":"Qiaozhi","family":"He","sequence":"additional","affiliation":[{"name":"Tencent, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5842-6501","authenticated-orcid":false,"given":"Tong","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Northeastern University and also NiuTrans Research, Shenyang, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9667-3071","authenticated-orcid":false,"given":"Jingbo","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Northeastern University and also NiuTrans Research, Shenyang, Liaoning, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,3,25]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1002\/047084535X.ch7"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.265960"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d18-1338"},{"key":"e_1_3_2_5_2","first-page":"549","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML\u201918),","volume":"80","author":"Bender Gabriel","year":"2018","unstructured":"Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc V. Le. 2018. Understanding and simplifying one-shot architecture search. In Proceedings of the 35th International Conference on Machine Learning (ICML\u201918), Proceedings of Machine Learning Research, Jennifer G. Dy and Andreas Krause (Eds.), Vol. 80. PMLR, 549\u2013558."},{"key":"e_1_3_2_6_2","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920)","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (NeurIPS\u201920), Hugo Larochelle, Marc\u2019Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)."},{"key":"e_1_3_2_7_2","first-page":"2787","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI\u201918), the 30th innovative Applications of Artificial Intelligence (IAAI\u201918), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI\u201918)","author":"Cai Han","year":"2018","unstructured":"Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang. 2018. Efficient architecture search by network transformation. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI\u201918), the 30th innovative Applications of Artificial Intelligence (IAAI\u201918), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI\u201918), Sheila A. McIlraith and Kilian Q. Weinberger (Eds.). AAAI Press, 2787\u20132794."},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)","author":"Cai Han","year":"2019","unstructured":"Han Cai, Ligeng Zhu, and Song Han. 2019. ProxylessNAS: Direct neural architecture search on target task and hardware. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919). OpenReview.net."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01396-x"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.5555\/2832581.2832731"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00186"},{"key":"e_1_3_2_13_2","author":"Elsken Thomas","unstructured":"Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2021. Neural architecture search: A survey. J. Mach. Learn. Res. 20, 1 (2021), 1997--2017.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI\u201922)","author":"Gao Jiahui","year":"2022","unstructured":"Jiahui Gao, Hang Xu, Han Shi, Xiaozhe Ren, Philip L. H. Yu, Xiaodan Liang, Xin Jiang, and Zhenguo Li. 2022. AutoBERT-Zero: Evolving BERT backbone from scratch. In Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI\u201922). AAAI Press."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-84858-7"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-05318-5","volume-title":"Automated Machine Learning: Methods, Systems, Challenges (1st ed.)","author":"Hutter Frank","year":"2019","unstructured":"Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren. 2019. Automated Machine Learning: Methods, Systems, Challenges (1st ed.). Springer Publishing Company, Incorporated."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1367"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/10.52325"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917)","author":"Li Lisha","year":"2017","unstructured":"Lisha Li, Kevin G. Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2017. Hyperband: Bandit-based configuration evaluation for hyperparameter optimization. In Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917). OpenReview.net."},{"key":"e_1_3_2_21_2","first-page":"129","volume-title":"Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI\u201919)","author":"Li Liam","year":"2019","unstructured":"Liam Li and Ameet Talwalkar. 2019. Random search and reproducibility for neural architecture search. In Proceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI\u201919). 129."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.592"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)","author":"Liu Hanxiao","year":"2019","unstructured":"Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2019. DARTS: Differentiable architecture search. In Proceedings of the 7th International Conference on Learning Representations (ICLR\u201919)."},{"key":"e_1_3_2_24_2","first-page":"7827","volume-title":"Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems (NeurIPS\u201918)","author":"Luo Renqian","year":"2018","unstructured":"Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, and Tie-Yan Liu. 2018. Neural architecture optimization. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems (NeurIPS\u201918), Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicol\u00f2 Cesa-Bianchi, and Roman Garnett (Eds.). 7827\u20137838."},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918)","author":"Merity Stephen","year":"2018","unstructured":"Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018. Regularizing and optimizing LSTM language models. In Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918). OpenReview.net."},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913), Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2010-343"},{"key":"e_1_3_2_28_2","first-page":"3111","volume-title":"Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems (NeurIPS\u201913)","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems (NeurIPS\u201913), Christopher J. C. Burges, L\u00e9on Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger (Eds.). 3111\u20133119."},{"key":"e_1_3_2_29_2","first-page":"4092","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML\u201918),","volume":"80","author":"Pham Hieu","year":"2018","unstructured":"Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean. 2018. Efficient neural architecture search via parameter sharing. In Proceedings of the 35th International Conference on Machine Learning (ICML\u201918), Proceedings of Machine Learning Research, Jennifer G. Dy and Andreas Krause (Eds.), Vol. 80. PMLR, 4092\u20134101."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1137\/0330046"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33014780"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00033"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1983.6313167"},{"issue":"3","key":"e_1_3_2_34_2","first-page":"35","article-title":"The curse of dimensionality in classification","volume":"21","author":"Spruyt Vincent","year":"2014","unstructured":"Vincent Spruyt. 2014. The curse of dimensionality in classification. Comput. Vis. Dummies 21, 3 (2014), 35\u201340.","journal-title":"Comput. Vis. Dummies"},{"key":"e_1_3_2_35_2","unstructured":"Kevin Swersky Jasper Snoek and Ryan Prescott Adams. 2014. Freeze-Thaw Bayesian optimization. arXiv:1406.3896. Retrieved from https:\/\/arxiv.org\/abs\/1406.3896."},{"key":"e_1_3_2_36_2","first-page":"1063","volume-title":"Proceedings of the 15th Annual Conference of the International Speech Communication Association (INTERSPEECH\u201914)","author":"Takeda Ryu","year":"2014","unstructured":"Ryu Takeda, Naoyuki Kanda, and Nobuo Nukaga. 2014. Boundary contraction training for acoustic models based on discrete deep neural networks. In Proceedings of the 15th Annual Conference of the International Speech Communication Association (INTERSPEECH\u201914), Haizhou Li, Helen M. Meng, Bin Ma, Engsiong Chng, and Lei Xie (Eds.). ISCA, 1063\u20131067."},{"key":"e_1_3_2_37_2","first-page":"6000","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 6000\u20136010."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p19-1176"},{"key":"e_1_3_2_39_2","series-title":"Proceedings of the 33nd International Conference on Machine Learning (ICML\u201916),","first-page":"564","volume":"48","author":"Wei Tao","year":"2016","unstructured":"Tao Wei, Changhu Wang, Yong Rui, and Chang Wen Chen. 2016. Network morphism. In Proceedings of the 33nd International Conference on Machine Learning (ICML\u201916), JMLR Workshop and Conference Proceedings, Maria-Florina Balcan and Kilian Q. Weinberger (Eds.), Vol. 48. JMLR.org, 564\u2013572."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.40"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.58337"},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Xu Yuhui","year":"2020","unstructured":"Yuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen, Guo-Jun Qi, Qi Tian, and Hongkai Xiong. 2020. {PC}-{DARTS}: Partial channel connections for memory-efficient architecture search. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_43_2","volume-title":"Proceedings of the 8th International Conference on Learning Representations (ICLR\u201920)","author":"Yu Kaicheng","year":"2020","unstructured":"Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat, and Mathieu Salzmann. 2020. Evaluating the search phase of neural architecture search. In Proceedings of the 8th International Conference on Learning Representations (ICLR\u201920). OpenReview.net."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2021.01.001"},{"key":"e_1_3_2_45_2","unstructured":"Julian G. Zilly Rupesh Kumar Srivastava Jan Koutn\u00edk and J\u00fcrgen Schmidhuber. 2016. Recurrent highway networks. arXiv:1607.03474. Retrieved from https:\/\/arxiv.org\/abs\/1607.03474."},{"key":"e_1_3_2_46_2","volume-title":"Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917)","author":"Zoph Barret","year":"2017","unstructured":"Barret Zoph and Quoc V. Le. 2017. Neural architecture search with reinforcement learning. In Proceedings of the 5th International Conference on Learning Representations (ICLR\u201917). OpenReview.net."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00907"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3578709","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3578709","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:50Z","timestamp":1750183730000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3578709"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,25]]},"references-count":46,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3578709"],"URL":"https:\/\/doi.org\/10.1145\/3578709","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,25]]},"assertion":[{"value":"2022-06-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-12-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}