{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:36:42Z","timestamp":1760060202767,"version":"build-2065373602"},"reference-count":38,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T00:00:00Z","timestamp":1754956800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Henan Province Philosophy and Social Sciences Planning Project: Research on Promoting People-Centered New Urbanization in Henan","award":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"],"award-info":[{"award-number":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"]}]},{"name":"Henan Provincial Science and Technology Research Program: Key Technologies for Crop Identification and Yield Estimation Using Remote Sensing Data","award":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"],"award-info":[{"award-number":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"]}]},{"name":"Foundation and Cutting-Edge Technologies Research Program of Henan Province","award":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"],"award-info":[{"award-number":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"]}]},{"name":"Research and Practice Project on Higher Education Teaching Reform in Henan Province","award":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"],"award-info":[{"award-number":["2022BJJ026","242102320345","252102211067","252102210064","252102210124","2024SJGLX0133"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>The remarkable success of autoregressive Large Language Models (LLMs) is predicated on the causal attention mechanism, which enforces a static and rigid form of informational asymmetry by permitting each token to attend only to its predecessors. While effective for sequential generation, this hard-coded unidirectional constraint fails to capture the more complex, dynamic, and nonlinear dependencies inherent in sophisticated reasoning, logical inference, and discourse. In this paper, we challenge this paradigm by introducing Dynamic Asymmetric Attention (DAA), a novel mechanism that replaces the static causal mask with a learnable context-aware guidance module. DAA dynamically generates a continuous-valued attention bias for each query\u2013key pair, effectively learning a \u201csoft\u201d information flow policy that guides rather than merely restricts the model\u2019s focus. Trained end-to-end, our DAA-augmented models demonstrate significant performance gains on a suite of benchmarks, including improvements in perplexity on language modeling and notable accuracy boosts on complex reasoning tasks such as code generation (HumanEval) and mathematical problem-solving (GSM8k). Crucially, DAA provides a new lens for model interpretability. By visualizing the learned asymmetric attention patterns, it is possible to uncover the implicit information flow graphs that the model constructs during inference. These visualizations reveal how the model dynamically prioritizes evidence and forges directed logical links in chain-of-thought reasoning, making its decision-making process more transparent. Our work demonstrates that transitioning from a static hard-wired asymmetry to a learned and dynamic one not only enhances model performance but also paves the way for a new class of more capable and profoundly more explainable LLMs.<\/jats:p>","DOI":"10.3390\/sym17081303","type":"journal-article","created":{"date-parts":[[2025,8,12]],"date-time":"2025-08-12T15:51:02Z","timestamp":1755013862000},"page":"1303","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Dynamic Asymmetric Attention for Enhanced Reasoning and Interpretability in LLMs"],"prefix":"10.3390","volume":"17","author":[{"given":"Feng","family":"Wen","sequence":"first","affiliation":[{"name":"School of Surveying and Urban Spatial Information, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoming","family":"Lu","sequence":"additional","affiliation":[{"name":"The Scientific Academy of Land and Resource of Henan Province, Zhengzhou 467300, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haikun","family":"Yu","sequence":"additional","affiliation":[{"name":"Henan Remote Sensing Institute, Henan Provincial Science and Technology Innovation Center for Natural Resources (Satellite Remote Sensing Research), Zhengzhou 467300, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chunyang","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Surveying and Urban Spatial Information, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huijie","family":"Li","sequence":"additional","affiliation":[{"name":"School of Surveying and Urban Spatial Information, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1539-712X","authenticated-orcid":false,"given":"Xiayang","family":"Shi","sequence":"additional","affiliation":[{"name":"College of Software, Zhengzhou University of Light Industry, Zhengzhou 450000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,12]]},"reference":[{"key":"ref_1","unstructured":"Zhu, S., Xu, S., Sun, H., Pan, L., Cui, M., Du, J., Jin, R., Branco, A., and Xiong, D. (2024). Multilingual Large Language Models: A Systematic Survey. arXiv."},{"key":"ref_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_3","unstructured":"Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H.W., Sutton, C., and Gehrmann, S. (2022). Palm: Scaling language modeling with pathways. arXiv."},{"key":"ref_4","unstructured":"Zhu, S., Cui, M., and Xiong, D. (2024, January 20\u201325). Towards robust in-context learning for machine translation with large language models. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy."},{"key":"ref_5","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2023). Attention is all you need. arXiv."},{"key":"ref_6","unstructured":"Dong, T., Li, B., Liu, J., Zhu, S., and Xiong, D. (August, January 27). MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine Translation. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv.","DOI":"10.1162\/tacl_a_00638"},{"key":"ref_8","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","first-page":"11809","article-title":"Tree of thoughts: Deliberate problem solving with large language models","volume":"36","author":"Yao","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_10","unstructured":"Child, R., Gray, S., Radford, A., and Sutskever, I. (2019). Generating long sequences with sparse transformers. arXiv."},{"key":"ref_11","unstructured":"Beltagy, I., Peters, M.E., and Cohan, A. (2020, January 5\u201310). Longformer: The long-document transformer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online."},{"key":"ref_12","unstructured":"Zaheer, M., Guruganesh, G., Dubey, A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., and Yang, L. (2020, January 6\u201312). Big bird: Transformers for longer sequences. Proceedings of the Advances in Neural Information Processing Systems, Virtual."},{"key":"ref_13","unstructured":"Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. (2021). Roformer: Enhanced transformer with rotary position embedding. arXiv."},{"key":"ref_14","unstructured":"Press, O., Smith, N., and Lewis, M. (2021). Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F. (2023, January 9\u201314). A Length-Extrapolatable Transformer. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.816"},{"key":"ref_16","unstructured":"Duan, S., Shi, Y., and Xu, W. (2023). From interpolation to extrapolation: Complete length generalization for arithmetic transformers. arXiv."},{"key":"ref_17","unstructured":"Veisi, A., Amirzadeh, H., and Mansourian, A. (2025). Context-aware Biases for Length Extrapolation. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"103825","DOI":"10.1016\/j.ipm.2024.103825","article-title":"FEDS-ICL: Enhancing translation ability and efficiency of large language model by optimizing demonstration selection","volume":"61","author":"Zhu","year":"2024","journal-title":"Inf. Process. Manag."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yao, Y., Li, Z., and Zhao, H. (2023). Beyond chain-of-thought, effective graph-of-thought reasoning in language models. arXiv.","DOI":"10.18653\/v1\/2024.findings-naacl.183"},{"key":"ref_20","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (May, January 1). Self-consistency improves chain of thought reasoning in language models. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_21","unstructured":"Zhang, Z., Zhang, A., Li, M., and Smola, A. (2022). Automatic chain of thought prompting in large language models. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Diao, S., Wang, P., Lin, Y., Pan, R., Liu, X., and Zhang, T. (2024, January 11\u201316). Active Prompting with Chain-of-Thought for Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.73"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhu, S., Pan, L., Li, B., and Xiong, D. (2024, January 11\u201316). LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.656"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"104078","DOI":"10.1016\/j.ipm.2025.104078","article-title":"Overcoming language barriers via machine translation with sparse Mixture-of-Experts fusion of large language models","volume":"62","author":"Zhu","year":"2025","journal-title":"Inf. Process. Manag."},{"key":"ref_25","unstructured":"Jain, S., and Wallace, B.C. (2019, January 2\u20137). Attention is not explanation. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wiegreffe, S., and Pinter, Y. (2019, January 3\u20137). Attention is not not explanation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1002"},{"key":"ref_27","unstructured":"He, G., Song, X., and Sun, A. (2025). Knowledge updating? no more model editing! just selective contextual reasoning. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Xu, D., Zhang, Z., Zhu, Z., Lin, Z., Liu, Q., Wu, X., Xu, T., Wang, W., Ye, Y., and Zhao, X. (2024, January 21\u201325). Editing factual knowledge and explanatory ability of medical large language models. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA.","DOI":"10.1145\/3627673.3679673"},{"key":"ref_29","first-page":"1113","article-title":"Probing classifiers: Promises, pitfalls, and a better way","volume":"10","author":"Belinkov","year":"2022","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_30","first-page":"17359","article-title":"Locating and editing factual associations in gpt","volume":"35","author":"Meng","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Sharma, A.S., Atkinson, D., and Bau, D. (2024). Locating and editing factual associations in mamba. arXiv."},{"key":"ref_32","unstructured":"Computer, T. (2025, March 11). RedPajama-1T: An Open, Reproducible, 1.2 Trillion Token Dataset for Training Large Language Models. Available online: https:\/\/simonwillison.net\/2023\/Apr\/17\/redpajama-data\/."},{"key":"ref_33","first-page":"5485","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_34","unstructured":"Chen, M., Tworek, J., Jun, H., Yuan, Q., Pires, H.P.d.O., Le, Q., Luan, Y., Jiang, H., Misra, I., and Krueger, G. (2021). Evaluating large language models trained on code. arXiv."},{"key":"ref_35","unstructured":"Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., and Nakano, R. (2021). Training verifiers to solve math word problems. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Pang, R.Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., and He, H. (2022, January 10\u201315). QuALITY: Question Answering with Long Input Texts, Yes!. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.naacl-main.391"},{"key":"ref_37","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_38","unstructured":"Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., and Saulnier, L. (2023). Mistral 7B. arXiv."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/8\/1303\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:25:38Z","timestamp":1760034338000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/8\/1303"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,12]]},"references-count":38,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8]]}},"alternative-id":["sym17081303"],"URL":"https:\/\/doi.org\/10.3390\/sym17081303","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2025,8,12]]}}}