{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T13:08:33Z","timestamp":1765285713305,"version":"3.46.0"},"reference-count":59,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T00:00:00Z","timestamp":1765238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"State Grid Corporation Headquarters Science and Technology Project","award":["5700-202458333A-2-1-ZX"],"award-info":[{"award-number":["5700-202458333A-2-1-ZX"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Chain-of-Thought (CoT) prompting has demonstrated strong effectiveness in improving the reasoning capabilities of Large Language Models (LLMs). However, existing CoT optimization approaches still lack systematic mechanisms for evaluating and refining prompts. To address this gap, we propose Adversarial Chain-of-Thought (adv-CoT), a framework that introduces adversarial learning into prompt optimization. Adv-CoT iteratively refines an initial prompt through generator\u2013discriminator interactions and integrates both feedback and verification mechanisms. This process enables more targeted and interpretable improvements to CoT instructions and demonstrations. We evaluate adv-CoT on twelve datasets across commonsense, factual, symbolic, and arithmetic reasoning. Across 12 reasoning datasets, adv-CoT yields an average improvement of 4.44% on GPT-3.5-turbo and 1.08% on GPT-4o-mini, with both gains being statistically significant (paired t-test, p &lt; 0.05). The experimental results show that the framework yields consistent but task-dependent gains, particularly on numerical and factual reasoning tasks, and maintains competitive performance on symbolic and commonsense benchmarks. Paired significance tests further indicate that improvements are statistically reliable on high-capacity proprietary models, while results on smaller open-source models exhibit greater variance. Although these findings demonstrate the promise of adversarial refinement for CoT prompting, the conclusions remain preliminary. The effectiveness of adv-CoT depends on the base model\u2019s reasoning capability, and the current evaluation is limited to four major categories of reasoning tasks. We will release the full implementation and prompts to support further investigation into broader applications and more generalizable prompt optimization strategies.<\/jats:p>","DOI":"10.3390\/info16121092","type":"journal-article","created":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T12:42:43Z","timestamp":1765284163000},"page":"1092","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Chain-of-Thought Prompt Optimization via Adversarial Learning"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-0624-3913","authenticated-orcid":false,"given":"Guang","family":"Yang","sequence":"first","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiantao","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1737-6912","authenticated-orcid":false,"given":"Shaohe","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Power Transmission and Transformation Technology, State Grid Zhejiang Electric Power Co., Ltd., Research Institute, Hangzhou 310014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3907-8820","authenticated-orcid":false,"given":"Juhua","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,9]]},"reference":[{"key":"ref_1","first-page":"1877","article-title":"Language Models are Few-Shot Learners","volume":"Volume 33","author":"Larochelle","year":"2020","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_2","unstructured":"Rae, J.W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., and Young, S. (2021). Scaling Language Models: Methods, Analysis & Insights from Training Gopher. arXiv."},{"key":"ref_3","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., and Azhar, F. (2023). Llama: Open and efficient foundation language models. arXiv."},{"key":"ref_4","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). Gpt-4 technical report. arXiv."},{"key":"ref_5","unstructured":"Xiao, T., and Zhu, J. (2025). Foundations of large language models. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H. (2022, January 10\u201315). Metaicl: Learning to learn in context. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.naacl-main.201"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., and Chang, B. (2024, January 12\u201316). A survey on in-context learning. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.emnlp-main.64"},{"key":"ref_8","unstructured":"Dherin, B., Munn, M., Mazzawi, H., Wunder, M., and Gonzalvo, J. (2025). Learning without training: The implicit dynamics of in-context learning. arXiv."},{"key":"ref_9","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_10","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_11","first-page":"11809","article-title":"Tree of thoughts: Deliberate problem solving with large language models","volume":"36","author":"Yao","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Weng, Y., Zhu, M., Xia, F., Li, B., He, S., Liu, S., Sun, B., Liu, K., and Zhao, J. (2023, January 6\u201310). Large language models are better reasoners with self-verification. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore.","DOI":"10.18653\/v1\/2023.findings-emnlp.167"},{"key":"ref_13","unstructured":"Zhang, Z., Zhang, A., Li, M., and Smola, A. (2023, January 1\u20135). Automatic Chain of Thought Prompting in Large Language Models. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3571730","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, J., Chen, J., Ren, R., Cheng, X., Zhao, W.X., Nie, J.Y., and Wen, J.R. (2024). The dawn after the dark: An empirical study on factuality hallucination in large language models. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.586"},{"key":"ref_16","unstructured":"Xu, Z., Jain, S., and Kankanhalli, M. (2024). Hallucination is inevitable: An innate limitation of large language models. arXiv."},{"key":"ref_17","unstructured":"Sun, Y., Yin, Z., Guo, Q., Wu, J., Qiu, X., and Zhao, H. (2024). Benchmarking hallucination in large language models based on unanswerable math word problem. arXiv."},{"key":"ref_18","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022). Self-consistency improves chain of thought reasoning in language models. arXiv."},{"key":"ref_19","first-page":"333","article-title":"Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs","volume":"Volume 37","author":"Globerson","year":"2024","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Diao, S., Wang, P., Lin, Y., Pan, R., Liu, X., and Zhang, T. (2024, January 11\u201316). Active prompting with chain-of-thought for large language models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.73"},{"key":"ref_21","first-page":"105345","article-title":"Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models","volume":"Volume 37","author":"Globerson","year":"2024","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Luo, W., Wang, W., Li, X., Zhou, W., Jia, P., and Zhao, X. (2025, January 6\u201311). TAPO: Task-Referenced Adaptation for Prompt Optimization. Proceedings of the ICASSP 2025\u20142025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India.","DOI":"10.1109\/ICASSP49660.2025.10888677"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Larionov, D., and Eger, S. (2024). Promptoptme: Error-aware prompt compression for llm-based mt evaluation metrics. arXiv.","DOI":"10.18653\/v1\/2025.naacl-long.592"},{"key":"ref_24","unstructured":"Guo, P.F., Tsai, Y.D., and Lin, S.D. (2024). Benchmarking large language model uncertainty for prompt optimization. arXiv."},{"key":"ref_25","unstructured":"Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Adv. Neural Inf. Process. Syst., 27, Available online: https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/hash\/f033ed80deb0234979a61f95710dbe25-Abstract.html."},{"key":"ref_26","unstructured":"Radford, A., Metz, L., and Chintala, S. (2015). Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv."},{"key":"ref_27","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013). Intriguing properties of neural networks. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Biggio, B., Corona, I., Maiorca, D., Nelson, B., \u0160rndi\u0107, N., Laskov, P., Giacinto, G., and Roli, F. (2013, January 22\u201326). Evasion attacks against machine learning at test time. Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Prague, Czech Republic.","DOI":"10.1007\/978-3-642-40994-3_25"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Long, D., Zhao, Y., Brown, H., Xie, Y., Zhao, J., Chen, N., Kawaguchi, K., Shieh, M., and He, J. (2024, January 11\u201316). Prompt optimization via adversarial in-context learning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.395"},{"key":"ref_30","unstructured":"Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. (2022, January 25\u201329). Large language models are human-level prompt engineers. Proceedings of the Eleventh International Conference on Learning Representations, Online."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Pryzant, R., Iter, D., Li, J., Lee, Y.T., Zhu, C., and Zeng, M. (2023). Automatic prompt optimization with \u201cgradient descent\u201d and beam search. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.494"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Cheng, J., Liu, X., Zheng, K., Ke, P., Wang, H., Dong, Y., Tang, J., and Huang, M. (2023). Black-box prompt optimization: Aligning large language models without model training. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.176"},{"key":"ref_33","unstructured":"Schneider, L., Wistuba, M., Klein, A., Golebiowski, J., Zappella, G., and Merra, F.A. (2024). Hyperband-based Bayesian optimization for black-box prompt selection. arXiv."},{"key":"ref_34","unstructured":"Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. (2023, January 1\u20135). Let\u2019s verify step by step. Proceedings of the Twelfth International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, M., Wang, W., Feng, F., Cao, Y., Zhang, J., and Chua, T.S. (2023). Robust prompt optimization for large language models against distribution shifts. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.95"},{"key":"ref_36","unstructured":"Ashok, D., and May, J. (2025). Language models can predict their own behavior. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Ledig, C., Theis, L., Husz\u00e1r, F., Caballero, J., Cunningham, A., Acosta, A., Aitken, A., Tejani, A., Totz, J., and Wang, Z. (2017, January 21\u201326). Photo-realistic single image super-resolution using a generative adversarial network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.19"},{"key":"ref_38","first-page":"1","article-title":"Domain-adversarial training of neural networks","volume":"17","author":"Ganin","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T. (2017, January 21\u201326). Adversarial discriminative domain adaptation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.316"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"100553","DOI":"10.1016\/j.cosrev.2023.100553","article-title":"A survey on GANs for computer vision: Recent research, analysis and taxonomy","volume":"48","author":"Iglesias","year":"2023","journal-title":"Comput. Sci. Rev."},{"key":"ref_41","unstructured":"Ren, D., Cai, Y., and Li, Q. (2023). Unlocking the Power of GANs in Non-Autoregressive Text Generation. arXiv."},{"key":"ref_42","unstructured":"Maus, N., Chao, P., Wong, E., and Gardner, J. (2023). Black box adversarial prompting for foundation models. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yu, L., Zhang, W., Wang, J., and Yu, Y. (2017, January 4\u20139). Seqgan: Sequence generative adversarial nets with policy gradient. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10804"},{"key":"ref_44","unstructured":"Nie, W., Narodytska, N., and Patel, A. (May, January 30). Relgan: Relational generative adversarial networks for text generation. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Guo, J., Lu, S., Cai, H., Zhang, W., Yu, Y., and Wang, J. (2018, January 2\u20137). Long text generation via adversarial training with leaked information. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11957"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhang, R., Chen, C., Gan, Z., Wang, W., Shen, D., Wang, G., Wen, Z., and Carin, L. (2020, January 5\u201310). Improving Adversarial Text Generation by Modeling the Distant Future. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.227"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"5135","DOI":"10.1007\/s10994-023-06367-0","article-title":"Robust generative adversarial network","volume":"112","author":"Zhang","year":"2023","journal-title":"Mach. Learn."},{"key":"ref_48","unstructured":"Wang, J., Sun, Q., Li, X., and Gao, M. (2023). Boosting language models reasoning with chain-of-knowledge prompting. arXiv."},{"key":"ref_49","unstructured":"Talmor, A., Herzig, J., Lourie, N., and Berant, J. (2019, January 2\u20137). COMMONSENSEQA: A Question Answering Challenge Targeting Commonsense Knowledge. Proceedings of the NAACL-HLT, Minneapolis, MN, USA."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1162\/tacl_a_00370","article-title":"Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies","volume":"9","author":"Geva","year":"2021","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018). Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv.","DOI":"10.18653\/v1\/D18-1260"},{"key":"ref_52","unstructured":"Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018). Think you have solved question answering? Try arc, the ai2 reasoning challenge. arXiv."},{"key":"ref_53","first-page":"1","article-title":"Beyond the imitation game: Quantifying and extrapolating the capabilities of language models","volume":"2023","author":"Srivastava","year":"2023","journal-title":"Trans. Mach. Learn. Res."},{"key":"ref_54","unstructured":"Clark, C., Lee, K., Chang, M.W., Kwiatkowski, T., Collins, M., and Toutanova, K. (2019). Boolq: Exploring the surprising difficulty of natural yes\/no questions. arXiv."},{"key":"ref_55","unstructured":"Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., and Nakano, R. (2021). Training verifiers to solve math word problems. arXiv."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Patel, A., Bhattamishra, S., and Goyal, N. (2021). Are NLP models really able to solve simple math word problems?. arXiv.","DOI":"10.18653\/v1\/2021.naacl-main.168"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Ling, W., Yogatama, D., Dyer, C., and Blunsom, P. (2017). Program induction by rationale generation: Learning to solve and explain algebraic word problems. arXiv.","DOI":"10.18653\/v1\/P17-1015"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Roy, S., and Roth, D. (2016). Solving general arithmetic word problems. arXiv.","DOI":"10.18653\/v1\/D15-1202"},{"key":"ref_59","unstructured":"Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., and Fan, A. (2024). The llama 3 herd of models. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1092\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T13:04:49Z","timestamp":1765285489000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1092"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,9]]},"references-count":59,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["info16121092"],"URL":"https:\/\/doi.org\/10.3390\/info16121092","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,9]]}}}