{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T13:27:07Z","timestamp":1783690027282,"version":"3.55.0"},"reference-count":50,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2024,6,6]],"date-time":"2024-06-06T00:00:00Z","timestamp":1717632000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"},{"start":{"date-parts":[[2024,6,6]],"date-time":"2024-06-06T00:00:00Z","timestamp":1717632000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Argument &amp; Computation"],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:p>\n                    In natural language understanding, a crucial goal is correctly interpreting open-textured phrases. In practice, disagreements over the meanings of open-textured phrases are often resolved through the generation and evaluation of\n                    <jats:italic>interpretive arguments<\/jats:italic>\n                    , arguments designed to support or attack a specific interpretation of an expression within a document. In this paper, we discuss some of our work towards the goal of automatically generating and evaluating interpretive arguments. We have curated a set of rules from the code of ethics of various professional organizations and a set of associated scenarios that are ambiguous with respect to some open-textured phrase within the rule. We collected and evaluated arguments from both human annotators and state-of-the-art generative language models in order to determine the relative quality and persuasiveness of both sets of arguments. Finally, we performed a Turing test-inspired study in order to assess whether human annotators can tell the difference between human arguments and machine-generated arguments. The results show that machine-generated arguments, when prompted a certain way, can be consistently rated as more convincing than human-generated arguments, and to the untrained eye, the machine-generated arguments can convincingly sound human-like.\n                  <\/jats:p>","DOI":"10.3233\/aac-230014","type":"journal-article","created":{"date-parts":[[2024,6,7]],"date-time":"2024-06-07T11:22:32Z","timestamp":1717759352000},"page":"362-404","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":2,"title":["Evaluating large language models\u2019 ability to generate interpretive arguments"],"prefix":"10.1177","volume":"16","author":[{"given":"Zaid","family":"Marji","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, University of South Florida, FL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"Licato","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, University of South Florida, FL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2024,6,6]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1177\/147078539703900202"},{"key":"e_1_3_3_3_2","author":"Bench-Capon T.","year":"2012","unstructured":"Bench-Capon T., Open texture and argumentation: What makes an argument persuasive?, in: Logic Programs, Norms and Action: Essays in Honor of Marek J. Sergot on the Occasion of His 60th Birthday, Artikis A., Craven R., \u00c7i\u00e7ekli N.K., Sadighi B., Stathis K., eds, Springer, 2012.","journal-title":"Logic Programs, Norms and Action: Essays in Honor of Marek J. Sergot on the Occasion of His 60th Birthday"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1093\/acref\/9780198735304.001.0001"},{"key":"e_1_3_3_5_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown T.","year":"2020","unstructured":"Brown T., Mann B., Ryder N., Subbiah M., Kaplan J.D., Dhariwal P., Neelakantan A., Shyam P., Sastry G., Askell A.et al., Language models are few-shot learners, Advances in neural information processing systems 33 (2020), 1877\u20131901.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_3_6_2","first-page":"4299","author":"Christiano P.F.","year":"2017","unstructured":"Christiano P.F., Leike J., Brown T., Martic M., Legg S., Amodei D., Deep reinforcement learning from human preferences, in: Advances in Neural Information Processing Systems, 2017, pp.\u00a04299\u20134307.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_7_2","author":"Fields L.","year":"2022","unstructured":"Fields L., Marji Z., Licato J., Of a different persuasion: Perception of minority status and persuasive impact, in: Proceedings of the Annual Meeting of the Cognitive Science Society, Vol.\u00a044, 2022.","journal-title":"Proceedings of the Annual Meeting of the Cognitive Science Society"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1093\/lpr\/mgs007"},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","unstructured":"Gao T. Fisch A. Chen D. Making pre-trained language models better few-shot learners 2020 arXiv preprint arXiv:2012.15723.","DOI":"10.18653\/v1\/2021.acl-long.295"},{"key":"e_1_3_3_10_2","author":"Hart H.L.A.","year":"1961","unstructured":"Hart H.L.A., The Concept of Law, Clarendon Press, 1961.","journal-title":"The Concept of Law"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.3233\/AAC-210026"},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","unstructured":"Huang J. Chang K.C.-C. Towards reasoning in large language models: A survey 2022 arXiv preprint arXiv:2212.10403.","DOI":"10.18653\/v1\/2023.findings-acl.67"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00324"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Jung J. Qin L. Welleck S. Brahman F. Bhagavatula C. Bras R.L. Choi Y. Maieutic prompting: Logically consistent reasoning with recursive explanations 2022 arXiv preprint arXiv:2205.11822.","DOI":"10.18653\/v1\/2022.emnlp-main.82"},{"key":"e_1_3_3_15_2","first-page":"22199","article-title":"Large language models are zero-shot reasoners","volume":"35","author":"Kojima T.","year":"2022","unstructured":"Kojima T., Gu S.S., Reid M., Matsuo Y., Iwasawa Y., Large language models are zero-shot reasoners, Advances in neural information processing systems 35 (2022), 22199\u201322213.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_3_16_2","author":"Licato J.","year":"2022","unstructured":"Licato J., Automated ethical reasoners must be interpretation-capable, in: Proceedings of the AAAI 2022 Spring Workshop on \u201cEthical Computing: Metrics for Measuring AI\u2019s Proficiency and Competency for Ethical Reasoning\u201d, 2022.","journal-title":"Proceedings of the AAAI 2022 Spring Workshop on \u201cEthical Computing: Metrics for Measuring AI\u2019s Proficiency and Competency for Ethical Reasoning\u201d"},{"key":"e_1_3_3_17_2","author":"Licato J.","year":"2022","unstructured":"Licato J., War-gaming needs argument-justified AI more than explainable AI, in: Proceedings of the 2022 Advances on Societal Digital Transformation (DIGITAL) Special Track on Explainable AI in Societal Games (XAISG), 2022.","journal-title":"Proceedings of the 2022 Advances on Societal Digital Transformation (DIGITAL) Special Track on Explainable AI in Societal Games (XAISG)"},{"key":"e_1_3_3_18_2","author":"Licato J.","year":"2022","unstructured":"Licato J., Fields L., Marji Z., Resolving open-textured rules with templated interpretive arguments, in: European Conference on Argumentation, 2022.","journal-title":"European Conference on Argumentation"},{"key":"e_1_3_3_19_2","author":"Licato J.","year":"2018","unstructured":"Licato J., Marji Z., Probing formal\/informal misalignment with the loophole task, in: Proceedings of the 2018 International Conference on Robot Ethics and Standards (ICRES 2018), 2018.","journal-title":"Proceedings of the 2018 International Conference on Robot Ethics and Standards (ICRES 2018)"},{"key":"e_1_3_3_20_2","author":"Licato J.","year":"2019","unstructured":"Licato J., Marji Z., Abraham S., Scenarios and recommendations for ethical interpretive AI, in: Proceedings of the AAAI 2019 Fall Symposium on Human-Centered AI, Arlington, VA, 2019.","journal-title":"Proceedings of the AAAI 2019 Fall Symposium on Human-Centered AI"},{"key":"e_1_3_3_21_2","unstructured":"Liu P. Yuan W. Fu J. Jiang Z. Hayashi H. Neubig G. Pre-train prompt and predict: A systematic survey of prompting methods in natural language processing 2021 arXiv preprint arXiv:2107.13586."},{"key":"e_1_3_3_22_2","unstructured":"Liu X. Zheng Y. Du Z. Ding M. Qian Y. Yang Z. Tang J. GPT understands too 2021 arXiv preprint arXiv:2103.10385."},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10982-017-9306-4"},{"key":"e_1_3_3_24_2","author":"MacCormick D.N.","year":"1991","unstructured":"MacCormick D.N., Summers R.S., Interpreting Statutes: A Comparative Study, Routledge, 1991.","journal-title":"Interpreting Statutes: A Comparative Study"},{"key":"e_1_3_3_25_2","author":"Marji Z.","year":"2021","unstructured":"Marji Z., Licato J., Aporia: The argumentation game, in: Proceedings of the Third Workshop on Argument Strength, 2021.","journal-title":"Proceedings of the Third Workshop on Argument Strength"},{"key":"e_1_3_3_26_2","unstructured":"OpenAI Introducing ChatGPT 2022."},{"key":"e_1_3_3_27_2","unstructured":"OpenAI GPT-4 technical report 2023."},{"key":"e_1_3_3_28_2","author":"Ouyang L.","year":"2022","unstructured":"Ouyang L., Wu J., Jiang X., Almeida D., Wainwright C.L., Mishkin P., Zhang C., Agarwal S., Slama K., Ray A., Schulman J., Hilton J., Kelton F., Miller L., Simens M., Askell A., Welinder P., Christiano P., Leike J., Lowe R., Training Language Models to Follow Instructions with Human Feedback, 2022.","journal-title":"Training Language Models to Follow Instructions with Human Feedback"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Petroni F. Rockt\u00e4schel T. Lewis P. Bakhtin A. Wu Y. Miller A.H. Riedel S. Language models as knowledge bases? 2019 arXiv preprint arXiv:1909.01066.","DOI":"10.18653\/v1\/D19-1250"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10506-017-9210-0"},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Press O. Zhang M. Min S. Schmidt L. Smith N.A. Lewis M. Measuring and narrowing the compositionality gap in language models 2022 arXiv preprint arXiv:2210.03350.","DOI":"10.18653\/v1\/2023.findings-emnlp.378"},{"key":"e_1_3_3_32_2","author":"Quandt R.","year":"2020","unstructured":"Quandt R., Licato J., Problems of autonomous agents following informal, open-textured rules, in: Human\u2013Machine Shared Contexts, Lawless W.F., Mittu R., Sofge D.A., eds, Academic Press, 2020.","journal-title":"Human\u2013Machine Shared Contexts"},{"issue":"8","key":"e_1_3_3_33_2","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford A.","year":"2019","unstructured":"Radford A., Wu J., Child R., Luan D., Amodei D., Sutskever I.et al., Language models are unsupervised multitask learners, OpenAI blog 1(8) (2019), 9.","journal-title":"OpenAI blog"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2746090.2746100"},{"key":"e_1_3_3_35_2","unstructured":"Sadasivan V.S. Kumar A. Balasubramanian S. Wang W. Feizi S. 2023 Can AI-generated text be reliably detected?"},{"key":"e_1_3_3_36_2","first-page":"21","author":"Sartor G.","year":"2014","unstructured":"Sartor G., Walton D., Macagno F., Rotolo A., Argumentation schemes for statutory interpretation: A logical analysis, in: Legal Knowledge and Information Systems. (Proceedings of JURIX, Vol.\u00a014, 2014, pp.\u00a021\u201328.","journal-title":"Legal Knowledge and Information Systems. (Proceedings of JURIX"},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","unstructured":"Schick T. Sch\u00fctze H. Exploiting cloze questions for few shot text classification and natural language inference 2020 arXiv preprint arXiv:2001.07676.","DOI":"10.18653\/v1\/2021.eacl-main.20"},{"key":"e_1_3_3_38_2","unstructured":"Singh V.P. Pal P. Survey of different types of CAPTCHA."},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1093\/mind\/LIX.236.433"},{"key":"e_1_3_3_40_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","unstructured":"Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser \u0141., Polosukhin I., Attention is all you need, Advances in neural information processing systems 30 (2017).","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1080\/0020174X.2020.1787222"},{"key":"e_1_3_3_42_2","author":"Waismann F.","year":"1965","unstructured":"Waismann F., The Principles of Linguistic Philosophy, St. Martins Press, 1965.","journal-title":"The Principles of Linguistic Philosophy"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","unstructured":"Wallace E. Feng S. Kandpal N. Gardner M. Singh S. Universal adversarial triggers for attacking and analyzing NLP 2019 arXiv preprint arXiv:1908.07125.","DOI":"10.18653\/v1\/D19-1221"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1017\/9781108554572"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10506-016-9179-0"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-90-481-9452-0_18"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3365000"},{"key":"e_1_3_3_48_2","unstructured":"Wei J. Bosma M. Zhao V.Y. Guu K. Yu A.W. Lester B. Du N. Dai A.M. Le Q.V. Finetuned language models are zero-shot learners 2021 arXiv preprint arXiv:2109.01652."},{"key":"e_1_3_3_49_2","unstructured":"Wei J. Wang X. Schuurmans D. Bosma M. Chi E. Le Q. Zhou D. Chain of thought prompting elicits reasoning in large language models 2022 arXiv preprint arXiv:2201.11903."},{"key":"e_1_3_3_50_2","unstructured":"Zhao W.X. Zhou K. Li J. Tang T. Wang X. Hou Y. Min Y. Zhang B. Zhang J. Dong Z.et al. A survey of large language models 2023 arXiv preprint arXiv:2303.18223."},{"key":"e_1_3_3_51_2","doi-asserted-by":"crossref","unstructured":"Zhong Z. Friedman D. Chen D. Factual probing is [MASK]: Learning vs. learning to recall 2021 arXiv preprint arXiv:2104.05240.","DOI":"10.18653\/v1\/2021.naacl-main.398"}],"container-title":["Argument &amp; Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/AAC-230014","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.3233\/AAC-230014","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.3233\/AAC-230014","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T11:53:30Z","timestamp":1777377210000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.3233\/AAC-230014"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,6]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["10.3233\/AAC-230014"],"URL":"https:\/\/doi.org\/10.3233\/aac-230014","relation":{},"ISSN":["1946-2166","1946-2174"],"issn-type":[{"value":"1946-2166","type":"print"},{"value":"1946-2174","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,6]]}}}