{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T15:07:55Z","timestamp":1783955275012,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":80,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,4,12]],"date-time":"2026-04-12T00:00:00Z","timestamp":1775952000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,12]]},"DOI":"10.1145\/3786580.3786972","type":"proceedings-article","created":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T14:37:06Z","timestamp":1783953426000},"page":"162-173","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Beyond Answer Engines: LLMs as Reasoning Partners in Data Structures and Algorithms Education"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-6522-8536","authenticated-orcid":false,"given":"Saad Zafar","family":"Khan","sequence":"first","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8329-0990","authenticated-orcid":false,"given":"Desiree","family":"Leal","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9044-104X","authenticated-orcid":false,"given":"Lucas","family":"Valen\u00e7a","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1863-9147","authenticated-orcid":false,"given":"Ahmad","family":"Abdellatif","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8400-9069","authenticated-orcid":false,"given":"Mea","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6098-4801","authenticated-orcid":false,"given":"Diwakar","family":"Krishnamurthy","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3235-6530","authenticated-orcid":false,"given":"Ronnie de Souza","family":"Santos","sequence":"additional","affiliation":[{"name":"University of Calgary, Calgary, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,13]]},"reference":[{"key":"e_1_3_3_2_2_2","volume-title":"LeetCode Weekly Contest 431","year":"2025","unstructured":"2025. LeetCode Weekly Contest 431. https:\/\/leetcode.com\/contest\/weekly-contest-431\/ Accessed 8 September 2025; contest date from standings resource."},{"key":"e_1_3_3_2_3_2","volume-title":"LeetCode Weekly Contest 452","year":"2025","unstructured":"2025. LeetCode Weekly Contest 452. https:\/\/leetcode.com\/contest\/weekly-contest-452\/ Accessed 8 September 2025; contest date from standings resource."},{"key":"e_1_3_3_2_4_2","unstructured":"Asma\u00a0Ben Abacha Wen wai Yim Yujuan Fu Zhaoyi Sun Meliha Yetisgen Fei Xia and Thomas Lin. 2025. MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes. arxiv:https:\/\/arXiv.org\/abs\/2412.19260\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2412.19260"},{"key":"e_1_3_3_2_5_2","unstructured":"Matin Amoozadeh David Daniels Daye Nam Aayush Kumar Stella Chen Michael Hilton Sruti\u00a0Srinivasa Ragavan and Mohammad\u00a0Amin Alipour. 2024. Trust in Generative AI among students: An Exploratory Study. arxiv:https:\/\/arXiv.org\/abs\/2310.04631\u00a0[cs.HC] https:\/\/arxiv.org\/abs\/2310.04631"},{"key":"e_1_3_3_2_6_2","volume-title":"Claude 3.7 Sonnet and Claude Code","year":"2025","unstructured":"Anthropic. 2025. Claude 3.7 Sonnet and Claude Code. https:\/\/www.anthropic.com\/news\/claude-3-7-sonnet"},{"key":"e_1_3_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Mark Ardis David Budgen Gregory\u00a0W. Hislop Jeff Offutt Mark Sebern and Willem Visser. 2015. SE 2014: Curriculum Guidelines for Undergraduate Degree Programs in Software Engineering. Computer 48 11 (2015) 106\u2013109. doi:10.1109\/MC.2015.345","DOI":"10.1109\/MC.2015.345"},{"key":"e_1_3_3_2_8_2","doi-asserted-by":"crossref","unstructured":"R. Azoulay T. Hirst and S. Reches. 2025. Large Language Models in Computer Science Classrooms: Ethical Challenges and Strategic Solutions. Applied Sciences 15 4 (2025) 1793. doi:10.3390\/app15041793","DOI":"10.3390\/app15041793"},{"key":"e_1_3_3_2_9_2","unstructured":"Steve\u00a0Olusegun Bada. 2015. Constructivism Learning Theory: A Paradigm for Teaching and Learning. IOSR Journal of Research & Method in Education (IOSR-JRME) 5 6 Ver. I (2015) 66\u201370. doi:10.9790\/7388-05616670"},{"key":"e_1_3_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Yoav Benjamini and Yosef Hochberg. 1995. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological) 57 1 (1995) 289\u2013300. doi:10.1111\/j.2517-6161.1995.tb02031.x","DOI":"10.1111\/j.2517-6161.1995.tb02031.x"},{"key":"e_1_3_3_2_11_2","volume-title":"Taxonomy of Educational Objectives: Handbook I, Cognitive Domain","author":"Bloom Benjamin\u00a0S.","year":"1956","unstructured":"Benjamin\u00a0S. Bloom, Max\u00a0D. Engelhart, Edward\u00a0J. Furst, Walter\u00a0H. Hill, and David\u00a0R. Krathwohl. 1956. Taxonomy of Educational Objectives: Handbook I, Cognitive Domain. David McKay Company, New York. Handbook I of *Handbook: The Classification of Educational Goals*."},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3636243.3636257"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"crossref","unstructured":"C.K.Y. Chan. 2023. A Comprehensive AI Policy Education Framework for University Teaching and Learning. International Journal of Educational Technology in Higher Education 20 1 (2023) 38. doi:10.1186\/s41239-023-00408-3","DOI":"10.1186\/s41239-023-00408-3"},{"key":"e_1_3_3_2_14_2","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde\u00a0de Oliveira\u00a0Pinto Jared Kaplan Harri Edwards et\u00a0al. 2021. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2107.03374 (2021)."},{"key":"e_1_3_3_2_15_2","unstructured":"Yucheng Chu Hang Li Kaiqi Yang Harry Shomer Hui Liu Yasemin Copur-Gencturk and Jiliang Tang. 2025. A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization. arxiv:https:\/\/arXiv.org\/abs\/2410.02165\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2410.02165"},{"key":"e_1_3_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Norman Cliff. 1993. Dominance Statistics: Ordinal Analyses to Answer Ordinal Questions. Psychological Bulletin 114 3 (1993) 494\u2013509.","DOI":"10.1037\/0033-2909.114.3.494"},{"key":"e_1_3_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Jacob Cohen. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20 1 (1960) 37\u201346. doi:10.1177\/001316446002000104","DOI":"10.1177\/001316446002000104"},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3661167.3661221"},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"crossref","unstructured":"D. Coleman D. Ash B. Lowther and P. Oman. 1994. Using metrics to evaluate software system maintainability. Computer 27 8 (1994) 44\u201349. doi:10.1109\/2.303623","DOI":"10.1109\/2.303623"},{"key":"e_1_3_3_2_20_2","unstructured":"Giuseppe Crupi Rosalia Tufano Alejandro Velasco Antonio Mastropaolo Denys Poshyvanyk and Gabriele Bavota. 2025. On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization. arxiv:https:\/\/arXiv.org\/abs\/2507.16587\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2507.16587"},{"key":"e_1_3_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Paul Denny James Prather Brett\u00a0A. Becker James Finnie-Ansley Arto Hellas Juho Leinonen Andrew Luxton-Reilly Brent\u00a0N. Reeves Eddie\u00a0Antonio Santos and Sami Sarsa. 2024. Computing Education in the Era of Generative AI. Commun. ACM 67 2 (Jan. 2024) 56\u201367. doi:10.1145\/3624720","DOI":"10.1145\/3624720"},{"key":"e_1_3_3_2_22_2","unstructured":"Aaron\u00a0Grattafiori et al.2024. The Llama 3 Herd of Models. arxiv:https:\/\/arXiv.org\/abs\/2407.21783\u00a0[cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_3_2_23_2","unstructured":"DeepSeek-AI et al.2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arxiv:https:\/\/arXiv.org\/abs\/2501.12948\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2501.12948"},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Guangrui Fan Dandan Liu Rui Zhang and Lihu Pan. 2025. The impact of AI-assisted pair programming on student motivation programming anxiety collaborative learning and programming performance: a comparative study with traditional pair programming and individual approaches. International Journal of STEM Education 12 1 (2025) 16. doi:10.1186\/s40594-025-00537-3","DOI":"10.1186\/s40594-025-00537-3"},{"key":"e_1_3_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Sue Fitzgerald Gary Lewandowski Ren\u00e9e McCauley Laurie Murphy Beth Simon Lynda Thomas and Carol Zander. 2008. Debugging: finding fixing and flailing a multi-institutional study of novice debuggers. Computer Science Education 18 2 (2008) 93\u2013116. doi:10.1080\/08993400802114508","DOI":"10.1080\/08993400802114508"},{"key":"e_1_3_3_2_26_2","first-page":"231","volume-title":"The Nature of Intelligence","author":"Flavell John\u00a0H.","year":"1976","unstructured":"John\u00a0H. Flavell. 1976. Metacognitive aspects of problem solving. In The Nature of Intelligence, Lauren\u00a0B. Resnick (Ed.). Lawrence Erlbaum Associates, 231\u2013235."},{"key":"e_1_3_3_2_27_2","series-title":"(SIGCSE \u201915)","doi-asserted-by":"crossref","first-page":"452","DOI":"10.1145\/2676723.2677311","volume-title":"Proceedings of the 46th ACM Technical Symposium on Computer Science Education","author":"Ginat David","year":"2015","unstructured":"David Ginat and Eti Menashe. 2015. SOLO Taxonomy for Assessing Novices\u2019 Algorithmic Design. In Proceedings of the 46th ACM Technical Symposium on Computer Science Education (Kansas City, Missouri, USA) (SIGCSE \u201915). Association for Computing Machinery, New York, NY, USA, 452\u2013457. doi:10.1145\/2676723.2677311"},{"key":"e_1_3_3_2_28_2","volume-title":"Anthropic Education Report: How University Students Use Claude","author":"Handa Kunal","year":"2025","unstructured":"Kunal Handa, Drew Bent, Alex Tamkin, Miles McCain, Esin Durmus, Michael Stern, Mike Schiraldi, Saffron Huang, Stuart Ritchie, Steven Syverud, Kamya Jagadish, Margaret Vo, Matt Bell, and Deep Ganguli. 2025. Anthropic Education Report: How University Students Use Claude. https:\/\/www.anthropic.com\/news\/anthropic-education-report-how-university-students-use-claude"},{"key":"e_1_3_3_2_29_2","unstructured":"Kaivalya Hariharan Uzay Girit Atticus Wang and Jacob Andreas. 2025. Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents. arxiv:https:\/\/arXiv.org\/abs\/2506.00172\u00a0[cs.LG] https:\/\/arxiv.org\/abs\/2506.00172"},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.5555\/1995312"},{"key":"e_1_3_3_2_31_2","unstructured":"Dan Hendrycks Steven Basart Saurav Kadavath Mantas Mazeika Akul Arora Ethan Guo Collin Burns Samir Puranik Horace He Dawn Song and Jacob Steinhardt. 2021. Measuring Coding Challenge Competence With APPS. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2105.09938 (2021)."},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Tja\u0161a Heri\u010dko and Bo\u0161tjan \u0160umak. 2023. Exploring Maintainability Index Variants for Software Maintainability Measurement in Object-Oriented Systems. Applied Sciences 13 5 (2023). doi:10.3390\/app13052972","DOI":"10.3390\/app13052972"},{"key":"e_1_3_3_2_33_2","unstructured":"Irene Hou Sophia Metille Zhuo Li Owen Man Cynthia Zastudil and Stephen MacNeil. 2024. The Effects of Generative AI on Computing Students\u2019 Help-Seeking Preferences. arxiv:https:\/\/arXiv.org\/abs\/2401.02262\u00a0[cs.HC] https:\/\/arxiv.org\/abs\/2401.02262"},{"key":"e_1_3_3_2_34_2","unstructured":"Naman Jain King Han Alex Gu Wen-Ding Li Fanjia Yan Tianjun Zhang Sida Wang Armando Solar-Lezama Koushik Sen and Ion Stoica. 2024. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. arxiv:https:\/\/arXiv.org\/abs\/2403.07974\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2403.07974"},{"key":"e_1_3_3_2_35_2","unstructured":"Hongchao Jiang Yiming Chen Yushi Cao Hung yi Lee and Robby\u00a0T. Tan. 2025. CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks. arxiv:https:\/\/arXiv.org\/abs\/2507.10535\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2507.10535"},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"crossref","unstructured":"E. Kasneci K. Sessler S. K\u00fcchemann M. Bannert D. Dementieva F. Fischer et\u00a0al. 2023. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences (2023). doi:10.1016\/j.lindif.2023.102274","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_3_3_2_37_2","volume-title":"CHI \u201924","author":"Kazemitabaar Majeed","year":"2024","unstructured":"Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, et\u00a0al. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. In CHI \u201924. doi:10.1145\/3613904.3642773"},{"key":"e_1_3_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Hieke Keuning Isaac Alpizar-Chacon Ioanna Lykourentzou Lauren Beehler Christian K\u00f6ppe Imke de Jong and Sergey Sosnovsky. 2024. Students\u2019 Perceptions and Use of Generative AI Tools for Programming Across Different Computing Courses. arxiv:https:\/\/arXiv.org\/abs\/2410.06865\u00a0[cs.CY] https:\/\/arxiv.org\/abs\/2410.06865","DOI":"10.1145\/3699538.3699546"},{"key":"e_1_3_3_2_39_2","unstructured":"Saad\u00a0Zafar Khan Desiree Leal Lucas Valen\u00e7a Ahmad Abdellatif Mea Wang Diwakar Krishnamurthy and Ronnie de Souza\u00a0Santos. 2025. Replication Package for: Beyond Answer Engines: LLMs as Reasoning Partners in Data Structures and Algorithms Education. https:\/\/zenodo.org\/records\/17229516"},{"key":"e_1_3_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664191"},{"key":"e_1_3_3_2_41_2","unstructured":"Michele Lacchia. 2023. Radon: Code Metrics in Python (v6.0.1). https:\/\/pypi.org\/project\/radon\/. Python package MIT License computes code metrics including cyclomatic complexity raw metrics Halstead metrics and Maintainability Index."},{"key":"e_1_3_3_2_42_2","doi-asserted-by":"crossref","unstructured":"K.\u00a0Y. Lau and S. Sotiriadis. 2023. Learning to program with large language models: A case study with ChatGPT. Proceedings of the 28th ACM Conference on Innovation and Technology in Computer Science Education 1 (2023). doi:10.1145\/3587102.3588830","DOI":"10.1145\/3587102.3588830"},{"key":"e_1_3_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3568813.3600138"},{"key":"e_1_3_3_2_44_2","doi-asserted-by":"crossref","unstructured":"D. Lee and E. Palmer. 2025. Prompt engineering in higher education: a systematic review to help inform curricula. International Journal of Educational Technology in Higher Education 22 7 (2025). doi:10.1186\/s41239-025-00503-7","DOI":"10.1186\/s41239-025-00503-7"},{"key":"e_1_3_3_2_45_2","unstructured":"LeetCode. 2025. LeetCode Contest. https:\/\/leetcode.com\/contest\/. Accessed: 2025-02-28."},{"key":"e_1_3_3_2_46_2","unstructured":"LeetCode. 2025. Top Interview 150. https:\/\/leetcode.com\/studyplan\/top-interview-150\/. Accessed: 2025-02-28."},{"key":"e_1_3_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3587102.3588785"},{"key":"e_1_3_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641554.3701946"},{"key":"e_1_3_3_2_49_2","doi-asserted-by":"crossref","unstructured":"Yi Liu et\u00a0al. 2025. Evaluating LLMs for Automated Scoring in Formative Assessments of a Programming Course. Applied Sciences 15 5 (2025) 2787. doi:10.3390\/app15052787","DOI":"10.3390\/app15052787"},{"key":"e_1_3_3_2_50_2","doi-asserted-by":"crossref","unstructured":"Wenhan Lyu Yimeng Wang Tingting\u00a0(Rachel) Chung Yifan Sun and Yixuan Zhang. 2024. Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study. 63\u201374\u00a0pages. doi:10.1145\/3657604.3662036","DOI":"10.1145\/3657604.3662036"},{"key":"e_1_3_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Lauren\u00a0E. Margulieux and Richard Catrambone. 2016. Improving problem solving with subgoal labels in expository text and worked examples. Learning and Instruction 42 (2016) 58\u201371. doi:10.1016\/j.learninstruc.2015.12.002","DOI":"10.1016\/j.learninstruc.2015.12.002"},{"key":"e_1_3_3_2_52_2","doi-asserted-by":"crossref","unstructured":"Thomas\u00a0J. McCabe. 1976. A Complexity Measure. IEEE Transactions on Software Engineering SE-2 4 (1976) 308\u2013320. doi:10.1109\/TSE.1976.233837","DOI":"10.1109\/TSE.1976.233837"},{"key":"e_1_3_3_2_53_2","doi-asserted-by":"crossref","unstructured":"R. McCauley S. Fitzgerald G. Lewandowski L. Murphy B. Simon L. Thomas and C. Zander. 2008. Teaching debugging skills in the 21st century. Proceedings of the 13th annual conference on Innovation and technology in computer science education (2008). doi:10.1145\/1384271.1384387","DOI":"10.1145\/1384271.1384387"},{"key":"e_1_3_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Quinn McNemar. 1947. Note on the Sampling Error of the Difference between Correlated Proportions or Percentages. Psychometrika 12 2 (1947) 153\u2013157. doi:10.1007\/BF02295996","DOI":"10.1007\/BF02295996"},{"key":"e_1_3_3_2_55_2","volume-title":"Assigning AI: Seven Approaches for Students, with Prompts","author":"Mollick Ethan\u00a0R.","year":"2023","unstructured":"Ethan\u00a0R. Mollick and Lilach Mollick. 2023. Assigning AI: Seven Approaches for Students, with Prompts. Technical Report. The Wharton School Research Paper. doi:10.2139\/ssrn.4475995Available at SSRN: https:\/\/ssrn.com\/abstract=4475995 or http:\/\/dx.doi.org\/10.2139\/ssrn.4475995."},{"key":"e_1_3_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSM.1992.242525"},{"key":"e_1_3_3_2_57_2","series-title":"(ISSTA 2024)","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1145\/3650212.3652140","volume-title":"Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis","author":"Ouyang Yicheng","year":"2024","unstructured":"Yicheng Ouyang, Jun Yang, and Lingming Zhang. 2024. Benchmarking Automated Program Repair: An Extensive Study on Both Real-World and Artificial Bugs. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (Vienna, Austria) (ISSTA 2024). Association for Computing Machinery, New York, NY, USA, 440\u2013452. doi:10.1145\/3650212.3652140"},{"key":"e_1_3_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3702652.3744220"},{"key":"e_1_3_3_2_59_2","doi-asserted-by":"crossref","unstructured":"Haritz Puerto Martin Tutek Somak Aditya Xiaodan Zhu and Iryna Gurevych. 2024. Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.10065 (2024).","DOI":"10.18653\/v1\/2024.emnlp-main.629"},{"key":"e_1_3_3_2_60_2","doi-asserted-by":"crossref","unstructured":"Alexander Renkl. 2014. Toward an instructionally oriented theory of example-based learning. Cognitive Science 38 1 (2014) 1\u201337. doi:10.1111\/cogs.12086","DOI":"10.1111\/cogs.12086"},{"key":"e_1_3_3_2_61_2","doi-asserted-by":"crossref","unstructured":"Anthony Robins Jennifer Rountree and Nathan Rountree. 2003. Learning and teaching programming: A review and discussion. Computer Science Education 13 2 (2003) 137\u2013172. doi:10.1076\/csed.13.2.137.14200","DOI":"10.1076\/csed.13.2.137.14200"},{"key":"e_1_3_3_2_62_2","doi-asserted-by":"crossref","first-page":"10776","DOI":"10.18653\/v1\/2023.findings-emnlp.722","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Sainz Oscar","year":"2023","unstructured":"Oscar Sainz, Jon Campos, Iker Garc\u00eda-Ferrero, Julen Etxaniz, Oier\u00a0Lopez de Lacalle, and Eneko Agirre. 2023. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 10776\u201310787. doi:10.18653\/v1\/2023.findings-emnlp.722"},{"key":"e_1_3_3_2_63_2","doi-asserted-by":"crossref","unstructured":"Ana Stojanov Qian Liu and Joyce Hwee\u00a0Ling Koh. 2024. University students\u2019 self-reported reliance on ChatGPT for learning: A latent profile analysis. Computers and Education: Artificial Intelligence 6 (2024) 100243. doi:10.1016\/j.caeai.2024.100243","DOI":"10.1016\/j.caeai.2024.100243"},{"key":"e_1_3_3_2_64_2","doi-asserted-by":"crossref","unstructured":"Marielle\u00a0Justine Sumilong. 2025. Instructional affect and learner motivation in generative AI-restrictive and permissive classrooms. Frontiers in Education Volume 10 - 2025 (2025). doi:10.3389\/feduc.2025.1626802","DOI":"10.3389\/feduc.2025.1626802"},{"key":"e_1_3_3_2_65_2","unstructured":"Wannita Takerngsaksiri Cleshan Warusavitarne Christian Yaacoub Matthew Hee\u00a0Keng Hou and Chakkrit Tantithamthavorn. 2024. Students\u2019 Perspective on AI Code Completion: Benefits and Challenges. arxiv:https:\/\/arXiv.org\/abs\/2311.00177\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2311.00177"},{"key":"e_1_3_3_2_66_2","unstructured":"Lun Wang Chuanqi Shi Shaoshui Du Yiyi Tao Yixian Shen Hang Zheng Yanxin Shen and Xinyu Qiu. 2025. Performance Review on LLM for solving leetcode problems. arxiv:https:\/\/arXiv.org\/abs\/2502.15770\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2502.15770"},{"key":"e_1_3_3_2_67_2","unstructured":"Shen Wang Tianlong Xu Hang Li Chaoli Zhang Joleen Liang Jiliang Tang Philip\u00a0S. Yu and Qingsong Wen. 2024. Large Language Models for Education: A Survey and Outlook. arxiv:https:\/\/arXiv.org\/abs\/2403.18105\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2403.18105"},{"key":"e_1_3_3_2_68_2","unstructured":"Xuezhi Wang Jason Wei Dale Schuurmans et\u00a0al. 2022. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2203.11171 (2022)."},{"key":"e_1_3_3_2_69_2","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans et\u00a0al. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2201.11903 (2022)."},{"key":"e_1_3_3_2_70_2","unstructured":"Colin White Samuel Dooley Manley Roberts Arka Pal Ben Feuer Siddhartha Jain Ravid Shwartz-Ziv Neel Jain Khalid Saifullah Sreemanti Dey Shubh-Agrawal Sandeep\u00a0Singh Sandha Siddartha Naidu Chinmay Hegde Yann LeCun Tom Goldstein Willie Neiswanger and Micah Goldblum. 2025. LiveBench: A Challenging Contamination-Limited LLM Benchmark. arxiv:https:\/\/arXiv.org\/abs\/2406.19314\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2406.19314"},{"key":"e_1_3_3_2_71_2","doi-asserted-by":"crossref","unstructured":"Frank Wilcoxon. 1945. Individual Comparisons by Ranking Methods. Biometrics Bulletin 1 6 (1945) 80\u201383.","DOI":"10.2307\/3001968"},{"key":"e_1_3_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00129"},{"key":"e_1_3_3_2_73_2","unstructured":"Yunhui Xia Wei Shen Yan Wang Jason\u00a0Klein Liu Huifeng Sun Siyue Wu Jian Hu and Xiaolong Xu. 2025. LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs. arxiv:https:\/\/arXiv.org\/abs\/2504.14655\u00a0[cs.LG]"},{"key":"e_1_3_3_2_74_2","unstructured":"Jessica Xing. 2021. Here\u2019s what job seekers need to know about LeetCode the coding-skills platform millions of developers use to ace the notoriously difficult technical interviews at firms such as Apple Amazon and Google. Business Insider (Nov. 2021). https:\/\/www.businessinsider.com\/leetcode-coding-test-apple-amazon-google-technical-interview-prep-job-2021-11"},{"key":"e_1_3_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639474.3640076"},{"key":"e_1_3_3_2_76_2","volume-title":"ICER \u201924","author":"Yang Stephanie","year":"2024","unstructured":"Stephanie Yang, Hanzhang Zhao, Yudian Xu, Karen Brennan, and Bertrand Schneider. 2024. Debugging with an AI Tutor: Investigating Novice Help-seeking Behaviors and Perceived Learning. In ICER \u201924. doi:10.1145\/3632620.3671092"},{"key":"e_1_3_3_2_77_2","unstructured":"Shunyu Yao Dian Bosma Jeffrey Zhao et\u00a0al. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2305.10601 (2023)."},{"key":"e_1_3_3_2_78_2","volume-title":"Proceedings of ICSME 2025, NIER Track","author":"Yuen Adam","year":"2025","unstructured":"Adam Yuen, John Pangas, Md\u00a0Mainul\u00a0Hasan Polash, and Ahmad Abdellatif. 2025. Prompting Matters: Assessing the Effect of Prompting Techniques on LLM\u2010Generated Class Code. In Proceedings of ICSME 2025, NIER Track. Case Room 3, ICSME 2025."},{"key":"e_1_3_3_2_79_2","doi-asserted-by":"crossref","unstructured":"Chunpeng Zhai Santoso Wibowo and Lily\u00a0D. Li. 2024. The effects of over-reliance on AI dialogue systems on students\u2019 cognitive abilities: a systematic review. Smart Learning Environments 11 (2024). doi:10.1186\/s40561-024-00316-7","DOI":"10.1186\/s40561-024-00316-7"},{"key":"e_1_3_3_2_80_2","unstructured":"Kyrie\u00a0Zhixuan Zhou Zachary Kilhoffer Madelyn\u00a0Rose Sanfilippo Ted Underwood Ece Gumusel Mengyi Wei Abhinav Choudhry and Jinjun Xiong. 2024. The teachers are confused as well: A Multiple-Stakeholder Ethics Discussion on Large Language Models in Computing Education. arxiv:https:\/\/arXiv.org\/abs\/2401.12453https:\/\/arxiv.org\/abs\/2401.12453"},{"key":"e_1_3_3_2_81_2","unstructured":"Terry\u00a0Yue Zhuo Minh\u00a0Chien Vu Jenny Chim Han Hu Wenhao Yu and et al.2025. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.15877 (2025)."}],"event":{"name":"ICSE-SEET '26: 2026 IEEE\/ACM 48th International Conference on Software Engineering","location":"Rio de Janeiro Brazil","acronym":"ICSE-SEET '26","sponsor":["SIGSOFT ACM Special Interest Group on Software Engineering"]},"container-title":["Proceedings of the IEEE\/ACM 48th International Conference on Software Engineering"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786580.3786972","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T14:37:36Z","timestamp":1783953456000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786580.3786972"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,12]]},"references-count":80,"alternative-id":["10.1145\/3786580.3786972","10.1145\/3786580"],"URL":"https:\/\/doi.org\/10.1145\/3786580.3786972","relation":{},"subject":[],"published":{"date-parts":[[2026,4,12]]},"assertion":[{"value":"2026-07-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}