{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:00:17Z","timestamp":1784300417593,"version":"3.55.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"name":"Hong Kong SAR Research Grant Council\/General Research","award":["16205722"],"award-info":[{"award-number":["16205722"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    Differential testing offers a promising strategy to alleviate the test oracle problem by comparing the test results between alternative implementations. However, existing differential testing techniques for deep learning (DL) libraries are limited by the key challenges of finding alternative implementations (called\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(counterparts\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ) for a given API and subsequently generating diverse test inputs. To address the two challenges, this article introduces\n                    <jats:sc>DLLens<\/jats:sc>\n                    , a large language model (LLM)-enhanced differential testing technique for DL libraries. The first challenge is addressed by an observation that DL libraries are commonly designed to support the computation of a similar set of DL algorithms. Therefore, the counterpart of a given API\u2019s computation could be successfully synthesized through certain composition and adaptation of the APIs from another DL library.\n                    <jats:sc>DLLens<\/jats:sc>\n                    incorporates a novel counterpart synthesis workflow, leveraging a LLM to search for valid counterparts for differential testing. To address the second challenge,\n                    <jats:sc>DLLens<\/jats:sc>\n                    incorporates a static analysis technique that extracts the path constraints from the implementations of a given API and its counterpart to guide diverse test input generation. The extraction is facilitated by LLM\u2019s knowledge of the concerned DL library and its upstream libraries.\n                    <jats:sc>DLLens<\/jats:sc>\n                    incorporates validation mechanisms to manage the LLM\u2019s hallucinations in counterpart synthesis and path constraint extraction.\n                  <\/jats:p>\n                  <jats:p>\n                    We evaluate\n                    <jats:sc>DLLens<\/jats:sc>\n                    on two popular DL libraries, TensorFlow and PyTorch. Our evaluation shows that\n                    <jats:sc>DLLens<\/jats:sc>\n                    synthesizes counterparts for 1.84 times as many APIs as those found by state-of-the-art techniques on these libraries. Moreover, under the same time budget,\n                    <jats:sc>DLLens<\/jats:sc>\n                    covers 7.23% more branches and detects 1.88 times as many bugs as state-of-the-art techniques on 200 randomly sampled APIs.\n                    <jats:sc>DLLens<\/jats:sc>\n                    has successfully detected 71 bugs in recent TensorFlow and PyTorch libraries. Among them, 59 are confirmed by developers, including 46 confirmed as previously unknown bugs, and 10 of these previously unknown bugs have been fixed in the latest version of TensorFlow and PyTorch.\n                  <\/jats:p>","DOI":"10.1145\/3735637","type":"journal-article","created":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T12:24:27Z","timestamp":1747225467000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Enhancing Differential Testing with LLMs for Testing Deep Learning Libraries"],"prefix":"10.1145","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5947-4030","authenticated-orcid":false,"given":"Meiziniu","family":"Li","sequence":"first","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6482-8509","authenticated-orcid":false,"given":"Dongze","family":"Li","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-8423-7781","authenticated-orcid":false,"given":"Jianmeng","family":"Liu","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4892-6294","authenticated-orcid":false,"given":"Jialun","family":"Cao","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1644-2965","authenticated-orcid":false,"given":"Yongqiang","family":"Tian","sequence":"additional","affiliation":[{"name":"Monash University, Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3508-7172","authenticated-orcid":false,"given":"Shing-Chi","family":"Cheung","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,11]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI \u201916)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI \u201916). USENIX Association, 265\u2013283."},{"key":"e_1_3_3_3_2","unstructured":"ACETest. 2024. A GitHub Repository Maintaining the Source Code of ACETest. Retrieved October 2024 from https:\/\/github.com\/shijy16\/ACETest\/tree\/main"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/229867"},{"key":"e_1_3_3_5_2","unstructured":"Ned Batchelder. 2021. coverage.py. Retrieved from https:\/\/coverage.readthedocs.io\/en\/coverage-5.5\/"},{"key":"e_1_3_3_6_2","unstructured":"Max Brunsfeld. 2024. Tree-Sitter: A Parser Generator Tool and an Incremental Parsing Library. Retrieved from https:\/\/github.com\/tree-sitter\/tree-sitter"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.312"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","unstructured":"Junjie Chen Yihua Liang Qingchao Shen Jiajun Jiang and Shuochuan Li. 2023. Toward understanding deep learning framework bugs. ACM Transactions on Software Engineering and Methodology 32 6 Article 135 (Nov. 2023) 31 pages. DOI: 10.1145\/3587155","DOI":"10.1145\/3587155"},{"key":"e_1_3_3_9_2","unstructured":"Xinyun Chen Maxwell Lin Nathanael Sch\u00e4rli and Denny Zhou. 2023. Teaching large language models to self-debug. arXiv:2304.05128. Retrieved from https:\/\/arxiv.org\/abs\/2304.05128"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICC.2018.8422832"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598067"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597503.3623343"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549085"},{"key":"e_1_3_3_14_2","first-page":"7393","volume-title":"Proceedings of the 32nd USENIX Security Symposium (USENIX Security \u201923)","author":"Deng Zizhuang","year":"2023","unstructured":"Zizhuang Deng, Guozhu Meng, Kai Chen, Tong Liu, Lu Xiang, and Chunyang Chen. 2023. Differential testing of cross deep learning framework APIs: Revealing inconsistencies and vulnerabilities. In Proceedings of the 32nd USENIX Security Symposium (USENIX Security \u201923). USENIX Association, Anaheim, CA, 7393\u20137410. Retrieved from https:\/\/www.usenix.org\/conference\/usenixsecurity23\/presentation\/deng-zizhuang"},{"key":"e_1_3_3_15_2","unstructured":"DLLens. 2024. DLLens. Retrieved October 2024 from https:\/\/github.com\/maybeLee\/DLLens"},{"key":"e_1_3_3_16_2","unstructured":"DocTer. 2024. A GitHub Repository Maintaining the Source Code of DocTer. Retrieved October 2024 from https:\/\/github.com\/lin-tan\/DocTer\/tree\/main"},{"key":"e_1_3_3_17_2","unstructured":"Example. 2024. Example of False Positive. Retrieved March 2024 from https:\/\/github.com\/pytorch\/pytorch\/issues\/122426"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCOMM.2018.2878025"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510092"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416571"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"e_1_3_3_22_2","unstructured":"Gabor Horvath Reka Kovacs and Zoltan Porkolab. 2024. Scaling symbolic execution to large software systems. arXiv:2408.01909. Retrieved from https:\/\/arxiv.org\/abs\/2408.01909"},{"key":"e_1_3_3_23_2","unstructured":"Joern. 2024. The Bug Hunter\u2019s Workbench. Retrieved March 2024 from https:\/\/joern.io"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-011-4368-7"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649828"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","unstructured":"Meiziniu Li Jialun Cao Yongqiang Tian Tsz On Li Ming Wen and Shing-Chi Cheung. 2023. COMET: Coverage-guided model generation for deep learning library testing. ACM Transactions on Software Engineering and Methodology 32 5 Article 127 (July 2023) 34 pages. DOI: 10.1145\/3583566","DOI":"10.1145\/3583566"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.media.2017.07.005"},{"key":"e_1_3_3_28_2","unstructured":"Fang Liu Yang Liu Lin Shi Houkun Huang Ruifeng Wang Zhen Yang and Li Zhang. 2024. Exploring and evaluating hallucinations in LLM-powered code generation. arXiv:2404.00971. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:268819908"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","unstructured":"Jiawei Liu Yuheng Huang Zhijie Wang Lei Ma Chunrong Fang Mingzheng Gu Xufan Zhang and Zhenyu Chen. 2024. Generation-based differential fuzzing for deep learning libraries. ACM Transactions on Software Engineering and Methodology 33 2 Article 50 (Dec. 2023) 28 pages. DOI: 10.1145\/3628159","DOI":"10.1145\/3628159"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00037"},{"key":"e_1_3_3_31_2","unstructured":"Peter Oberparleiter. 2024. lcov\u2014A Graphical GCOV Front-End. Retrieved from https:\/\/manpages.ubuntu.com\/manpages\/xenial\/man1\/lcov.1.html"},{"key":"e_1_3_3_32_2","unstructured":"Tensorflow Onnx. 2022. tf2onnx\u2014Convert TensorFlow Keras Tensorflow.js and Tflite Models to ONNX. Retrieved from https:\/\/github.com\/onnx\/tensorflow-onnx"},{"key":"e_1_3_3_33_2","volume-title":"PyTorch: An Imperative Style, High-Performance Deep Learning Library","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. Curran Associates Inc., Red Hook, NY."},{"key":"e_1_3_3_34_2","first-page":"1027","volume-title":"Proceedings of the 2019 IEEE\/ACM 41st International Conference on Software Engineering (ICSE)","author":"Viet Pham Hung","year":"2019","unstructured":"Hung Viet Pham, Thibaud Lutellier, Weizhen Qi, and Lin Tan. 2019. CRADLE: Cross-backend validation to detect and localize bugs in deep learning libraries. In Proceedings of the 2019 IEEE\/ACM 41st International Conference on Software Engineering (ICSE). IEEE, 1027\u20131038."},{"key":"e_1_3_3_35_2","unstructured":"Python. 2024. Python AST Package. Retrieved March 2024 from https:\/\/docs.python.org\/3\/library\/ast.html"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00146"},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","unstructured":"Ahmad El Sallab Mohammed Abdou Etienne Perot and Senthil Yogamani. 2017. Deep reinforcement learning framework for autonomous driving. arXiv:1704.02532. Retrieved from https:\/\/arxiv.org\/abs\/1704.02532","DOI":"10.2352\/ISSN.2470-1173.2017.19.AVM-023"},{"key":"e_1_3_3_38_2","unstructured":"Shai Shalev-Shwartz Shaked Shammah and Amnon Shashua. 2016. Safe multi-agent reinforcement learning for autonomous driving. arXiv:1610.03295. Retrieved from https:\/\/arxiv.org\/abs\/1610.03295"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597926.3598088"},{"key":"e_1_3_3_40_2","unstructured":"stack-overflow. 2024. How Can I Know if a List Is Decreasing? (Python). Retrieved March 2024 from https:\/\/stackoverflow.com\/questions\/69576011\/how-can-i-know-if-a-list-is-decreasing-python"},{"key":"e_1_3_3_41_2","unstructured":"Florian Tambon Amin Nikanjam Le An Foutse Khomh and Giuliano Antoniol. 2021. Silent bugs in deep learning frameworks: An empirical study of Keras and TensorFlow. arXiv:2112.13314. Retrieved from https:\/\/arxiv.org\/abs\/2112.13314"},{"key":"e_1_3_3_42_2","unstructured":"TensorFlow. 2024. tf.math.is_non_decreasing Outputs Incorrect Result When Input Is an Uint Tensor. Retrieved March 2024 from https:\/\/github.com\/tensorflow\/tensorflow\/issues\/62072"},{"key":"e_1_3_3_43_2","unstructured":"TensorFlow. 2024. tf.tensor_scatter_nd_update Lead to a Program Abortion When Receiving a 3D Indices. Retrieved March 2024 from https:\/\/github.com\/tensorflow\/tensorflow\/issues\/63575"},{"key":"e_1_3_3_44_2","unstructured":"TensorFlow. 2024. tf.raw_ops.ApproximateEqual\/Erfinv Fail Support a Few Data Types Inconsistent with the Documentation. Retrieved October 2024 from https:\/\/github.com\/tensorflow\/tensorflow\/issues\/64256"},{"key":"e_1_3_3_45_2","unstructured":"TensorFlow. 2024. When Receiving Infinity Complex Tensor tf.math.sigmoid Outputs NaN. Retrieved October 2024 from https:\/\/github.com\/tensorflow\/tensorflow\/issues\/63051"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3395363.3397380"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510165"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3409761"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510041"},{"key":"e_1_3_3_50_2","unstructured":"Yixuan Weng Minjun Zhu Fei Xia Bin Li Shizhu He Shengping Liu Bin Sun Kang Liu and Jun Zhao. 2022. Large language models are better reasoners with self-verification. arXiv:2212.09561. Retrieved from https:\/\/arxiv.org\/abs\/2212.09561"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533767.3534220"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689736"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE48619.2023.00105"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2022.107004"},{"key":"e_1_3_3_55_2","unstructured":"z3py. 2024. z3py Namespace Reference. Retrieved September 2024 from https:\/\/z3prover.github.io\/api\/html\/namespacez3py.html"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460319.3464843"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3213846.3213866"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3735637","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T16:26:36Z","timestamp":1773246396000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3735637"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,11]]},"references-count":56,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3735637"],"URL":"https:\/\/doi.org\/10.1145\/3735637","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,11]]},"assertion":[{"value":"2024-10-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}