{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T09:05:07Z","timestamp":1784797507927,"version":"3.55.0"},"publisher-location":"Cham","reference-count":26,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032325365","type":"print"},{"value":"9783032325372","type":"electronic"}],"license":[{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T00:00:00Z","timestamp":1784851200000},"content-version":"vor","delay-in-days":204,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Machine Learning (ML) inference is shifting from using pre-developed static, CUDA C++, GPU kernel libraries to using MLIR-based graph compilers that perform advanced optimizations and generate custom kernels. This paradigm shift reimagines how we achieve ML inference in safety-critical domains such as automotive applications. Traditional approaches relied on qualifying static kernel libraries\u2014pre-built for fixed input shapes and parameter ranges\u2014according to the ISO 26262 standard. However, the demanding performance requirements of diverse ML models and rapidly evolving hardware accelerators necessitate generating optimized kernels on the fly, which only ML graph compilers can provide.<\/jats:p>\n                  <jats:p>This paper presents an industrial experience report on a comprehensive verification framework for ML inference in automotive applications. We describe the transition from static kernels to dynamic ML graph compilation and introduce two complementary verification strategies: (1) formal methods targeting memory safety and concurrency properties in CUDA kernels and MLIR-based compiler; and (2) AI-driven testing for functional correctness. Our experience over multiple years of production use demonstrates that validating ML graph compiler output can satisfy the ISO 26262 ASIL B requirements - without requiring compiler tool qualification - while enabling performance and flexibility benefits. We discuss remaining challenges including scalability of formal verification and adapting to evolving compilers and hardware platforms.<\/jats:p>","DOI":"10.1007\/978-3-032-32537-2_26","type":"book-chapter","created":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:41:58Z","timestamp":1784796118000},"page":"550-564","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Ensuring Safety in\u00a0Automotive Machine Learning Inference: From Pre-validated Static Kernels to\u00a0Machine Learning Graph Compilation"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-5559-1672","authenticated-orcid":false,"given":"Jelena","family":"Frtunikj","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9342-2211","authenticated-orcid":false,"given":"Alex","family":"Latz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ajit","family":"Mistry","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthew","family":"Propp","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-5468-7899","authenticated-orcid":false,"given":"Vasu","family":"Singh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Suresh","family":"Talapaneni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4076-7750","authenticated-orcid":false,"given":"Amanda","family":"Tang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3197-8736","authenticated-orcid":false,"given":"Damien","family":"Zufferey","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,24]]},"reference":[{"key":"26_CR1","volume-title":"SPARK: The Proven Approach to High Integrity Software","author":"J Barnes","year":"2012","unstructured":"Barnes, J.: SPARK: The Proven Approach to High Integrity Software. Altran Praxis, London, GBR (2012)"},{"key":"26_CR2","doi-asserted-by":"publisher","unstructured":"Betts, A., Chong, N., Donaldson, A.F., Qadeer, S., Thomson, P.: GPUVerify: a verifier for GPU kernels. In: Leavens, G.T., Dwyer, M.B. (eds.) Proceedings of the 27th Annual ACM SIGPLAN Conference on Object-Oriented Programming, Systems, Languages, and Applications, OOPSLA 2012, part of SPLASH 2012, Tucson, AZ, USA, October 21-25, 2012, pp. 113\u2013132. ACM (2012). https:\/\/doi.org\/10.1145\/2384616.2384625","DOI":"10.1145\/2384616.2384625"},{"key":"26_CR3","doi-asserted-by":"publisher","unstructured":"Blanchet, B., Cousot, P., Cousot, R., Feret, J., Mauborgne, L., Min\u00e9, A., Monniaux, D., Rival, X.: A static analyzer for large safety-critical software. In: Cytron, R., Gupta, R. (eds.) Proceedings of the ACM SIGPLAN 2003 Conference on Programming Language Design and Implementation 2003, San Diego, California, USA, June 9-11, 2003, pp. 196\u2013207. ACM (2003). https:\/\/doi.org\/10.1145\/781131.781153","DOI":"10.1145\/781131.781153"},{"key":"26_CR4","doi-asserted-by":"publisher","unstructured":"Chakraborty, S., Krishna, S., Pavlogiannis, A., Tuppe, O.: GPUMC: a stateless model checker for GPU weak memory concurrency. In: Piskac, R., Rakamaric, Z. (eds.) Computer Aided Verification - 37th International Conference, CAV 2025, Zagreb, Croatia, July 23-25, 2025, Proceedings, Part III. Lecture Notes in Computer Science, vol. 15933, pp. 321\u2013346. Springer (2025). https:\/\/doi.org\/10.1007\/978-3-031-98682-6_17, https:\/\/doi.org\/10.1007\/978-3-031-98682-6_17","DOI":"10.1007\/978-3-031-98682-6_17"},{"key":"26_CR5","unstructured":"Chen, T., et al.: TVM: an automated end-to-end optimizing compiler for deep learning. In: Arpaci-Dusseau, A.C., Voelker, G. (eds.) 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 2018, pp. 578\u2013594. USENIX Association (2018). https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen"},{"key":"26_CR6","unstructured":"Dao, T.: FlashAttention-2: faster attention with better parallelism and work partitioning. In: The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net (2024). https:\/\/openreview.net\/forum?id=mZn2Xyh9Ec"},{"key":"26_CR7","unstructured":"Dao, T., Fu, D.Y., Ermon, S., Rudra, A., R\u00e9, C.: FlashAttention: fast and memory-efficient exact attention with io-awareness. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 (2022). http:\/\/papers.nips.cc\/paper_files\/paper\/2022\/hash\/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html"},{"key":"26_CR8","unstructured":"Hazy Research: No Bubbles: Eliminating Pipeline Bubbles with Megakernels. Blog post (2025). https:\/\/hazyresearch.stanford.edu\/blog\/2025-05-27-no-bubbles. Accessed 24 April 2026"},{"key":"26_CR9","doi-asserted-by":"publisher","unstructured":"Kuppe, M.A., Lamport, L., Ricketts, D.: The TLA+ toolbox. In: Monahan, R., Prevosto, V., Proen\u00e7a, J. (eds.) Proceedings Fifth Workshop on Formal Integrated Development Environment, F-IDE@FM 2019, Porto, Portugal, 7th October 2019. EPTCS, vol.\u00a0310, pp. 50\u201362 (2019). https:\/\/doi.org\/10.4204\/EPTCS.310.6, https:\/\/doi.org\/10.4204\/EPTCS.310.6","DOI":"10.4204\/EPTCS.310.6"},{"key":"26_CR10","unstructured":"Lamport, L.: Specifying Systems, The TLA+ Language and Tools for Hardware and Software Engineers. Addison-Wesley (2002). http:\/\/research.microsoft.com\/users\/lamport\/tla\/book.html"},{"key":"26_CR11","unstructured":"Lattner, C., et al.: MLIR: a compiler infrastructure for the end of Moore\u2019s law. CoRR abs\/2002.11054 (2020), https:\/\/arxiv.org\/abs\/2002.11054"},{"key":"26_CR12","doi-asserted-by":"crossref","unstructured":"Leroy, X.: Formal verification of a realistic compiler. Commun. ACM 52(7), 107\u2013115 (2009)","DOI":"10.1145\/1538788.1538814"},{"key":"26_CR13","doi-asserted-by":"publisher","unstructured":"Li, G., Li, P., Sawaya, G., Gopalakrishnan, G., Ghosh, I., Rajan, S.P.: GKLEE: concolic verification and test generation for GPUs. In: Ramanujam, J., Sadayappan, P. (eds.) Proceedings of the 17th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, PPOPP 2012, New Orleans, LA, USA, February 25-29, 2012, pp. 215\u2013224. ACM (2012). https:\/\/doi.org\/10.1145\/2145816.2145844","DOI":"10.1145\/2145816.2145844"},{"key":"26_CR14","doi-asserted-by":"crossref","unstructured":"Liew, D., Cogumbreiro, T., Lange, J.: Sound and partially-complete static analysis of data-races in GPU programs. Proc. ACM Program. Lang. 8(OOPSLA2), 2434\u20132461 (2024)","DOI":"10.1145\/3689797"},{"key":"26_CR15","doi-asserted-by":"publisher","unstructured":"Lustig, D., Cooksey, S., Giroux, O.: Mixed-proxy extensions for the NVIDIA PTX memory consistency model: industrial product. In: Salapura, V., Zahran, M., Chong, F., Tang, L. (eds.) ISCA \u201922: The 49th Annual International Symposium on Computer Architecture, New York, New York, USA, June 18 - 22, 2022, pp. 1058\u20131070. ACM (2022). https:\/\/doi.org\/10.1145\/3470496.3533045","DOI":"10.1145\/3470496.3533045"},{"key":"26_CR16","doi-asserted-by":"publisher","unstructured":"Lustig, D., Sahasrabuddhe, S., Giroux, O.: A formal analysis of the NVIDIA PTX memory consistency model. In: Bahar, I., Herlihy, M., Witchel, E., Lebeck, A.R. (eds.) Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2019, Providence, RI, USA, April 13-17, 2019, pp. 257\u2013270. ACM (2019). https:\/\/doi.org\/10.1145\/3297858.3304043","DOI":"10.1145\/3297858.3304043"},{"key":"26_CR17","doi-asserted-by":"publisher","unstructured":"de\u00a0Moura, L.M., Bj\u00f8rner, N.S.: Z3: an efficient SMT solver. In: Ramakrishnan, C.R., Rehof, J. (eds.) Tools and Algorithms for the Construction and Analysis of Systems, 14th International Conference, TACAS 2008, Held as Part of the Joint European Conferences on Theory and Practice of Software, ETAPS 2008, Budapest, Hungary, March 29-April 6, 2008. Proceedings. Lecture Notes in Computer Science, vol.\u00a04963, pp. 337\u2013340. Springer (2008). https:\/\/doi.org\/10.1007\/978-3-540-78800-3_24","DOI":"10.1007\/978-3-540-78800-3_24"},{"key":"26_CR18","unstructured":"NVIDIA Corporation: Compute Sanitizer. https:\/\/docs.nvidia.com\/compute-sanitizer\/ComputeSanitizer\/index.html. Accessed Oct 2025"},{"key":"26_CR19","unstructured":"NVIDIA Corporation: Blackwell Cluster Launch Control (2022). https:\/\/docs.nvidia.com\/cutlass\/media\/docs\/cpp\/blackwell_cluster_launch_control.html. Accessed 23 Oct 2025"},{"key":"26_CR20","unstructured":"NVIDIA Corporation: NVIDIA CUTLASS Documentation (2025). https:\/\/docs.nvidia.com\/cutlass\/latest\/. Accessed 23 Oct 2025"},{"key":"26_CR21","unstructured":"OpenXLA: OpenXLA GPU Architecture Overview (2026). https:\/\/openxla.org\/xla\/gpu_architecture. Accessed 23 April 2026"},{"key":"26_CR22","doi-asserted-by":"publisher","unstructured":"Peng, Y., Grover, V., Devietti, J.: CURD: a dynamic CUDA race detector. In: Foster, J.S., Grossman, D. (eds.) Proceedings of the 39th ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI 2018, Philadelphia, PA, USA, June 18-22, 2018, pp. 390\u2013403. ACM (2018). https:\/\/doi.org\/10.1145\/3192366.3192368","DOI":"10.1145\/3192366.3192368"},{"key":"26_CR23","unstructured":"Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., Dao, T.: FlashAttention-3: fast and accurate attention with asynchrony and low-precision. In: Globersons, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J.M., Zhang, C. (eds.) Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024 (2024). http:\/\/papers.nips.cc\/paper_files\/paper\/2024\/hash\/7ede97c3e082c6df10a8d6103a2eebd2-Abstract-Conference.html"},{"key":"26_CR24","doi-asserted-by":"publisher","unstructured":"Sharma, R., Bauer, M., Aiken, A.: Verification of producer-consumer synchronization in GPU programs. In: Grove, D., Blackburn, S.M. (eds.) Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Implementation, Portland, OR, USA, June 15-17, 2015, pp. 88\u201398. ACM (2015). https:\/\/doi.org\/10.1145\/2737924.2737962","DOI":"10.1145\/2737924.2737962"},{"key":"26_CR25","doi-asserted-by":"publisher","unstructured":"Tillet, P., Kung, H.T., Cox, D.: Triton: an intermediate language and compiler for tiled neural network computations. In: Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pp. 10\u201319. MAPL 2019, Association for Computing Machinery, New York, NY, USA (2019). https:\/\/doi.org\/10.1145\/3315508.3329973","DOI":"10.1145\/3315508.3329973"},{"key":"26_CR26","unstructured":"Vaswani, A., et al.: Attention is all you need. CoRR abs\/1706.03762 (2017). http:\/\/arxiv.org\/abs\/1706.03762"}],"container-title":["Lecture Notes in Computer Science","Computer Aided Verification"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-32537-2_26","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T08:42:09Z","timestamp":1784796129000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-32537-2_26"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"ISBN":["9783032325365","9783032325372"],"references-count":26,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-32537-2_26","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026]]},"assertion":[{"value":"24 July 2026","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"All authors are employed by NVIDIA or one of its affiliated entities. The work reported in this paper was conducted as part of that employment. The authors have no other competing interests to declare.","order":1,"name":"Ethics","label":"Disclosure of Interests","group":{"name":"EthicsHeading","label":"Ethics"}},{"value":"CAV","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Computer Aided Verification","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Lisbon","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Portugal","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2026","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"26 July 2026","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"29 July 2026","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"38","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"cav2026","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/www.floc26.org\/program","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}