{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T11:58:35Z","timestamp":1774353515501,"version":"3.50.1"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"Dutch Research Council (NWO) in the framework of the NWA ORC Call","award":["NWA.1160.18.316"],"award-info":[{"award-number":["NWA.1160.18.316"]}]},{"name":"ESiWACE3 is funded by EuroHPC JU and national co-funding bodies","award":["101093054"],"award-info":[{"award-number":["101093054"]}]},{"DOI":"10.13039\/100013407","name":"Netherlands eScience Center","doi-asserted-by":"crossref","award":["NLESC.OEC.2022.001"],"award-info":[{"award-number":["NLESC.OEC.2022.001"]}],"id":[{"id":"10.13039\/100013407","id-type":"DOI","asserted-by":"crossref"}]},{"name":"EuroHPC supercomputer LUMI was through SURF","award":["EINF-11263"],"award-info":[{"award-number":["EINF-11263"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>\n                    Modern GPUs feature specialized hardware for low-precision floating-point arithmetic to accelerate compute-intensive workloads that do not require high numerical accuracy, such as those from artificial intelligence. However, despite the significant gains in computational throughput, memory bandwidth utilization, and energy efficiency, integrating low-precision formats into scientific applications remains difficult. We introduce\n                    <jats:italic toggle=\"yes\">Kernel Float<\/jats:italic>\n                    , a header-only C++ library that simplifies the development of portable mixed-precision GPU kernels. Kernel Float provides a generic vector type, a unified interface for common mathematical operations, and fast approximations for low-precision transcendental functions that lack native hardware support. To demonstrate the potential of mixed-precision computing unlocked by our library, we integrated Kernel Float into nine GPU kernels from various domains. Our evaluation on Nvidia A100 and AMD MI250X GPUs shows performance improvements of up to\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(12\\times\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    over double precision, while reducing source code length by up to 50% compared to handwritten kernels and having negligible runtime overhead. Our results further show that mixed-precision performance depends not only on choosing appropriate data types but also on tuning traditional optimization parameters (e.g., block size and vector width) and, when relevant, even domain-specific parameters.\n                  <\/jats:p>","DOI":"10.1145\/3779120","type":"journal-article","created":{"date-parts":[[2025,12,4]],"date-time":"2025-12-04T14:55:18Z","timestamp":1764860118000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Kernel Float: Unlocking Mixed-Precision GPU Programming"],"prefix":"10.1145","volume":"52","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8792-6305","authenticated-orcid":false,"given":"Stijn","family":"Heldens","sequence":"first","affiliation":[{"name":"Netherlands eScience Center, Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7508-3272","authenticated-orcid":false,"given":"Ben","family":"van Werkhoven","sequence":"additional","affiliation":[{"name":"Leiden University, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,17]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"1","volume-title":"Proceedings of the 2022 IEEE High Performance Extreme Computing Conference (HPEC)","author":"Abdelkhalik Hamdy","year":"2022","unstructured":"Hamdy Abdelkhalik, Yehia Arafa, Nandakishore Santhi, and Abdel-Hameed A. Badawy. 2022. Demystifying the Nvidia Ampere architecture through microbenchmarking and instruction-level analysis. In Proceedings of the 2022 IEEE High Performance Extreme Computing Conference (HPEC), 1\u20138. DOI: 10.1109\/HPEC55821.2022.9926299"},{"key":"e_1_3_2_3_2","volume-title":"Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables","author":"Abramowitz Milton","year":"1948","unstructured":"Milton Abramowitz and Irene A. Stegun. 1948. Handbook of Mathematical Functions with Formulas, Graphs, and Mathematical Tables, Vol. 55. US Government printing office."},{"key":"e_1_3_2_4_2","unstructured":"AMD. 2020. Introducing CDNA Architecture. Retrieved June 2021 from https:\/\/www.amd.com\/system\/files\/documents\/amd-cdna-whitepaper.pdf"},{"key":"e_1_3_2_5_2","unstructured":"Patrick Amestoy Antoine Jego Jean-Yves L\u2019Excellent Th\u00e9o Mary and Gr\u00e9goire Pichon. 2025. BLAS-based Block Memory Accessor with Applications to Mixed Precision Sparse Direct Solvers. Retrieved from https:\/\/hal.science\/hal-05019106"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3151032"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","unstructured":"Hartwig Anzt Terry Cojean Goran Flegar Fritz G\u00f6bel Thomas Gr\u00fctzmacher Pratik Nayak Tobias Ribizel Yuhsiang Mike Tsai and Enrique S. Quintana-Ort\u00ed. 2022. Ginkgo: A modern linear operator algebra framework for high performance computing. ACM Transactions on Mathematical Software 48 1 Article 2 (Feb. 2022) 1\u201333. DOI: 10.1145\/3480935","DOI":"10.1145\/3480935"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2016.127"},{"key":"e_1_3_2_9_2","volume-title":"Introduction to Bessel Functions","author":"Bowman F.","year":"1958","unstructured":"F. Bowman. 1958. Introduction to Bessel Functions. Dover Publications."},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1109\/ARITH.2019.00022","volume-title":"Proceedings of the 2019 IEEE 26th Symposium on Computer Arithmetic (ARITH)","author":"Burgess Neil","year":"2019","unstructured":"Neil Burgess, Jelena Milanovic, Nigel Stephens, Konstantinos Monachopoulos, and David Mansell. 2019. Bfloat16 processing for neural networks. In Proceedings of the 2019 IEEE 26th Symposium on Computer Arithmetic (ARITH). IEEE, 88\u201391."},{"key":"e_1_3_2_11_2","first-page":"44","volume-title":"Proceedings of the 2009 IEEE International Symposium on Workload Characterization (IISWC)","author":"Che Shuai","year":"2009","unstructured":"Shuai Che, Michael Boyer, Jiayuan Meng, David Tarjan, Jeremy W. Sheaffer, Sang-Ha Lee, and Kevin Skadron. 2009. Rodinia: A benchmark suite for heterogeneous computing. In Proceedings of the 2009 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 44\u201354."},{"key":"e_1_3_2_12_2","unstructured":"LUMI Consortium. 2024. LUMI Supercomputer. Retrieved from https:\/\/www.lumi-supercomputer.eu\/"},{"key":"e_1_3_2_13_2","unstructured":"Matthieu Courbariaux Yoshua Bengio and Jean-Pierre David. 2014. Training deep neural networks with low precision multiplications. arXiv:1412.7024. Retrieved from https:\/\/arxiv.org\/abs\/1412.7024"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.5194\/gmd-10-2221-2017"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Massimiliano Fasi and Mantas Mikaitis. 2023. CPFloat: A C library for simulating low-precision arithmetic. ACM Transactions on Mathematical Software 49 2 Article 18 (June 2023) 32 pages. DOI: 10.1145\/3585515","DOI":"10.1145\/3585515"},{"issue":"4","key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3368086","article-title":"FloatX: A C++ library for customized floating-point arithmetic","volume":"45","author":"Flegar Goran","year":"2019","unstructured":"Goran Flegar, Florian Scheidegger, Vedran Novakovi\u0107, Giovani Mariani, Andr\u00e9s E. Tom\u00e1s, A. Cristiano I. Malossi, and Enrique S. Quintana-Ort\u00ed. 2019. FloatX: A C++ library for customized floating-point arithmetic. ACM Transactions on Mathematical Software 45, 4 (2019), 1\u201323.","journal-title":"ACM Transactions on Mathematical Software"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Michael J. Flynn. 1970. On division by functional iteration. IEEE Transactions on Computers 100 8 (1970) 702\u2013706.","DOI":"10.1109\/T-C.1970.223019"},{"issue":"2","key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1145\/1236463.1236468","article-title":"MPFR: A multiple-precision binary floating-point library with correct rounding","volume":"33","author":"Fousse Laurent","year":"2007","unstructured":"Laurent Fousse, Guillaume Hanrot, Vincent Lef\u00e8vre, Patrick P\u00e9lissier, and Paul Zimmermann. 2007. MPFR: A multiple-precision binary floating-point library with correct rounding. ACM Transactions on Mathematical Software 33, 2 (2007), 13\u2013es.","journal-title":"ACM Transactions on Mathematical Software"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1002\/spe.3041"},{"key":"e_1_3_2_20_2","first-page":"1737","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Gupta Suyog","year":"2015","unstructured":"Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep learning with limited numerical precision. In Proceedings of the International Conference on Machine Learning. PMLR, 1737\u20131746."},{"key":"e_1_3_2_21_2","first-page":"551","volume-title":"Proceedings of the 2022 IEEE 18th International Conference on e-Science (e-Science)","author":"Harvey Evan","year":"2022","unstructured":"Evan Harvey, Reed Milewicz, Christian Trott, Luc Berger-Vergiat, and Siva Rajamanickam. 2022. Half-precision scalar support in Kokkos and Kokkos kernels: An engineering study and experience report. In Proceedings of the 2022 IEEE 18th International Conference on e-Science (e-Science). IEEE, 551\u2013560."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372390"},{"issue":"5","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"C585","DOI":"10.1137\/19M1251308","article-title":"Simulating low precision floating-point arithmetic","volume":"41","author":"Higham Nicholas J.","year":"2019","unstructured":"Nicholas J. Higham and Srikara Pranesh. 2019. Simulating low precision floating-point arithmetic. SIAM Journal on Scientific Computing 41, 5 (2019), C585\u2013C602.","journal-title":"SIAM Journal on Scientific Computing"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Nhut-Minh Ho Himeshi De Silva and Weng-Fai Wong. 2021. GRAM: A framework for dynamically mixing precisions in GPU applications. ACM Transactions on Architecture and Code Optimization 18 2 Article 19 (Feb. 2021) 24 pages. DOI: 10.1145\/3441830","DOI":"10.1145\/3441830"},{"key":"e_1_3_2_25_2","first-page":"1","volume-title":"Proceedings of the 2017 IEEE High Performance Extreme Computing Conference (HPEC)","author":"Ho Nhut-Minh","year":"2017","unstructured":"Nhut-Minh Ho and Weng-Fai Wong. 2017. Exploiting half precision arithmetic in Nvidia GPUs. In Proceedings of the 2017 IEEE High Performance Extreme Computing Conference (HPEC), 1\u20137. DOI: 10.1109\/HPEC.2017.8091072"},{"key":"e_1_3_2_26_2","article-title":"Anatomy of AMD\u2019s TeraScale microarchitecture","year":"2008","unstructured":"Mike Houston. 2008. Anatomy of AMD\u2019s TeraScale microarchitecture. In Proceedings of the SIGGRAPH 2008.","journal-title":"Proceedings of the SIGGRAPH 2008"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"IEEE. 2019. IEEE Standard for Floating-Point Arithmetic (IEEE Std 754-2019). In IEEE Std 754-2019 (Revision of IEEE 754-2008) 1\u201384. DOI: 10.1109\/IEEESTD.2019.8766229","DOI":"10.1109\/IEEESTD.2019.8766229"},{"key":"e_1_3_2_28_2","unstructured":"Aditya Kashi Hao Lu Wesley Brewer David Rogers Michael Matheson Mallikarjun Shankar and Feiyi Wang. 2024. Mixed-precision numerics in scientific applications: Survey and perspectives. arXiv:2412.19322. Retrieved from https:\/\/arxiv.org\/abs\/2412.19322"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1145\/3330345.3330360","volume-title":"Proceedings of the ACM International Conference on Supercomputing","author":"Kotipalli Pradeep V.","year":"2019","unstructured":"Pradeep V. Kotipalli, Ranvijay Singh, Paul Wood, Ignacio Laguna, and Saurabh Bagchi. 2019. AMPT-GA: Automatic mixed precision floating point tuning for GPU applications. In Proceedings of the ACM International Conference on Supercomputing, 160\u2013170."},{"key":"e_1_3_2_30_2","first-page":"227","volume-title":"Proceedings of the 34th International Conference on High Performance Computing (ISC High Performance \u201919)","author":"Laguna Ignacio","year":"2019","unstructured":"Ignacio Laguna, Paul C. Wood, Ranvijay Singh, and Saurabh Bagchi. 2019. GPUMixer: Performance-driven floating-point tuning for GPU scientific applications. In Proceedings of the 34th International Conference on High Performance Computing (ISC High Performance \u201919). Springer, 227\u2013246."},{"key":"e_1_3_2_31_2","first-page":"27","volume-title":"Proceedings of the 2019 IEEE\/ACM 3rd International Workshop on Software Correctness for HPC Applications (Correctness)","author":"Lam Michael O.","year":"2019","unstructured":"Michael O. Lam, Tristan Vanderbruggen, Harshitha Menon, and Markus Schordan. 2019. Tool integration for source-level mixed precision. In Proceedings of the 2019 IEEE\/ACM 3rd International Workshop on Software Correctness for HPC Applications (Correctness), 27\u201335. DOI: 10.1109\/Correctness49594.2019.00009"},{"issue":"7553","key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun Yann","year":"2015","unstructured":"Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. Nature 521, 7553 (2015), 436\u2013444.","journal-title":"Nature"},{"issue":"2","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1109\/MM.2008.31","article-title":"NVIDIA tesla: A unified graphics and computing architecture","volume":"28","author":"Lindholm Erik","year":"2008","unstructured":"Erik Lindholm, John Nickolls, Stuart Oberman, and John Montrym. 2008. NVIDIA tesla: A unified graphics and computing architecture. IEEE Micro 28, 2 (2008), 39\u201355.","journal-title":"IEEE Micro"},{"key":"e_1_3_2_34_2","unstructured":"Chris Lomont. 2003. Fast Inverse Square Root. Technical Report 32."},{"key":"e_1_3_2_35_2","unstructured":"Charles McEniry. 2007. The Mathematics Behind the Fast Inverse Square Root Function Code. Technical Report."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/2893356"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1145\/3183767.3183776","volume-title":"Proceedings of the 9th Workshop and 7th Workshop on Parallel Programming and RunTime Management Techniques for Manycore Architectures and Design Tools and Architectures for Multicore Embedded Computing Platforms","author":"Nobre Ricardo","year":"2018","unstructured":"Ricardo Nobre, Lu\u00eds Reis, Jo\u00e3o Bispo, Tiago Carvalho, Jo\u00e3o M. P. Cardoso, Stefano Cherubin, and Giovanni Agosta. 2018. Aspect-driven mixed-precision tuning targeting GPUs. In Proceedings of the 9th Workshop and 7th Workshop on Parallel Programming and RunTime Management Techniques for Manycore Architectures and Design Tools and Architectures for Multicore Embedded Computing Platforms, 26\u201331."},{"key":"e_1_3_2_38_2","unstructured":"Nvidia. 2010. NVIDIA\u2019s Next Generation CUDA Compute Architecture: Fermi. Retrieved June 2021 from https:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/NVIDIA_Fermi_Compute_Architecture_Whitepaper.pdf"},{"key":"e_1_3_2_39_2","unstructured":"Nvidia. 2016. Nvidia Tesla P100 (Whitepaper). Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf"},{"key":"e_1_3_2_40_2","unstructured":"Nvidia. 2017. Nvidia Tesla V100 Whitepaper. Retrieved June 2021 from https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_3_2_41_2","unstructured":"Nvidia. 2021. Nvidia Ampere GA-102 GPU Architecture (Whitepaper). Retrieved from https:\/\/www.nvidia.com\/content\/PDF\/nvidia-ampere-ga-102-GPU-architecture-whitepaper-v2.pdf"},{"key":"e_1_3_2_42_2","unstructured":"Nvidia. 2022. NVIDIA H100 Tensor Core GPU Architecture (Whitepaper). Retrieved from https:\/\/resources.nvidia.com\/en-us-tensor-core\/gtc22-whitepaper-hopper"},{"key":"e_1_3_2_43_2","unstructured":"Nvidia. 2024. CUDA Samples. Retrieved from https:\/\/github.com\/NVIDIA\/cuda-samples"},{"key":"e_1_3_2_44_2","unstructured":"Nvidia. 2024. NVIDIA Blackwell Architecture Technical Brief (Whitepaper). Retrieved from https:\/\/resources.nvidia.com\/en-us-blackwell-architecture\/blackwell-architecture-technical-brief"},{"key":"e_1_3_2_45_2","unstructured":"Sivasankaran Rajamanickam Seher Acer Luc Berger-Vergiat Vinh Dang Nathan Ellingwood Evan Harvey Brian Kelley Christian R. Trott Jeremiah Wilke and Ichitaro Yamazaki. 2021. Kokkos kernels: Performance portable sparse\/dense linear algebra and graph kernels. arXiv:2103.11991. Retrieved from https:\/\/arxiv.org\/abs\/2103.11991"},{"issue":"196","key":"e_1_3_2_46_2","first-page":"41","article-title":"Sur la d\u00e9termination des polyn\u00f4mes d\u2019approximation de degr\u00e9 donn\u00e9e","volume":"10","author":"Remez Eugene Y.","year":"1934","unstructured":"Eugene Y. Remez. 1934. Sur la d\u00e9termination des polyn\u00f4mes d\u2019approximation de degr\u00e9 donn\u00e9e. Communications of the Kharkov Mathematical Society 10, 196 (1934), 41\u201363.","journal-title":"Communications of the Kharkov Mathematical Society"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3596218"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1162\/089976699300016467"},{"key":"e_1_3_2_49_2","unstructured":"Benjamin F. Spector Simran Arora Aaryan Singhal Daniel Y. Fu and Christopher R\u00e9. 2024. ThunderKittens: Simple fast and adorable AI kernels. arXiv:2410.20399. Retrieved from https:\/\/arxiv.org\/abs\/2410.20399"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","unstructured":"Stijn Heldens and Ben van Werkhoven. 2025. Kernel Float: Header-Only Library for Writing Mixed-Precision GPU Kernels. Version 0.3. Zenodo. DOI: 10.5281\/zenodo.15195615","DOI":"10.5281\/zenodo.15195615"},{"key":"e_1_3_2_51_2","article-title":"Parboil: A revised benchmark suite for scientific and commercial throughput computing","author":"Stratton John A.","year":"2012","unstructured":"John A. Stratton, Christopher Rodrigues, I-Jui Sung, Nady Obeid, Li-Wen Chang, Nasser Anssari, Geng Daniel Liu, and Wen-Mei W. Hwu. 2012. Parboil: A revised benchmark suite for scientific and commercial throughput computing. Center for Reliable and High-Performance Computing. Technical Report No. IMPACT-12-01, University of Illinois at Urbana-Champaign. Retrieved from https:\/\/impact.crhc.illinois.edu\/Shared\/Docs\/impact-12-01.parboil.pdf","journal-title":"Center for Reliable and High-Performance Computing. Technical Report No. IMPACT-12-01, University of Illinois at Urbana-Champaign"},{"issue":"1","key":"e_1_3_2_52_2","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1109\/TCAD.2018.2883902","article-title":"FlexFloat: A software library for transprecision computing","volume":"39","author":"Tagliavini Giuseppe","year":"2018","unstructured":"Giuseppe Tagliavini, Andrea Marongiu, and Luca Benini. 2018. FlexFloat: A software library for transprecision computing. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 39, 1 (2018), 145\u2013156.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems"},{"key":"e_1_3_2_53_2","first-page":"724","volume-title":"Proceedings of the 2023 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)","author":"T\u00f8rring Jacob O.","year":"2023","unstructured":"Jacob O. T\u00f8rring, Ben van Werkhoven, Filip Petrov\u010d, Floris-Jan Willemsen, Ji\u0159\u00ed Filipovi\u010d, and Anne C. Elster. 2023. Towards a benchmarking suite for kernel tuners. In Proceedings of the 2023 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE, 724\u2013733."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2018.08.004"},{"key":"e_1_3_2_55_2","first-page":"399","volume-title":"Proceedings of the International Conference on Computational Science","author":"van Werkhoven Ben","year":"2020","unstructured":"Ben van Werkhoven, Willem Jan Palenstijn, and Alessio Sclocco. 2020. Lessons learned in a decade of research software engineering GPU applications. In Proceedings of the International Conference on Computational Science. Springer, 399\u2013412. DOI: 10.1007\/978-3-030-50436-6_29"},{"key":"e_1_3_2_56_2","article-title":"Training deep neural networks with 8-bit floating point numbers","volume":"31","author":"Wang Naigang","year":"2018","unstructured":"Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan. 2018. Training deep neural networks with 8-bit floating point numbers. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 31.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"issue":"2","key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1109\/MDAT.2016.2630270","article-title":"AxBench: A multiplatform benchmark suite for approximate computing","volume":"34","author":"Yazdanbakhsh Amir","year":"2016","unstructured":"Amir Yazdanbakhsh, Divya Mahajan, Hadi Esmaeilzadeh, and Pejman Lotfi-Kamran. 2016. AxBench: A multiplatform benchmark suite for approximate computing. IEEE Design & Test 34, 2 (2016), 60\u201368.","journal-title":"IEEE Design & Test"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3779120","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T16:03:54Z","timestamp":1773763434000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3779120"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,17]]},"references-count":56,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3779120"],"URL":"https:\/\/doi.org\/10.1145\/3779120","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,17]]},"assertion":[{"value":"2025-04-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}