{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T11:05:29Z","timestamp":1784199929476,"version":"3.55.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"PLDI","license":[{"start":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T00:00:00Z","timestamp":1749772800000},"content-version":"vor","delay-in-days":3,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2110861;2312220"],"award-info":[{"award-number":["2110861;2312220"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Program. Lang."],"published-print":{"date-parts":[[2025,6,10]]},"abstract":"<jats:p>\n                    Our\n                    <jats:sc>RLibm<\/jats:sc>\n                    project has recently proposed methods to generate a single implementation for an elementary function that produces correctly rounded results for multiple rounding modes and representations with up to 32-bits. They are appealing for developing fast reference libraries without double rounding issues. The key insight is to build polynomial approximations that produce the correctly rounded result for a representation with two additional bits when compared to the largest target representation and with the \u201cnon-standard\u201d round-to-odd rounding mode, which makes double rounding the\n                    <jats:sc>RLibm<\/jats:sc>\n                    math library result to any smaller target representation innocuous. The resulting approximations generated by the\n                    <jats:sc>RLibm<\/jats:sc>\n                    approach are implemented with machine supported foating-point operations with the\n                    <jats:italic toggle=\"yes\">round-to-nearest<\/jats:italic>\n                    rounding mode. When an application uses a rounding mode other than the round-to-nearest mode, the\n                    <jats:sc>RLibm<\/jats:sc>\n                    math library saves the application\u2019s rounding mode, changes the system\u2019s rounding mode to round-to-nearest, computes the correctly rounded result, and restores the application\u2019s rounding mode. This frequent change of rounding modes has a performance cost.\n                  <\/jats:p>\n                  <jats:p>\n                    This paper proposes two new methods, which we call rounding-invariant outputs and rounding-invariant input bounds, to avoid the frequent changes to the rounding mode and the dependence on the round-to-nearest mode. First, our new rounding-invariant outputs method proposes using the round-to-zero rounding mode to implement\n                    <jats:sc>RLibm<\/jats:sc>\n                    \u2019s polynomial approximations. We propose fast, error-free transformations to emulate a round-to-zero result from any standard rounding mode without changing the rounding mode. Second, our rounding-invariant input bounds method factors any rounding error due to different rounding modes using interval bounds in the\n                    <jats:sc>RLibm<\/jats:sc>\n                    pipeline. Both methods make a different set of trade-offs and improve the performance of resulting libraries by more than 2\u00d7.\n                  <\/jats:p>","DOI":"10.1145\/3729332","type":"journal-article","created":{"date-parts":[[2025,6,13]],"date-time":"2025-06-13T16:02:27Z","timestamp":1749830547000},"page":"2032-2055","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Correctly Rounded Math Libraries without Worrying about the Application\u2019s Rounding Mode"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-1528-562X","authenticated-orcid":false,"given":"Sehyeok","family":"Park","sequence":"first","affiliation":[{"name":"Rutgers University, Piscataway, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1481-5019","authenticated-orcid":false,"given":"Justin","family":"Kim","sequence":"additional","affiliation":[{"name":"Rutgers University, Piscataway, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5048-8548","authenticated-orcid":false,"given":"Santosh","family":"Nagarakatte","sequence":"additional","affiliation":[{"name":"Rutgers University, Piscatway, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Mridul Aanjaneya Jay P. Lim and Santosh Nagarakatte. 2021. RLIBM-Prog: Progressive Polynomial Approximations for Correctly Rounded Math Libraries. arXiv:2111.12852 Rutgers Department of Computer Science Technical Report DCS-TR-758.","DOI":"10.1145\/3519939.3523447"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3519939.3523447"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579990.3580022"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656427"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","unstructured":"Sylvie Boldo Marc Daumas and Ren-Cang Li. 2009. Formally Verified Argument Reduction with a Fused Multiply-Add. In IEEE Transactions on Computers Vol. 58. 1139\u20131145. doi:10.1109\/TC.2008.216","DOI":"10.1109\/TC.2008.216"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3054947"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3632874"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","unstructured":"Nicolas Brisebarre and Sylvain Chevillard. 2007. Efficient polynomial L\u221e-approximations. In 18th IEEE Symposium on Computer Arithmetic (ARITH '07). doi:10.1109\/ARITH.2007.17","DOI":"10.1109\/ARITH.2007.17"},{"key":"e_1_3_2_10_2","unstructured":"Nicolas Brisebarre Guillaume Hanrot Jean-Michel Muller and Paul Zimmermann. 2024. Correctly-rounded evaluation of a function: why how and at what cost? (May 2024). https:\/\/hal.science\/hal-04474530 working paper or preprint."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","unstructured":"Sylvain Chevillard John Harrison Miaora Joldes and Christoph Lauter. 2011. Efficient and accurate computation of upper bounds of approximation errors. In Theoretical Computer Science Vol. 412. doi:10.1016\/j.tcs.2010.11.052","DOI":"10.1016\/j.tcs.2010.11.052"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15582-6_5"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Sylvain Chevillard and Christopher Lauter. 2007. A Certified Infinite Norm for the Implementation of Elementary Functions. In Seventh International Conference on Quality Software (QSIC 2007). 153\u2013160. doi:10.1109\/QSIC.2007.4385491","DOI":"10.1109\/QSIC.2007.4385491"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3563353"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"William J Cody and William M Waite. 1980. Software manual for the elementary functions. Prentice-Hall Englewood Cliffs NJ. doi:10.1137\/1024023","DOI":"10.1137\/1024023"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","unstructured":"Catherine Daramy David Defour Florent Dinechin and Jean-Michel Muller. 2003. CR-LIBM: A correctly rounded elementary function library. In Proceedings of SPIE Vol. 5205: Advanced Signal Processing Algorithms Architectures and Implementations XIII Vol. 5205. doi:10.1117\/12.505591","DOI":"10.1117\/12.505591"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Marc Daumas Guillaume Melquiond and Cesar Munoz. 2005. Guaranteed proofs using interval arithmetic. In 17th IEEE Symposium on Computer Arithmetic (ARITH\u201905). 188\u2013195. doi:10.1109\/ARITH.2005.25","DOI":"10.1109\/ARITH.2005.25"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","unstructured":"Florent de Dinechin Christopher Lauter and Guillaume Melquiond. 2011. Certifying the Floating-Point Implementation of an Elementary Function Using Gappa. In IEEE Transactions on Computers Vol. 60. 242\u2013253. doi:10.1109\/TC.2010.128","DOI":"10.1109\/TC.2010.128"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1141277.1141584"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01397083"},{"key":"e_1_3_2_21_2","unstructured":"Nestor Demeure. 2020. Compromise between precision and performance in high-performance computing. Ph.D. Dissertation. University Paris-Saclay. https:\/\/tel.archives-ouvertes.fr\/tel-03116750"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/1236463.1236468"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0000475"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"John Harrison. 1997. Verifying the Accuracy of Polynomial Approximations in HOL. In International Conference on Theorem Proving in Higher Order Logics. doi:10.1007\/BFb0028391","DOI":"10.1007\/BFb0028391"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642\u201303359-9_4"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/363707.363723"},{"key":"e_1_3_2_27_2","unstructured":"William Kahan. 2004. A Logarithm To Clever by Half. https:\/\/people.eecs.berkeley.edu\/~wkahan\/LOG10HAF.TXT."},{"key":"e_1_3_2_28_2","unstructured":"Donald E. Knuth. 1998. The Art of Computer Programming Volume 2: Seminumerical Algorithms. Addison-Wesley."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.4230\/DagSemProc.05391.3"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3158135"},{"key":"e_1_3_2_31_2","unstructured":"Jay P. Lim Mridul Aanjaneya John Gustafson and Santosh Nagarakatte. 2020. A Novel Approach to Generate Correctly Rounded Math Libraries for New Floating Point Representations. arXiv:2007.05344 Rutgers Department of Computer Science Technical Report DCS-TR-753."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3434310"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","unstructured":"Jay P. Lim and Santosh Nagarakatte. 2021. High Performance Correctly Rounded Math Libraries for 32-bit Floating Point Representations. In 42nd ACM SIGPLAN Conference on Programming Language Design and Implementation (HLD\u201921). doi:10.1145\/3453483.3454049","DOI":"10.1145\/3453483.3454049"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Jay P Lim and Santosh Nagarakatte. 2021. RLIBM-32: High Performance Correctly Rounded Math Libraries for 32-bit Floating Point Representations. arXiv:2104.04043 Rutgers Department of Computer Science Technical Report DCS-TR-754.","DOI":"10.1145\/3453483.3454049"},{"key":"e_1_3_2_35_2","unstructured":"Jay P. Lim and Santosh Nagarakatte. 2021. RLIBM-ALL: A Novel Polynomial Approximation Method to Produce Correctly Rounded Results for Multiple Representations and Rounding Modes. arXiv:2108.06756 Rutgers Department of Computer Science Technical Report DCS-TR-757."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3498664"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/1353445.1353446"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","unstructured":"Jean-Michel Muller. 2016. Elementary Functions: Algorithms and Implementation 3rd edition. Springer. doi:10.1007\/978-1-4899-7983-4","DOI":"10.1007\/978-1-4899-7983-4"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","unstructured":"Jean-Michel Muller Nicolas Brunie Florent de Dinechin Claude-Pierre Jeannerod Miaora Joldes Vincent Lefvre Guillaume Melquiond Nathalie Revel and Serge Torres. 2018. Handbook of Floating-Point Arithmetic (2nd ed.). Birkh\u00e4user Basel. doi:10.1007\/978-3-319-76526-6","DOI":"10.1007\/978-3-319-76526-6"},{"key":"e_1_3_2_40_2","unstructured":"Santosh Nagarakatte Sehyeok Park Mridul Aanjaneya and Jay P. Lim. 2024. The RLIBM Project. https:\/\/www.cs.rutgers.edu\/~santosh.nagarakatte\/rlibm\/."},{"key":"e_1_3_2_41_2","unstructured":"NVIDIA. 2020. TensorFloat-32 in the A100 GPU Accelerates AI Training HPC up to 20x. https:\/\/blogs.nvidia.com\/blog\/2020\/05\/14\/tensorfloat-32-precision-format\/."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","unstructured":"Michael Overton. 2001. Numerical computing with IEEE floating point arithmetic. SIAM Society for Industrial and Applied Mathematics. doi:10.1137\/1.9780898718072","DOI":"10.1137\/1.9780898718072"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","unstructured":"Sehyeok Park Justin Kim and Santosh Nagarakatte. 2025. Artifact for Correctly Rounded Math Libraries Without Worrying about the Application\u2019s Rounding Mode. doi:10.5281\/zenodo.1506682","DOI":"10.5281\/zenodo.1506682"},{"key":"e_1_3_2_44_2","doi-asserted-by":"crossref","unstructured":"Sehyeok Park Justin Kim and Santosh Nagarakatte. 2025. RLIBM-MultiRound: Correctly Rounded Math Libraries Without Worrying about the Application\u2019s Rounding Mode. arXiv:2504.07409 Rutgers Department of Computer Science Technical Report DCS-TR-759.","DOI":"10.1145\/3729332"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"Sehyeok Park and Santosh Nagarakatte. 2025. Fast Trigonometric Functions using the RLIBM Approach. In Proceedings of the International Workshop on Verification of Scientific Software (VSS 2025).","DOI":"10.4204\/EPTCS.432.9"},{"key":"e_1_3_2_46_2","unstructured":"Douglas M. Priest. 1992. On Properties of Floating Point Arithmetic: Numerical Stability and the Cost of Accurate Computations. Ph.D. Dissertation. USA. UMI Order No. GAX93-30692."},{"key":"e_1_3_2_47_2","unstructured":"Eugene Remes. 1934. Sur un proc\u00e9d\u00e9 convergent d\u2019approximations successives pour d\u00e9terminer les polyn\u00f4mes d\u2019approximation. Comptes rendus de l\u2019Acad\u00e9mie des Sciences 198 (1934) 2063\u20132065."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1137\/080738490"},{"key":"e_1_3_2_49_2","unstructured":"Jun Sawada. 2002. Formal verification of divide and square root algorithms using series calculation. In 3rd International Workshop on the ACL2 Theorem Prover and its Applications."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","unstructured":"Jonathan Shevchuk. 1996. Adaptive Precision Floating-Point Arithmetic and Fast Robust Geometric Predicates. Discrete and Computational Geometry 18 (Jul 1996). doi:10.1007\/PL00009321","DOI":"10.1007\/PL00009321"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","unstructured":"Alexei Sibidanov Paul Zimmermann and St\u00e9phane Glondlo. 2022. The CORE-MATH Project. In ARITH 2022 - 29th IEEE Symposium on Computer Arithmetic virtual France. https:\/\/hal.inria.fr\/hal-03721225","DOI":"10.1109\/ARITH54963.2022.00014"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","unstructured":"Shane Story and Ping Tak Peter Tang. 1999. New algorithms for improved transcendental functions on IA-64. In Proceedings 14th IEEE Symposium on Computer Arithmetic. 4\u201311. doi:10.1109\/ARITH.1999.762822","DOI":"10.1109\/ARITH.1999.762822"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/63522.214389"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/98267.98294"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","unstructured":"P. T. P. Tang. 1991. Table-lookup algorithms for elementary functions and their error analysis. In [1991] Proceedings 10th IEEE Symposium on Computer Arithmetic. 232\u2013236. doi:10.1109\/ARITH.1991.145565","DOI":"10.1109\/ARITH.1991.145565"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","unstructured":"Lloyd N. Trefethen. 2012. Approximation Theory and Approximation Practice (Other Titles in Applied Mathematics). Society for Industrial and Applied Mathematics USA. doi:10.1137\/1.9781611975949","DOI":"10.1137\/1.9781611975949"},{"key":"e_1_3_2_57_2","unstructured":"Shibo Wang and Pankaj Kanwar. 2019. BFloat16: The secret to high performance on Cloud TPUs. https:\/\/cloud.google.com\/blog\/products\/ai-machine-learning\/bfloat16-the-secret-to-high-performance-on-cloud-tpus."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3290369"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/114697.116813"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3371128"}],"container-title":["Proceedings of the ACM on Programming Languages"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729332","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3729332","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T10:05:39Z","timestamp":1784196339000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3729332"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":59,"journal-issue":{"issue":"PLDI","published-print":{"date-parts":[[2025,6,10]]}},"alternative-id":["10.1145\/3729332"],"URL":"https:\/\/doi.org\/10.1145\/3729332","relation":{},"ISSN":["2475-1421"],"issn-type":[{"value":"2475-1421","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,10]]},"assertion":[{"value":"2024-11-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}