{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T15:02:51Z","timestamp":1779894171597,"version":"3.53.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T00:00:00Z","timestamp":1779840000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Ministry of Culture and Science of the State of North Rhine-Westphalia","award":["NW21-059D (SAIL)"],"award-info":[{"award-number":["NW21-059D (SAIL)"]}]},{"name":"German Federal Ministry for the Environment, Nature Conservation, Nuclear Safety and Consumer Protection","award":["67KI32004A (eki)"],"award-info":[{"award-number":["67KI32004A (eki)"]}]},{"name":"AGH University of Krakow","award":["16.16.120.773"],"award-info":[{"award-number":["16.16.120.773"]}]},{"name":"Federal Ministry of Education and Research and the state governments"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    While neural network quantization effectively reduces the cost of matrix multiplications, aggressive quantization can expose non-matrix-multiply operations as significant performance and resource bottlenecks on embedded systems. Addressing such bottlenecks requires a comprehensive approach to tailoring the precision across operations in the inference computation. To this end, we introduce scaled-integer range analysis (\n                    <jats:sc>SIRA<\/jats:sc>\n                    ), a static analysis technique employing interval arithmetic to determine the range, scale, and bias for tensors in quantized neural networks. We show how this information can be exploited to reduce the resource footprint of FPGA dataflow neural network accelerators via tailored bitwidth adaptation for accumulators and downstream operations, aggregation of scales and biases, and conversion of consecutive elementwise operations to thresholding operations. We integrate\n                    <jats:sc>SIRA<\/jats:sc>\n                    -driven optimizations into the open source FINN framework, then evaluate their effectiveness across a range of quantized neural network workloads and compare implementation alternatives for non-matrix-multiply operations. We demonstrate an average reduction of 17% for LUTs, 66% for DSPs, and 22% for accumulator bitwidths with\n                    <jats:sc>SIRA<\/jats:sc>\n                    optimizations, providing detailed benchmark analysis and analytical models to guide the implementation style for non-matrix layers. Finally, we open source\n                    <jats:sc>SIRA<\/jats:sc>\n                    to facilitate community exploration of its benefits across various applications and hardware platforms.\n                  <\/jats:p>","DOI":"10.1145\/3807510","type":"journal-article","created":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T14:48:11Z","timestamp":1775832491000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["<scp>SIRA<\/scp>\n                    : Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3700-5935","authenticated-orcid":false,"given":"Yaman","family":"Umuroglu","sequence":"first","affiliation":[{"name":"AMD Research, Trondheim, Norway and Norwegian University of Science and Technology, Trondheim, Norway"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6956-0657","authenticated-orcid":false,"given":"Christoph","family":"Berganski","sequence":"additional","affiliation":[{"name":"Paderborn University, Paderborn, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4987-5708","authenticated-orcid":false,"given":"Felix","family":"Jentzsch","sequence":"additional","affiliation":[{"name":"Paderborn University, Paderborn, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8851-8186","authenticated-orcid":false,"given":"Michal","family":"Danilowicz","sequence":"additional","affiliation":[{"name":"AGH University of Krakow, Krakow, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6798-4444","authenticated-orcid":false,"given":"Tomasz","family":"Kryjak","sequence":"additional","affiliation":[{"name":"AGH University of Krakow, Krakow, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7905-8357","authenticated-orcid":false,"given":"Charalampos","family":"Bezaitis","sequence":"additional","affiliation":[{"name":"Norwegian University of Science and Technology, Trondheim, Norway"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4232-6976","authenticated-orcid":false,"given":"Magnus","family":"Sj\u00e4lander","sequence":"additional","affiliation":[{"name":"Norwegian University of Science and Technology, Trondheim, Norway"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1669-5519","authenticated-orcid":false,"given":"Ian","family":"Colbert","sequence":"additional","affiliation":[{"name":"AMD, San Jose, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3998-7896","authenticated-orcid":false,"given":"Thomas","family":"Preusser","sequence":"additional","affiliation":[{"name":"AMD Research, Dresden, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2457-7819","authenticated-orcid":false,"given":"Jakoba","family":"Petri-Koenig","sequence":"additional","affiliation":[{"name":"AMD Research, Dublin, Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7833-4057","authenticated-orcid":false,"given":"Michaela","family":"Blott","sequence":"additional","affiliation":[{"name":"AMD Research, Dublin, Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,27]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3547141"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470567"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","unstructured":"QONNX authors. 2022. fastmachinelearning\/qonnx. DOI: 10.5281\/zenodo.7622236","DOI":"10.5281\/zenodo.7622236"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICFPT64416.2024.11113391"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3242897"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3507570"},{"key":"e_1_3_2_8_2","unstructured":"Javier Campos Zhen Dong Javier Duarte Amir Gholami Michael W. Mahoney Jovan Mitrevski and Nhan Tran. 2023. End-to-end codesign of Hessian-aware quantized neural networks for FPGAs and ASICs. arXiv:2304.06745. Retrieved from https:\/\/arxiv.org\/abs\/2304.06745"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 10th International Workshop on Frontiers in Handwriting Recognition","author":"Chellapilla Kumar","year":"2006","unstructured":"Kumar Chellapilla, Sidd Puri, and Patrice Simard. 2006. High performance convolutional neural networks for document processing. In Proceedings of the 10th International Workshop on Frontiers in Handwriting Recognition. Suvisoft."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/SiPS52927.2021.00028"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-021-00356-5"},{"key":"e_1_3_2_12_2","unstructured":"Ian Colbert Fabian Grob Giuseppe Franco Jinjie Zhang and Rayan Saab. 2024. Accumulator-aware post-training quantization. arXiv:2409.17092. Retrieved from https:\/\/arxiv.org\/abs\/2409.17092"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01558"},{"key":"e_1_3_2_14_2","unstructured":"QONNX Community. 2025. QONNX Model Zoo. Retrieved May 16 2025 from https:\/\/github.com\/fastmachinelearning\/qonnx_model_zoo"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/1120725.1121055"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2003.818119"},{"key":"e_1_3_2_17_2","first-page":"873","volume-title":"Proceedings of Machine Learning and Systems","volume":"3","author":"Dai Steve","year":"2021","unstructured":"Steve Dai, Rangha Venkatesan, Mark Ren, Brian Zimmer, William Dally, and Brucek Khailany. 2021. VS-Quant: Per-vector scaled quantization for accurate low-precision neural network inference. In Proceedings of Machine Learning and Systems 3 (2021), 873\u2013884."},{"issue":"3","key":"e_1_3_2_18_2","first-page":"15","article-title":"A logical formalization of the notion of interval dependency","volume":"1","author":"Dawood Hend","year":"2019","unstructured":"Hend Dawood and Yasser Dawood. 2019. A logical formalization of the notion of interval dependency. Online Mathematics Journal 1, 3 (2019), 15\u201336.","journal-title":"Online Mathematics Journal"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1088\/1748-0221\/13\/07\/P07027"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.3333552"},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Frantar Elias","year":"2023","unstructured":"Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan-Adrian Alistarh. 2023. OPTQ: Accurate post-training quantization for generative pre-trained transformers. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASPDAC.2013.6509694"},{"key":"e_1_3_2_23_2","unstructured":"Sven Gowal Krishnamurthy Dvijotham Robert Stanforth Rudy Bunel Chongli Qin Jonathan Uesato Relja Arandjelovic Timothy Mann and Pushmeet Kohli. 2018. On the effectiveness of interval bound propagation for training verifiably robust models. arXiv:1810.12715. Retrieved from https:\/\/arxiv.org\/abs\/1810.12715"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Mathew Hall and Vaughn Betz. 2020. HPIPE: Heterogeneous layer-pipelined and sparse-aware CNN inference for FPGAs. arXiv:2007.10451. Retrieved from https:\/\/arxiv.org\/abs\/2007.10451","DOI":"10.1145\/3373087.3375380"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3236974"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2024.3440884"},{"key":"e_1_3_2_27_2","unstructured":"Benoit Jacob Skirmantas Kligys Bo Chen Menglong Zhu Matthew Tang Andrew G. Howard Hartwig Adam and Dmitry Kalenichenko. 2017. Quantization and training of neural networks for efficient integer-arithmetic-only inference. arXiv:1712.05877. Retrieved from http:\/\/arxiv.org\/abs\/1712.05877"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS64960.2025.00011"},{"key":"e_1_3_2_29_2","unstructured":"Raghuraman Krishnamoorthi. 2018. Quantizing deep convolutional networks for efficient inference: A whitepaper. arXiv:1806.08342. Retrieved from https:\/\/arxiv.org\/abs\/1806.08342"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2006.873887"},{"key":"e_1_3_2_31_2","unstructured":"Lu Lu Yeonjong Shin Yanhui Su and George Em Karniadakis. 2019. Dying ReLU and initialization: Theory and numerical examples. arXiv:1903.06733. Retrieved from https:\/\/arxiv.org\/abs\/1903.06733"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPA63168.2024.00143"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3608447"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898717716"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00141"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01514"},{"key":"e_1_3_2_37_2","volume-title":"Proceedings of the 4th Workshop on Accelerated Machine Learning (AccML) at HiPEAC 2022 Conference","author":"Pappalardo Alessandro","year":"2022","unstructured":"Alessandro Pappalardo, Yaman Umuroglu, Michaela Blott, Jovan Mitrevski, Ben Hawks, Nhan Tran, Vladimir Loncar, Sioni Summers, Hendrik Borras, Jules Muhizi, et al. 2022. QONNX: Representing arbitrary-precision quantized neural networks. In Proceedings of the 4th Workshop on Accelerated Machine Learning (AccML) at HiPEAC 2022 Conference. Retrieved from https:\/\/accml.dcs.gla.ac.uk\/papers\/2022\/4thAccML_paper_1(12).pdf"},{"key":"e_1_3_2_38_2","unstructured":"Magnus Sj\u00e4lander Magnus Jahre Gunnar Tufte and Nico Reissmann. 2019. EPIC: An energy-efficient high-performance GPGPU computing research infrastructure. arXiv:1912.05848. Retrieved from https:\/\/arxiv.org\/abs\/1912.05848"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3490422.3502364"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021744"},{"key":"e_1_3_2_41_2","unstructured":"Yaman Umuroglu and Magnus Jahre. 2017. Streamlined deployment for quantized neural networks. arXiv:1709.04060. Retrieved from https:\/\/arxiv.org\/abs\/1709.04060"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSD60849.2023.00032"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00030"},{"key":"e_1_3_2_44_2","first-page":"868","volume-title":"Proceedings of the 29th International Conference on International Joint Conferences on Artificial Intelligence","author":"Xie Hongwei","year":"2021","unstructured":"Hongwei Xie, Yafei Song, Ling Cai, and Mingyang Li. 2021. Overflow aware quantization: Accelerating neural network inference by low-bit multiply-accumulate operations. In Proceedings of the 29th International Conference on International Joint Conferences on Artificial Intelligence, 868\u2013875."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL53798.2021.00011"},{"key":"e_1_3_2_46_2","unstructured":"Hongyi Yao Pu Li Jian Cao Xiangcheng Liu Chenying Xie and Bingzhang Wang. 2022. RAPQ: Rescuing accuracy for power-of-two low-bit post-training quantization. arXiv:2204.12322. Retrieved from https:\/\/arxiv.org\/abs\/2204.12322"},{"key":"e_1_3_2_47_2","unstructured":"Aozhong Zhang Naigang Wang Yanxia Deng Xin Li Zi Yang and Penghang Yin. 2024. MagR: Weight magnitude reduction for enhancing post-training quantization. arXiv:2406.00800. Retrieved from https:\/\/arxiv.org\/abs\/2406.00800"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/app12157829"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807510","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T14:04:49Z","timestamp":1779890689000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807510"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,27]]},"references-count":47,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807510"],"URL":"https:\/\/doi.org\/10.1145\/3807510","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,27]]},"assertion":[{"value":"2025-05-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}