{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:11:38Z","timestamp":1750306298943,"version":"3.41.0"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2016,3,7]],"date-time":"2016-03-07T00:00:00Z","timestamp":1457308800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"ONR-G","award":["N62909-14-1-N072"],"award-info":[{"award-number":["N62909-14-1-N072"]}]},{"name":"E4Bio RTD project","award":["200021_159853"],"award-info":[{"award-number":["200021_159853"]}]},{"name":"Swiss NSF"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2016,7,21]]},"abstract":"<jats:p>The performance and the efficiency of recent computing platforms have been deeply influenced by the widespread adoption of hardware accelerators, such as graphics processing units (GPUs) or field-programmable gate arrays (FPGAs), which are often employed to support the tasks of general-purpose processors (GPPs). One of the main advantages of these accelerators over their sequential counterparts (GPPs) is their ability to perform massive parallel computation. However, to exploit this competitive edge, it is necessary to extract the parallelism from the target algorithm to be executed, which generally is a very challenging task.<\/jats:p>\n          <jats:p>This concept is demonstrated, for instance, by the poor performance achieved on relevant multimedia algorithms, such as Chambolle, which is a well-known algorithm employed for the optical flow estimation. The implementations of this algorithm that can be found in the state of the art are generally based on GPUs but barely improve the performance that can be obtained with a powerful GPP. In this article, we propose a novel approach to extract the parallelism from computation-intensive multimedia algorithms, which includes an analysis of their dependency schema and an assessment of their data reuse. We then perform a thorough analysis of the Chambolle algorithm, providing a formal proof of its inner data dependencies and locality properties. Then, we exploit the considerations drawn from this analysis by proposing an architectural template that takes advantage of the fine-grained parallelism of FPGA devices. Moreover, since the proposed template can be instantiated with different parameters, we also propose a design metric, the expansion rate, to help the designer in the estimation of the efficiency and performance of the different instances, making it possible to select the right one before the implementation phase. We finally show, by means of experimental results, how the proposed analysis and parallelization approach leads to the design of efficient and high-performance FPGA-based implementations that are orders of magnitude faster than the state-of-the-art ones.<\/jats:p>","DOI":"10.1145\/2851497","type":"journal-article","created":{"date-parts":[[2016,3,8]],"date-time":"2016-03-08T13:33:07Z","timestamp":1457443987000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Parallelizing the Chambolle Algorithm for Performance-Optimized Mapping on FPGA Devices"],"prefix":"10.1145","volume":"15","author":[{"given":"Ivan","family":"Beretta","sequence":"first","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne, Lausanne, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vincenzo","family":"Rana","sequence":"additional","affiliation":[{"name":"Politecnico di Milano, Milano"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abdulkadir","family":"Akin","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne, Lausanne, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alessandro Antonio","family":"Nacci","sequence":"additional","affiliation":[{"name":"Politecnico di Milano, Milano"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donatella","family":"Sciuto","sequence":"additional","affiliation":[{"name":"Politecnico di Milano, Milano"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Atienza","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne, Lausanne, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,3,7]]},"reference":[{"volume-title":"Proceedings of the National Radio Science Conference (NRSC\u201909)","author":"Abutaleb M. M.","key":"e_1_2_1_1_1","unstructured":"M. M. Abutaleb , A. Hamdy , M. E. Abuelwafa , and E. M. Saad . 2009. A reliable FPGA-based real-time optical-flow estimation . In Proceedings of the National Radio Science Conference (NRSC\u201909) . 1--8. M. M. Abutaleb, A. Hamdy, M. E. Abuelwafa, and E. M. Saad. 2009. A reliable FPGA-based real-time optical-flow estimation. In Proceedings of the National Radio Science Conference (NRSC\u201909). 1--8."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/DATE.2011.5763232"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ReConFig.2014.7032547"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0036139998340170"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5539932"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSPIT.2007.4458079"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 4th International Conference on Computer Vision. 231--236","author":"Black M. J.","year":"1993","unstructured":"M. J. Black and P. Anandan . 1993. A framework for the robust estimation of optical flow . In Proceedings of the 4th International Conference on Computer Vision. 231--236 . DOI:http:\/\/dx.doi.org\/10.1109\/ICCV. 1993 .378214 10.1109\/ICCV.1993.378214 M. J. Black and P. Anandan. 1993. A framework for the robust estimation of optical flow. In Proceedings of the 4th International Conference on Computer Vision. 231--236. DOI:http:\/\/dx.doi.org\/10.1109\/ICCV.1993.378214"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1754386.1754387"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:JMIV.0000011325.36760.1e"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2012.14"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2555729.2555733"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 2014 World Congress on Computer Applications and Systems (WCCAIS\u201914)","author":"Ghodhbani R.","year":"2014","unstructured":"R. Ghodhbani , T. Saidani , L. Horrigue , and M. Atri . 2014. Analysis and implementation of parallel causal bit plane coding in JPEG2000 standard . In Proceedings of the 2014 World Congress on Computer Applications and Systems (WCCAIS\u201914) . 1--6. DOI:http:\/\/dx.doi.org\/10.1109\/WCCAIS. 2014 .6916602 10.1109\/WCCAIS.2014.6916602 R. Ghodhbani, T. Saidani, L. Horrigue, and M. Atri. 2014. Analysis and implementation of parallel causal bit plane coding in JPEG2000 standard. In Proceedings of the 2014 World Congress on Computer Applications and Systems (WCCAIS\u201914). 1--6. DOI:http:\/\/dx.doi.org\/10.1109\/WCCAIS.2014.6916602"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 2nd International Symposium on Image and Signal Processing and Analysis (ISPA\u201901)","author":"Jamro E.","year":"2001","unstructured":"E. Jamro and K. Wiatr . 2001. Convolution operation implemented in FPGA structures for real-time image processing . In Proceedings of the 2nd International Symposium on Image and Signal Processing and Analysis (ISPA\u201901) . 417--422. DOI:http:\/\/dx.doi.org\/10.1109\/ISPA. 2001 .938666 10.1109\/ISPA.2001.938666 E. Jamro and K. Wiatr. 2001. Convolution operation implemented in FPGA structures for real-time image processing. In Proceedings of the 2nd International Symposium on Image and Signal Processing and Analysis (ISPA\u201901). 417--422. DOI:http:\/\/dx.doi.org\/10.1109\/ISPA.2001.938666"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLDI-DAT.2013.6533845"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the International Conference on Control, Automation, and Systems (ICCAS\u201907)","author":"Kim Sungbok","year":"2007","unstructured":"Sungbok Kim , Ilhwa Jeong , and Sanghyup Lee . 2007 . Mobile robot velocity estimation using an array of optical flow sensors . In Proceedings of the International Conference on Control, Automation, and Systems (ICCAS\u201907) . 616--621. DOI:http:\/\/dx.doi.org\/10.1109\/ICCAS.2007.4407097 10.1109\/ICCAS.2007.4407097 Sungbok Kim, Ilhwa Jeong, and Sanghyup Lee. 2007. Mobile robot velocity estimation using an array of optical flow sensors. In Proceedings of the International Conference on Control, Automation, and Systems (ICCAS\u201907). 616--621. DOI:http:\/\/dx.doi.org\/10.1109\/ICCAS.2007.4407097"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 5th Annual IEEE Symposium on Field-Programmable Custom Computing Machines. 226--232","author":"Li Yamin","year":"1997","unstructured":"Yamin Li and Wanming Chu . 1997 . Implementation of single precision floating point square root on FPGAs . In Proceedings of the 5th Annual IEEE Symposium on Field-Programmable Custom Computing Machines. 226--232 . DOI:http:\/\/dx.doi.org\/10.1109\/FPGA.1997.624623 10.1109\/FPGA.1997.624623 Yamin Li and Wanming Chu. 1997. Implementation of single precision floating point square root on FPGAs. In Proceedings of the 5th Annual IEEE Symposium on Field-Programmable Custom Computing Machines. 226--232. DOI:http:\/\/dx.doi.org\/10.1109\/FPGA.1997.624623"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/839276.839928"},{"key":"e_1_2_1_19_1","unstructured":"NVIDIA. 2007. NVIDIA CUDA Compute Unified Device Architecture Programming Guide. Available at http:\/\/www.nvidia.com.  NVIDIA. 2007. NVIDIA CUDA Compute Unified Device Architecture Programming Guide. Available at http:\/\/www.nvidia.com."},{"key":"e_1_2_1_20_1","unstructured":"NVIDIA. 2009. NVIDIA Next Generation CUDA Compute Architecture: Fermi. Available at http:\/\/www.nvidia.com.  NVIDIA. 2009. NVIDIA Next Generation CUDA Compute Architecture: Fermi. Available at http:\/\/www.nvidia.com."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-005-3960-y"},{"key":"e_1_2_1_22_1","first-page":"511","article-title":"A duality based algorithm for TV-L1-optical-flow image registration","volume":"10","author":"Pock T.","year":"2007","unstructured":"T. Pock , M. Urschler , C. Zach , R. Beichel , and H. Bischof . 2007 . A duality based algorithm for TV-L1-optical-flow image registration . Med Image Computing and Computer Assisted Intervention 10 , 511 -- 518 . http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/18044607. T. Pock, M. Urschler, C. Zach, R. Beichel, and H. Bischof. 2007. A duality based algorithm for TV-L1-optical-flow image registration. Med Image Computing and Computer Assisted Intervention 10, 511--518. http:\/\/www.ncbi.nlm.nih.gov\/pubmed\/18044607.","journal-title":"Med Image Computing and Computer Assisted Intervention"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008193018623"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/0167-2789(92)90242-F"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAE.2010.5452039"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479894270427"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 2000 International Conference on Image Processing","volume":"1","author":"Sun S.","year":"2000","unstructured":"S. Sun , D. Haynor , and Yongmin Kim . 2000 . Motion estimation based on optical flow with adaptive gradients . In Proceedings of the 2000 International Conference on Image Processing , Vol. 1 . 852--855. DOI:http:\/\/dx.doi.org\/10.1109\/ICIP.2000.901093 10.1109\/ICIP.2000.901093 S. Sun, D. Haynor, and Yongmin Kim. 2000. Motion estimation based on optical flow with adaptive gradients. In Proceedings of the 2000 International Conference on Image Processing, Vol. 1. 852--855. DOI:http:\/\/dx.doi.org\/10.1109\/ICIP.2000.901093"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.24781"},{"key":"e_1_2_1_30_1","unstructured":"Xilinx. 2009. Virtex-5 Family Overview DS100 (v5.0). Available at http:\/\/www.xilinx.com.  Xilinx. 2009. Virtex-5 Family Overview DS100 (v5.0). Available at http:\/\/www.xilinx.com."},{"volume-title":"Proceedings of the 29th DAGM Conference on Pattern Recognition. 214--223","author":"Zach C.","key":"e_1_2_1_31_1","unstructured":"C. Zach , T. Pock , and H. Bischof . 2007. A duality based approach for realtime TV-L1 optical flow . In Proceedings of the 29th DAGM Conference on Pattern Recognition. 214--223 . http:\/\/dl.acm.org\/citation.cfm?id&equals;1771530.1771554 C. Zach, T. Pock, and H. Bischof. 2007. A duality based approach for realtime TV-L1 optical flow. In Proceedings of the 29th DAGM Conference on Pattern Recognition. 214--223. http:\/\/dl.acm.org\/citation.cfm?id&equals;1771530.1771554"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2851497","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2851497","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:39:15Z","timestamp":1750221555000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2851497"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,3,7]]},"references-count":30,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2016,7,21]]}},"alternative-id":["10.1145\/2851497"],"URL":"https:\/\/doi.org\/10.1145\/2851497","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2016,3,7]]},"assertion":[{"value":"2015-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-03-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}