{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T21:57:24Z","timestamp":1766267844800,"version":"3.41.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,7,18]],"date-time":"2023-07-18T00:00:00Z","timestamp":1689638400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"European Union\u2019s Horizon 2020"},{"name":"Marie Sk\u0142odowska-Curie","award":["765866"],"award-info":[{"award-number":["765866"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Model. Perform. Eval. Comput. Syst."],"published-print":{"date-parts":[[2023,9,30]]},"abstract":"<jats:p>\n            The introduction of deep learning algorithms, such as Convolutional Neural Networks (CNNs) in many near-sensor embedded systems, opens new challenges in terms of energy efficiency and hardware performance. An emerging solution to address these challenges is to use tailored heterogeneous hardware accelerators combining processing elements of different architectural natures such as Central Processing Unit (CPU), Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), or Application Specific Integrated Circuit (ASIC). To progress towards heterogeneity, a great asset would be an automated design space exploration tool that chooses, for each accelerated partition of a CNN, the most appropriate architecture considering available resources. To feed such a design space exploration process, models are required that provide very fast yet precise evaluations of alternative architectures or alternative forms of CNNs. Quick configuration estimation could be achieved with few parameters from representative input sequences. This article studies a solution called\n            <jats:italic>flydeling<\/jats:italic>\n            (as a contraction of\n            <jats:italic>fly<\/jats:italic>\n            weight mo\n            <jats:italic>deling<\/jats:italic>\n            ) for obtaining these models by inspiring from the black-box System Identification (SI) domain. We refer to models derived using the proposed approach as\n            <jats:italic>fly<\/jats:italic>\n            weight mo\n            <jats:italic>dels<\/jats:italic>\n            (\n            <jats:italic>flydels<\/jats:italic>\n            ).\n          <\/jats:p>\n          <jats:p>\n            A methodology is proposed to generate these\n            <jats:italic>flydels<\/jats:italic>\n            , using CNN properties as predictor features together with SI techniques with a stochastic excitation input at a feature map dimensions level. For an embedded CPU-FPGA-GPU heterogeneous platform, it is demonstrated that it is possible to learn these Key Performance Indicators (KPIs)\n            <jats:italic>flydels<\/jats:italic>\n            at an early design stage and from high-level application features. For latency, energy, and resource utilization,\n            <jats:italic>flydels<\/jats:italic>\n            obtain estimation errors varying between 5% and 10% with less model parameters compared to state-of-the-art solutions and are built automatically from platform measurements.\n          <\/jats:p>","DOI":"10.1145\/3594870","type":"journal-article","created":{"date-parts":[[2023,5,12]],"date-time":"2023-05-12T11:52:01Z","timestamp":1683892321000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Flydeling: Streamlined Performance Models for Hardware Acceleration of CNNs through System Identification"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2765-0228","authenticated-orcid":false,"given":"Walther","family":"Carballo-Hern\u00e1ndez","sequence":"first","affiliation":[{"name":"Universit\u00e9 Clermont Auvergne, CNRS, Clermont Auvergne INP, Institut Pascal, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1158-0915","authenticated-orcid":false,"given":"Maxime","family":"Pelcat","sequence":"additional","affiliation":[{"name":"IETR, UMR CNRS 6164, Institut Pascal, UMR CNRS 6602, Univ Rennes, INSA Rennes, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7719-1106","authenticated-orcid":false,"given":"Shuvra S.","family":"Bhattacharyya","sequence":"additional","affiliation":[{"name":"University of Maryland, College Park, USA and IETR, UMR CNRS 6164, Rennes, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4230-3988","authenticated-orcid":false,"given":"Ricardo Carmona","family":"Gal\u00e1n","sequence":"additional","affiliation":[{"name":"Instituto de Microelectr\u00f3nica de Sevilla (IMSE-CNM), CSIC-Universidad de Sevilla, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5899-4672","authenticated-orcid":false,"given":"Fran\u00e7ois","family":"Berry","sequence":"additional","affiliation":[{"name":"Universit\u00e9 Clermont Auvergne, CNRS, Clermont Auvergne INP, Institut Pascal, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,7,18]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/LES.2017.2743247"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/1465482.1465560"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2018.01.002"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3457388.3458666"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2962131"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1908.09791"},{"key":"e_1_3_1_8_2","unstructured":"Walther Carballo-Hern\u00e1ndez Maxime Pelcat and Fran\u00e7ois Berry. 2021. Why is FPGA-GPU heterogeneity the best option for embedded deep neural networks?arxiv:2102.01343 [cs.AR]."},{"key":"e_1_3_1_9_2","volume-title":"12th International Workshop on Worst-Case Execution Time Analysis","author":"Cassez Franck","year":"2012","unstructured":"Franck Cassez, Ren\u00e9 Rydhof Hansen, and Mads Chr Olesen. 2012. What is a timing anomaly? In 12th International Workshop on Worst-Case Execution Time Analysis. Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik."},{"key":"e_1_3_1_10_2","unstructured":"Sharan Chetlur Cliff Woolley Philippe Vandermersch Jonathan Cohen John Tran Bryan Catanzaro and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arxiv:1410.0759 [cs.NE]."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-11179-7_36"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2009.5206848"},{"key":"e_1_3_1_13_2","unstructured":"Nicolas Derumigny Fabian Gruber Th\u00e9ophile Bastian Christophe Guillon Louis-Noel Pouchet and Fabrice Rastello. 2020. From micro-OPs to abstract resources: Constructing a simpler CPU performance model through microbenchmarking. arxiv:2012.11473 [cs.AR]."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2007.08668"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.2200\/s00273ed1v01y201006cac010"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-36187-1_45"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2019.07.007"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.5812\/ijem.3505"},{"key":"e_1_3_1_19_2","article-title":"INFER: INterFerence-aware estimation of runtime for concurrent CNN execution on DPUs","author":"Goel S.","year":"2020","unstructured":"S. Goel, R. Kedia, M. Balakrishnan, and R. Sen. 2020. INFER: INterFerence-aware estimation of runtime for concurrent CNN execution on DPUs. In International Conference on Field Programmable Technology (FPT). Retrieved from https:\/\/www.cse.iitd.ac.in\/kedia\/research\/DPU_runtime_FPT_short_paper.pdf.","journal-title":"International Conference on Field Programmable Technology (FPT)"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-018-6861-0"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1512.03385"},{"key":"e_1_3_1_22_2","article-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size","author":"Iandola F. N.","year":"2016","unstructured":"F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size. In Conference on Computer Vision and Pattern Recognition (CVPR\u201916). arXiv:1602.07360v4.","journal-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201916)"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/iccad.2010.5653959"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60939-9_3"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-85729-522-4"},{"key":"e_1_3_1_26_2","volume-title":"Learning Multiple Layers of Features from Tiny Images","author":"Krizhevsky Alex","year":"2009","unstructured":"Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report. Massachusetts Institute of Technology (MIT) and New York University (NYU). Retrieved from https:\/\/www.cs.toronto.edu\/kriz\/learning-features-2009-TR.pdf."},{"key":"e_1_3_1_27_2","article-title":"ImageNet classification with deep convolutional neural networks","author":"Krizhevsky A.","year":"2012","unstructured":"A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Conference on Neural Information Processing Systems (NIPS).","journal-title":"Conference on Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-46805-6_19"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/1168917.1168881"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2106.08630"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1090\/qam\/10666"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/tcad.2019.2897634"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/fpl.2019.00069"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1137\/0111030"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/2379776.2379786"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/greencomp.2010.5598315"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/tcad.2020.3003276"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTC.2018.8539528"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/isvlsi.2018.00143"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062257"},{"key":"e_1_3_1_41_2","article-title":"Automatic differentiation in PyTorch","author":"Paszke A.","year":"2017","unstructured":"A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. 2017. Automatic differentiation in PyTorch. In NIPS 2017 Workshop. Retrieved from https:\/\/openreview.net\/forum?id=BJJsrmfCZ.","journal-title":"NIPS 2017 Workshop"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2017.2774822"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/mdat.2016.2626445"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1.1.135.865"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-019-00645-y"},{"key":"e_1_3_1_46_2","first-page":"448","article-title":"Batch normalization: Accelerating deep network training by reducing internal covariate shift","author":"Segey I.","year":"2015","unstructured":"I. Segey and C. Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In 32nd International Conference on International Conference on Machine Learning (ICML\u201915). 448\u2013456. Retrieved from https:\/\/arxiv.org\/abs\/1502.03167v3.","journal-title":"32nd International Conference on International Conference on Machine Learning (ICML\u201915)"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-25636-4_5"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1409.1556"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1512.00567"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1807.11626"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2019.00040"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICACCS48705.2020.9074444"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/jiot.2020.2981684"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/icnnb.2005.1614615"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_1_56_2","unstructured":"Nan Wu Yuan Xie and Cong Hao. 2021. IronMan: GNN-assisted design space exploration in high-level synthesis via reinforcement learning. arxiv:2102.08138 [cs.AR]."},{"key":"e_1_3_1_57_2","article-title":"Designing energy-efficient convolutional neural networks using energy-aware pruning","author":"Yang Tien-Ju","year":"2017","unstructured":"Tien-Ju Yang, Yu-Hsin Chen, and Vivienne Sze. 2017. Designing energy-efficient convolutional neural networks using energy-aware pruning. In Conference on Computer Vision and Pattern Recognition (CVPR\u201917). arxiv:1611.05128 [cs.CV].","journal-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201917)"},{"key":"e_1_3_1_58_2","volume-title":"2nd USENIX Workshop on Hot Topics in Edge Computing (HotEdge\u201919)","author":"Zhou L.","year":"2019","unstructured":"L. Zhou, H. Wen, R. Teodorescu, and David H. C. Du. 2019. Distributing deep neural networks with containerized partitions at the edge. In 2nd USENIX Workshop on Hot Topics in Edge Computing (HotEdge\u201919). USENIX Association, Renton, WA. Retrieved from https:\/\/www.usenix.org\/conference\/hotedge19\/presentation\/zhou."},{"key":"e_1_3_1_59_2","volume-title":"System Identification","author":"\u00c5str\u00f6m K. J.","year":"1970","unstructured":"K. J. \u00c5str\u00f6m and E. Pieter. 1970. System Identification. Retrieved from https:\/\/portal.research.lu.se\/portal\/files\/48192977\/TFRT_7011.pdf."}],"container-title":["ACM Transactions on Modeling and Performance Evaluation of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594870","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3594870","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:08Z","timestamp":1750182548000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594870"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,18]]},"references-count":58,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,9,30]]}},"alternative-id":["10.1145\/3594870"],"URL":"https:\/\/doi.org\/10.1145\/3594870","relation":{},"ISSN":["2376-3639","2376-3647"],"issn-type":[{"type":"print","value":"2376-3639"},{"type":"electronic","value":"2376-3647"}],"subject":[],"published":{"date-parts":[[2023,7,18]]},"assertion":[{"value":"2022-05-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-04","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}