{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:23:08Z","timestamp":1759332188901,"version":"3.41.0"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2024,2,14]],"date-time":"2024-02-14T00:00:00Z","timestamp":1707868800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Recently Processing-in-Memory (PIM) has become a promising solution to achieve energy-efficient computation in data-intensive applications by placing computation near or inside the memory. In most Deep Learning (DL) frameworks, a user manually partitions a model\u2019s computational graph (CG) onto the computing devices by considering the devices\u2019 capability and the data transfer. The Deep Neural Network (DNN) models become increasingly complex for improving accuracy; thus, it is exceptionally challenging to partition the execution to achieve the best performance, especially on a PIM-based platform requiring frequent offloading of large amounts of data.<\/jats:p>\n          <jats:p>This article proposes two novel algorithms for DL inference to resolve the challenge: low-overhead profiling and optimal model partitioning. First, we reconstruct CG by considering the devices\u2019 capability to represent all the possible scheduling paths. Second, we develop a profiling algorithm to find the required minimum profiling paths to measure all the node and edge costs of the reconstructed CG. Finally, we devise the model partitioning algorithm to get the optimal minimum execution time using the dynamic programming technique with the profiled data. We evaluated our work by executing the BERT, RoBERTa, and GPT-2 models on the ARM multicores with the PIM-modeled FPGA platform with various sequence lengths. For three computing devices in the platform, i.e., CPU serial\/parallel and PIM executions, we could find all the costs only in four profile runs, three for node costs and one for edge costs. Also, our model partitioning algorithm achieved the highest performance in all the experiments over the execution with manually assigned device priority and the state-of-the-art greedy approach.<\/jats:p>","DOI":"10.1145\/3628599","type":"journal-article","created":{"date-parts":[[2023,11,2]],"date-time":"2023-11-02T04:39:35Z","timestamp":1698899975000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Optimal Model Partitioning with Low-Overhead Profiling on the PIM-based Platform for Deep Learning Inference"],"prefix":"10.1145","volume":"29","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6259-9413","authenticated-orcid":false,"given":"Seok Young","family":"Kim","sequence":"first","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1179-4730","authenticated-orcid":false,"given":"Jaewook","family":"Lee","sequence":"additional","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8294-1079","authenticated-orcid":false,"given":"Yoonah","family":"Paik","sequence":"additional","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6224-7074","authenticated-orcid":false,"given":"Chang Hyun","family":"Kim","sequence":"additional","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6555-1741","authenticated-orcid":false,"given":"Won Jun","family":"Lee","sequence":"additional","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6161-0871","authenticated-orcid":false,"given":"Seon Wook","family":"Kim","sequence":"additional","affiliation":[{"name":"Korea University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,2,14]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation.","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zhen. 2016. Tensorflow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation. Savannah, GA, 265\u2013283."},{"key":"e_1_3_1_3_2","unstructured":"Open AI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPPW.2012.14"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400716"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2482767.2482794"},{"key":"e_1_3_1_7_2","volume-title":"Introduction to Algorithms (2nd. ed.)","author":"Cormen Thomas H.","year":"2001","unstructured":"Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. 2001. Introduction to Algorithms (2nd. ed.). MIT Press."},{"key":"e_1_3_1_8_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370866"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/2856636.2856639"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19861-8_16"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/54.232470"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00040"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15291-7_23"},{"key":"e_1_3_1_15_2","article-title":"HTG-Z920","author":"Global Hitech","year":"2017","unstructured":"Hitech Global. 2017. HTG-Z920. Retrieved August 07, 2022 from https:\/\/www.xilinx.com\/products\/boards-and-kits\/1-qwrzuv.html","journal-title":"R"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_17_2","volume-title":"Recurrent Neural Networks: Design and Applications (1st ed.)","author":"Jain Lakhmi","year":"1999","unstructured":"Lakhmi Jain and Larry Medsker. 1999. Recurrent Neural Networks: Design and Applications (1st ed.). CRC Press, Inc."},{"key":"e_1_3_1_18_2","unstructured":"JEDEC Standard: DDR4 SDRAM JESD79-4B. 2012. Retrieved from https:\/\/www.jedec.org\/standards-documents\/docs\/jesd79-4a"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3097700"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3065365"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465007"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.14"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00013"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00013"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2928289"},{"key":"e_1_3_1_26_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669121"},{"key":"e_1_3_1_28_2","article-title":"Direct Memory Access (DMA) (Part III)","year":"2008","unstructured":"Microchip. 2008. Direct Memory Access (DMA) (Part III). Retrieved January 02, 2022 from http:\/\/ww1.microchip.com\/downloads\/en\/devicedoc\/70215c.pdf","journal-title":"R"},{"key":"e_1_3_1_29_2","unstructured":"NVIDIA Corporation. 2020. CUDA. Retrieved January 02 2022 from https:\/\/developer.nvidia.com\/cuda-toolkit accessed: 2022-01-02."},{"key":"e_1_3_1_30_2","unstructured":"ONNX Developers. 2019. Open Neural Network Exchange (ONNX). Retrieved December 22 2021 from https:\/\/onnx.ai\/"},{"key":"e_1_3_1_31_2","unstructured":"ONNX Runtime Developers. 2021. ONNX Runtime. Retrieved December 22 2021 from https:\/\/onnxruntime.ai\/"},{"key":"e_1_3_1_32_2","first-page":"8026","volume-title":"Proceedings of the 32th Neural Information Processing Systems.","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the 32th Neural Information Processing Systems. Vancouver, Canada, 8026\u20138037."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/2744769.2744819"},{"key":"e_1_3_1_34_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.231340"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2509972"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2304576.2304625"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2010.69"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2002.1176259"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.18"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00024"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3628599","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3628599","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:44Z","timestamp":1750291424000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3628599"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,14]]},"references-count":40,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3628599"],"URL":"https:\/\/doi.org\/10.1145\/3628599","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2024,2,14]]},"assertion":[{"value":"2022-08-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}