{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T11:59:06Z","timestamp":1784894346508,"version":"3.55.0"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,5,25]],"date-time":"2022-05-25T00:00:00Z","timestamp":1653436800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Science and Technology","award":["MOST-110-2221-E-001-007-MY2, MOST-110-2221-E-002-072-MY2, and MOST-110-2221-E-001-004"],"award-info":[{"award-number":["MOST-110-2221-E-001-007-MY2, MOST-110-2221-E-002-072-MY2, and MOST-110-2221-E-001-004"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>\n            Video captioning is a core technology to many important applications, such as AI-assisted medical diagnosis, video question answering, storytelling through videos, and lip-reading. Video captioning employs a hybrid CNN + RNN model. Accelerating such a hybrid model on a heterogeneous system is challenging for two reasons. First, CNN and RNN exhibit very different computing behaviors, making the mapping between computation and heterogeneous devices difficult. Second, data dependency exists between the CNN and RNN within a video frame and between adjacent RNNs across video frames. These data dependencies prohibit the full parallelization of the hybrid model. The issues also include the utilization of accelerator resources, which is critical to maximizing the performance. In this work, we propose a fine-grained scheduling scheme for mapping computation and devices within a video frame, and a pipeline scheduling scheme for exploiting maximum parallelism between the execution of the video frames. In addition, we propose two capacity-guided scheduling methods. On the server, the concurrent kernel execution mechanism is exploited for improving GPU utilization. On the edge platform, we rearrange CNN computation among the CPU and EdgeTPUs guided by the EdgeTPU\u2019s SRAM capacity so that balanced computation is achieved and off-chip memory overhead is minimized. Experimental results show that our scheduling scheme improves video captioning performance by up to 3.24\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( \\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            with CPU + GPU collaboration over the GPU-only execution. On an edge platform with an ARM CPU and two EdgeTPUs, our CPU + EdgeTPU scheduling exhibits outstanding performance, which achieves up to 54.9\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( \\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            speedup compared to using ARM CPU only and can perform video captioning of 59 frames per second.\n          <\/jats:p>","DOI":"10.1145\/3527609","type":"journal-article","created":{"date-parts":[[2022,4,1]],"date-time":"2022-04-01T11:39:12Z","timestamp":1648813152000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Accelerating Video Captioning on Heterogeneous System Architectures"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4209-5258","authenticated-orcid":false,"given":"Horng-Ruey","family":"Huang","sequence":"first","affiliation":[{"name":"Academia Sinica, Nankang, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7649-7581","authenticated-orcid":false,"given":"Ding-Yong","family":"Hong","sequence":"additional","affiliation":[{"name":"Academia Sinica, Nankang, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jan-Jan","family":"Wu","sequence":"additional","affiliation":[{"name":"Academia Sinica, Nankang, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kung-Fu","family":"Chen","sequence":"additional","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pangfeng","family":"Liu","sequence":"additional","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei-Chung","family":"Hsu","sequence":"additional","affiliation":[{"name":"National Taiwan University, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,5,25]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"265","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, et\u00a0al. 2016. TensorFlow: A system for large-scale machine learning. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation. 265\u2013283."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783725"},{"key":"e_1_3_2_4_2","first-page":"340","volume-title":"Proceedings of the International Symposium on Code Generation and Optimization","author":"Anderson Andrew","year":"2018","unstructured":"Andrew Anderson and David Gregg. 2018. Optimal DNN primitive selection with partitioned Boolean quadratic programming. In Proceedings of the International Symposium on Code Generation and Optimization. 340\u2013351."},{"key":"e_1_3_2_5_2","article-title":"Optimizing performance of recurrent neural networks on GPUs","author":"Appleyard Jeremy","year":"2016","unstructured":"Jeremy Appleyard, Tom\u00e1s Kocisk\u00fd, and Phil Blunsom. 2016. Optimizing performance of recurrent neural networks on GPUs. arXiv:1604.01946 (2016).","journal-title":"arXiv:1604.01946"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/SLT.2018.8639690"},{"key":"e_1_3_2_7_2","first-page":"513","volume-title":"Proceedings of the 4th Machine Learning for Healthcare Conference","author":"Biswal Siddharth","year":"2019","unstructured":"Siddharth Biswal, Cao Xiao, M. Brandon Westover, and Jimeng Sun. 2019. EEGtoText: Learning to write medical reports from EEG recordings. In Proceedings of the 4th Machine Learning for Healthcare Conference. 513\u2013531."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001177"},{"key":"e_1_3_2_9_2","article-title":"cuDNN: Efficient primitives for deep learning","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arXiv:1410.0759 (2014).","journal-title":"arXiv:1410.0759"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-4012"},{"key":"e_1_3_2_11_2","unstructured":"Coral. 2020. Coral Dev Board. Retrieved April 8 2022 from https:\/\/coral.ai\/products\/dev-board\/."},{"key":"e_1_3_2_12_2","unstructured":"Coral. 2021. Operations Supported by the EdgeTPU. Retrieved April 8 2022 from https:\/\/coral.ai\/docs\/edgetpu\/models-intro#supported-operations."},{"key":"e_1_3_2_13_2","first-page":"1676","volume-title":"Proceedings of the 35th International Conference on Machine Learning","author":"Gao Yuanxiang","year":"2018","unstructured":"Yuanxiang Gao, Li Chen, and Baochun Li. 2018. Spotlight: Optimizing device placement for training deep neural networks. In Proceedings of the 35th International Conference on Machine Learning. 1676\u20131684."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_16_2","article-title":"Exploring edge TPU for network intrusion detection in IoT","author":"Hosseininoorbin Seyedehfaezeh","year":"2021","unstructured":"Seyedehfaezeh Hosseininoorbin, Siamak Layeghy, Mohanad Sarhan, Raja Jurdak, and Marius Portmann. 2021. Exploring edge TPU for network intrusion detection in IoT. arXiv:2103.16295 (2021).","journal-title":"arXiv:2103.16295"},{"key":"e_1_3_2_17_2","article-title":"MobileNets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861 (2017).","journal-title":"arXiv:1704.04861"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358263"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00112"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1007\/978-3-030-27562-4_4","volume-title":"Embedded Computer Systems: Architectures, Modeling, and Simulation","author":"Huang Zi Xuan","year":"2019","unstructured":"Zi Xuan Huang, Shen Yu Fu, and Wei Chung Hsu. 2019. Efficient dynamic device placement for deep neural network training on heterogeneous systems. In Embedded Computer Systems: Architectures, Modeling, and Simulation. Springer, 51\u201364."},{"key":"e_1_3_2_22_2","unstructured":"Intel. 2017. Neural Compute Stick. Retrieved April 8 2022 from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/intel-movidius-neural-compute-stick.html."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01189-x"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3012722"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS47924.2020.00042"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/398"},{"key":"e_1_3_2_27_2","first-page":"31","volume-title":"Proceedings of the Workshop on Mobile Edge Communications","author":"Li En","year":"2018","unstructured":"En Li, Zhi Zhou, and Xu Chen. 2018. Edge intelligence: On-demand deep learning model co-inference with device-edge synergy. In Proceedings of the Workshop on Mobile Edge Communications. 31\u201336."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CANDARW.2019.00033"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001179"},{"key":"e_1_3_2_30_2","first-page":"Lecture Notes i","volume-title":"Neural Information Processing","author":"Liu Shuanglong","year":"2017","unstructured":"Shuanglong Liu, Chao Zhang, and Jinwen Ma. 2017. CNN-LSTM neural network model for quantitative strategy analysis in stock markets. In Neural Information Processing. Lecture Notes in Computer Science, Vol. 10635. Springer, 198\u2013206."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5945"},{"key":"e_1_3_2_33_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Mirhoseini Azalia","year":"2018","unstructured":"Azalia Mirhoseini, Anna Goldie, Hieu Pham, Benoit Steiner, Quoc V. Le, and Jeff Dean. 2018. A hierarchical model for device placement. In Proceedings of the International Conference on Learning Representations. 1\u201311."},{"key":"e_1_3_2_34_2","unstructured":"Nvidia. 2021. CUDA C++ Programming Guide v11.4.0. Retrieved April 8 2022 from https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00185"},{"key":"e_1_3_2_36_2","first-page":"311","volume-title":"Proceedings of the Annual Meeting on Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting on Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00063"},{"key":"e_1_3_2_38_2","article-title":"Real-time mask detection on Google edge TPU","author":"Park Keondo","year":"2020","unstructured":"Keondo Park, Wonyoung Jang, Woochul Lee, Kisung Nam, Kihong Seong, Kyuwook Chai, and Wen-Syan Li. 2020. Real-time mask detection on Google edge TPU. arXiv:2010.04427 (2020).","journal-title":"arXiv:2010.04427"},{"key":"e_1_3_2_39_2","first-page":"Lecture Notes i","volume-title":"KI 2020: Advances in Artificial Intelligence","author":"Puchtler Pascal","year":"2020","unstructured":"Pascal Puchtler and Ren\u00e9 Peinl. 2020. Evaluation of deep learning accelerators for object detection at the edge. In KI 2020: Advances in Artificial Intelligence. Lecture Notes in Computer Science, Vol. 12325. Springer, 320\u2013326."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2897028"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001165"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-019-07788-7"},{"key":"e_1_3_2_44_2","first-page":"91","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems, Vol. 1. 91\u201399."},{"key":"e_1_3_2_45_2","first-page":"234","volume-title":"Medical Image Computing and Computer-Assisted Intervention","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention. Lecture Notes in Computer Science, Vol. 9351. Springer, 234\u2013241."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080221"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2851077"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00015"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.114"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/2786984.2786995"},{"key":"e_1_3_2_52_2","article-title":"Mutual inclusivity of the critical path and its partial schedule on heterogeneous systems","author":"Vasudevan Aravind","year":"2017","unstructured":"Aravind Vasudevan and David Gregg. 2017. Mutual inclusivity of the critical path and its partial schedule on heterogeneous systems. arXiv:1701.08800 (2017).","journal-title":"arXiv:1701.08800"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1173"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3349614.3356027"},{"key":"e_1_3_2_56_2","unstructured":"Ronald J. Williams and David Zipser. 1995. Gradient-based learning algorithms for recurrent networks and their computational complexity. In Backpropagation: Theory Architectures and Applications Yves Chauvin and David E. Rumelhart (Eds.). Lawrence Erlbaum Associates Hillsdale NJ 433\u2013486."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCIS201949030.2019.8959857"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/AICAS48895.2020.9073977"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS46320.2019.00042"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CISP-BMEI.2017.8302078"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/tip.2018.2855422"},{"key":"e_1_3_2_62_2","article-title":"An evaluation of edge TPU accelerators for convolutional neural networks","author":"Yazdanbakhsh Amir","year":"2021","unstructured":"Amir Yazdanbakhsh, Kiran Seshadri, Berkin Akin, James Laudon, and Ravi Narayanaswami. 2021. An evaluation of edge TPU accelerators for convolutional neural networks. arXiv:2102.10423 (2021).","journal-title":"arXiv:2102.10423"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018754"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401319"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00629"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.496"},{"key":"e_1_3_2_67_2","first-page":"951","volume-title":"Proceedings of the USENIX Annual Technical Conference","author":"Zhang Minjia","year":"2018","unstructured":"Minjia Zhang, Samyam Rajbhandari, Wenhan Wang, and Yuxiong He. 2018. DeepCPU: Serving RNN-based deep learning models 10x faster. In Proceedings of the USENIX Annual Technical Conference. 951\u2013965."},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2018.8573476"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/SIU.2019.8806555"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527609","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3527609","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:53Z","timestamp":1750191533000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3527609"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,25]]},"references-count":68,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3527609"],"URL":"https:\/\/doi.org\/10.1145\/3527609","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,25]]},"assertion":[{"value":"2021-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}