{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T16:00:32Z","timestamp":1779292832933,"version":"3.51.4"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2018,11,16]],"date-time":"2018-11-16T00:00:00Z","timestamp":1542326400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004663","name":"Ministry of Science and Technology of Taiwan","doi-asserted-by":"crossref","award":["MOST-106-2218-E-002-040 and MOST-107-2221-E-001-002"],"award-info":[{"award-number":["MOST-106-2218-E-002-040 and MOST-107-2221-E-001-002"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2018,12,31]]},"abstract":"<jats:p>Region formation is an important step in dynamic binary translation to select hot code regions for translation and optimization. The quality of the formed regions determines the extent of optimizations and thus determines the final execution performance. Moreover, the overall performance is very sensitive to the formation overhead, because region formation can have a non-trivial cost. For addressing the dual issues of region quality and region formation overhead, this article presents a lightweight region formation method guided by processor tracing, e.g., Intel PT. We leverage the branch history information stored in the processor to reconstruct the program execution profile and effectively form high-quality regions with low cost. Furthermore, we present the designs of lightweight hardware performance monitoring sampling and the branch instruction decode cache to minimize region formation overhead. Using ARM64 to x86-64 translations, the experiment results show that our method achieves a performance speedup of up to 1.53\u00d7 (1.16\u00d7 on average) for SPEC CPU2006 benchmarks with reference inputs, compared to the well-known software-based trace formation method, Next Executing Tail (NET). The performance results of x86-64 to ARM64 translations also show a speedup of up to 1.25\u00d7 over NET for CINT2006 benchmarks with reference inputs. The comparison with a relaxed NETPlus region formation method further demonstrates that our method achieves the best performance and lowest compilation overhead.<\/jats:p>","DOI":"10.1145\/3281664","type":"journal-article","created":{"date-parts":[[2018,11,16]],"date-time":"2018-11-16T13:08:54Z","timestamp":1542373734000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Processor-Tracing Guided Region Formation in Dynamic Binary Translation"],"prefix":"10.1145","volume":"15","author":[{"given":"Ding-Yong","family":"Hong","sequence":"first","affiliation":[{"name":"Institute of Information Science, Academia Sinica, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan-Jan","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Academia Sinica, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu-Ping","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sheng-Yu","family":"Fu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei-Chung","family":"Hsu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Information Engineering, National Taiwan University, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,11,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1147\/sj.391.0211"},{"key":"e_1_2_1_2_1","unstructured":"ARM. 2012. CoreSight Components Technical Reference Manual. ARM.  ARM. 2012. CoreSight Components Technical Reference Manual. ARM."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/378795.378832"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/349299.349303"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/183432.183527"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 29th Annual ACM\/IEEE International Symposium on Microarchitecture. 46--57","author":"Ball Thomas"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956550"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/1247360.1247401"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993508"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772954.1772959"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/776261.776290"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1297027.1297068"},{"key":"e_1_2_1_13_1","unstructured":"J. G. Castanos H. Hayashizaki H. Inoue M. J. Serrano and P. Wu. 2014. Adaptive next-executing-cycle trace selection for trace-driven code optimizers. http:\/\/www.google.com\/patents\/US8756581 US Patent 8 756 581.  J. G. Castanos H. Hayashizaki H. Inoue M. J. Serrano and P. Wu. 2014. Adaptive next-executing-cycle trace selection for trace-driven code optimizers. http:\/\/www.google.com\/patents\/US8756581 US Patent 8 756 581."},{"key":"e_1_2_1_14_1","volume-title":"ACM Workshop on Feedback-Directed and Dynamic Optimization. 81--90","author":"Chen Wen-Ke"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3062341.3062371"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the ASPLOS Workshop on Runtime Environments\/Systems, Layering, and Virtualized Environments.","author":"Derek"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/776261.776263"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/378993.379241"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542476.1542528"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/800230.806987"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950365.1950412"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2005.22"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 4th ACM Workshop on Feedback-Directed and Dynamic Optimization.","author":"Hirzel Martin","year":"2001"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259030"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451512.2451519"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2190025.2190071"},{"key":"e_1_2_1_27_1","unstructured":"Intel Corporation 2018. Intel(R) 64 and IA-32 Architectures Software Developer\u2019s Manual: Volume 3. Intel Corporation.  Intel Corporation 2018. Intel(R) 64 and IA-32 Architectures Software Developer\u2019s Manual: Volume 3. Intel Corporation."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-92990-1_6"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_2_1_30_1","unstructured":"Linaro. 2018. OpenCSD library. Retrieved from https:\/\/github.com\/Linaro\/OpenCSD.  Linaro. 2018. OpenCSD library. Retrieved from https:\/\/github.com\/Linaro\/OpenCSD."},{"key":"e_1_2_1_31_1","unstructured":"Linaro ToolChain. 2017. Linaro ARM GCC toolchain. Retrieved from http:\/\/www.linaro.org\/downloads\/.  Linaro ToolChain. 2017. Linaro ARM GCC toolchain. Retrieved from http:\/\/www.linaro.org\/downloads\/."},{"key":"e_1_2_1_32_1","first-page":"1","article-title":"Design and implementation of a lightweight dynamic optimization system","volume":"6","author":"Lu Jiwei","year":"2004","journal-title":"J. Instruct.-Level Parall."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065010.1065034"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250734.1250746"},{"key":"e_1_2_1_35_1","unstructured":"Andreas Neustifter. 2010. Efficient Profiling in the LLVM Compiler. Master\u2019s thesis. Vienna University of Technology.  Andreas Neustifter. 2010. Efficient Profiling in the LLVM Compiler. Master\u2019s thesis. Vienna University of Technology."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2006.16"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/566172.566186"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/2392163.2392166"},{"key":"e_1_2_1_40_1","unstructured":"C. Wang B. Zheng H. S. Kim M. Breternitz and Y. Wu. 2010. Two-pass MRET trace selection for dynamic optimization. http:\/\/www.google.com\/patents\/US7694281 US Patent 7 694 281.  C. Wang B. Zheng H. S. Kim M. Breternitz and Y. Wu. 2010. Two-pass MRET trace selection for dynamic optimization. http:\/\/www.google.com\/patents\/US7694281 US Patent 7 694 281."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/337449.337483"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2048066.2048127"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3281664","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3281664","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:02:10Z","timestamp":1750208530000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3281664"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,16]]},"references-count":41,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,12,31]]}},"alternative-id":["10.1145\/3281664"],"URL":"https:\/\/doi.org\/10.1145\/3281664","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,11,16]]},"assertion":[{"value":"2018-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-11-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}