{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T15:38:44Z","timestamp":1762875524898,"version":"3.41.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2013,5,1]],"date-time":"2013-05-01T00:00:00Z","timestamp":1367366400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61272131, 61202053"],"award-info":[{"award-number":["61272131, 61202053"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["BH0110000014"],"award-info":[{"award-number":["BH0110000014"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004608","name":"National Natural Science Foundation of Jiangsu Province","doi-asserted-by":"publisher","award":["SBK201240198"],"award-info":[{"award-number":["SBK201240198"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2013,5]]},"abstract":"<jats:p>This article presents MP-Tomasulo, a dependency-aware automatic parallel task execution engine for sequential programs. Applying the instruction-level Tomasulo algorithm to MPSoC environments, MP-Tomasulo detects and eliminates Write-After-Write (WAW) and Write-After-Read (WAR) inter-task dependencies in the dataflow execution, therefore to operate out-of-order task execution on heterogeneous units. We implemented the prototype system within a single FPGA. Experimental results on EEMBC applications demonstrate that MP-Tomasulo can execute the tasks out-of-order to achieve as high as 93.6% to 97.6% of ideal peak speedup. A comparative study against a state-of-the-art dataflow execution scheme is illustrated with a classic JPEG application. The promising results show MP-Tomasulo enables programmers to uncover more task-level parallelism on heterogeneous systems, as well as to ease the burden of programmers.<\/jats:p>","DOI":"10.1145\/2459316.2459320","type":"journal-article","created":{"date-parts":[[2013,5,3]],"date-time":"2013-05-03T12:30:25Z","timestamp":1367584225000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":24,"title":["MP-Tomasulo"],"prefix":"10.1145","volume":"10","author":[{"given":"Chao","family":"Wang","sequence":"first","affiliation":[{"name":"University of Science and Technology of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Li","sequence":"additional","affiliation":[{"name":"Suzhou Institute for University of Science and Technology of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junneng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Suzhou Institute for University of Science and Technology of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuehai","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoning","family":"Nie","sequence":"additional","affiliation":[{"name":"Intel"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,5]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504190"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1188455.1188546"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-007-0138-6"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/209937.209958"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941487.1941507"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.17"},{"key":"e_1_2_1_7_1","unstructured":"EEMBC. 2010. The embedded microprocessor benchmark consortium. http:\/\/www.eembc.org\/.  EEMBC. 2010. The embedded microprocessor benchmark consortium. http:\/\/www.eembc.org\/."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.13"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2002.804107"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155628"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.848474"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/1148882.1148891"},{"volume-title":"Proceedings of the Workshop on Application Specific Processors.","author":"Karim F.","key":"e_1_2_1_13_1","unstructured":"Karim , F. , Mellan , A. , Aydonat , U. , Abdelrahman , T. S. , Stramm , B. , and Nguyen , A . 2003. The Hyperprocessor: A template system-on-chip architecture for embedded multimedia applications . In Proceedings of the Workshop on Application Specific Processors. Karim, F., Mellan, A., Aydonat, U., Abdelrahman, T. S., Stramm, B., and Nguyen, A. 2003. The Hyperprocessor: A template system-on-chip architecture for embedded multimedia applications. In Proceedings of the Workshop on Application Specific Processors."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2004.1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.27"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250662.1250683"},{"volume-title":"Proceedings of the 12th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201904)","author":"Kuzmanov G.","key":"e_1_2_1_17_1","unstructured":"Kuzmanov , G. , Gaydadjiev , G. , and Vassiliadis , S . 2004. The MOLEN processor prototype . In Proceedings of the 12th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201904) . Kuzmanov, G., Gaydadjiev, G., and Vassiliadis, S. 2004. The MOLEN processor prototype. In Proceedings of the 12th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201904)."},{"volume-title":"Proceedings of the 46th Design Automation Conference (DAC\u201909)","author":"Limberg T.","key":"e_1_2_1_18_1","unstructured":"Limberg , T. , Winter , M. , Bimberg , M. , Klemm , R. , and Fettweis , G . 2009. A heterogeneous MPSoC with hardware supported dynamic task scheduling for software defined radio . In Proceedings of the 46th Design Automation Conference (DAC\u201909) . Limberg, T., Winter, M., Bimberg, M., Klemm, R., and Fettweis, G. 2009. A heterogeneous MPSoC with hardware supported dynamic task scheduling for software defined radio. In Proceedings of the 46th Design Automation Conference (DAC\u201909)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/18927.18916"},{"volume-title":"Proceedings of the 30th IEEE\/ACM International Symposium on Microarchitecture (MICRO). 138--148","author":"Rotenberg E.","key":"e_1_2_1_20_1","unstructured":"Rotenberg , E. , Jacobson , Q. , Sazeides , Y. , and Smith , J. E . 1997. Trace processors . In Proceedings of the 30th IEEE\/ACM International Symposium on Microarchitecture (MICRO). 138--148 . Rotenberg, E., Jacobson, Q., Sazeides, Y., and Smith, J. E. 1997. Trace processors. In Proceedings of the 30th IEEE\/ACM International Symposium on Microarchitecture (MICRO). 138--148."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.9"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2006.19"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.224451"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555815.1555763"},{"volume-title":"WaveScalar. In Proceedings of the 36th International Symposium on Microarchitecture.","author":"Swanson S.","key":"e_1_2_1_25_1","unstructured":"Swanson , S. , Michelson , K. , Schwerin , A. , and Oskin , M . 2003 . WaveScalar. In Proceedings of the 36th International Symposium on Microarchitecture. Swanson, S., Michelson, K., Schwerin, A., and Oskin, M. 2003. WaveScalar. In Proceedings of the 36th International Symposium on Microarchitecture."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2002.997877"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.111.0025"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2012.62"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/SCC.2011.26"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2007.39"},{"key":"e_1_2_1_31_1","unstructured":"Xilinx. 2009. Fast simplex link (FSL) specification. http:\/\/www.xilinx.com\/products\/ipcenter\/FSL.htm.  Xilinx. 2009. Fast simplex link (FSL) specification. http:\/\/www.xilinx.com\/products\/ipcenter\/FSL.htm."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2459316.2459320","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2459316.2459320","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:18:38Z","timestamp":1750234718000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2459316.2459320"}},"subtitle":["A Dependency-Aware Automatic Parallel Execution Engine for Sequential Programs"],"short-title":[],"issued":{"date-parts":[[2013,5]]},"references-count":31,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2013,5]]}},"alternative-id":["10.1145\/2459316.2459320"],"URL":"https:\/\/doi.org\/10.1145\/2459316.2459320","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2013,5]]},"assertion":[{"value":"2012-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-05-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}