{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:49:26Z","timestamp":1750308566382,"version":"3.41.0"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2015,9,9]],"date-time":"2015-09-09T00:00:00Z","timestamp":1441756800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003968","name":"Iranian National Science Foundation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003968","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2015,12,8]]},"abstract":"<jats:p>In this article, a heuristic custom instruction (CI) selection algorithm is presented. The proposed algorithm, which is called OPLE for \u201cOptimization based on Partitioning and Local Exploration,\u201d uses a combination of greedy and optimal optimization methods. It searches for the near-optimal solution by reducing the search space based on partitioning the identified CI set. The partitioning of the identified set guarantees the success of the algorithm independent of the size of the identified set. First, the algorithm finds the near-optimal CIs from the candidate CIs for each part. Next, the suggested CIs from different parts are combined to determine the final selected CI set. To improve the set of the selected CIs, the solution is evolved by calling the algorithm iteratively. The efficacy of the algorithm is assessed by comparing its performance to those of optimal and nonoptimal methods. A comparative study is performed for a number of benchmarks under different area budgets and I\/O constraints. The results reveal higher speedups for the OPLE algorithm, especially for larger identified candidate sets and\/or small area budgets compared to those of the nonoptimal solutions. Compared to the nonoptimal techniques, the proposed algorithm provides 30% higher speedup improvement on average. The maximum improvement is 117%. The results also demonstrate that in many cases OPLE is able to find the optimal solution.<\/jats:p>","DOI":"10.1145\/2764458","type":"journal-article","created":{"date-parts":[[2015,9,15]],"date-time":"2015-09-15T12:09:15Z","timestamp":1442318955000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["OPLE"],"prefix":"10.1145","volume":"14","author":[{"given":"Mehdi","family":"Kamal","sequence":"first","affiliation":[{"name":"University of Tehran, Tehran, Iran"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ali","family":"Afzali-Kusha","sequence":"additional","affiliation":[{"name":"University of Tehran, Tehran, Iran"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saeed","family":"Safari","sequence":"additional","affiliation":[{"name":"University of Tehran, Tehran, Iran"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Massoud","family":"Pedram","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,9,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSI.Design.2010.68"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2008.915536"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2010.2090543"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2006.890582"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/1266366.1266657"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2008.2001863"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1176760.1176779"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.156"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/800263.809204"},{"key":"e_1_2_1_10_1","unstructured":"FreePDK. 2010. A Free OpenAccess 45nm PDK and Cell Library for university. http:\/\/www.eda.ncsu.edu."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1968502.1968509"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.848473"},{"key":"e_1_2_1_13_1","unstructured":"Gurobi. 2015. Gurobi Optimization. http:\/\/www.gurobi.com\/."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1128020.1128563"},{"volume-title":"Proceedings of the 21st IEEE International Conference on Application-Specific Systems, Architectures and Processors (ASAP\u201910)","author":"Kamal M.","key":"e_1_2_1_15_1","unstructured":"M. Kamal, N. Kazemian-Amiri, A. Kamran, S. A. Hoseini, M. Dehyadegari, and H. Noori. 2010. Dual-purpose custom instruction identification algorithm based on particle swarm optimization. In Proceedings of the 21st IEEE International Conference on Application-Specific Systems, Architectures and Processors (ASAP\u201910). 159--166."},{"volume-title":"Proceedings of the Design, Automation and Test in Europe (DATE\u201911)","author":"Kamal M.","key":"e_1_2_1_16_1","unstructured":"M. Kamal, A. Afzali-Kusha, and M. Pedram. 2011. Timing variation-aware custom instruction extension technique. In Proceedings of the Design, Automation and Test in Europe (DATE\u201911). 1517--1520."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/1326073.1326108"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/846216.846940"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2009.2017091"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/266800.266832"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIS.2009.108"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/266021.266046"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1049\/iet-cdt:20070104"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/996566.996764"},{"volume-title":"Proceedings of the Application-Specific Systems, Architectures, and Processors (ASAP\u201903)","author":"Peymandoust A.","key":"e_1_2_1_25_1","unstructured":"A. Peymandoust, L. Pozzi, P. Ienne, and G. De Micheli. 2003. Automatic instruction set extension and utilization for embedded processors. In Proceedings of the Application-Specific Systems, Architectures, and Processors (ASAP\u201903). 108--118."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/VLSI.2008.93"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2005.855950"},{"volume-title":"Proceedings of the IEEE International Workshop on Workload Characterization. 42--50","author":"Ramaswamy R.","key":"e_1_2_1_28_1","unstructured":"R. Ramaswamy and T. Wolf. 2003. PacketBench: A tool for workload characterization of network processing. In Proceedings of the IEEE International Workshop on Workload Characterization. 42--50."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2011.2173221"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10617-011-9080-8"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/321958.321963"},{"key":"e_1_2_1_32_1","unstructured":"SNU. 2015. SNU Real Time Benchmarks. http:\/\/www.cprover.org\/goto-cc\/examples\/snu.html."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1289881.1289905"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2010.2041849"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1973009.1973047"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2764458","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2764458","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:04:19Z","timestamp":1750273459000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2764458"}},"subtitle":["A Heuristic Custom Instruction Selection Algorithm Based on Partitioning and Local Exploration of Application Dataflow Graphs"],"short-title":[],"issued":{"date-parts":[[2015,9,9]]},"references-count":35,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,12,8]]}},"alternative-id":["10.1145\/2764458"],"URL":"https:\/\/doi.org\/10.1145\/2764458","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2015,9,9]]},"assertion":[{"value":"2014-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-09-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}