{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,3,29]],"date-time":"2022-03-29T23:37:03Z","timestamp":1648597023602},"reference-count":9,"publisher":"World Scientific Pub Co Pte Lt","issue":"08","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J CIRCUIT SYST COMP"],"published-print":{"date-parts":[[2015,9]]},"abstract":"<jats:p> This paper proposes a simultaneous multithreaded matrix processor (SMMP) to improve the performance of data-parallel applications by exploiting instruction-level parallelism (ILP) data-level parallelism (DLP) and thread-level parallelism (TLP). In SMMP, the well-known five-stage pipeline (baseline scalar processor) is extended to execute multi-scalar\/vector\/matrix instructions on unified parallel execution datapaths. SMMP can issue four scalar instructions from two threads each cycle or four vector\/matrix operations from one thread, where the execution of vector\/matrix instructions in threads is done in round-robin fashion. Moreover, this paper presents the implementation of our proposed SMMP using VHDL targeting FPGA Virtex-6. In addition, the performance of SMMP is evaluated on some kernels from the basic linear algebra subprograms (BLAS). Our results show that, the hardware complexity of SMMP is 5.68 times higher than the baseline scalar processor. However, speedups of 4.9, 6.09, 6.98, 8.2, 8.25, 8.72, 9.36, 11.84 and 21.57 are achieved on BLAS kernels of applying Givens rotation, scalar times vector plus another, vector addition, vector scaling, setting up Givens rotation, dot-product, matrix\u2013vector multiplication, Euclidean length, and matrix\u2013matrix multiplications, respectively. The average speedup over the baseline is 9.55 and the average speedup over complexity is 1.68. Comparing with Xilinx MicroBlaze, the complexity of SMMP is 6.36 times higher, however, its speedup ranges from 6.87 to 12.07 on vector\/matrix kernels, which is 9.46 in average. <\/jats:p>","DOI":"10.1142\/s0218126615501145","type":"journal-article","created":{"date-parts":[[2015,5,29]],"date-time":"2015-05-29T05:50:54Z","timestamp":1432878654000},"page":"1550114","source":"Crossref","is-referenced-by-count":0,"title":["Simultaneous Multithreaded Matrix Processor"],"prefix":"10.1142","volume":"24","author":[{"given":"Mostafa I.","family":"Soliman","sequence":"first","affiliation":[{"name":"Computer Science and Information Department, Community College, Taibah University, Al-Madinah Al-Munawwarah 2898, Saudi Arabia"},{"name":"Computer and System Section, Electrical Engineering Department, Faculty of Engineering, Aswan University, Aswan 81542, Egypt"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elsayed A.","family":"Elsayed","sequence":"additional","affiliation":[{"name":"Computer and System Section, Electrical Engineering Department, Faculty of Engineering, Aswan University, Aswan 81542, Egypt"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2015,8,12]]},"reference":[{"key":"rf1","volume-title":"Computer Architecture: A Quantitative Approach","author":"Hennessay J.","year":"2011"},{"key":"rf3","doi-asserted-by":"publisher","DOI":"10.1109\/40.621209"},{"key":"rf4","doi-asserted-by":"publisher","DOI":"10.1109\/5.476078"},{"key":"rf6","volume-title":"Embedded Computing: A VLIW Approach to Architecture, Compilers and Tools","author":"Fisher J.","year":"2004"},{"key":"rf9","doi-asserted-by":"publisher","DOI":"10.1145\/2518037.2491464"},{"key":"rf10","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2004.10.002"},{"key":"rf18","first-page":"224","volume":"2","author":"Kumar D.","year":"2013","journal-title":"Int. J. Eng. Res."},{"key":"rf19","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2013.03.004"},{"key":"rf20","volume-title":"Computer Organization and Design: The Hardware\/Software Interface","author":"Patterson D.","year":"2013"}],"container-title":["Journal of Circuits, Systems and Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218126615501145","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T12:14:51Z","timestamp":1565093691000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0218126615501145"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,8,12]]},"references-count":9,"journal-issue":{"issue":"08","published-online":{"date-parts":[[2015,8,12]]},"published-print":{"date-parts":[[2015,9]]}},"alternative-id":["10.1142\/S0218126615501145"],"URL":"https:\/\/doi.org\/10.1142\/s0218126615501145","relation":{},"ISSN":["0218-1266","1793-6454"],"issn-type":[{"value":"0218-1266","type":"print"},{"value":"1793-6454","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,8,12]]}}}