{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T16:30:09Z","timestamp":1779381009096,"version":"3.53.1"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2009,10,1]],"date-time":"2009-10-01T00:00:00Z","timestamp":1254355200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000149","name":"Division of Engineering Education and Centers","doi-asserted-by":"publisher","award":["EEC-9986821"],"award-info":[{"award-number":["EEC-9986821"]}],"id":[{"id":"10.13039\/100000149","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2009,10]]},"abstract":"<jats:p>We have implemented a two-dimensional systolic array QR decomposition on a Xilinx Virtex5 FPGA using the Givens rotation algorithm. QR decomposition is a key step in many DSP applications including sonar beamforming, channel equalization, and 3G wireless communication. Compared to previous work that implements Givens rotations using a one-dimensional systolic array, our implementation uses a truly two-dimensional systolic array architecture. As a result, latency scales well for larger matrices. In addition, prior work avoids divide and square root operations in the Givens rotation algorithm by using special operations such as CORDIC or special number systems such as the logarithmic number system (LNS). In contrast, our design uses straightforward floating-point divide and square root implementations, which makes it easier to be used within a larger system. In our design, the input matrix size can be configured at compile time to many different sizes, making it easily scalable to future large FPGAs or over multiple FPGAs. The QR module is fully pipelined with a throughput of over 130MHz for the IEEE single-precision floating-point format. The peak performance for a 12 \u00d7 12 input matrix is approximately 35 GFLOPs.<\/jats:p>","DOI":"10.1145\/1596532.1596535","type":"journal-article","created":{"date-parts":[[2009,10,27]],"date-time":"2009-10-27T13:28:14Z","timestamp":1256650094000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":23,"title":["A truly two-dimensional systolic array FPGA implementation of QR decomposition"],"prefix":"10.1145","volume":"9","author":[{"given":"Xiaojun","family":"Wang","sequence":"first","affiliation":[{"name":"Airvana"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Miriam","family":"Leeser","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2009,10,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Accel-dsp. AccelWare DSP IP Toolkits. http:\/\/www.xilinx.com\/ise\/dspdesignprod\/acceldsp\/accelware\/.  Accel-dsp. AccelWare DSP IP Toolkits. http:\/\/www.xilinx.com\/ise\/dspdesignprod\/acceldsp\/accelware\/."},{"key":"e_1_2_1_2_1","unstructured":"Altera-cordic. ALTERA CORDIC reference design. http:\/\/www.altera.com\/literature\/an\/an263.pdf.  Altera-cordic. ALTERA CORDIC reference design. http:\/\/www.altera.com\/literature\/an\/an263.pdf."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 12th IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM'04)","author":"Boppana D."},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Dohler R. 1991. Squared Givens rotations. IMA J. Numer. Anal. 1--5.  Dohler R. 1991. Squared Givens rotations. IMA J. Numer. Anal. 1--5.","DOI":"10.1093\/imanum\/11.1.1"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.863031"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the Software Defined Radio Technical Conference. ACM","author":"Fitton M. P."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/4.4.332"},{"key":"e_1_2_1_8_1","unstructured":"Gay M. and Phillips I. 2005. Real-time adaptive beam forming: FPGA implementation using QR decomposition. J. Electron. Defense.  Gay M. and Phillips I. 2005. Real-time adaptive beam forming: FPGA implementation using QR decomposition. J. Electron. Defense."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of Real-Time Signal Processing IV, SPIE 298","author":"Gentleman M. W."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1137\/0106004"},{"key":"e_1_2_1_11_1","unstructured":"Golub G. H. and Van Loan C. F. 1996. Matrix Computations. Johns Hopkins University Press Baltimore MD.  Golub G. H. and Van Loan C. F. 1996. Matrix Computations. Johns Hopkins University Press Baltimore MD."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 12th European Signal Processing Conference. ACM","author":"Hermanek J. A."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/320941.320947"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646146"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 39rd Asilomar Conference on Signals, Systems and Computers. IEEE","author":"Karkooti M."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/55364.55427"},{"key":"e_1_2_1_17_1","volume-title":"Lecture Notes in Computer Science","volume":"2438","author":"Matousek"},{"key":"e_1_2_1_18_1","unstructured":"Moon T. K. and Stirling W. C. 2000. Mathematical Methods and Algorithms for Signal Processing. Prentice Hall International Upper Saddle River NJ 285--300.  Moon T. K. and Stirling W. C. 2000. Mathematical Methods and Algorithms for Signal Processing. Prentice Hall International Upper Saddle River NJ 285--300."},{"key":"e_1_2_1_19_1","unstructured":"Northeastern University Floating-Point Library. Northeastern University variable precision floating-point modules. http:\/\/www.ece.neu.edu\/groups\/rcl\/projects\/floatingpoint\/.  Northeastern University Floating-Point Library. Northeastern University variable precision floating-point modules. http:\/\/www.ece.neu.edu\/groups\/rcl\/projects\/floatingpoint\/."},{"key":"e_1_2_1_20_1","unstructured":"Quixilica-qr. Quixilica Floating-Point QR Processor Core. http:\/\/www.eonic.co.kr\/data\/datasheet\/transtech\/FPGA\/qx qr.pdf.  Quixilica-qr. Quixilica Floating-Point QR Processor Core. http:\/\/www.eonic.co.kr\/data\/datasheet\/transtech\/FPGA\/qx qr.pdf."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-1966-0192673-4"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the 14th International Conference on Field-Programmable Logic and Applications (FPL'04)","author":"Schier J."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the 6th Baiona Workshop on Signal Processing in Communications. ACM","author":"Schier J."},{"key":"e_1_2_1_24_1","volume-title":"Lecture Notes in Computer Science","volume":"2328","author":"Sergyienko A."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-006-0004-y"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEC.1959.5222693"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 33rd Asilomar Conference on Signals, Systems and Computers.","volume":"2","author":"Walke R. L., M."},{"key":"e_1_2_1_28_1","first-page":"300","article-title":"20 GFLOPS QR processor on a Xilinx Virtex-E FPGA. In Proceedings of the Advanced Signal Processing Algorithms, Architectures, and Implementations X","volume":"4116","author":"Walke R. L., M.","year":"2000","journal-title":"SPIE."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2006.21"},{"key":"e_1_2_1_30_1","unstructured":"Xilinx. XILINX LogiCORE floating-point operator v3.0. http:\/\/www.xilinx.com\/bvdocs\/ipcenter\/data sheet\/floatingpointds335.pdf.  Xilinx. XILINX LogiCORE floating-point operator v3.0. http:\/\/www.xilinx.com\/bvdocs\/ipcenter\/data sheet\/floatingpointds335.pdf."},{"key":"e_1_2_1_31_1","unstructured":"Xilinx-cordic. XILINX LogiCORE CORDIC v3.0. http:\/\/www.xilinx.com\/bvdocs\/ipcenter\/datasheet\/cordic.pdf.  Xilinx-cordic. XILINX LogiCORE CORDIC v3.0. http:\/\/www.xilinx.com\/bvdocs\/ipcenter\/datasheet\/cordic.pdf."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1596532.1596535","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1596532.1596535","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T12:23:32Z","timestamp":1750249412000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1596532.1596535"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,10]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2009,10]]}},"alternative-id":["10.1145\/1596532.1596535"],"URL":"https:\/\/doi.org\/10.1145\/1596532.1596535","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,10]]},"assertion":[{"value":"2008-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-10-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}