{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T00:29:33Z","timestamp":1782174573720,"version":"3.54.5"},"reference-count":26,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2016,6,30]],"date-time":"2016-06-30T00:00:00Z","timestamp":1467244800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2018,5]]},"abstract":"<jats:p>The well-known Smith\u2013Waterman algorithm is a high-sensitivity method for local sequence alignment. Unfortunately, the Smith\u2013Waterman algorithm has quadratic time complexity, which makes it computationally demanding for large protein databases. In this paper, we present OSWALD, a portable, fully functional and general implementation to accelerate Smith\u2013Waterman database searches in heterogeneous platforms based on Altera\u2019s FPGA. OSWALD exploits OpenMP multithreading and SIMD computing through SSE and AVX2 extensions on the host while taking advantage of pipeline and vectorial parallelism by way of OpenCL on the FPGAs. Performance evaluations on two different heterogeneous architectures with real amino acid datasets show that OSWALD is competitive in comparison with other top-performing Smith\u2013Waterman implementations, attaining up to 442 GCUPS peak with the best GCUPS\/watts ratio.<\/jats:p>","DOI":"10.1177\/1094342016654215","type":"journal-article","created":{"date-parts":[[2016,7,1]],"date-time":"2016-07-01T20:38:35Z","timestamp":1467405515000},"page":"337-350","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":36,"title":["OSWALD"],"prefix":"10.1177","volume":"32","author":[{"given":"Enzo","family":"Rucci","sequence":"first","affiliation":[{"name":"Instituto de Investigacion en Informatica LIDI, Universidad Nacional de La Plata, Argentina"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Carlos","family":"Garcia","sequence":"additional","affiliation":[{"name":"Depto. Arquitectura de Computadores y Automatica, Universidad Complutense de Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guillermo","family":"Botella","sequence":"additional","affiliation":[{"name":"Depto. Arquitectura de Computadores y Automatica, Universidad Complutense de Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Armando E","family":"De Giusti","sequence":"additional","affiliation":[{"name":"Instituto de Investigacion en Informatica LIDI, Universidad Nacional de La Plata, Argentina"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marcelo","family":"Naiouf","sequence":"additional","affiliation":[{"name":"Instituto de Investigacion en Informatica LIDI, Universidad Nacional de La Plata, Argentina"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manuel","family":"Prieto-Matias","sequence":"additional","affiliation":[{"name":"Depto. Arquitectura de Computadores y Automatica, Universidad Complutense de Madrid, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2016,6,30]]},"reference":[{"key":"bibr1-1094342016654215","volume-title":"Altera SDK for OpenCL Programming Guide","author":"Altera","year":"2014"},{"key":"bibr2-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-2836(05)80360-2"},{"key":"bibr3-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30117-2_5"},{"key":"bibr4-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btl582"},{"key":"bibr5-1094342016654215","unstructured":"Farrar MS (2008) Optimizing Smith\u2013Waterman for the cell broad-band engine, http:\/\/cudasw.sourceforge.net\/sw-cellbe.pdf."},{"key":"bibr6-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1016\/0022-2836(82)90398-9"},{"key":"bibr7-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-014-0266-8"},{"key":"bibr8-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1109\/AHS.2011.5963957"},{"key":"bibr9-1094342016654215","volume-title":"The OpenCL Specification","author":"Howes L","year":"2014"},{"key":"bibr10-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-8-185"},{"key":"bibr11-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1126\/science.2983426"},{"key":"bibr12-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP.2014.6868657"},{"key":"bibr13-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1756-0500-2-73"},{"key":"bibr14-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1756-0500-3-93"},{"key":"bibr15-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2014.6968772"},{"key":"bibr16-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-14-117"},{"key":"bibr17-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-11-S12-S3"},{"key":"bibr18-1094342016654215","first-page":"xli","volume-title":"High Performance Parallelism Pearls, Volume 1: Multicore and Many-Core Programming Approaches","author":"Reinders J","year":"2014"},{"key":"bibr19-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-12-221"},{"key":"bibr20-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2014.6968784"},{"key":"bibr21-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1109\/Trustcom.2015.634"},{"key":"bibr22-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3598"},{"key":"bibr23-1094342016654215","first-page":"1","volume-title":"IEEE High Performance Extreme Computing Conference","author":"Settle SO","year":"2014"},{"key":"bibr24-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1016\/0022-2836(81)90087-5"},{"key":"bibr25-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1145\/611817.611845"},{"key":"bibr26-1094342016654215","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-19475-7_20"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342016654215","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342016654215","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342016654215","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:32Z","timestamp":1777450532000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342016654215"}},"subtitle":["<i>O<\/i>\n                    penCL\n                    <i>S<\/i>\n                    mith\u2013\n                    <i>W<\/i>\n                    aterman on\n                    <i>A<\/i>\n                    ltera\u2019s FPGA for\n                    <i>L<\/i>\n                    arge Protein\n                    <i>D<\/i>\n                    atabases"],"short-title":[],"issued":{"date-parts":[[2016,6,30]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,5]]}},"alternative-id":["10.1177\/1094342016654215"],"URL":"https:\/\/doi.org\/10.1177\/1094342016654215","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,6,30]]}}}