{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,27]],"date-time":"2025-09-27T13:48:07Z","timestamp":1758980887174,"version":"3.41.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2017,12,15]],"date-time":"2017-12-15T00:00:00Z","timestamp":1513296000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Australian Research Councils Linkage Projects","award":["LP130101034"],"award-info":[{"award-number":["LP130101034"]}]},{"name":"Zomojo Pty Ltd"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2017,12,31]]},"abstract":"<jats:p>Kernel adaptive filters (KAFs) are online machine learning algorithms which are amenable to highly efficient streaming implementations. They require only a single pass through the data and can act as universal approximators, i.e. approximate any continuous function with arbitrary accuracy. KAFs are members of a family of kernel methods which apply an implicit non-linear mapping of input data to a high dimensional feature space, permitting learning algorithms to be expressed entirely as inner products. Such an approach avoids explicit projection into the feature space, enabling computational efficiency. In this paper, we propose the first fully pipelined implementation of the kernel normalised least mean squares algorithm for regression. Independent training tasks necessary for hyperparameter optimisation fill pipeline stages, so no stall cycles to resolve dependencies are required. Together with other optimisations to reduce resource utilisation and latency, our core achieves 161 GFLOPS on a Virtex 7 XC7VX485T FPGA for a floating point implementation and 211 GOPS for fixed point. Our PCI Express based floating-point system implementation achieves 80% of the core\u2019s speed, this being a speedup of 10\u00d7 over an optimised implementation on a desktop processor and 2.66\u00d7 over a GPU.<\/jats:p>","DOI":"10.1145\/3106744","type":"journal-article","created":{"date-parts":[[2017,12,18]],"date-time":"2017-12-18T13:22:44Z","timestamp":1513603364000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["FPGA Implementations of Kernel Normalised Least Mean Squares Processors"],"prefix":"10.1145","volume":"10","author":[{"given":"Nicholas J.","family":"Fraser","sequence":"first","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junkyu","family":"Lee","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Duncan J. M.","family":"Moss","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Julian","family":"Faraone","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stephen","family":"Tridgell","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Craig T.","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philip H. W.","family":"Leong","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,12,15]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1987535.1987578"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0218126611007244"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2188385.2188395"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2011.2178446"},{"volume-title":"The 2013 International Joint Conference on Neural Networks (IJCNN\u201913)","author":"Chen Badong","key":"e_1_2_1_5_1"},{"volume-title":"The XI Metaheuristics International Conference (MIC\u201915)","year":"2015","author":"Claesen Marc","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2011.5771385"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2005.1568520"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.661345"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2015.7293952"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2012.44"},{"volume-title":"International Conference on Field Programmable Logic and Applications (FPL\u201907)","author":"Jamro E.","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2004.830991"},{"volume-title":"Proceedings of the 16th Annual Conference on Neural Information Processing Systems. 609--616","year":"2003","author":"Lawrence Neil","key":"e_1_2_1_14_1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"e_1_2_1_16_1","unstructured":"Weifeng Liu Jos\u00e9 C. Pr\u00edncipe and Simon Haykin. 2011. Kernel Adaptive Filtering: A Comprehensive Introduction. Vol. 57. John Wiley 8 Sons Hoboken NJ. Weifeng Liu Jos\u00e9 C. Pr\u00edncipe and Simon Haykin. 2011. Kernel Adaptive Filtering: A Comprehensive Introduction. Vol. 57. John Wiley 8 Sons Hoboken NJ."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/29.31293"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2133382.2133388"},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Yeyong Pang Shaojun Wang Yu Peng Nicholas J. Fraser and Philip H. W. Leong. 2013. A low latency kernel recursive least squares processor using FPGA technology. In FPT. 144--151. Yeyong Pang Shaojun Wang Yu Peng Nicholas J. Fraser and Philip H. W. Leong. 2013. A low latency kernel recursive least squares processor using FPGA technology. In FPT. 144--151.","DOI":"10.1109\/FPT.2013.6718345"},{"volume-title":"International Conference on ICECE Technology (FPT\u201908)","author":"Papadonikolakis M.","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1000579"},{"key":"e_1_2_1_22_1","unstructured":"John Platt and others. 1998. Sequential minimal optimization: A fast algorithm for training support vector machines. https:\/\/www.microsoft.com\/en-us\/research\/publication\/sequential-minimal-optimization-a-fast-algorithm-for-training-support-vector-machines\/. John Platt and others. 1998. Sequential minimal optimization: A fast algorithm for training support vector machines. https:\/\/www.microsoft.com\/en-us\/research\/publication\/sequential-minimal-optimization-a-fast-algorithm-for-training-support-vector-machines\/."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/97.475854"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1508128.1508198"},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Carl E. Rasmussen and Christoper K. I. Williams. 2006. Gaussian Processes for Machine Learning. MIT Press Cambridge MA. Carl E. Rasmussen and Christoper K. I. Williams. 2006. Gaussian Processes for Machine Learning. MIT Press Cambridge MA.","DOI":"10.7551\/mitpress\/3206.001.0001"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2014.6889689"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2008.2009895"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Bernhard Scholkopf and Alexander J. Smola. 2001. Learning with Kernels: Support Vector Machines Regularization Optimization and Beyond. MIT Press Cambridge MA. Bernhard Scholkopf and Alexander J. Smola. 2001. Learning with Kernels: Support Vector Machines Regularization Optimization and Beyond. MIT Press Cambridge MA.","DOI":"10.7551\/mitpress\/4175.001.0001"},{"key":"e_1_2_1_29_1","unstructured":"Matthias Seeger. 2000. Relationships between Gaussian processes support vector machines and smoothing splines. Machine Learning). Matthias Seeger. 2000. Relationships between Gaussian processes support vector machines and smoothing splines. Machine Learning)."},{"volume-title":"Proceedings of the International Conference on Field Programmable Technology (FPT\u201915)","author":"Tridgell Stephen","key":"e_1_2_1_30_1"},{"key":"e_1_2_1_31_1","unstructured":"Steven Van Vaerenbergh. 2012. Kernel Methods Toolbox KAFBOX: a Matlab benchmarking toolbox for kernel adaptive filtering. Retrieved October 1 2017 at http:\/\/sourceforge.net\/p\/kafbox. Steven Van Vaerenbergh. 2012. Kernel Methods Toolbox KAFBOX: a Matlab benchmarking toolbox for kernel adaptive filtering. Retrieved October 1 2017 at http:\/\/sourceforge.net\/p\/kafbox."},{"volume":"5","volume-title":"2006 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201906)","author":"Vaerenbergh S. Van","key":"e_1_2_1_32_1"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.v35:2"},{"volume-title":"IRE WESCON Convention Record. 96--104","author":"Widrow B.","key":"e_1_2_1_34_1"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-78610-8_28"},{"key":"e_1_2_1_36_1","unstructured":"James H. Wilkinson. 1994. Rounding Errors in Algebraic Processes. Dover Publications Incorporated Mineola NY. James H. Wilkinson. 1994. Rounding Errors in Algebraic Processes. Dover Publications Incorporated Mineola NY."},{"key":"e_1_2_1_37_1","unstructured":"Zhang Xianyi Wang Qian and Zaheer Chothia. 2014. Openblas. Retrieved October 1 2017 from http:\/\/xianyi.github.io\/OpenBLAS. Zhang Xianyi Wang Qian and Zaheer Chothia. 2014. Openblas. Retrieved October 1 2017 from http:\/\/xianyi.github.io\/OpenBLAS."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VLSI.0000047275.54691.be"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2012.2200889"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3106744","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3106744","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:20Z","timestamp":1750217420000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3106744"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,12,15]]},"references-count":39,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,12,31]]}},"alternative-id":["10.1145\/3106744"],"URL":"https:\/\/doi.org\/10.1145\/3106744","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2017,12,15]]},"assertion":[{"value":"2016-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-12-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}