{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:28:28Z","timestamp":1750307308158,"version":"3.41.0"},"reference-count":17,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2011,8,1]],"date-time":"2011-08-01T00:00:00Z","timestamp":1312156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2011,8]]},"abstract":"<jats:p>In modern Web search engines, Neural Network (NN)-based learning to rank algorithms is intensively used to increase the quality of search results. LambdaRank is one such algorithm. However, it is hard to be efficiently accelerated by computer clusters or GPUs, because: (i) the cost function for the ranking problem is much more complex than that of traditional Back-Propagation(BP) NNs, and (ii) no coarse-grained parallelism exists in the algorithm. This article presents an FPGA-based accelerator solution to provide high computing performance with low power consumption. A compact deep pipeline is proposed to handle the complex computing in the batch updating. The area scales linearly with the number of hidden nodes in the algorithm. We also carefully design a data format to enable streaming consumption of the training data from the host computer. The accelerator shows up to 15.3X (with PCIe x4) and 23.9X (with PCIe x8) speedup compared with the pure software implementation on datasets from a commercial search engine.<\/jats:p>","DOI":"10.1145\/2000832.2000837","type":"journal-article","created":{"date-parts":[[2011,8,30]],"date-time":"2011-08-30T13:30:18Z","timestamp":1314711018000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["An FPGA-based accelerator for LambdaRank in Web search engines"],"prefix":"10.1145","volume":"4","author":[{"given":"Jing","family":"Yan","sequence":"first","affiliation":[{"name":"Microsoft Research Asia and Tsinghua University, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning-YI","family":"Xu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiong-FEI","family":"Cai","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Gao","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Wang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rong","family":"Luo","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng-HSIUNG","family":"Hsu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,8,22]]},"reference":[{"volume-title":"Modern Information Retrieval","author":"Baeza-Yates R.","key":"e_1_2_1_1_1","unstructured":"Baeza-Yates , R. and Ribeiro-Neto , B. 1999. Modern Information Retrieval . Addison Wesley . Baeza-Yates, R. and Ribeiro-Neto, B. 1999. Modern Information Retrieval. Addison Wesley."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7552(98)00110-X"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the NIPS Workshop on Machine Learning. 7--11","author":"Burges C.","year":"2005","unstructured":"Burges , C. 2005 . Ranking as learning structured outputs . In Proceedings of the NIPS Workshop on Machine Learning. 7--11 . Burges, C. 2005. Ranking as learning structured outputs. In Proceedings of the NIPS Workshop on Machine Learning. 7--11."},{"volume-title":"Proceedings of the Conferences on Advances in Neural Information Processing Systems (NIPS). 193--200","author":"Burges C.","key":"e_1_2_1_4_1","unstructured":"Burges , C. , Ragno , R. , and Le , Q. V . 2006. Learning to rank with nonsmooth cost functions . In Proceedings of the Conferences on Advances in Neural Information Processing Systems (NIPS). 193--200 . Burges, C., Ragno, R., and Le, Q. V. 2006. Learning to rank with nonsmooth cost functions. In Proceedings of the Conferences on Advances in Neural Information Processing Systems (NIPS). 193--200."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102363"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/345508.345545"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 25th Annual International ACM SIGIR Workshop on Mathematical\/Formal Methods in Information Retrieval (SIGIR '02)","author":"Joachims T.","year":"2002","unstructured":"Joachims , T. 2002 . Evaluating retrieval performance using clickthrough data . In Proceedings of the 25th Annual International ACM SIGIR Workshop on Mathematical\/Formal Methods in Information Retrieval (SIGIR '02) . Joachims, T. 2002. Evaluating retrieval performance using clickthrough data. In Proceedings of the 25th Annual International ACM SIGIR Workshop on Mathematical\/Formal Methods in Information Retrieval (SIGIR '02)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1328964.1328974"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2003.1220582"},{"key":"e_1_2_1_10_1","unstructured":"Microsoft. 2011 http:\/\/research.microsoft.com\/en-us\/um\/beijing\/projects\/letor\/.  Microsoft. 2011 http:\/\/research.microsoft.com\/en-us\/um\/beijing\/projects\/letor\/."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/0743-7315(92)90068-X"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Omondi A. R. and Rajapakse J. C. 2006. FPGA Implementations of Neural Networks. Birkhauser.   Omondi A. R. and Rajapakse J. C. 2006. FPGA Implementations of Neural Networks. Birkhauser.","DOI":"10.1007\/0-387-28487-7"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2006.883002"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1049\/ip-cdt:20030965"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1462586.1462588"},{"volume-title":"Proceedings of the 19th International Conference on Field Programmable Logic and Applications (FPL'09)","author":"Yan J.","key":"e_1_2_1_16_1","unstructured":"Yan , J. , Xu , N.-Y. , Cai , X. F. , Gao , R. , Wang , Y. , Luo , R. , and Hsu , F. H . 2009. Fpga based acceleration of neural network for ranking in web search engine with a streaming architecture . In Proceedings of the 19th International Conference on Field Programmable Logic and Applications (FPL'09) . Yan, J., Xu, N.-Y., Cai, X. F., Gao, R., Wang, Y., Luo, R., and Hsu, F. H. 2009. Fpga based acceleration of neural network for ranking in web search engine with a streaming architecture. In Proceedings of the 19th International Conference on Field Programmable Logic and Applications (FPL'09)."},{"volume-title":"Proceedings of the 13th International Conference on Field Programmable Logic and Applications (FPL'03)","author":"Zhu J.","key":"e_1_2_1_17_1","unstructured":"Zhu , J. and Sutton , P . 2003. Fpga implementations of neural networks\u2014A survey of a decade of progress . In Proceedings of the 13th International Conference on Field Programmable Logic and Applications (FPL'03) . 1062--1066. Zhu, J. and Sutton, P. 2003. Fpga implementations of neural networks\u2014A survey of a decade of progress. In Proceedings of the 13th International Conference on Field Programmable Logic and Applications (FPL'03). 1062--1066."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2000832.2000837","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2000832.2000837","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T11:00:03Z","timestamp":1750244403000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2000832.2000837"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,8]]},"references-count":17,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,8]]}},"alternative-id":["10.1145\/2000832.2000837"],"URL":"https:\/\/doi.org\/10.1145\/2000832.2000837","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2011,8]]},"assertion":[{"value":"2010-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-08-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}