{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T10:46:47Z","timestamp":1781779607057,"version":"3.54.5"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,6,26]],"date-time":"2019-06-26T00:00:00Z","timestamp":1561507200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Basic Research Program","award":["2015CB358700"],"award-info":[{"award-number":["2015CB358700"]}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61622208, 61732008, 61472206"],"award-info":[{"award-number":["61622208, 61732008, 61472206"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2019,7,31]]},"abstract":"<jats:p>Relevance estimation is among the most important tasks in the ranking of search results. Current methodologies mainly concentrate on text matching, link analysis, and user behavior models. However, users judge the relevance of search results directly from Search Engine Result Pages (SERPs), which provide valuable signals for reranking. In this article, we propose two different approaches to aggregate the visual, structure, as well as textual information sources of search results in relevance estimation. The first one is a late-fusion framework named Joint Relevance Estimation model (JRE). JRE estimates the relevance independently from screenshots, textual contents, and HTML source codes of search results and jointly makes the final decision through an inter-modality attention mechanism. The second one is an early-fusion framework named Tree-based Deep Neural Network (TreeNN), which embeds the texts and images into the HTML parse tree through a recursive process. To evaluate the performance of the proposed models, we construct a large-scale practical Search Result Relevance (SRR) dataset that consists of multiple information sources and relevance labels of over 60,000 search results. Experimental results show that the proposed two models achieve better performance than state-of-the-art ranking solutions as well as the original rankings of commercial search engines.<\/jats:p>","DOI":"10.1145\/3329188","type":"journal-article","created":{"date-parts":[[2019,6,26]],"date-time":"2019-06-26T12:36:24Z","timestamp":1561552584000},"page":"1-38","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Search Result Reranking with Visual and Structure Information Sources"],"prefix":"10.1145","volume":"37","author":[{"given":"Yiqun","family":"Liu","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junqi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaxin","family":"Mao","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shaoping","family":"Ma","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Tian","sequence":"additional","affiliation":[{"name":"Huawei Noah's Ark Lab, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanxiong","family":"Lu","sequence":"additional","affiliation":[{"name":"Tencent, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Leyu","family":"Lin","sequence":"additional","affiliation":[{"name":"Tencent, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,6,26]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2187980.2188075"},{"key":"e_1_2_1_2_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473.  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/1766091.1766143"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526711"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2124295.2124351"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.657"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339652"},{"key":"e_1_2_1_8_1","volume-title":"Dzmitry Bahdanau, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho , Bart Van Merrienboer , Dzmitry Bahdanau, and Yoshua Bengio. 2014 . On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259. Kyunghyun Cho, Bart Van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the Proceedings-International Joint Conference on Artificial Intelligence (IJCAI\u201911)","volume":"22","author":"Ciresan Dan C.","year":"2011","unstructured":"Dan C. Ciresan , Ueli Meier , Jonathan Masci , Luca Maria Gambardella , and J\u00fcrgen Schmidhuber . 2011 . Flexible, high performance convolutional neural networks for image classification . In Proceedings of the Proceedings-International Joint Conference on Artificial Intelligence (IJCAI\u201911) , Vol. 22 , 1237. Dan C. Ciresan, Ueli Meier, Jonathan Masci, Luca Maria Gambardella, and J\u00fcrgen Schmidhuber. 2011. Flexible, high performance convolutional neural networks for image classification. In Proceedings of the Proceedings-International Joint Conference on Artificial Intelligence (IJCAI\u201911), Vol. 22, 1237."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0026256"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390392"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-014-0733-5"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3132943"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/72.712151"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 15th International Conference on Machine Learning. 170--178","author":"Freund Yoav","year":"1998","unstructured":"Yoav Freund , Raj D. Iyer , Robert E. Schapire , and Yoram Singer . 1998 . An efficient boosting algorithm for combining preferences . In Proceedings of the 15th International Conference on Machine Learning. 170--178 . Yoav Freund, Raj D. Iyer, Robert E. Schapire, and Yoram Singer. 1998. An efficient boosting algorithm for combining preferences. In Proceedings of the 15th International Conference on Machine Learning. 170--178."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1013203451"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICNN.1996.548916"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498759.1498818"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(90)90004-J"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems. 2042--2050","author":"Hu Baotian","year":"2014","unstructured":"Baotian Hu , Zhengdong Lu , Hang Li , and Qingcai Chen . 2014 . Convolutional neural network architectures for matching natural language sentences . In Proceedings of the International Conference on Neural Information Processing Systems. 2042--2050 . Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen. 2014. Convolutional neural network architectures for matching natural language sentences. In Proceedings of the International Conference on Neural Information Processing Systems. 2042--2050."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775067"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_2_1_26_1","volume-title":"Hinton","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E . Hinton . 2012 . Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems . 1097--1105. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097--1105."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_2_1_28_1","volume-title":"Knowing when to look: Adaptive attention via A visual sentinel for image captioning. arXiv preprint arXiv:1612.01887","author":"Lu Jiasen","year":"2016","unstructured":"Jiasen Lu , Caiming Xiong , Devi Parikh , and Richard Socher . 2016. Knowing when to look: Adaptive attention via A visual sentinel for image captioning. arXiv preprint arXiv:1612.01887 ( 2016 ). Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2016. Knowing when to look: Adaptive attention via A visual sentinel for image captioning. arXiv preprint arXiv:1612.01887 (2016)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080795"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2488198"},{"key":"e_1_2_1_31_1","volume-title":"Manning","author":"Luong Minh Thang","year":"2015","unstructured":"Minh Thang Luong , Hieu Pham , and Christopher D . Manning . 2015 . Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025. Minh Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1149941.1149957"},{"key":"e_1_2_1_33_1","first-page":"824","article-title":"Introduction to information retrieval","volume":"43","author":"Raghavan Schutze Manning","year":"2008","unstructured":"Schutze Manning Raghavan . 2008 . Introduction to information retrieval . J. Am. Soc. Inf. Sci. Technol. 43 , 3 (2008), 824 -- 825 . Schutze Manning Raghavan. 2008. Introduction to information retrieval. J. Am. Soc. Inf. Sci. Technol. 43, 3 (2008), 824--825.","journal-title":"J. Am. Soc. Inf. Sci. Technol."},{"key":"e_1_2_1_34_1","first-page":"3111","article-title":"Distributed representations of words and phrases and their compositionality","volume":"26","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov , Ilya Sutskever , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013 . Distributed representations of words and phrases and their compositionality . Adv. Neural Inf. Process. Syst. 26 (2013), 3111 -- 3119 . Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. Adv. Neural Inf. Process. Syst. 26 (2013), 3111--3119.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_1_35_1","first-page":"1","article-title":"The PageRank citation ranking: Bringing order to the web","volume":"9","author":"Page L.","year":"1999","unstructured":"L. Page . 1999 . The PageRank citation ranking: Bringing order to the web . Stanf. Dig. Libr. Work. Pap. 9 , 1 (1999), 1 -- 14 . L. Page. 1999. The PageRank citation ranking: Bringing order to the web. Stanf. Dig. Libr. Work. Pap. 9, 1 (1999), 1--14.","journal-title":"Stanf. Dig. Libr. Work. Pap."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the Association for the Advancement of Artificial Intelligence Conference (AAAI\u201916)","author":"Pang Liang","year":"2016","unstructured":"Liang Pang , Yanyan Lan , Jiafeng Guo , Jun Xu , Shengxian Wan , and Xueqi Cheng . 2016 . Text matching as image recognition . In Proceedings of the Association for the Advancement of Artificial Intelligence Conference (AAAI\u201916) . 2793--2799. Liang Pang, Yanyan Lan, Jiafeng Guo, Jun Xu, Shengxian Wan, and Xueqi Cheng. 2016. Text matching as image recognition. In Proceedings of the Association for the Advancement of Artificial Intelligence Conference (AAAI\u201916). 2793--2799."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(90)90005-K"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-009-9123-y"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000019"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/582415.582418"},{"key":"e_1_2_1_42_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 151--161","author":"Socher Richard","unstructured":"Richard Socher , Jeffrey Pennington , Eric H. Huang , Andrew Y. Ng , and Christopher D. Manning . 2011. Semi-supervised recursive autoencoders for predicting sentiment distributions . In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 151--161 . Richard Socher, Jeffrey Pennington, Eric H. Huang, Andrew Y. Ng, and Christopher D. Manning. 2011. Semi-supervised recursive autoencoders for predicting sentiment distributions. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 151--161."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/988672.988700"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1046456.1046459"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/72.572108"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_48_1","volume-title":"Manning","author":"Tai Kai Sheng","year":"2015","unstructured":"Kai Sheng Tai , Richard Socher , and Christopher D . Manning . 2015 . Improved semantic representations from tree-structured long short-term memory networks. arXiv preprint arXiv:1503.00075 (2015). Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015. Improved semantic representations from tree-structured long short-term memory networks. arXiv preprint arXiv:1503.00075 (2015)."},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of theThirtieth AAAI Conference on Artificial Intelligence. 2835--2841","author":"Wan Shengxian","year":"2015","unstructured":"Shengxian Wan , Yanyan Lan , Jiafeng Guo , Jun Xu , Liang Pang , and Xueqi Cheng . 2015 . A deep architecture for semantic matching with multiple positional sentence representations . In Proceedings of theThirtieth AAAI Conference on Artificial Intelligence. 2835--2841 . Shengxian Wan, Yanyan Lan, Jiafeng Guo, Jun Xu, Liang Pang, and Xueqi Cheng. 2015. A deep architecture for semantic matching with multiple positional sentence representations. In Proceedings of theThirtieth AAAI Conference on Artificial Intelligence. 2835--2841."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766462.2767712"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2484028.2484036"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-009-9112-1"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277809"},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the International Conference on Machine Learning. 2048--2057","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhutdinov , Richard Zemel , and Yoshua Bengio . 2015 . Show, attend and tell: Neural image caption generation with visual attention . In Proceedings of the International Conference on Machine Learning. 2048--2057 . Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning. 2048--2057."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939677"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271673"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2911451.2911500"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3329188","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3329188","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:41Z","timestamp":1750204481000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3329188"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,26]]},"references-count":57,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,7,31]]}},"alternative-id":["10.1145\/3329188"],"URL":"https:\/\/doi.org\/10.1145\/3329188","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,26]]},"assertion":[{"value":"2018-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-06-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}