{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T08:34:13Z","timestamp":1780389253321,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,4,19]]},"DOI":"10.1145\/3442381.3449988","type":"proceedings-article","created":{"date-parts":[[2021,6,3]],"date-time":"2021-06-03T19:00:27Z","timestamp":1622746827000},"page":"2466-2475","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":33,"title":["Using Prior Knowledge to Guide BERT\u2019s Attention in Semantic Textual Matching Tasks"],"prefix":"10.1145","author":[{"given":"Tingyu","family":"Xia","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence and Jilin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yue","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information and Library Science and University of North Carolina at Chapel Hill, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuan","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Jilin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Chang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Jilin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,6,3]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Islam Beltagy Stephen Roller Pengxiang Cheng Katrin Erk and Raymond\u00a0J. Mooney. 2015. Representing Meaning with a Combination of Logical Form and Vectors. CoRR abs\/1505.06816(2015).  Islam Beltagy Stephen Roller Pengxiang Cheng Katrin Erk and Raymond\u00a0J. Mooney. 2015. Representing Meaning with a Combination of Logical Form and Vectors. CoRR abs\/1505.06816(2015)."},{"key":"e_1_3_2_1_2_1","volume-title":"The unified medical language system (UMLS): integrating biomedical terminology. Nucleic acids research 32, suppl_1","author":"Bodenreider Olivier","year":"2004","unstructured":"Olivier Bodenreider . 2004. The unified medical language system (UMLS): integrating biomedical terminology. Nucleic acids research 32, suppl_1 ( 2004 ), D267\u2013D270. Olivier Bodenreider. 2004. The unified medical language system (UMLS): integrating biomedical terminology. Nucleic acids research 32, suppl_1 (2004), D267\u2013D270."},{"key":"e_1_3_2_1_3_1","unstructured":"Daniel Cer Mona Diab Eneko Agirre Inigo Lopez-Gazpio and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055(2017).  Daniel Cer Mona Diab Eneko Agirre Inigo Lopez-Gazpio and Lucia Specia. 2017. Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055(2017)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33016252"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1152"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1224"},{"key":"e_1_3_2_1_7_1","unstructured":"Yun-Nung Chen Dilek Hakkani-Tur Gokhan Tur Asli Celikyilmaz Jianfeng Gao and Li Deng. 2016. Knowledge as a teacher: Knowledge-guided structural attention networks. arXiv preprint arXiv:1609.03286(2016).  Yun-Nung Chen Dilek Hakkani-Tur Gokhan Tur Asli Celikyilmaz Jianfeng Gao and Li Deng. 2016. Knowledge as a teacher: Knowledge-guided structural attention networks. arXiv preprint arXiv:1609.03286(2016)."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Kevin Clark Urvashi Khandelwal Omer Levy and Christopher\u00a0D. Manning. 2019. What Does BERT Look At? An Analysis of BERT\u2019s Attention. In BlackBoxNLP@ACL.  Kevin Clark Urvashi Khandelwal Omer Levy and Christopher\u00a0D. Manning. 2019. What Does BERT Look At? An Analysis of BERT\u2019s Attention. In BlackBoxNLP@ACL.","DOI":"10.18653\/v1\/W19-4828"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.3115\/1687878.1687944"},{"key":"e_1_3_2_1_10_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805(2018)."},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of the Third International Workshop on Paraphrasing (IWP2005)","author":"Dolan B","year":"2005","unstructured":"William\u00a0 B Dolan and Chris Brockett . 2005 . Automatically Constructing a Corpus of Sentential Paraphrases . In Proceedings of the Third International Workshop on Paraphrasing (IWP2005) . William\u00a0B Dolan and Chris Brockett. 2005. Automatically Constructing a Corpus of Sentential Paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00298"},{"key":"e_1_3_2_1_13_1","volume-title":"Proceedings of the 11th annual research colloquium of the UK special interest group for computational linguistics. 45\u201352","author":"Fernando Samuel","year":"2008","unstructured":"Samuel Fernando and Mark Stevenson . 2008 . A semantic similarity approach to paraphrase detection . In Proceedings of the 11th annual research colloquium of the UK special interest group for computational linguistics. 45\u201352 . Samuel Fernando and Mark Stevenson. 2008. A semantic similarity approach to paraphrase detection. In Proceedings of the 11th annual research colloquium of the UK special interest group for computational linguistics. 45\u201352."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.740"},{"key":"e_1_3_2_1_15_1","unstructured":"Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2020. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arxiv:2006.03654\u00a0[cs.CL]  Pengcheng He Xiaodong Liu Jianfeng Gao and Weizhu Chen. 2020. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arxiv:2006.03654\u00a0[cs.CL]"},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Hewitt John","year":"2019","unstructured":"John Hewitt and Christopher\u00a0 D Manning . 2019 . A structural probe for finding syntax in word representations . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Volume 1 (Long and Short Papers). 4129\u20134138. John Hewitt and Christopher\u00a0D Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4129\u20134138."},{"key":"e_1_3_2_1_17_1","volume-title":"Long short-term memory. Neural computation 9, 8","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber . 1997. Long short-term memory. Neural computation 9, 8 ( 1997 ), 1735\u20131780. Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735\u20131780."},{"key":"e_1_3_2_1_18_1","unstructured":"Baotian Hu Zhengdong Lu Hang Li and Qingcai Chen. 2014. Convolutional neural network architectures for matching natural language sentences. In Advances in neural information processing systems. 2042\u20132050.  Baotian Hu Zhengdong Lu Hang Li and Qingcai Chen. 2014. Convolutional neural network architectures for matching natural language sentences. In Advances in neural information processing systems. 2042\u20132050."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.3115\/1654536.1654562"},{"key":"e_1_3_2_1_20_1","unstructured":"Shankar Iyer Nikhil Dandekar and Korn\u00e9l Csernai. 2017. First quora dataset release: Question pairs. URL https:\/\/data.quora.com\/First-Quora-Dataset-Release-Question-Pairs(2017).  Shankar Iyer Nikhil Dandekar and Korn\u00e9l Csernai. 2017. First quora dataset release: Question pairs. URL https:\/\/data.quora.com\/First-Quora-Dataset-Release-Question-Pairs(2017)."},{"key":"e_1_3_2_1_21_1","unstructured":"Mandar Joshi Danqi Chen Yinhan Liu Daniel\u00a0S. Weld Luke Zettlemoyer and Omer Levy. 2019. SpanBERT: Improving Pre-training by Representing and Predicting Spans. arXiv preprint arXiv:1907.10529(2019).  Mandar Joshi Danqi Chen Yinhan Liu Daniel\u00a0S. Weld Luke Zettlemoyer and Omer Levy. 2019. SpanBERT: Improving Pre-training by Representing and Predicting Spans. arXiv preprint arXiv:1907.10529(2019)."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1126"},{"key":"e_1_3_2_1_23_1","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics, 3890\u20133902","author":"Lan Wuwei","year":"2018","unstructured":"Wuwei Lan and Wei Xu . 2018 . Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering . In Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics, 3890\u20133902 . Wuwei Lan and Wei Xu. 2018. Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering. In Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics, 3890\u20133902."},{"key":"e_1_3_2_1_24_1","volume-title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations.","author":"Lan Zhenzhong","year":"2019","unstructured":"Zhenzhong Lan , Mingda Chen , Sebastian Goodman , Kevin Gimpel , Piyush Sharma , and Radu Soricut . 2019 . ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations. Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_25_1","unstructured":"Guanyu Li Pengfei Zhang and Caiyan Jia. 2018. Attention Boosted Sequential Inference Model. CoRR abs\/1812.01840(2018).  Guanyu Li Pengfei Zhang and Caiyan Jia. 2018. Attention Boosted Sequential Inference Model. CoRR abs\/1812.01840(2018)."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i03.5681"},{"key":"e_1_3_2_1_27_1","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692(2019).  Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692(2019)."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_2_1_29_1","volume-title":"A Decomposable Attention Model for Natural Language Inference","author":"Parikh P.","unstructured":"Ankur\u00a0 P. Parikh , Oscar T\u00e4ckstr\u00f6m , Dipanjan Das , and Jakob Uszkoreit . 2016. A Decomposable Attention Model for Natural Language Inference . In EMNLP. The Association for Computational Linguistics , 2249\u20132255. Ankur\u00a0P. Parikh, Oscar T\u00e4ckstr\u00f6m, Dipanjan Das, and Jakob Uszkoreit. 2016. A Decomposable Attention Model for Natural Language Inference. In EMNLP. The Association for Computational Linguistics, 2249\u20132255."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_1_32_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.  Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Anna Rogers Olga Kovaleva and Anna Rumshisky. 2020. A primer in bertology: What we know about how bert works. arXiv preprint arXiv:2002.12327(2020).  Anna Rogers Olga Kovaleva and Anna Rumshisky. 2020. A primer in bertology: What we know about how bert works. arXiv preprint arXiv:2002.12327(2020).","DOI":"10.1162\/tacl_a_00349"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00178"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S15-2027"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1452"},{"key":"e_1_3_2_1_37_1","unstructured":"Ian Tenney Patrick Xia Berlin Chen Alex Wang Adam Poliak R\u00a0Thomas McCoy Najoung Kim Benjamin Van\u00a0Durme Samuel\u00a0R Bowman Dipanjan Das 2019. What do you learn from context? probing for sentence structure in contextualized word representations. arXiv preprint arXiv:1905.06316(2019).  Ian Tenney Patrick Xia Berlin Chen Alex Wang Adam Poliak R\u00a0Thomas McCoy Najoung Kim Benjamin Van\u00a0Durme Samuel\u00a0R Bowman Dipanjan Das 2019. What do you learn from context? probing for sentence structure in contextualized word representations. arXiv preprint arXiv:1905.06316(2019)."},{"key":"e_1_3_2_1_38_1","volume-title":"SWCN@EMNLP","author":"Tomar Gaurav\u00a0Singh","unstructured":"Gaurav\u00a0Singh Tomar , Thyago Duque , Oscar T\u00e4ckstr\u00f6m , Jakob Uszkoreit , and Dipanjan Das . 2017. Neural Paraphrase Identification of Questions with Noisy Pretraining . In SWCN@EMNLP . Association for Computational Linguistics , 142\u2013147. Gaurav\u00a0Singh Tomar, Thyago Duque, Oscar T\u00e4ckstr\u00f6m, Jakob Uszkoreit, and Dipanjan Das. 2017. Neural Paraphrase Identification of Questions with Noisy Pretraining. In SWCN@EMNLP. Association for Computational Linguistics, 142\u2013147."},{"key":"e_1_3_2_1_39_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008."},{"key":"e_1_3_2_1_40_1","volume-title":"GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In the Proceedings of ICLR.","author":"Wang Alex","year":"2019","unstructured":"Alex Wang , Amanpreet Singh , Julian Michael , Felix Hill , Omer Levy , and Samuel\u00a0 R. Bowman . 2019 . GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In the Proceedings of ICLR. Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel\u00a0R. Bowman. 2019. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In the Proceedings of ICLR."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/579"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.3115\/981732.981751"},{"key":"e_1_3_2_1_43_1","unstructured":"Qizhe Xie Zihang Dai Eduard Hovy Minh-Thang Luong and Quoc\u00a0V Le. 2019. Unsupervised data augmentation for consistency training. arXiv preprint arXiv:1904.12848(2019).  Qizhe Xie Zihang Dai Eduard Hovy Minh-Thang Luong and Quoc\u00a0V Le. 2019. Unsupervised data augmentation for consistency training. arXiv preprint arXiv:1904.12848(2019)."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00194"},{"key":"e_1_3_2_1_45_1","volume-title":"Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems. 5753\u20135763.","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang , Zihang Dai , Yiming Yang , Jaime Carbonell , Russ\u00a0 R Salakhutdinov , and Quoc\u00a0 V Le . 2019 . Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems. 5753\u20135763. Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ\u00a0R Salakhutdinov, and Quoc\u00a0V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in neural information processing systems. 5753\u20135763."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00097"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1139"},{"key":"e_1_3_2_1_48_1","volume-title":"Semantics-Aware BERT for Language Understanding","author":"Zhang Zhuosheng","unstructured":"Zhuosheng Zhang , Yuwei Wu , Hai Zhao , Zuchao Li , Shuailiang Zhang , Xi Zhou , and Xiang Zhou . 2020. Semantics-Aware BERT for Language Understanding . In AAAI. AAAI Press , 9628\u20139635. Zhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li, Shuailiang Zhang, Xi Zhou, and Xiang Zhou. 2020. Semantics-Aware BERT for Language Understanding. In AAAI. AAAI Press, 9628\u20139635."}],"event":{"name":"WWW '21: The Web Conference 2021","location":"Ljubljana Slovenia","acronym":"WWW '21","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the Web Conference 2021"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449988","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3442381.3449988","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:45Z","timestamp":1750195485000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3442381.3449988"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,19]]},"references-count":48,"alternative-id":["10.1145\/3442381.3449988","10.1145\/3442381"],"URL":"https:\/\/doi.org\/10.1145\/3442381.3449988","relation":{},"subject":[],"published":{"date-parts":[[2021,4,19]]},"assertion":[{"value":"2021-06-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}