{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T16:42:28Z","timestamp":1777567348746,"version":"3.51.4"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,2,13]],"date-time":"2022-02-13T00:00:00Z","timestamp":1644710400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,2,13]],"date-time":"2022-02-13T00:00:00Z","timestamp":1644710400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100008354","name":"Callaghan Innovation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100008354","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100008205","name":"Auckland University of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100008205","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"published-print":{"date-parts":[[2022,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Attention mechanisms have been incorporated into many neural network-based natural language processing (NLP) models. They enhance the ability of these models to learn and reason with long input texts. A critical part of such mechanisms is the computation of attention similarity scores between two elements of the texts using a similarity score function. Given that these models have different architectures, it is difficult to comparatively evaluate the effectiveness of different similarity score functions. In this paper, we proposed a baseline model that captures the common components of recurrent neural network-based Question Answering (QA) systems found in the literature. By isolating the attention function, this baseline model allows us to study the effects of different similarity score functions on the performance of such systems. Experimental results show that a trilinear function produced the best results among the commonly used functions. Based on these insights, a new T-trilinear similarity function is proposed which achieved the higher predictive EM and F1 scores than these existing functions. A heatmap visualization of the attention score matrix explains why this T-trilinear function is effective.<\/jats:p>","DOI":"10.1007\/s11063-021-10730-4","type":"journal-article","created":{"date-parts":[[2022,2,13]],"date-time":"2022-02-13T10:02:21Z","timestamp":1644746541000},"page":"2283-2302","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Effects of Similarity Score Functions in Attention Mechanisms on the Performance of Neural Question Answering Systems"],"prefix":"10.1007","volume":"54","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0968-4472","authenticated-orcid":false,"given":"Yuanyuan","family":"Shen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9159-3718","authenticated-orcid":false,"given":"Edmund M.-K.","family":"Lai","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2228-8300","authenticated-orcid":false,"given":"Mahsa","family":"Mohaghegh","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,2,13]]},"reference":[{"key":"10730_CR1","doi-asserted-by":"publisher","unstructured":"Cho K, Van Merri\u00ebnboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using RNN encoder-decoder for statistical machine translation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). Doha, Qatar. https:\/\/doi.org\/10.3115\/v1\/D14-1179","DOI":"10.3115\/v1\/D14-1179"},{"key":"10730_CR2","unstructured":"Sutskever I, Vinyals O, Le QV (2014) Sequence to sequence learning with neural networks. In: Proceedings of the 27th international conference on neural information processing systems (NIPS\u201914). Montreal, Canada"},{"key":"10730_CR3","doi-asserted-by":"publisher","unstructured":"Cho K, van Merrienboer B, Bahdanau D, Bengio Y (2014) On the properties of neural machine translation: Encoder-decoder approaches. In: Proceedings of the 8th workshop on syntax, semantics and structure in statistical translation (SSST-8). Doha, Qatar. https:\/\/doi.org\/10.3115\/v1\/W14-4012","DOI":"10.3115\/v1\/W14-4012"},{"key":"10730_CR4","doi-asserted-by":"publisher","unstructured":"Luong M-T, Pham H, Manning C D (2015) Effective approaches to attention-based neural machine translation. In: Proceedings of the 2015 conference on empirical methods in natural language processing (EMNLP). Lisbon, Portugal. https:\/\/doi.org\/10.18653\/v1\/D15-1166","DOI":"10.18653\/v1\/D15-1166"},{"key":"10730_CR5","doi-asserted-by":"publisher","first-page":"4688","DOI":"10.1109\/TNNLS.2019.2957276","volume":"31","author":"B Zhang","year":"2020","unstructured":"Zhang B, Xiong D, Xie J, Su J (2020) Neural machine translation with gru-gated attention model. IEEE Trans Neural Netw Learn Syst 31:4688\u20134698. https:\/\/doi.org\/10.1109\/TNNLS.2019.2957276","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"10730_CR6","unstructured":"Seo M, Kembhavi A, Farhadi A, Hajishirzi H (2017) Bidirectional attention flow for machine comprehension. In: Proceedings of the 5th international conference on learning representations (ICLR). Toulon, France"},{"key":"10730_CR7","unstructured":"Xiong C, Zhong V, Socher R (2017) Dynamic coattention networks for question answering. In: Proceedings of the 5th international conference on learning representations (ICLR). Toulon, France"},{"key":"10730_CR8","doi-asserted-by":"publisher","unstructured":"Wang W, Yang N, Wei F, Chang B, Zhou M (2017) Gated self-matching networks for reading comprehension and question answering. In: Proceedings of the 55th annual meeting of the association for computational linguistics (Volume 1: Long Papers). Vancouver, Canada. https:\/\/doi.org\/10.18653\/v1\/P17-1018","DOI":"10.18653\/v1\/P17-1018"},{"key":"10730_CR9","unstructured":"Wang S, Jiang J (2016) Machine comprehension using match-lstm and answer pointer. In: Proceedings of the 5th international conference on learning representations (ICLR). Toulon, France"},{"key":"10730_CR10","doi-asserted-by":"publisher","unstructured":"Wang J, Sun C, Li S, Liu X, Si L, Zhang M, Zhou G (2019) Aspect sentiment classification towards question-answering with reinforced bidirectional attention network. In: Proceedings of the 57th annual meeting of the association for computational linguistics. Florence, Italy. https:\/\/doi.org\/10.18653\/v1\/P19-1345","DOI":"10.18653\/v1\/P19-1345"},{"key":"10730_CR11","doi-asserted-by":"publisher","first-page":"2745","DOI":"10.1007\/s11063-019-10049-1","volume":"50","author":"H Sadr","year":"2019","unstructured":"Sadr H, Pedram MM, Teshnehlab M (2019) A robust sentiment analysis method based on sequential combination of convolutional and recursive neural networks. Neural Process Lett 50:2745\u20132761. https:\/\/doi.org\/10.1007\/s11063-019-10049-1","journal-title":"Neural Process Lett"},{"key":"10730_CR12","unstructured":"Kumar A, Irsoy O, Ondruska P, Iyyer M, Bradbury J, Gulrajani I, Zhong V et al (2016) Ask me anything: dynamic memory networks for natural language processing. In: Proceedings of the 33rd international conference on machine learning. New York, USA"},{"key":"10730_CR13","unstructured":"Xiong C, Merity S, Socher R (2016) Dynamic memory networks for visual and textual question answering. In: Proceedings of the 33rd international conference on machine learning. New York, USA"},{"key":"10730_CR14","unstructured":"Yu AW, Dohan D, Luong M-T, Zhao R, Chen K, Norouzi M, Le QV (2018) QANet: combining local convolution with global self-attention for reading comprehension. In: Proceedings of the 6th international conference on learning representations (ICLR). Vancouver, Canada"},{"key":"10730_CR15","doi-asserted-by":"publisher","unstructured":"Feldman Y, El-Yaniv R (2019) Multi-hop paragraph retrieval for open-domain question answering. In: Proceedings of the 57th annual meeting of the association for computational linguistics. Florence, Italy. https:\/\/doi.org\/10.18653\/v1\/P19-1222","DOI":"10.18653\/v1\/P19-1222"},{"key":"10730_CR16","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1016\/j.neucom.2020.01.056","volume":"391","author":"R Li","year":"2020","unstructured":"Li R, Jiang Z, Wang L, Lu X, Zhao M (2020) Directional attention weaving for text-grounded conversational question answering. Neurocomputing 391:13\u201324. https:\/\/doi.org\/10.1016\/j.neucom.2020.01.056","journal-title":"Neurocomputing"},{"key":"10730_CR17","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser \u0141 et al (2017) Attention is all you need. In: Proceedings of the 31st international conference on neural information processing systems (NIPS\u201917). Long Beach, California, USA"},{"key":"10730_CR18","unstructured":"Radford A, Narasimhan K, Salimans T, Sutskever I (2018) Improving language understanding by generative pre-training. OpenAI blog"},{"key":"10730_CR19","unstructured":"Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I (2019) Language models are unsupervised multitask learners. OpenAI blog"},{"key":"10730_CR20","unstructured":"Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A et al (2020) Language models are few-shot learners. arXiv preprint arXiv:200514165"},{"key":"10730_CR21","doi-asserted-by":"publisher","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2019) Bert: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 annual conference of the North American chapter of the association for computational linguistics (NAACL-HLT). Minneapolis, Minnesota, USA. https:\/\/doi.org\/10.18653\/v1\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"key":"10730_CR22","doi-asserted-by":"publisher","unstructured":"Dai Z, Yang Z, Yang Y, Carbonell J, Le Q, Salakhutdinov R (2019) Transformer-xl: attentive language models beyond a fixed-length context. In: Proceedings of the 57th annual meeting of the association for computational linguistics. Florence, Italy. https:\/\/doi.org\/10.18653\/v1\/P19-1285","DOI":"10.18653\/v1\/P19-1285"},{"key":"10730_CR23","doi-asserted-by":"publisher","unstructured":"Weissenborn D, Wiese G, Seiffe L (2017) Making neural QA as simple as possible but not simpler. In: Proceedings of the 21st conference on computational natural language learning (CoNLL 2017). Vancouver, Canada. https:\/\/doi.org\/10.18653\/v1\/K17-1028","DOI":"10.18653\/v1\/K17-1028"},{"key":"10730_CR24","doi-asserted-by":"publisher","unstructured":"Geng X, Wang L, Wang X, Qin B, Liu T, Tu Z (2020) How does selective mechanism improve self-attention networks?. In: Proceedings of the 58th annual meeting of the association for computational linguistics. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.269","DOI":"10.18653\/v1\/2020.acl-main.269"},{"key":"10730_CR25","doi-asserted-by":"publisher","unstructured":"Britz D, Goldie A, Luong M-T, Le Q (2017) Massive exploration of neural machine translation architectures. In: Proceedings of the 2017 conference on empirical methods in natural language processing (EMNLP). Copenhagen, Denmark. https:\/\/doi.org\/10.18653\/v1\/D17-1151","DOI":"10.18653\/v1\/D17-1151"},{"key":"10730_CR26","unstructured":"Graves A, Wayne G, Danihelka I (2014) Neural turing machines. arXiv preprint arXiv:14105401"},{"key":"10730_CR27","doi-asserted-by":"publisher","unstructured":"Liu X, Shen Y, Duh K, Gao J (2018) Stochastic answer networks for machine reading comprehension. In: Proceedings of the 56th annual meeting of the association for computational linguistics (Volume 1: Long Papers). Melbourne, Australia. https:\/\/doi.org\/10.18653\/v1\/P18-1157","DOI":"10.18653\/v1\/P18-1157"},{"key":"10730_CR28","unstructured":"Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A et al (2020) Language models are few-shot learners. CoRR abs\/2005.14165"},{"key":"10730_CR29","doi-asserted-by":"publisher","unstructured":"Clark C, Gardner M (2018) Simple and effective multi-paragraph reading comprehension. In: Proceedings of the 56th annual meeting of the association for computational linguistics (Volume 1: Long Papers). Melbourne, Australia. https:\/\/doi.org\/10.18653\/v1\/P18-1078","DOI":"10.18653\/v1\/P18-1078"},{"key":"10730_CR30","doi-asserted-by":"publisher","unstructured":"Chen D, Fisch A, Weston J, Bordes A (2017) Reading wikipedia to answer open-domain questions. In: Proceedings of the 55th annual meeting of the association for computational linguistics (Volume 1: Long Papers). Vancouver, Canada. https:\/\/doi.org\/10.18653\/v1\/P17-1171","DOI":"10.18653\/v1\/P17-1171"},{"key":"10730_CR31","unstructured":"Huang H-Y, Zhu C, Shen Y, Chen W (2018) Fusionnet: fusing via fully-aware attention with application to machine comprehension. In: Proceedings of the 6th international conference on learning representations (ICLR). Vancouver, Canada"},{"key":"10730_CR32","unstructured":"Fedus W, Zoph B, Shazeer N (2021) Switch transformers: scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:210103961"},{"key":"10730_CR33","unstructured":"Ramesh A, Pavlov M, Goh G, Gray S, Voss C, Radford A, Chen M et al (2021) Zero-shot text-to-image generation. arXiv preprint arXiv:210212092"},{"key":"10730_CR34","unstructured":"Romero A (2021) Gpt-3 scared you? Meet wu dao 2.0: a monster of 1.75 trillion parameters"},{"key":"10730_CR35","doi-asserted-by":"publisher","unstructured":"Kobayashi S, Tian R, Okazaki N, Inui K (2016) Dynamic entity representation with max-pooling improves machine reading. In: Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies. San Diego, California, USA. https:\/\/doi.org\/10.18653\/v1\/N16-1099","DOI":"10.18653\/v1\/N16-1099"},{"key":"10730_CR36","unstructured":"Sukhbaatar S, Weston J, Fergus R (2015) End-to-end memory networks. In: Proceedings of the 28th international conference on neural information processing systems (NIPS\u201915). Montreal, Canada"},{"key":"10730_CR37","unstructured":"Xiong C, Zhong V, Socher R (2018) DCN+: mixed objective and deep residual coattention for question answering. In: Proceedings of the 6th international conference on learning representations (ICLR). Vancouver, Canada"},{"key":"10730_CR38","doi-asserted-by":"publisher","unstructured":"Wang Y, Liu K, Liu J, He W, Lyu Y, Wu H, Li S et al (2018) Multi-passage machine reading comprehension with cross-passage answer verification. In: Proceedings of the 56th annual meeting of the association for computational linguistics (Volume 1: Long Papers). Melbourne, Australia. https:\/\/doi.org\/10.18653\/v1\/P18-1178","DOI":"10.18653\/v1\/P18-1178"},{"key":"10730_CR39","unstructured":"Bahdanau D, Cho K, Bengio Y (2015) Neural machine translation by jointly learning to align and translate. In: Proceedings of the 3rd international conference on learning representations (ICLR). San Diego, CA, USA"},{"key":"10730_CR40","doi-asserted-by":"publisher","unstructured":"Rajpurkar P, Zhang J, Lopyrev K, Liang P (2016) Squad: 100,000+ questions for machine comprehension of text. In: Proceedings of the 2016 conference on empirical methods in natural language processing. Austin, Texas, USA. https:\/\/doi.org\/10.18653\/v1\/D16-1264","DOI":"10.18653\/v1\/D16-1264"},{"key":"10730_CR41","doi-asserted-by":"publisher","unstructured":"Pennington J, Socher R, Manning C (2014) Glove: global vectors for word representation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). Doha, Qatar. https:\/\/doi.org\/10.3115\/v1\/D14-1162","DOI":"10.3115\/v1\/D14-1162"},{"key":"10730_CR42","unstructured":"Zeiler MD (2012) Adadelta: an adaptive learning rate method. arXiv preprint arXiv:12125701"},{"key":"10730_CR43","doi-asserted-by":"crossref","unstructured":"Radosavovic I, Johnson J, Xie S, Lo W-Y, Doll\u00e1r P (2019) On network design spaces for visual recognition. In: Proceedings of the IEEE\/CVF international conference on computer vision","DOI":"10.1109\/ICCV.2019.00197"},{"key":"10730_CR44","doi-asserted-by":"crossref","unstructured":"Radosavovic I, Kosaraju RP, Girshick R, He K, Doll\u00e1r P (2020) Designing network design spaces. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR42600.2020.01044"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-021-10730-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-021-10730-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-021-10730-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,5,28]],"date-time":"2022-05-28T13:24:52Z","timestamp":1653744292000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-021-10730-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,13]]},"references-count":44,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,6]]}},"alternative-id":["10730"],"URL":"https:\/\/doi.org\/10.1007\/s11063-021-10730-4","relation":{},"ISSN":["1370-4621","1573-773X"],"issn-type":[{"value":"1370-4621","type":"print"},{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,13]]},"assertion":[{"value":"20 December 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 February 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}