{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T01:53:22Z","timestamp":1760234002091,"version":"build-2065373602"},"reference-count":33,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2021,3,13]],"date-time":"2021-03-13T00:00:00Z","timestamp":1615593600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Using the single premise entailment (SPE) model to accomplish the multi-premise entailment (MPE) task can alleviate the problem that the neural network cannot be effectively trained due to the lack of labeled multi-premise training data. Moreover, the abundant judgment methods for the relationship between sentence pairs can also be applied in this task. However, the single-premise pre-trained model does not have a structure for processing multi-premise relationships, and this structure is a crucial technique for solving MPE problems. This paper proposes adding a multi-premise relationship processing module based on not changing the structure of the pre-trained model to compensate for this deficiency. Moreover, we proposed a three-step training method combining this module, which ensures that the module focuses on dealing with the multi-premise relationship during matching, thus applying the single-premise model to multi-premise tasks. Besides, this paper also proposes a specific structure of the relationship processing module, i.e., we call it the attention-backtracking mechanism. Experiments show that this structure can fully consider the context of multi-premise, and the structure combined with the three-step training can achieve better accuracy on the MPE test set than other transfer methods.<\/jats:p>","DOI":"10.3390\/fi13030071","type":"journal-article","created":{"date-parts":[[2021,3,14]],"date-time":"2021-03-14T22:13:10Z","timestamp":1615759990000},"page":"71","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Transfer Learning for Multi-Premise Entailment with Relationship Processing Module"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1411-1188","authenticated-orcid":false,"given":"Pin","family":"Wu","sequence":"first","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3603-1573","authenticated-orcid":false,"given":"Rukang","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3158-5393","authenticated-orcid":false,"given":"Zhidan","family":"Lei","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,13]]},"reference":[{"key":"ref_1","unstructured":"Lai, A., Bisk, Y., and Hockenmaier, J. (2017). Natural language inference from multiple premises. arXiv."},{"key":"ref_2","first-page":"26","article-title":"Probabilistic textual entailment: Generic applied modeling of language variability","volume":"2004","author":"Dagan","year":"2004","journal-title":"Learn. Methods Text Underst. Min."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"763","DOI":"10.1002\/asi.24007","article-title":"Defining textual entailment","volume":"69","author":"Korman","year":"2018","journal-title":"J. Assoc. Inf. Sci. Technol."},{"key":"ref_4","unstructured":"Liu, Y., Sun, C., Lin, L., and Wang, X. (2016). Learning natural language inference using bidirectional LSTM model and inner-attention. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Shen, T., Zhou, T., Long, G., Jiang, J., Pan, S., and Zhang, C. (2018, January 2\u20137). Disan: Directional self-attention network for rnn\/cnn-free language understanding. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11941"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ghaeini, R., Hasan, S.A., Datla, V., Liu, J., Lee, K., Qadir, A., Ling, Y., Prakash, A., Fern, X.Z., and Farri, O. (2018). Dr-bilstm: Dependent reading bidirectional lstm for natural language inference. arXiv.","DOI":"10.18653\/v1\/N18-1132"},{"key":"ref_7","unstructured":"Gong, Y., Luo, H., and Zhang, J. (2017). Natural language inference over interaction space. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Parikh, A.P., T\u00e4ckstr\u00f6m, O., Das, D., and Uszkoreit, J. (2016). A decomposable attention model for natural language inference. arXiv.","DOI":"10.18653\/v1\/D16-1244"},{"key":"ref_9","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_10","unstructured":"Rockt\u00e4schel, T., Grefenstette, E., Hermann, K.M., Ko\u010disk\u1ef3, T., and Blunsom, P. (2015). Reasoning about entailment with neural attention. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Williams, A., Nangia, N., and Bowman, S.R. (2017). A broad-coverage challenge corpus for sentence understanding through inference. arXiv.","DOI":"10.18653\/v1\/N18-1101"},{"key":"ref_12","unstructured":"Huh, M., Agrawal, P., and Efros, A.A. (2016). What makes ImageNet good for transfer learning?. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., and Fergus, R. (2014). Visualizing and understanding convolutional networks. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"ref_14","unstructured":"Bengio, Y., Laufer, E., Alain, G., and Yosinski, J. (2014, January 21\u201326). Deep generative stochastic networks trainable by backprop. Proceedings of the International Conference on Machine Learning, Beijing, China."},{"key":"ref_15","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_16","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_17","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2021, March 03). Improving Language Understanding by Generative Pre-Training. Available online: https:\/\/www.cs.ubc.ca\/~amuham01\/LING530\/papers\/radford2018improving.pdf."},{"key":"ref_18","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, X., He, P., Chen, W., and Gao, J. (2019). Multi-task deep neural networks for natural language understanding. arXiv.","DOI":"10.18653\/v1\/P19-1441"},{"key":"ref_20","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"8","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Mou, L., Meng, Z., Yan, R., Li, G., Xu, Y., Zhang, L., and Jin, Z. (2016). How transferable are neural networks in nlp applications?. arXiv.","DOI":"10.18653\/v1\/D16-1046"},{"key":"ref_22","unstructured":"Bromley, J., Guyon, I., LeCun, Y., S\u00e4ckinger, E., and Shah, R. (December, January 28). Signature verification using a \u201csiamese\u201d time delay neural network. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Peters, M.E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv.","DOI":"10.18653\/v1\/N18-1202"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, C., Jiang, S., Yu, H., and Yu, D. (2018). Multi-turn inference matching network for natural language inference. CCF International Conference on Natural Language Processing and Chinese Computing, Springer.","DOI":"10.1007\/978-3-319-99501-4_11"},{"key":"ref_27","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Boureau, Y.L., Bach, F., LeCun, Y., and Ponce, J. (2010, January 13\u201318). Learning mid-level features for recognition. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539963"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Bowman, S.R., Angeli, G., Potts, C., and Manning, C.D. (2015). A large annotated corpus for learning natural language inference. arXiv.","DOI":"10.18653\/v1\/D15-1075"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_31","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_32","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Howard, J., and Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv.","DOI":"10.18653\/v1\/P18-1031"}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/13\/3\/71\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:35:11Z","timestamp":1760160911000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/13\/3\/71"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,13]]},"references-count":33,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["fi13030071"],"URL":"https:\/\/doi.org\/10.3390\/fi13030071","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2021,3,13]]}}}