{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T15:33:58Z","timestamp":1772120038291,"version":"3.50.1"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"5-6","license":[{"start":{"date-parts":[[2024,8,3]],"date-time":"2024-08-03T00:00:00Z","timestamp":1722643200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2024,8,3]],"date-time":"2024-08-03T00:00:00Z","timestamp":1722643200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Evol. Intel."],"published-print":{"date-parts":[[2024,10]]},"DOI":"10.1007\/s12065-024-00954-3","type":"journal-article","created":{"date-parts":[[2024,8,3]],"date-time":"2024-08-03T10:01:53Z","timestamp":1722679313000},"page":"3723-3744","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["The triple attention transformer: advancing contextual coherence in transformer models"],"prefix":"10.1007","volume":"17","author":[{"given":"Shadi","family":"Ghaith","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,8,3]]},"reference":[{"key":"954_CR1","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2018) Bert: pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805"},{"issue":"8","key":"954_CR2","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I (2019) Language models are unsupervised multitask learners. OpenAI Blog 1(8):9","journal-title":"OpenAI Blog"},{"key":"954_CR3","unstructured":"Chen M, Tworek J, Jun H, Yuan Q, Pinto H, Kaplan J, Edwards H, Burda Y, Joseph N, Brockman G, et\u00a0al. (2021) Evaluating large language models trained on code. In: Proceedings of the 38th international conference on machine learning"},{"key":"954_CR4","unstructured":"Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et\u00a0al. (2020) Language models are few-shot learners. arXiv:2005.14165"},{"key":"954_CR5","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. In: Advances in neural information processing systems. Curran Associates, Inc"},{"key":"954_CR6","unstructured":"Liu Y, Bi J-Y, Fan Z-P (2020) Deep learning for natural language processing: advantages and challenges. In: National CCF conference on natural language processing and Chinese computing. Springer, pp 3\u201314"},{"key":"954_CR7","unstructured":"Kitaev N, Kaiser \u0141, Levskaya A (2020) Reformer: the efficient transformer. arXiv:2001.04451"},{"key":"954_CR8","unstructured":"Beltagy I, Peters ME, Cohan A (2020) Longformer: the long-document transformer. arXiv:2004.05150"},{"key":"954_CR9","doi-asserted-by":"crossref","unstructured":"Zhang Y, Sun S, Galley M, Chen Y-C, Brockett C, Gao X, Gao J, Liu J, Dolan B (2020) Dialogpt: large-scale generative pre-training for conversational response generation. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 270\u2013278","DOI":"10.18653\/v1\/2020.acl-demos.30"},{"key":"954_CR10","doi-asserted-by":"crossref","unstructured":"Feng Z, Guo D, Tang D, Duan N, Feng X, Gong M, Shou L, Qin B, Liu T, Jiang D, et\u00a0al. (2020) Codebert: a pre-trained model for programming and natural languages. In: Proceedings of the 58th annual meeting of the association for computational linguistics, pp 1536\u20131547","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"954_CR11","doi-asserted-by":"crossref","unstructured":"Rastogi A, Zang X, Sunkara S, Gupta R, Khaitan P (2019) Towards scalable multi-domain conversational agents: the schema-guided dialogue dataset. In: The 8th international conference on learning representations","DOI":"10.1609\/aaai.v34i05.6394"},{"key":"954_CR12","doi-asserted-by":"crossref","unstructured":"Papineni K, Roukos S, Ward T, Zhu W-J (2002) Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th annual meeting of the association for computational linguistics, pp 311\u2013318","DOI":"10.3115\/1073083.1073135"},{"key":"954_CR13","unstructured":"Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. arXiv:1409.0473"},{"key":"954_CR14","unstructured":"Mehri S, Eric M, Hakkani-Tur D (2020) Dialoglue: a benchmark and analysis platform for natural language understanding. In: Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP)"},{"key":"954_CR15","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) Roberta: a robustly optimized bert pretraining approach. arXiv:1907.11692"},{"key":"954_CR16","doi-asserted-by":"crossref","unstructured":"See A, Liu PJ, Manning CD (2017) Get to the point: summarization with pointer-generator networks. In: Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: long papers), pp 1073\u20131083","DOI":"10.18653\/v1\/P17-1099"},{"key":"954_CR17","unstructured":"Socher R, Perelygin A, Wu JY, Chuang J, Manning CD, Ng A, Potts C (2013) Recursive deep models for semantic compositionality over a sentiment treebank. In: Proceedings of the 2013 conference on empirical methods in natural language processing, pp 1631\u20131642"},{"key":"954_CR18","first-page":"5776","volume":"33","author":"W Wang","year":"2020","unstructured":"Wang W, Bao H, Li S, Xia M, Jian W, Liu W (2020) Minilm: deep self-attention distillation for task-agnostic compression of pre-trained transformers. Adv Neural Inf Process Syst 33:5776\u20135788","journal-title":"Adv Neural Inf Process Syst"},{"key":"954_CR19","doi-asserted-by":"crossref","unstructured":"Sun Z, Yu H, Song X, Liu R, Yang Y, Zhou D (2020) Mobilebert: a compact task-agnostic bert for resource-limited devices. In: Annual meeting of the association for computational linguistics","DOI":"10.18653\/v1\/2020.acl-main.195"},{"key":"954_CR20","unstructured":"Tang R, Lin J (2020) Xtremedistil: multi-stage distillation for massive multilingual models. In: Proceedings of the 58th annual meeting of the association for computational linguistics. Association for Computational Linguistics, pp 2631\u20132637"},{"key":"954_CR21","unstructured":"Jaiswal A, Milios E (2023) Breaking the token barrier: chunking and convolution for efficient long text classification with bert."},{"issue":"1","key":"954_CR22","doi-asserted-by":"publisher","first-page":"20349","DOI":"10.1038\/s41598-022-24787-1","volume":"12","author":"Y Yang","year":"2022","unstructured":"Yang Y, Cao J, Wen Y, Zhang P (2022) Multiturn dialogue generation by modeling sentence-level and discourse-level contexts. Sci Rep 12(1):20349","journal-title":"Sci Rep"},{"issue":"3","key":"954_CR23","first-page":"1054","volume":"20","author":"N Zeng","year":"2024","unstructured":"Zeng N, Peishu W, Zhang Y, Li H, Mao J, Wang Z (2024) Dpmsn: a dual-pathway multiscale network for image forgery detection. IEEE Trans Ind Inform 20(3):1054\u20131065","journal-title":"IEEE Trans Ind Inform"},{"key":"954_CR24","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.121305","volume":"237","author":"W Peishu","year":"2024","unstructured":"Peishu W, Wang Z, Li H, Zeng N (2024) Kd-par: a knowledge distillation-based pedestrian attribute recognition model with multi-label mixed feature learning network. Expert Syst Appl 237:121305","journal-title":"Expert Syst Appl"},{"key":"954_CR25","doi-asserted-by":"crossref","unstructured":"Ni J, Young T, Pandelea V, Xue F, Cambria E (2021) Recent advances in deep learning based dialogue systems: a systematic survey. arXiv:2105.04387","DOI":"10.1007\/s10462-022-10248-8"},{"key":"954_CR26","doi-asserted-by":"crossref","unstructured":"Santra B, Anusha P, Goyal P (2021) Hierarchical transformer for task oriented dialog systems. In: Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies, page [pagination], Online. Association for Computational Linguistics","DOI":"10.18653\/v1\/2021.naacl-main.449"},{"key":"954_CR27","unstructured":"Morris A, Maier V, Green P (2004) The use of character error rate for machine translation evaluation. In: Proceedings of the 10th international conference on theoretical and methodological issues in machine translation"},{"key":"954_CR28","unstructured":"Pallett DS, Fiscus JG, Fisher WM (1994) A look at nist\u2019s benchmark asr tests: past, present, and future. In: IEEE aerospace and electronic systems magazine, 9. IEEE, 14\u201320"},{"issue":"05","key":"954_CR29","first-page":"8689","volume":"34","author":"A Rastogi","year":"2020","unstructured":"Rastogi A, Zang X, Sunkara S, Gupta R, Khaitan P (2020) Towards scalable multi-domain conversational agents: the schema-guided dialogue dataset. Proc AAAI Conf Artif Intell 34(05):8689\u20138696","journal-title":"Proc AAAI Conf Artif Intell"}],"container-title":["Evolutionary Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12065-024-00954-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12065-024-00954-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12065-024-00954-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,19]],"date-time":"2024-10-19T02:19:18Z","timestamp":1729304358000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12065-024-00954-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,3]]},"references-count":29,"journal-issue":{"issue":"5-6","published-print":{"date-parts":[[2024,10]]}},"alternative-id":["954"],"URL":"https:\/\/doi.org\/10.1007\/s12065-024-00954-3","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3916608\/v1","asserted-by":"object"}]},"ISSN":["1864-5909","1864-5917"],"issn-type":[{"value":"1864-5909","type":"print"},{"value":"1864-5917","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,3]]},"assertion":[{"value":"1 February 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 May 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 June 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 August 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The research in this paper used the Schema-Guided Dialogue dataset from the 8th Dialogue System Technology Challenge (DSTC8), available at\n                      \n                      . This publicly accessible dataset, created for research, complies with ethical research data standards. As the dataset does not contain any sensitive or personal data and is designed for academic research in dialogue systems, no ethical approval was necessary. The dataset usage adhered to the terms of the Hugging Face platform.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval and informed consent"}},{"value":"This paper is published with the statement that there are no conflicting interests. The research work, including the development and assessment of the Triple Attention Transformer model, was conducted independently, without any specific financial support from public, private, or non-profit organizations. The choice of software and tools mentioned in the paper is solely for their relevance to the study and does not imply any endorsement. The research approach remained unbiased, free from any external influences from software or hardware companies.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}