{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T16:17:03Z","timestamp":1772122623518,"version":"3.50.1"},"reference-count":57,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2021,11,5]],"date-time":"2021-11-05T00:00:00Z","timestamp":1636070400000},"content-version":"vor","delay-in-days":308,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,27]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Dialog acts can be interpreted as the atomic units of a conversation, more fine-grained than utterances, characterized by a specific communicative function. The ability to structure a conversational transcript as a sequence of dialog acts\u2014dialog act recognition, including the segmentation\u2014is critical for understanding dialog. We apply two pre-trained transformer models, XLNet and Longformer, to this task in English and achieve strong results on Switchboard Dialog Act and Meeting Recorder Dialog Act corpora with dialog act segmentation error rates (DSER) of 8.4% and 14.2%. To understand the key factors affecting dialog act recognition, we perform a comparative analysis of models trained under different conditions. We find that the inclusion of a broader conversational context helps disambiguate many dialog act classes, especially those infrequent in the training data. The presence of punctuation in the transcripts has a massive effect on the models\u2019 performance, and a detailed analysis reveals specific segmentation patterns observed in its absence. Finally, we find that the label set specificity does not affect dialog act segmentation performance. These findings have significant practical implications for spoken language understanding applications that depend heavily on a good-quality segmentation being available.<\/jats:p>","DOI":"10.1162\/tacl_a_00420","type":"journal-article","created":{"date-parts":[[2021,11,5]],"date-time":"2021-11-05T12:56:11Z","timestamp":1636116971000},"page":"1163-1179","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":7,"title":["What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition"],"prefix":"10.1162","volume":"9","author":[{"given":"Piotr","family":"\u017belasko","sequence":"first","affiliation":[{"name":"Center of Language and Speech Processing"},{"name":"Human Language Technology Center of Excellence, Johns Hopkins University, Baltimore, MD, USA. piotr.andrzej.zelasko@gmail.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raghavendra","family":"Pappagari","sequence":"additional","affiliation":[{"name":"Center of Language and Speech Processing"},{"name":"Human Language Technology Center of Excellence, Johns Hopkins University, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Najim","family":"Dehak","sequence":"additional","affiliation":[{"name":"Center of Language and Speech Processing"},{"name":"Human Language Technology Center of Excellence, Johns Hopkins University, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2021,10,27]]},"reference":[{"key":"2021110920491891600_bib1","first-page":"I","article-title":"Automatic dialog act segmentation and classification in multiparty meetings","volume-title":"Proceedings. (ICASSP\u201905). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005.","author":"Ang","year":"2005"},{"key":"2021110920491891600_bib2","volume-title":"How to Do Things with Words","author":"Austin","year":"1962"},{"key":"2021110920491891600_bib3","article-title":"Longformer: The long-document transformer","author":"Iz","year":"2020","journal-title":"arXiv preprint arXiv:2004.05150 [v1]"},{"key":"2021110920491891600_bib4","doi-asserted-by":"publisher","first-page":"2349","DOI":"10.1145\/1518701.1519060","article-title":"Conversation clusters: Grouping conversation topics through human-computer dialog","volume-title":"Proceedings of the SIGCHI Conference on Human Factors in Computing Systems","author":"Bergstrom","year":"2009"},{"key":"2021110920491891600_bib5","article-title":"A context-based approach for dialogue act recognition using simple recurrent neural networks","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)","author":"Bothe","year":"2018"},{"key":"2021110920491891600_bib6","first-page":"430","article-title":"ISO 24617-2: A semantically-based standard for dialogue annotation.","volume-title":"LREC","author":"Bunt","year":"2012"},{"key":"2021110920491891600_bib7","article-title":"Revisiting the ISO standard for dialogue act annotation","volume-title":"Proceedings of the 13th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (ISA-13)","author":"Bunt","year":"2017"},{"key":"2021110920491891600_bib8","first-page":"549","article-title":"The ISO standard for dialogue act annotation","volume-title":"Proceedings of the 12th Language Resources and Evaluation Conference","author":"Bunt","year":"2020"},{"key":"2021110920491891600_bib9","volume-title":"The British National Corpus Users Reference Guide","author":"Burnard","year":"2000"},{"key":"2021110920491891600_bib10","doi-asserted-by":"publisher","DOI":"10.3115\/1073336.1073352","article-title":"Edit detection and parsing for transcribed speech","volume-title":"Second Meeting of the North American Chapter of the Association for Computational Linguistics","author":"Charniak","year":"2001"},{"issue":"1","key":"2021110920491891600_bib11","doi-asserted-by":"publisher","first-page":"74","DOI":"10.1109\/TAFFC.2015.2444846","article-title":"Sentiment analysis: From opinion mining to human- agent interaction","volume":"7","author":"Clavel","year":"2015","journal-title":"IEEE Transactions on Affective Computing"},{"key":"2021110920491891600_bib12","doi-asserted-by":"publisher","first-page":"7594","DOI":"10.1609\/aaai.v34i05.6259","article-title":"Guiding attention in sequence- to-sequence models for dialogue act prediction.","volume-title":"AAAI","author":"Colombo","year":"2020"},{"key":"2021110920491891600_bib13","first-page":"28","article-title":"Coding dialogs with the DAMSL annotation scheme","volume-title":"AAAI Fall Symposium on Communicative Action in Humans and Machines","author":"Core","year":"1997"},{"key":"2021110920491891600_bib14","article-title":"Local contextual attention with hierarchical structure for dialogue act recognition","author":"Dai","year":"2020","journal-title":"arXiv preprint arXiv:2003. 06044 [v1]"},{"key":"2021110920491891600_bib15","doi-asserted-by":"crossref","first-page":"2978","DOI":"10.18653\/v1\/P19-1285","article-title":"Transformer-XL: Attentive language models beyond a fixed-length context","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Dai","year":"2019"},{"key":"2021110920491891600_bib16","doi-asserted-by":"crossref","first-page":"3910","DOI":"10.21437\/Interspeech.2020-1062","article-title":"End-to-end speech-to-dialog-act recognition","author":"Dang","year":"2020","journal-title":"Proceedings of Interspeech 2020"},{"key":"2021110920491891600_bib17","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2021110920491891600_bib18","first-page":"261","article-title":"Feedback in conversation as incremental semantic update","volume-title":"Proceedings of the 11th International Conference on Computational Semantics","author":"Eshghi","year":"2015"},{"issue":"2","key":"2021110920491891600_bib19","first-page":"23","article-title":"A new algorithm for data compression","volume":"12","author":"Gage","year":"1994","journal-title":"The C Users Journal"},{"key":"2021110920491891600_bib20","article-title":"Two-level transformer and auxiliary coherence modeling for improved text segmentation","author":"Glavas","year":"2020","journal-title":"ArXiv"},{"key":"2021110920491891600_bib21","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2010-765","article-title":"Dialogue act tagging and segmentation with a single perceptron","volume-title":"Eleventh Annual Conference of the International Speech Communication Association","author":"Granell","year":"2010"},{"key":"2021110920491891600_bib22","volume-title":"Switchboard SWBD-DAMSL Labeling Project Coder\u2019s Manual","author":"Jurafsky","year":"1997"},{"key":"2021110920491891600_bib23","unstructured":"Daniel Jurafsky , RebeccaBates, NoahCoccaro, RachelMartin, MarieMeteer, KlausRies, ElizabethShriberg, AndreasStolcke, PaulTaylor, and CarolVan Ess-Dykema DoD. 1998. Johns Hopkins LVCSR Workshop-97 Switchboard Discourse Language Modeling Project Final Report."},{"key":"2021110920491891600_bib24","first-page":"5156","article-title":"Transformers are RNNs: Fast autoregressive transformers with linear attention","volume-title":"International Conference on Machine Learning","author":"Katharopoulos","year":"2020"},{"issue":"3\u20134","key":"2021110920491891600_bib25","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1515\/tl-2016-0011","article-title":"Language as mechanisms for interaction","volume":"42","author":"Kempson","year":"2016","journal-title":"Theoretical Linguistics"},{"key":"2021110920491891600_bib26","volume-title":"Dynamic Syntax: The Flow of Language Understanding","author":"Kempson","year":"2000"},{"key":"2021110920491891600_bib27","article-title":"Reformer: The efficient transformer","author":"Kitaev","year":"2020","journal-title":"Proceedings of International Conference on Learning Representations (ICLR)"},{"key":"2021110920491891600_bib28","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/D18-2012","article-title":"SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Kudo","year":"2018"},{"key":"2021110920491891600_bib29","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.11701","article-title":"Dialogue act sequence labeling using hierarchical encoder with crf","volume-title":"Thirty-Second AAAI Conference on Artificial Intelligence","author":"Kumar","year":"2018"},{"key":"2021110920491891600_bib30","doi-asserted-by":"crossref","first-page":"383","DOI":"10.18653\/v1\/K19-1036","article-title":"A dual-attention hierarchical recurrent neural network for dialogue act classification","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Li","year":"2019"},{"key":"2021110920491891600_bib31","doi-asserted-by":"publisher","first-page":"2170","DOI":"10.18653\/v1\/D17-1231","article-title":"Using context information for dialog act classification in DNN framework","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Liu","year":"2017"},{"key":"2021110920491891600_bib32","doi-asserted-by":"publisher","first-page":"2170","DOI":"10.18653\/v1\/D17-1231","article-title":"Using context information for dialog act classification in dnn framework","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Liu","year":"2017"},{"key":"2021110920491891600_bib33","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019","journal-title":"arXiv preprint arXiv:1907.11692 [v1]"},{"key":"2021110920491891600_bib34","doi-asserted-by":"publisher","first-page":"247","DOI":"10.18653\/v1\/W17-5530","article-title":"Neural-based context representation learning for dialog act classification","volume-title":"Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue","author":"Ortega","year":"2017"},{"key":"2021110920491891600_bib35","doi-asserted-by":"publisher","first-page":"6194","DOI":"10.1109\/ICASSP.2018.8461371","article-title":"Lexico-acoustic neural-based models for dialog act classification","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Ortega","year":"2018"},{"key":"2021110920491891600_bib36","article-title":"Dialog intent structure: A hierarchical schema of linked dialog acts","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)","author":"Pareti","year":"2018"},{"issue":"2","key":"2021110920491891600_bib37","doi-asserted-by":"publisher","first-page":"425","DOI":"10.1111\/tops.12324","article-title":"Computational models of miscommunication phenomena","volume":"10","author":"Purver","year":"2018","journal-title":"Topics in Cognitive Science"},{"key":"2021110920491891600_bib38","doi-asserted-by":"publisher","first-page":"262","DOI":"10.3115\/1708376.1708413","article-title":"Split utterances in dialogue: A corpus study","volume-title":"Proceedings of the SIGDIAL 2009 Conference","author":"Purver","year":"2009"},{"key":"2021110920491891600_bib39","doi-asserted-by":"publisher","first-page":"5596","DOI":"10.1109\/ICASSP.2011.5947628","article-title":"Simultaneous dialog act segmentation and classification from human-human spoken conversations","volume-title":"2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Quarteroni","year":"2011"},{"key":"2021110920491891600_bib40","first-page":"3727","article-title":"Dialogue act classification with context-aware self-attention","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Raheja","year":"2019"},{"key":"2021110920491891600_bib41","doi-asserted-by":"publisher","DOI":"10.3115\/980432.980757","article-title":"Dialogue act tagging with transformation-based learning","volume-title":"COLING 1998 Volume 2: The 17th International Conference on Computational Linguistics","author":"Samuel","year":"1998"},{"key":"2021110920491891600_bib42","doi-asserted-by":"publisher","first-page":"1715","DOI":"10.18653\/v1\/P16-1162","article-title":"Neural machine translation of rare words with subword units","volume-title":"Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sennrich","year":"2016"},{"key":"2021110920491891600_bib43","article-title":"Multi-task learning for domain-general spoken disfluency detection in dialogue systems","author":"Shalyminov","year":"2018","journal-title":"The 22nd Workshop on the Semantics and Pragmatics of Dialogue SEMDIAL"},{"key":"2021110920491891600_bib44","doi-asserted-by":"publisher","first-page":"3531","DOI":"10.1109\/WACV48630.2021.00357","article-title":"Efficient attention: Attention with linear complexities","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Shen","year":"2021"},{"key":"2021110920491891600_bib45","doi-asserted-by":"crossref","unstructured":"Elizabeth Shriberg , RajDhillon, SonaliBhagat, JeremyAng, and HannahCarvey. 2004. The ICSI meeting recorder dialog act (mrda) corpus, International Computer Science Inst, Berkeley, CA. 10.21236\/ADA460980","DOI":"10.21236\/ADA460980"},{"key":"2021110920491891600_bib46","doi-asserted-by":"publisher","first-page":"7994","DOI":"10.1109\/ICASSP40776.2020.9053061","article-title":"A hierarchical model for dialog act recognition considering acoustic and lexical context information","volume-title":"ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Si","year":"2020"},{"issue":"1","key":"2021110920491891600_bib47","article-title":"How do we speak with Alexa: Subjective and objective assessments of changes in speaking style between HC and HH conversations","volume":"2018","author":"Siegert","year":"2018","journal-title":"Kognitive Systeme"},{"issue":"3","key":"2021110920491891600_bib48","doi-asserted-by":"publisher","first-page":"339","DOI":"10.1162\/089120100561737","article-title":"Dialogue act modeling for automatic tagging and recognition of conversational speech","volume":"26","author":"Stolcke","year":"2000","journal-title":"Computational Linguistics"},{"key":"2021110920491891600_bib49","first-page":"9438","article-title":"Sparse sinkhorn attention","volume-title":"International Conference on Machine Learning","author":"Yi","year":"2020"},{"key":"2021110920491891600_bib50","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"key":"2021110920491891600_bib51","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1109\/SLT.2006.326819","article-title":"Dialogue-act tagging using smart feature selection; results on multiple corpora","volume-title":"2006 IEEE Spoken Language Technology Workshop","author":"Verbree","year":"2006"},{"key":"2021110920491891600_bib52","article-title":"Linformer: Self-attention with linear complexity","author":"Wang","year":"2020","journal-title":"arXiv preprint arXiv:2006.04768 [v3]"},{"key":"2021110920491891600_bib53","article-title":"Lite transformer with long- short range attention","author":"Zhanghao","year":"2020","journal-title":"Proceedings of International Conference on Learning Representations (ICLR)"},{"key":"2021110920491891600_bib54","first-page":"5754","article-title":"XLNet: Generalized autoregressive pretraining for language understanding","volume-title":"Advances in Neural Information Processing Systems","author":"Yang","year":"2019"},{"key":"2021110920491891600_bib55","article-title":"BP- Transformer: Modelling long-range context via binary partitioning","author":"Ye","year":"2019","journal-title":"arXiv preprint arXiv: 1911.04070 [v1]"},{"key":"2021110920491891600_bib56","article-title":"Big bird: Transformers for longer sequences","volume":"33","author":"Zaheer","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2021110920491891600_bib57","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1016\/j.csl.2019.03.001","article-title":"Joint dialog act segmentation and recognition in human conversations using attention to dialog context","volume":"57","author":"Zhao","year":"2019","journal-title":"Computer Speech & Language"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00420\/1971801\/tacl_a_00420.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00420\/1971801\/tacl_a_00420.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,14]],"date-time":"2023-01-14T14:47:43Z","timestamp":1673707663000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00420\/107831\/What-Helps-Transformers-Recognize-Conversational"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021]]},"references-count":57,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00420","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021]]},"published":{"date-parts":[[2021]]}}}