{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,27]],"date-time":"2026-04-27T22:53:05Z","timestamp":1777330385536,"version":"3.51.4"},"reference-count":33,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T00:00:00Z","timestamp":1721088000000},"content-version":"vor","delay-in-days":197,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,7,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Authorship verification is the task of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text. In this paper, we explore the attribution of transcribed speech, which poses novel challenges. The main challenge is that many stylistic features, such as punctuation and capitalization, are not informative in this setting. On the other hand, transcribed speech exhibits other patterns, such as filler words and backchannels (e.g., um, uh-huh), which may be characteristic of different speakers. We propose a new benchmark for speaker attribution focused on human-transcribed conversational speech transcripts. To limit spurious associations of speakers with topic, we employ both conversation prompts and speakers participating in the same conversation to construct verification trials of varying difficulties. We establish the state of the art on this new benchmark by comparing a suite of neural and non-neural baselines, finding that although written text attribution models achieve surprisingly good performance in certain settings, they perform markedly worse as conversational topic is increasingly controlled. We present analyses of the impact of transcription style on performance as well as the ability of fine-tuning on speech transcripts to improve performance.1<\/jats:p>","DOI":"10.1162\/tacl_a_00678","type":"journal-article","created":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T18:36:04Z","timestamp":1721154964000},"page":"875-891","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":5,"title":["Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?"],"prefix":"10.1162","volume":"12","author":[{"given":"Cristina","family":"Aggazzotti","sequence":"first","affiliation":[{"name":"Johns Hopkins University, USA. caggazz1@jhu.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicholas","family":"Andrews","sequence":"additional","affiliation":[{"name":"Johns Hopkins University, USA. noa@jhu.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elizabeth Allyn","family":"Smith","sequence":"additional","affiliation":[{"name":"Universit\u00e9 du Qu\u00e9bec \u00e0 Montr\u00e9al, Canada. smith.elizabeth_allyn@uqam.ca"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2024,7,15]]},"reference":[{"key":"2024071618354074900_bib1","first-page":"69","article-title":"An experiment in authorship attribution","volume-title":"6es Journ\u00e9es Internationales d\u2019Analyse Statistique des Donn\u00e9es Textuelles (JADT)","author":"Baayen","year":"2002"},{"key":"2024071618354074900_bib2","doi-asserted-by":"publisher","first-page":"372","DOI":"10.1007\/978-3-030-58219-7_25","article-title":"Overview of PAN 2020: Authorship verification, celebrity profiling, profiling fake news spreaders on Twitter, and style change detection","volume-title":"Experimental IR Meets Multilinguality, Multimodality, and Interaction","author":"Bevendorff","year":"2020"},{"key":"2024071618354074900_bib3","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1109\/BigData47090.2019.9005650","article-title":"Explainable authorship verification in social media via attention-based similarity learning","volume-title":"IEEE International Conference on Big Data (Big Data)","author":"Boenninghoff","year":"2019"},{"key":"2024071618354074900_bib4","doi-asserted-by":"publisher","DOI":"10.35111\/w4bk-9b14","article-title":"The Fisher Corpus: A resource for the next generations of speech-to-text","author":"Cieri","year":"2004"},{"key":"2024071618354074900_bib5","doi-asserted-by":"publisher","first-page":"745","DOI":"10.1145\/1963405.1963509","article-title":"Mark my words! Linguistic style accommodation in social media","volume-title":"Proceedings of the 20th International Conference on World Wide Web","author":"Danescu-Niculescu-Mizil","year":"2011"},{"key":"2024071618354074900_bib6","article-title":"509 U.S. 579","author":"Daubert v. Merrell Dow Pharmaceuticals, Inc.","year":"1993"},{"issue":"1","key":"2024071618354074900_bib7","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1109\/TCYB.2017.2766189","article-title":"Learning stylometric representations for authorship analysis","volume":"49","author":"Ding","year":"2019","journal-title":"IEEE Transactions on Cybernetics"},{"issue":"2","key":"2024071618354074900_bib8","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1017\/S0047404500004322","article-title":"On the structure of speaker-auditor interaction during speaking turns","volume":"3","author":"Duncan","year":"1974","journal-title":"Language in Society"},{"key":"2024071618354074900_bib9","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2019-28","article-title":"Speaker anonymization using x-vector and neural waveform models","author":"Fang","year":"2019"},{"key":"2024071618354074900_bib10","doi-asserted-by":"publisher","DOI":"10.1016\/j.langsci.2023.101571","article-title":"Communication accommodation theory: Past accomplishments, current trends, and future prospects","volume":"99","author":"Giles","year":"2023","journal-title":"Language Sciences"},{"key":"2024071618354074900_bib11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1558\/ijsll.38028","article-title":"International practices in forensic speaker comparisons: Second survey","volume":"26","author":"Gold","year":"2019","journal-title":"International Journal of Speech, Language and the Law"},{"key":"2024071618354074900_bib12","doi-asserted-by":"publisher","first-page":"336","DOI":"10.3115\/1609067.1609104","article-title":"Person identification from text and speech genre samples","volume-title":"Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009)","author":"Goldstein-Stewart","year":"2009"},{"key":"2024071618354074900_bib13","first-page":"1312","article-title":"Identification of speakers in novels","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"He","year":"2013"},{"issue":"3","key":"2024071618354074900_bib14","doi-asserted-by":"publisher","first-page":"340","DOI":"10.1080\/0013838X.2012.668793","article-title":"Cross-genre authorship verification using unmasking","volume":"93","author":"Kestemont","year":"2012","journal-title":"English Studies"},{"key":"2024071618354074900_bib15","doi-asserted-by":"publisher","first-page":"5275","DOI":"10.18653\/v1\/2021.naacl-main.415","article-title":"A deep metric learning approach to account linking","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Khan","year":"2021"},{"key":"2024071618354074900_bib16","article-title":"Quick transcription of Fisher data with WordWave","author":"Kimball","year":".."},{"key":"2024071618354074900_bib17","article-title":"Reuters-21578 text categorization test collection, Distribution 1.0","author":"Lewis","year":"1997"},{"key":"2024071618354074900_bib18","volume-title":"Inference and Disputed Authorship: The Federalist","author":"Mosteller","year":"1964"},{"key":"2024071618354074900_bib19","article-title":"Text-to-text transformer in authorship verification via stylistic and semantical analysis","volume-title":"Notebook for PAN at CLEF 2022","author":"Najafi","year":"2022"},{"key":"2024071618354074900_bib20","doi-asserted-by":"publisher","DOI":"10.1016\/j.wocn.2022.101196","article-title":"Vocal accommodation in speech communication","volume":"95","author":"Pardo","year":"2022","journal-title":"Journal of Phonetics"},{"key":"2024071618354074900_bib21","doi-asserted-by":"publisher","first-page":"3982","DOI":"10.18653\/v1\/D19-1410","article-title":"Sentence-BERT: Sentence embeddings using Siamese BERT-Networks","author":"Reimers","year":"2019"},{"key":"2024071618354074900_bib22","doi-asserted-by":"publisher","first-page":"913","DOI":"10.18653\/v1\/2021.emnlp-main.70","article-title":"Learning universal authorship representations","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Rivera-Soto","year":"2021"},{"key":"2024071618354074900_bib23","volume-title":"Lectures on Conversation","author":"Sacks","year":"1992"},{"key":"2024071618354074900_bib24","unstructured":"Nelleke\n              Scheijen\n            \n          . 2020. Forensic speaker recognition: Based on text analysis of transcribed speech fragments. Master\u2019s thesis, Delft University of Technology."},{"key":"2024071618354074900_bib25","doi-asserted-by":"publisher","first-page":"132","DOI":"10.1109\/TASLP.2020.3038524","article-title":"An overview of voice conversion and its challenges: From statistical modeling to deep learning","volume":"29","author":"Sisman","year":"2021","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2024071618354074900_bib26","doi-asserted-by":"publisher","first-page":"5329","DOI":"10.1109\/ICASSP.2018.8461375","article-title":"X-vectors: Robust DNN embeddings for speaker recognition","volume-title":"2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Snyder","year":"2018"},{"issue":"3","key":"2024071618354074900_bib27","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1002\/asi.23968","article-title":"Masking topic-related information to enhance authorship attribution","volume":"69","author":"Stamatatos","year":"2018","journal-title":"Journal of the Association for Information Science and Technology"},{"key":"2024071618354074900_bib28","article-title":"Overview of the authorship verification task at PAN 2023","volume-title":"CLEF 2023: Conference and Labs of the Evaluation Forum","author":"Stamatatos","year":"2023"},{"key":"2024071618354074900_bib29","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.916","article-title":"HANSEN: Human and AI spoken text benchmark for authorship analysis","author":"Tripto","year":"2023"},{"key":"2024071618354074900_bib30","doi-asserted-by":"publisher","first-page":"1416","DOI":"10.1162\/tacl_a_00610","article-title":"Can authorship representation learning capture stylistic features?","volume":"11","author":"Wang","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2024071618354074900_bib31","doi-asserted-by":"publisher","DOI":"10.4324\/9780429030581-32","article-title":"Forensic phonetics and automatic speaker recognition: The complementarity of human- and machine-based forensic speaker comparison","volume-title":"The Routledge Handbook of Forensic Linguistics","author":"Watt","year":"2020"},{"key":"2024071618354074900_bib32","doi-asserted-by":"publisher","first-page":"249","DOI":"10.18653\/v1\/2022.repl4nlp-1.26","article-title":"Same author or just same topic? Towards content-independent style representations","volume-title":"Proceedings of the 7th Workshop on Representation Learning for NLP","author":"Wegmann","year":"2022"},{"key":"2024071618354074900_bib33","doi-asserted-by":"publisher","first-page":"279","DOI":"10.18653\/v1\/2021.emnlp-main.25","article-title":"Idiosyncratic but not arbitrary: Learning idiolects in online registers reveals distinctive yet consistent individual styles","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Zhu","year":"2021"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00678\/2461933\/tacl_a_00678.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00678\/2461933\/tacl_a_00678.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T18:36:13Z","timestamp":1721154973000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00678\/123650\/Can-Authorship-Attribution-Models-Distinguish"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":33,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00678","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}