{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T23:30:13Z","timestamp":1763767813875},"reference-count":31,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"vor","delay-in-days":330,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,11,16]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Automatically disentangling an author\u2019s style from the content of their writing is a longstanding and possibly insurmountable problem in computational linguistics. At the same time, the availability of large text corpora furnished with author labels has recently enabled learning authorship representations in a purely data-driven manner for authorship attribution, a task that ostensibly depends to a greater extent on encoding writing style than encoding content. However, success on this surrogate task does not ensure that such representations capture writing style since authorship could also be correlated with other latent variables, such as topic. In an effort to better understand the nature of the information these representations convey, and specifically to validate the hypothesis that they chiefly encode writing style, we systematically probe these representations through a series of targeted experiments. The results of these experiments suggest that representations learned for the surrogate authorship prediction task are indeed sensitive to writing style. As a consequence, authorship representations may be expected to be robust to certain kinds of data shift, such as topic drift over time. Additionally, our findings may open the door to downstream applications that require stylistic representations, such as style transfer.<\/jats:p>","DOI":"10.1162\/tacl_a_00610","type":"journal-article","created":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T16:48:02Z","timestamp":1701103682000},"page":"1416-1431","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Can Authorship Representation Learning Capture Stylistic Features?"],"prefix":"10.1162","volume":"11","author":[{"given":"Andrew","family":"Wang","sequence":"first","affiliation":[{"name":"Johns Hopkins University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cristina","family":"Aggazzotti","sequence":"additional","affiliation":[{"name":"Johns Hopkins University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rebecca","family":"Kotula","sequence":"additional","affiliation":[{"name":"Department of Defense, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rafael Rivera","family":"Soto","sequence":"additional","affiliation":[{"name":"Lawrence Livermore National Laboratory, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcus","family":"Bishop","sequence":"additional","affiliation":[{"name":"Department of Defense, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicholas","family":"Andrews","sequence":"additional","affiliation":[{"name":"Johns Hopkins University, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,11,16]]},"reference":[{"key":"2023112716473327300_bib1","doi-asserted-by":"publisher","first-page":"1684","DOI":"10.18653\/v1\/D19-1178","article-title":"Learning invariant representations of social media users","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Andrews","year":"2019"},{"key":"2023112716473327300_bib2","doi-asserted-by":"publisher","first-page":"830","DOI":"10.1609\/icwsm.v14i1.7347","article-title":"The Pushshift Reddit dataset","volume-title":"Proceedings of the 14th International AAAI Conference on Web and Social Media (ICWSM)","author":"Baumgartner","year":"2020"},{"key":"2023112716473327300_bib3","doi-asserted-by":"publisher","first-page":"372","DOI":"10.1007\/978-3-030-58219-7_25","article-title":"Overview of PAN 2020: Authorship verification, celebrity profiling, profiling fake news spreaders on Twitter, and style change detection","volume-title":"Experimental IR Meets Multilinguality, Multimodality, and Interaction","author":"Bevendorff","year":"2020"},{"key":"2023112716473327300_bib4","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1109\/BigData47090.2019.9005650","article-title":"Explainable authorship verification in social media via attention-based similarity learning","volume-title":"2019 IEEE International Conference on Big Data (Big Data)","author":"Boenninghoff","year":"2019"},{"key":"2023112716473327300_bib5","doi-asserted-by":"publisher","first-page":"3504","DOI":"10.1109\/TASLP.2021.3124365","article-title":"Pre-training with whole word masking for Chinese BERT","volume":"29","author":"Cui","year":"2021","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2023112716473327300_bib6","article-title":"Automatically constructing a corpus of sentential paraphrases","volume-title":"Third International Workshop on Paraphrasing (IWP2005)","author":"Dolan","year":"2005"},{"key":"2023112716473327300_bib7","doi-asserted-by":"publisher","first-page":"232","DOI":"10.18653\/v1\/2020.wnut-1.30","article-title":"Representation learning of writing style","volume-title":"Proceedings of the 6th Workshop on Noisy User-generated Text (W-NUT 2020)","author":"Hay","year":"2020"},{"key":"2023112716473327300_bib8","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2209.15373","article-title":"PART: Pre-trained Authorship Representation Transformer","volume":"cs.CL\/2209.15373v1","author":"Huertas-Tato","year":"2022"},{"key":"2023112716473327300_bib9","doi-asserted-by":"publisher","first-page":"2169","DOI":"10.18653\/v1\/2020.coling-main.197","article-title":"Style versus content: A distinction without a (learnable) difference?","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Jafaritazehjani","year":"2020"},{"issue":"1","key":"2023112716473327300_bib10","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1162\/coli_a_00426","article-title":"Deep learning for text style transfer: A survey","volume":"48","author":"Di","year":"2022","journal-title":"Computational Linguistics"},{"key":"2023112716473327300_bib11","article-title":"Supervised contrastive learning","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS)","author":"Khosla","year":"2020"},{"key":"2023112716473327300_bib12","first-page":"5530","article-title":"Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Kim","year":"2021"},{"key":"2023112716473327300_bib13","doi-asserted-by":"publisher","first-page":"737","DOI":"10.18653\/v1\/2020.emnlp-main.55","article-title":"Reformulating unsupervised style transfer as paraphrase generation","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Krishna","year":"2020"},{"key":"2023112716473327300_bib14","article-title":"Reuters-21578 text categorization test collection, Distribution 1.0","author":"Lewis","year":"1997"},{"key":"2023112716473327300_bib15","doi-asserted-by":"publisher","first-page":"1865","DOI":"10.18653\/v1\/N18-1169","article-title":"Delete, retrieve, generate: A simple approach to sentiment and style transfer","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Li","year":"2018"},{"key":"2023112716473327300_bib16","doi-asserted-by":"publisher","first-page":"1869","DOI":"10.18653\/v1\/2020.acl-main.169","article-title":"Politeness transfer: A tag and generate approach","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Madaan","year":"2020"},{"key":"2023112716473327300_bib17","volume-title":"Inference and Disputed Authorship: The Federalist","author":"Mosteller","year":"1964"},{"key":"2023112716473327300_bib18","doi-asserted-by":"publisher","first-page":"188","DOI":"10.18653\/v1\/D19-1018","article-title":"Justifying recommendations using distantly-labeled reviews and fine-grained aspects","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Ni","year":"2019"},{"key":"2023112716473327300_bib19","doi-asserted-by":"publisher","first-page":"4296","DOI":"10.18653\/v1\/2020.acl-main.396","article-title":"Toxicity detection: Does context really matter?","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Pavlopoulos","year":"2020"},{"key":"2023112716473327300_bib20","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1162\/tacl_a_00465","article-title":"Evaluating explanations: How much do explanations from the teacher aid students?","volume":"10","author":"Pruthi","year":"2022","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2023112716473327300_bib21","doi-asserted-by":"publisher","first-page":"101","DOI":"10.18653\/v1\/2020.acl-demos.14","article-title":"Stanza: A Python natural language processing toolkit for many human languages","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations","author":"Qi","year":"2020"},{"key":"2023112716473327300_bib22","doi-asserted-by":"publisher","first-page":"129","DOI":"10.18653\/v1\/N18-1012","article-title":"Dear Sir or Madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Rao","year":"2018"},{"key":"2023112716473327300_bib23","doi-asserted-by":"publisher","first-page":"913","DOI":"10.18653\/v1\/2021.emnlp-main.70","article-title":"Learning universal authorship representations","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Rivera-Soto","year":"2021"},{"key":"2023112716473327300_bib24","doi-asserted-by":"publisher","first-page":"101241","DOI":"10.1016\/j.csl.2021.101241","article-title":"Siamese networks for large-scale author identification","volume":"70","author":"Saedi","year":"2021","journal-title":"Computer Speech & Language"},{"key":"2023112716473327300_bib25","first-page":"343","article-title":"Topic or style? Exploring the most useful features for authorship attribution","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Sari","year":"2018"},{"issue":"3","key":"2023112716473327300_bib26","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1002\/asi.23968","article-title":"Masking topic-related information to enhance authorship attribution","volume":"69","author":"Stamatatos","year":"2018","journal-title":"Journal of the Association for Information Science and Technology"},{"key":"2023112716473327300_bib27","first-page":"3319","article-title":"Axiomatic attribution for deep networks","volume-title":"Proceedings of the 34th International Conference on Machine Learning (Volume 70)","author":"Sundararajan","year":"2017"},{"key":"2023112716473327300_bib28","doi-asserted-by":"publisher","first-page":"7109","DOI":"10.18653\/v1\/2021.emnlp-main.569","article-title":"Does it capture STEL? A modular, similarity-based linguistic style evaluation framework","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Wegmann","year":"2021"},{"key":"2023112716473327300_bib29","doi-asserted-by":"publisher","first-page":"249","DOI":"10.18653\/v1\/2022.repl4nlp-1.26","article-title":"Same author or just same topic? Towards content-independent style representations","volume-title":"Proceedings of the 7th Workshop on Representation Learning for NLP","author":"Wegmann","year":"2022"},{"key":"2023112716473327300_bib30","doi-asserted-by":"publisher","first-page":"451","DOI":"10.18653\/v1\/P18-1042","article-title":"ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1 (Long Papers)","author":"Wieting","year":"2018"},{"key":"2023112716473327300_bib31","article-title":"BERTScore: Evaluating text generation with BERT","author":"Zhang","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00610\/2184071\/tacl_a_00610.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00610\/2184071\/tacl_a_00610.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T16:48:45Z","timestamp":1701103725000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00610\/118299\/Can-Authorship-Representation-Learning-Capture"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":31,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00610","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}