{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T12:43:12Z","timestamp":1761396192781,"version":"3.41.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T00:00:00Z","timestamp":1641600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Crystal Photonics Inc","award":["1063271"],"award-info":[{"award-number":["1063271"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,8,31]]},"abstract":"<jats:p>\n            The syntactic structure of sentences in a document substantially informs about its authorial writing style. Sentence representation learning has been widely explored in recent years and it has been shown that it improves the generalization of different downstream tasks across many domains. Even though utilizing probing methods in several studies suggests that these learned contextual representations implicitly encode some amount of syntax, explicit syntactic information further improves the performance of deep neural models in the domain of authorship attribution. These observations have motivated us to investigate the explicit representation learning of syntactic structure of sentences. In this article, we propose a self-supervised framework for learning structural representations of sentences. The self-supervised network contains two components; a lexical sub-network and a syntactic sub-network which take the sequence of words and their corresponding structural labels as the input, respectively. Due to the\n            <jats:italic>n<\/jats:italic>\n            -to-1 mapping of words to their structural labels, each word will be embedded into a vector representation which mainly carries structural information. We evaluate the learned structural representations of sentences using different probing tasks, and subsequently utilize them in the authorship attribution task. Our experimental results indicate that the structural embeddings significantly improve the classification tasks when concatenated with the existing pre-trained word embeddings.\n          <\/jats:p>","DOI":"10.1145\/3491203","type":"journal-article","created":{"date-parts":[[2022,1,8]],"date-time":"2022-01-08T20:51:00Z","timestamp":1641675060000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["A Self-Supervised Representation Learning of Sentence Structure for Authorship Attribution"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7836-3771","authenticated-orcid":false,"given":"Fereshteh","family":"Jafariakinabad","sequence":"first","affiliation":[{"name":"University of Central Florida, Orlando, Florida, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kien A.","family":"Hua","sequence":"additional","affiliation":[{"name":"University of Central Florida, Orlando, Florida, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,1,8]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2012.34"},{"key":"e_1_3_2_3_2","first-page":"1","volume-title":"Proceedings of the AAAI Workshop on Text Categorization","author":"Argamon-Engelson Shlomo","year":"1998","unstructured":"Shlomo Argamon-Engelson, Moshe Koppel, and Galit Avneri. 1998. Style-based text categorization: What newspaper am I reading. In Proceedings of the AAAI Workshop on Text Categorization. 1\u20134."},{"key":"e_1_3_2_4_2","article-title":"Authorship clustering using multi-headed recurrent neural networks","author":"Bagnall Douglas","year":"2016","unstructured":"Douglas Bagnall. 2016. Authorship clustering using multi-headed recurrent neural networks. arXiv:1608.04485 . Retrieved from https:\/\/arxiv.org\/abs\/1608.04485","journal-title":"arXiv:1608.04485"},{"key":"e_1_3_2_5_2","article-title":"Neural machine translation by jointly learning to align and translate","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473. Retrieved from https:\/\/arxiv.org\/abs\/1409.0473","journal-title":"arXiv:1409.0473"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1602"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-2003"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2382448.2382450"},{"key":"e_1_3_2_10_2","article-title":"Universal sentence encoder","author":"Cer Daniel","year":"2018","unstructured":"Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018. Universal sentence encoder. arXiv:1803.11175 . Retrieved from https:\/\/arxiv.org\/abs\/1803.11175","journal-title":"arXiv:1803.11175"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.202"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1070"},{"key":"e_1_3_2_13_2","first-page":"2126","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Conneau Alexis","year":"2018","unstructured":"Alexis Conneau, German Kruszewski, Guillaume Lample, Lo\u00efc Barrault, and Marco Baroni. 2018. What you can cram into a single \\backslash \\&\\!\\#\\* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. Vol. 1, Association for Computational Linguistics, 2126\u20132136."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-37256-8_37"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/3016387.3016522"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1115"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/UEMCON51285.2020.9298158"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDMW51313.2020.00071"},{"key":"e_1_3_2_19_2","first-page":"4129","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Hewitt John","year":"2019","unstructured":"John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4129\u20134138."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W17-4907"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/SSCI.2016.7849940"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2019.00061"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the 33rd International Flairs Conference","author":"Jafariakinabad Fereshteh","year":"2020","unstructured":"Fereshteh Jafariakinabad, Sansiri Tarnpradab, and Kien A. Hua. 2020. Syntactic neural model for authorship attribution. In Proceedings of the 33rd International Flairs Conference."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-6010"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3070645"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-40943-2_17"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969607"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2556325.2567881"},{"key":"e_1_3_2_29_2","first-page":"191","volume-title":"Proceedings of the 5th Workshop on NLP for Similar Languages, Varieties and Dialects","author":"Kreutz Tim","year":"2018","unstructured":"Tim Kreutz and Walter Daelemans. 2018. Exploring classifier combinations for language variety identification. In Proceedings of the 5th Workshop on NLP for Similar Languages, Varieties and Dialects. 191\u2013198."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1132"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1064"},{"key":"e_1_3_2_32_2","article-title":"A structured self-attentive sentence embedding","author":"Lin Zhouhan","year":"2017","unstructured":"Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017. A structured self-attentive sentence embedding. arXiv:1703.03130. Retrieved from https:\/\/arxiv.org\/abs\/1703.03130.","journal-title":"arXiv:1703.03130"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1085"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.5555\/2999792.2999959"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1177\/0146167203029005010"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372923.3404790"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1144"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1037\/0022-3514.77.6.1296"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_41_2","article-title":"Evaluation of sentence embeddings in downstream and linguistic probing tasks","author":"Perone Christian S.","year":"2018","unstructured":"Christian S. Perone, Roberto Silveira, and Thomas S. Paula. 2018. Evaluation of sentence embeddings in downstream and linguistic probing tasks. arXiv:1806.06259. Retrieved from https:\/\/arxiv.org\/abs\/1806.06259.","journal-title":"arXiv:1806.06259"},{"key":"e_1_3_2_42_2","article-title":"Syntactic n-grams as features for the author profiling task","author":"Posadas-Dur\u00e1n Juan-Pablo","year":"2015","unstructured":"Juan-Pablo Posadas-Dur\u00e1n, Ilia Markov, Helena G\u00f3mez-Adorno, Grigori Sidorov, Ildar Batyrshin, Alexander Gelbukh, and Obdulia Pichardo-Lagunas. 2015. Syntactic n-grams as features for the author profiling task. In Proceedings of the Working Notes Papers of the CLEF.","journal-title":"Proceedings of the Working Notes Papers of the CLEF"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.5555\/1858842.1858850"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2043"},{"key":"e_1_3_2_45_2","first-page":"199","volume-title":"Proceedings of the AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs","volume":"6","author":"Schler Jonathan","year":"2006","unstructured":"Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James W. Pennebaker. 2006. Effects of age and gender on blogging. In Proceedings of the AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs. Vol. 6, 199\u2013205."},{"key":"e_1_3_2_46_2","article-title":"The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task","author":"Schwartz Roy","year":"2017","unstructured":"Roy Schwartz, Maarten Sap, Ioannis Konstas, Li Zilles, Yejin Choi, and Noah A. Smith. 2017. The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task. arXiv:1702.01841. Retrieved from https:\/\/arxiv.org\/abs\/1702.01841.","journal-title":"arXiv:1702.01841"},{"key":"e_1_3_2_47_2","article-title":"Neural language modeling by jointly learning syntax and lexicon","author":"Shen Yikang","year":"2017","unstructured":"Yikang Shen, Zhouhan Lin, Chin-Wei Huang, and Aaron Courville. 2017. Neural language modeling by jointly learning syntax and lexicon. arXiv:1711.02013. Retrieved from https:\/\/arxiv.org\/abs\/1711.02013.","journal-title":"arXiv:1711.02013"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2106"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2108"},{"key":"e_1_3_2_50_2","first-page":"1717","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Song Kaiqiang","year":"2018","unstructured":"Kaiqiang Song, Lin Zhao, and Fei Liu. 2018. Structure-infused copy mechanisms for abstractive summarization. In Proceedings of the 27th International Conference on Computational Linguistics. 1717\u20131729."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2007.05.012"},{"key":"e_1_3_2_52_2","first-page":"2814","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Sundararajan Kalaivani","year":"2018","unstructured":"Kalaivani Sundararajan and Damon Woodard. 2018. What represents \u201cstyle\u201d in authorship attribution? In Proceedings of the 27th International Conference on Computational Linguistics. 2814\u20132822."},{"key":"e_1_3_2_53_2","article-title":"What do you learn from context? Probing for sentence structure in contextualized word representations","author":"Tenney Ian","year":"2019","unstructured":"Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019. What do you learn from context? Probing for sentence structure in contextualized word representations. arXiv:1905.06316. Retrieved from https:\/\/arxiv.org\/abs\/1905.06316.","journal-title":"arXiv:1905.06316"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.18653\/v1\/W17-1224","volume-title":"Proceedings of the 4th Workshop on NLP for Similar Languages, Varieties and Dialects","author":"Lee Chris van der","year":"2017","unstructured":"Chris van der Lee and Antal van den Bosch. 2017. Exploring lexical and syntactic features for language variety identification. In Proceedings of the 4th Workshop on NLP for Similar Languages, Varieties and Dialects. 190\u2013199."},{"key":"e_1_3_2_55_2","article-title":"\u201cLiar, Liar Pants on Fire\u201d: A new benchmark dataset for fake news detection","author":"Wang William Yang","year":"2017","unstructured":"William Yang Wang. 2017. \u201cLiar, Liar Pants on Fire\u201d: A new benchmark dataset for fake news detection. arXiv:1705.00648. Retrieved from https:\/\/arxiv.org\/abs\/1705.00648.","journal-title":"arXiv:1705.00648"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/2747880"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1118"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1294"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491203","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3491203","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:56Z","timestamp":1750188656000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491203"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,8]]},"references-count":57,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,8,31]]}},"alternative-id":["10.1145\/3491203"],"URL":"https:\/\/doi.org\/10.1145\/3491203","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2022,1,8]]},"assertion":[{"value":"2020-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}