{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T06:12:43Z","timestamp":1780380763582,"version":"3.54.1"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,6,19]],"date-time":"2024-06-19T00:00:00Z","timestamp":1718755200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSF","award":["2113850"],"award-info":[{"award-number":["2113850"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2024,8,31]]},"abstract":"<jats:p>In the financial sphere, there is a wealth of accumulated unstructured financial data, such as the textual disclosure documents that companies submit on a regular basis to regulatory agencies, such as the Securities and Exchange Commission. These documents are typically very long and tend to contain valuable soft information about a company\u2019s performance that is not present in quantitative predictors. It is therefore of great interest to learn predictive models from these long textual documents, especially for forecasting numerical key performance indicators. In recent years, there has been great progress in natural language processing via pre-trained language models (LMs) learned from large corpora of textual data. This prompts the important question of whether they can be used effectively to produce representations for long documents, as well as how we can evaluate the effectiveness of representations produced by various LMs. Our work focuses on answering this critical question, namely, the evaluation of the efficacy of various LMs in extracting useful soft information from long textual documents for prediction tasks. In this article, we propose and implement a deep learning evaluation framework that utilizes a sequential chunking approach combined with an attention mechanism. We perform an extensive set of experiments on a collection of 10-K reports submitted annually by U.S. banks, and another dataset of reports submitted by U.S. companies, to investigate thoroughly the performance of different types of language models. Overall, our framework using LMs outperforms strong baseline methods for textual modeling as well as for numerical regression. Our work provides better insights into how utilizing pre-trained domain-specific and fine-tuned long-input LMs for representing long documents can improve the quality of representation of textual data and, therefore, help in improving predictive analyses.<\/jats:p>","DOI":"10.1145\/3657299","type":"journal-article","created":{"date-parts":[[2024,4,10]],"date-time":"2024-04-10T12:25:43Z","timestamp":1712751943000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["FETILDA: Evaluation Framework for Effective Representations of Long Financial Documents"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6841-9017","authenticated-orcid":false,"given":"Bolun (Namir)","family":"Xia","sequence":"first","affiliation":[{"name":"Rensselaer Polytechnic Institute School of Science, Troy, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4355-1393","authenticated-orcid":false,"given":"Vipula","family":"Rawte","sequence":"additional","affiliation":[{"name":"Rensselaer Polytechnic Institute School of Science, Troy, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5275-7756","authenticated-orcid":false,"given":"Aparna","family":"Gupta","sequence":"additional","affiliation":[{"name":"Rensselaer Polytechnic Institute, Troy, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4711-0234","authenticated-orcid":false,"given":"Mohammed","family":"Zaki","sequence":"additional","affiliation":[{"name":"Rensselaer Polytechnic Institute, Troy, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,6,19]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.19"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","unstructured":"Amir Amel-Zadeh and Jonathan Faasse. 2016. The information content of 10-K narratives: comparing MD&A and footnotes disclosures. Retrieved Nov. 2 2019 from 10.2139\/ssrn.2807546","DOI":"10.2139\/ssrn.2807546"},{"key":"e_1_3_1_4_2","volume-title":"FinBERT: Financial Sentiment Analysis with Pre-trained Language Models","author":"Araci Dogu","year":"2019","unstructured":"Dogu Araci. 2019. FinBERT: Financial Sentiment Analysis with Pre-trained Language Models. Master\u2019s thesis. University of Amsterdam."},{"key":"e_1_3_1_5_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The long-document transformer. Retrieved from https:\/\/arXiv:2004.05150"},{"key":"e_1_3_1_6_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, 1877\u20131901. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf"},{"key":"e_1_3_1_7_2","first-page":"3216","volume-title":"Proceedings of the 26th International Conference on Computer Linguistics (COLING\u201916)","author":"Chang Ching Yun","year":"2016","unstructured":"Ching Yun Chang, Yue Zhang, Zhiyang Teng, Zahn Bozanic, and Bin Ke. 2016. Measuring the information content of financial news. In Proceedings of the 26th International Conference on Computer Linguistics (COLING\u201916). 3216\u20133225."},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201920)","author":"Clark Kevin","year":"2020","unstructured":"Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In Proceedings of the International Conference on Learning Representations (ICLR\u201920). Retrieved from https:\/\/openreview.net\/pdf?id=r1xMH1BtvB"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1561\/0500000045"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Vinicio Desola Kevin Hanna and Pri Nonis. 2019. FinBERT: Pre-trained model on SEC filings for financial natural language tasks. Retrieved from 10.13140\/RG.2.2.19153.89442","DOI":"10.13140\/RG.2.2.19153.89442"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jacceco.2017.07.002"},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the 8th International Conference on Economics and Finance Research (ICEFR\u201919)","author":"Emerson Sophie","year":"2019","unstructured":"Sophie Emerson, Ruair\u00ed Kennedy, Luke O\u2019Shea, and John O\u2019Brien. 2019. Trends and applications of machine learning in quantitative finance. In Proceedings of the 8th International Conference on Economics and Finance Research (ICEFR\u201919)."},{"key":"e_1_3_1_14_2","unstructured":"Jason Fernando. 2021. How return on equity (ROE) works. Retrieved from https:\/\/www.investopedia.com\/terms\/r\/returnonequity.asp"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.603"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eacl-main.154"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1111\/1911-3846.12832"},{"key":"e_1_3_1_18_2","volume-title":"Speech and Language Processing (3rd ed.)","author":"Jurafsky Dan","year":"2021","unstructured":"Dan Jurafsky and James H. Martin. 2021. Speech and Language Processing (3rd ed.). Retrieved from https:\/\/web.stanford.edu\/jurafsky\/slp3\/"},{"key":"e_1_3_1_19_2","unstructured":"Will Kenton. 2021. What you should know About 10-KS. Retrieved from https:\/\/www.investopedia.com\/terms\/1\/10-k.asp"},{"key":"e_1_3_1_20_2","unstructured":"Luckyson Khaidem Snehanshu Saha and Sudeepa Roy Dey. 2016. Predicting the direction of stock market prices using random forest. Retrieved from https:\/\/arXiv:1605.00003"},{"key":"e_1_3_1_21_2","unstructured":"Nikita Kitaev \u0141ukasz Kaiser and Anselm Levskaya. 2020. Reformer: The efficient transformer. Retrieved from https:\/\/arXiv:2001.04451"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.3115\/1620754.1620794"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983328"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/622"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1540-6261.2010.01625.x"},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913), Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_1_27_2","unstructured":"Vipal Monga and Emily Chasan. 2015. The 109 894-word annual report. Retrieved from https:\/\/www.wsj.com\/articles\/the-109-894-word-annual-report-1433203762"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2916793"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_30_2","article-title":"Improving language understanding by generative pre-training","author":"Radford Alec","year":"2018","unstructured":"Alec Radford and Karthik Narasimhan. 2018. Improving language understanding by generative pre-training. OpenAI Blog (June 11, 2018).","journal-title":"OpenAI Blog"},{"key":"e_1_3_1_31_2","article-title":"Language models are unsupervised multitask learners","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog (Feb. 14, 2019).","journal-title":"OpenAI Blog"},{"issue":"140","key":"e_1_3_1_32_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140 (2020), 1\u201367. http:\/\/jmlr.org\/papers\/v21\/20-074.html","journal-title":"J. Mach. Learn. Res."},{"issue":"1","key":"e_1_3_1_33_2","first-page":"17","article-title":"Use advances in data science and computing power to invest in stock market","volume":"2","author":"Sakarwala Mustafa A.","year":"2019","unstructured":"Mustafa A. Sakarwala and Anthony Tanaydin. 2019. Use advances in data science and computing power to invest in stock market. SMU Data Sci. Rev. 2, 1 (2019), 17.","journal-title":"SMU Data Sci. Rev."},{"key":"e_1_3_1_34_2","unstructured":"Marcelo Sardelich and Suresh Manandhar. 2018. Multimodal deep learning for short-term stock volatility prediction. Retrieved from https:\/\/arXiv:1812.10479"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1034"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1540-6261.2008.01362.x"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-3104"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343039"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ejor.2016.06.069"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2948072"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1002\/0471746193"},{"key":"e_1_3_1_42_2","first-page":"5998","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917).","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917).5998\u20136008."},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Yan Wang and Xuelei Sherry Ni. 2019. A XGBoost risk model via feature selection and Bayesian hyper-parameter optimization. Retrieved from https:\/\/arXiv:1901.08433","DOI":"10.5121\/ijdms.2019.11101"},{"key":"e_1_3_1_44_2","unstructured":"Han Xiao. 2018. bert-as-service. Retrieved Mar. 3 2020 from https:\/\/github.com\/hanxiao\/bert-as-service"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i16.17664"},{"key":"e_1_3_1_46_2","article-title":"FinGPT: Open-source financial large language models","author":"Yang Hongyang","year":"2023","unstructured":"Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang. 2023. FinGPT: Open-source financial large language models. In Proceedings of the FinLLM Symposium at the International Joint Conference on Artificial Intelligence (IJCAI\u201923).","journal-title":"Proceedings of the FinLLM Symposium at the International Joint Conference on Artificial Intelligence (IJCAI\u201923)"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3411908"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCIS.2018.8691233"},{"key":"e_1_3_1_49_2","first-page":"17283","volume-title":"Advances in Neural Information Processing Systems","author":"Zaheer Manzil","year":"2020","unstructured":"Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big bird: Transformers for longer sequences. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, 17283\u201317297. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/c8512d142a2d849725f31a9a7a361ab9-Paper.pdf"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1017\/9781108564175"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3657299","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3657299","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:39Z","timestamp":1750295859000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3657299"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,19]]},"references-count":49,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,8,31]]}},"alternative-id":["10.1145\/3657299"],"URL":"https:\/\/doi.org\/10.1145\/3657299","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,19]]},"assertion":[{"value":"2022-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}