{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T18:01:24Z","timestamp":1771264884494,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,2,13]],"date-time":"2024-02-13T00:00:00Z","timestamp":1707782400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U21B2038 and U19B2039"],"award-info":[{"award-number":["U21B2038 and U19B2039"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021ZD0111902"],"award-info":[{"award-number":["2021ZD0111902"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"R&D Program of Beijing Municipal Education Commission","award":["KZ202210005008"],"award-info":[{"award-number":["KZ202210005008"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>Long Document Classification (LDC) has attracted great attention in Natural Language Processing and achieved considerable progress owing to the large-scale pre-trained language models. In spite of this, as a different problem from the traditional text classification, LDC is far from being settled. Long documents, such as news and articles, generally have more than thousands of words with complex structures. Moreover, compared with flat text, long documents usually contain multi-modal content of images, which provide rich information but not yet being utilized for classification. In this article, we propose a novel cross-modal method for long document classification, in which multiple granularity feature shifting networks are proposed to integrate the multi-scale text and visual features of long documents adaptively. Additionally, a multi-modal collaborative pooling block is proposed to eliminate redundant fine-grained text features and simultaneously reduce the computational complexity. To verify the effectiveness of the proposed model, we conduct experiments on the Food101 dataset and two constructed multi-modal long document datasets. The experimental results show that the proposed cross-modal method outperforms the single-modal text methods and defeats the state-of-the-art related multi-modal baselines.<\/jats:p>","DOI":"10.1145\/3631711","type":"journal-article","created":{"date-parts":[[2023,11,6]],"date-time":"2023-11-06T11:27:36Z","timestamp":1699270056000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Cross-modal Multiple Granularity Interactive Fusion Network for Long Document Classification"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2739-5220","authenticated-orcid":false,"given":"Tengfei","family":"Liu","sequence":"first","affiliation":[{"name":"Beijing University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0440-438X","authenticated-orcid":false,"given":"Yongli","family":"Hu","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9803-0256","authenticated-orcid":false,"given":"Junbin","family":"Gao","sequence":"additional","affiliation":[{"name":"The University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0872-384X","authenticated-orcid":false,"given":"Yanfeng","family":"Sun","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4164-6647","authenticated-orcid":false,"given":"Baocai","family":"Yin","sequence":"additional","affiliation":[{"name":"Beijing University of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,2,13]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies. 4046\u20134051","author":"Adhikari Ashutosh","unstructured":"Ashutosh Adhikari, Achyudh Ram, Raphael Tang, and Jimmy J. Lin. 2019. Rethinking complex neural network architectures for document classification. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies. 4046\u20134051."},{"key":"e_1_3_2_3_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 268\u2013284","author":"Ainslie Joshua","year":"2020","unstructured":"Joshua Ainslie, Santiago Onta\u00f1\u00f3n, Chris Alberti, Vaclav Cvicek, Zachary Kenneth Fisher, Philip Pham, Anirudh Ravula, Sumit K. Sanghai, Qifan Wang, and Li Yang. 2020. ETC: Encoding long and structured inputs in transformers. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 268\u2013284."},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies","volume":"3","author":"Ammar Waleed","year":"2018","unstructured":"Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu A. Ha, Rodney Michael Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler C. Murray, Hsu-Han Ooi, Matthew E. Peters, Joanna L. Power, Sam Skjonsberg, Lucy Lu Wang, Christopher Wilhelm, Zheng Yuan, Madeleine van Zuylen, and Oren Etzioni. 2018. Construction of the literature graph in semantic scholar. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, Vol. 3. 84\u201391."},{"key":"e_1_3_2_5_2","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Arevalo John","unstructured":"John Arevalo, Thamar Solorio, Manuel Montes y G\u00f3mez, and Fabio A. Gonz\u00e1lez. 2017. Gated multimodal units for information fusion. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_6_2","volume-title":"Longformer: The long-document transformer. arXiv:2004.05150.","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.3316\/QRJ0902027"},{"key":"e_1_3_2_8_2","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33. 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the Conference of North American Chapter of the Association for Computational Linguistics. 5881\u20135891","author":"Cui Peng","year":"2021","unstructured":"Peng Cui and Le Hu. 2021. Sliding selector network with dynamic memory for extractive summarization of long documents. In Proceedings of the Conference of North American Chapter of the Association for Computational Linguistics. 5881\u20135891."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_12_2","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies, Vol. 1. 4171\u20134186."},{"key":"e_1_3_2_13_2","first-page":"12792","article-title":"CogLTX: Applying BERT to long texts","volume":"33","author":"Ding Ming","year":"2020","unstructured":"Ming Ding, Chang Zhou, Hongxia Yang, and Jie Tang. 2020. CogLTX: Applying BERT to long texts. In Advances in Neural Information Processing Systems, Vol. 33. 12792\u201312804.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.710"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772731"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3003648"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.723"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413678"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2872427.2883037"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.670"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"32","author":"Kiela Douwe","year":"2018","unstructured":"Douwe Kiela, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2018. Efficient large-scale multi-modal classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32. 5198\u20135204."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_24_2","volume-title":"Proceedings of the International Conference on Learning Representations. 1\u201312","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. In Proceedings of the International Conference on Learning Representations. 1\u201312."},{"key":"e_1_3_2_25_2","first-page":"1","article-title":"DiMBERT: Learning vision-language grounded representations with disentangled multimodal-attention","volume":"16","author":"Liu Fenglin","year":"2021","unstructured":"Fenglin Liu, Xian Wu, Shen Ge, Xuancheng Ren, Wei Fan, Xu Sun, and Yuexian Zou. 2021. DiMBERT: Learning vision-language grounded representations with disentangled multimodal-attention. ACM Trans. Knowl. Discov. Data 16, 1 (2021), 1\u201319.","journal-title":"ACM Trans. Knowl. Discov. Data"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080834"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2976493"},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Liu Peter J.","unstructured":"Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam M. Shazeer. 2018a. Generating wikipedia by summarizing long sequences. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1209"},{"key":"e_1_3_2_30_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 3219\u20133232","author":"Luan Yi","year":"2018","unstructured":"Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 3219\u20133232."},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics. 142\u2013150","author":"Maas Andrew L.","year":"2011","unstructured":"Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, A. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics. 142\u2013150."},{"key":"e_1_3_2_32_2","volume-title":"Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU \u201919)","author":"Pappagari R.","year":"2019","unstructured":"R. Pappagari, Piotr \u017belasko, Jes\u00fas Villalba, Yishay Carmiel, and Najim Dehak. 2019. Hierarchical transformers for long document classification. In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU \u201919). 838\u2013844."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 2555\u20132565","author":"Qiu Jiezhong","year":"2020","unstructured":"Jiezhong Qiu, Hao Ma, Omer Levy, Scott Yih, Sinong Wang, and Jie Tang. 2020. Blockwise self-attention for long document understanding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 2555\u20132565."},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Rae Jack W.","unstructured":"Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap. 2020. Compressive transformers for long-range sequence modelling. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00353"},{"key":"e_1_3_2_37_2","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 305\u2013312","author":"Truong Quoc-Tuan","unstructured":"Quoc-Tuan Truong and Hady W. Lauw. 2019. VistaNet: Visual aspect attention network for multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence. 305\u2013312."},{"key":"e_1_3_2_39_2","unstructured":"Ashish Vaswani Noam M. Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the International Conference on Multimedia & Expo Workshops (ICMEW \u201915)","author":"Wang Xin","year":"2015","unstructured":"Xin Wang, Devinder Kumar, Nicolas Thome, Matthieu Cord, and Fr\u00e9d\u00e9ric Precioso. 2015. Recipe recognition with large multimodal food dataset. In Proceedings of the International Conference on Multimedia & Expo Workshops (ICMEW \u201915). 1\u20136."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.107"},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 3597\u20133606","author":"Wu Fangzhao","unstructured":"Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and M. Zhou. 2020. MIND: A large-scale dataset for news recommendation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 3597\u20133606."},{"key":"e_1_3_2_43_2","volume-title":"Pinheiro","author":"Xing Chen","year":"2019","unstructured":"Chen Xing, Negar Rostamzadeh, Boris N. Oreshkin, and Pedro H. O. Pinheiro. 2019. Adaptive cross-modal few-shot learning. In Advances in Neural Information Processing Systems. 4847\u20134857."},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the International Conference on Emerging Internetworking, Data & Web Technologies. Springer, 457\u2013466","author":"Xiong Caiquan","year":"2017","unstructured":"Caiquan Xiong, Yuan Li, and Ke Lv. 2017. Multi-documents summarization based on the textrank and its application in argumentation system. In Proceedings of the International Conference on Emerging Internetworking, Data & Web Technologies. Springer, 457\u2013466."},{"key":"e_1_3_2_45_2","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics. 3915\u20133926","author":"Yang Pengcheng","year":"2018","unstructured":"Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018. SGM: Sequence generation model for multi-label classification. In Proceedings of the 27th International Conference on Computational Linguistics. 3915\u20133926."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1257"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3035277"},{"key":"e_1_3_2_48_2","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1480\u20131489","author":"Yang Zichao","unstructured":"Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard H. Hovy. 2016. Hierarchical attention networks for document classification. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1480\u20131489."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i12.17289"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_2_51_2","volume-title":"Joshua Ainslie, Chris Alberti, Santiago Onta\u00f1\u00f3n","author":"Zaheer Manzil","year":"2020","unstructured":"Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Onta\u00f1\u00f3n, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big bird: Transformers for longer sequences. In Advances in Neural Information Processing Systems. 15."},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 5059\u20135069","author":"Zhang Xingxing","unstructured":"Xingxing Zhang, Furu Wei, and M. Zhou. 2019. HIBERT: Document level pre-training of hierarchical bidirectional transformers for document summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 5059\u20135069."},{"key":"e_1_3_2_53_2","first-page":"649","article-title":"Character-level convolutional networks for text classification","volume":"28","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in Neural Information Processing Systems, Vol. 28. 649\u2013657.","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631711","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3631711","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:43Z","timestamp":1750178143000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3631711"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,13]]},"references-count":52,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3631711"],"URL":"https:\/\/doi.org\/10.1145\/3631711","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,13]]},"assertion":[{"value":"2022-07-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-27","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}