{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T14:25:53Z","timestamp":1782397553477,"version":"3.54.5"},"reference-count":30,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T00:00:00Z","timestamp":1773792000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Research on Knowledge Representation Model and Smart System Construction of Multimodal Data of Cultural Heritage","award":["23BTQ088"],"award-info":[{"award-number":["23BTQ088"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Publications"],"abstract":"<jats:p>To address the efficiency and cost limitations of traditional manual cataloging, this study proposes a large language model-driven automated cataloging workflow in which the Metadata Extraction Agent (MEA), Description Cataloging Agent (DCA), Subject Analysis &amp; Indexing Agent (SAIA), and Quality Control Agent (QCA) collaborate to perform cataloging tasks. Experiments are conducted using a dataset of over 33,000 CNMARC bibliographic records from a University Library, together with data from the Chinese Library Classification (5th edition). Meanwhile, the agent-based workflow framework directly employs large language models without additional enhancement techniques, thereby providing a useful experimental benchmark for evaluating future AI-assisted cataloging systems. The results show that the framework performs well in metadata recognition, bibliographic description, and macro-level classification tasks, and can relatively stably generate standardized records. However, limitations remain in fine-grained semantic indexing and the interpretation of complex contexts. Therefore, in light of the capability limitations revealed by the experimental results, the study argues that fully automated end-to-end cataloging relying solely on generative AI is not yet entirely feasible. Future improvements should integrate techniques such as retrieval-augmented generation, supervised fine-tuning, and structured reasoning prompts, while establishing traceable mechanisms to enhance the reliability of intelligent cataloging.<\/jats:p>","DOI":"10.3390\/publications14010019","type":"journal-article","created":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T17:12:26Z","timestamp":1773853946000},"page":"19","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Research on Large Language Model-Based Bibliographic Cataloging Agent in the CNMARC Context"],"prefix":"10.3390","volume":"14","author":[{"given":"Zhuoxi","family":"Tan","sequence":"first","affiliation":[{"name":"School of Information Management, Sun Yat-sen University, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Information Resource Management, Renmin University of China, Beijing 100872, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qinyu","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Information Management, Sun Yat-sen University, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tao","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Information Management, Sun Yat-sen University, Guangzhou 510006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Alyafeai, Z., Al-Shaibani, M. S., and Ghanem, B. (2025). MOLE: Metadata extraction and validation in scientific papers using LLMs. arXiv.","DOI":"10.18653\/v1\/2025.findings-emnlp.655"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"775","DOI":"10.1080\/01639374.2021.1998283","article-title":"Kratt: Developing an automatic subject indexing tool for the national library of estonia","volume":"59","author":"Asula","year":"2021","journal-title":"Cataloging & Classification Quarterly"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"34","DOI":"10.6017\/ital.v33i1.5378","article-title":"assignFAST: An autosuggest-based tool for FAST subject assignment","volume":"33","author":"Bennett","year":"2014","journal-title":"Information Technology and Libraries"},{"key":"ref_4","unstructured":"Bodenhamer, J. (2026, February 27). Reliability and usability of ChatGPT for library metadata, Available online: https:\/\/openresearch.okstate.edu\/entities\/publication\/98b121d2-1f87-4824-b5c9-d204dfe87ced."},{"key":"ref_5","first-page":"147+180","article-title":"Exploration on the realization mechanism of linked data of special collections in digital humanities practice: Case study of local chronicles","volume":"45","author":"Chen","year":"2022","journal-title":"Information Studies: Theory & Application"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1080\/01639374.2024.2394516","article-title":"An experiment with the use of chatgpt for lcsh subject assignment on electronic theses and dissertations","volume":"62","author":"Chow","year":"2024","journal-title":"Cataloging & Classification Quarterly"},{"key":"ref_7","first-page":"1","article-title":"AI chatbots and subject cataloging: A performance test","volume":"69","author":"Dobreski","year":"2025","journal-title":"Library Resources & Technical Services"},{"key":"ref_8","unstructured":"D\u2019Souza, J., Sadruddin, S., and Israel, H. (2025). SemEval-2025 task 5: LLMs4Subjects\u2014LLM-based automated subject tagging for a national technical library\u2019s open-access catalog. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Facklam, D., Sweeney, S. J., Majumdar, S., and Dmello, K. R. N. (2025). Describing archival photographs using multimodal LLMs: A case study on evaluating vision-language model performance for creating descriptive metadata. The Electronic Library, Advance online publication.","DOI":"10.1108\/EL-06-2025-0270"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"38394.","DOI":"10.1038\/s41598-025-22272-z","article-title":"AI-powered knowledge organization: A next-generation approach to library classification using DeepSeek-R1","volume":"15","author":"Feng","year":"2025","journal-title":"Scientific Reports"},{"key":"ref_11","unstructured":"Greenly, G. A. (2025, December 31). CatalogerGPT, Available online: https:\/\/glengreenly.wixsite.com\/catalogergpt."},{"key":"ref_12","first-page":"49","article-title":"Cataloging: From digitization to datafication","volume":"45","author":"Hu","year":"2019","journal-title":"Journal of Library Science in China"},{"key":"ref_13","first-page":"425","article-title":"Generative and hierarchical classification of literature based on fine-tuned large language models","volume":"44","author":"Hu","year":"2025","journal-title":"Journal of the China Society for Scientific and Technical Information"},{"key":"ref_14","unstructured":"Isabel, B. (2025, December 31). Could artificial intelligence help catalog thousands of digital library books? An interview with Abigail Potter and Caroline Saccucci. Library of Congress Blogs, Available online: https:\/\/labs.loc.gov\/work\/experiments\/ECD\/?loclr=blogsig."},{"key":"ref_15","unstructured":"ISO (2008). Information and documentation\u2014Format for information exchange (Standard No. ISO Standard No. ISO 2709:2008)."},{"key":"ref_16","unstructured":"Joshi, B., Symeonidou, A., and Danish, S. M. (, January November). An end-to-end pipeline for bibliography extraction from scientific articles [Conference paper]. 2nd Workshop on Information Extraction from Scientific Publications, Bali, Indonesia."},{"key":"ref_17","first-page":"2","article-title":"Tesseract: An open-source optical character recognition engine","volume":"2007","author":"Kay","year":"2007","journal-title":"Linux Journal"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Koh\u00fat, J., Do\u010dekal, M., Hradi\u0161, M., and Va\u0161ko, M. (2025). BiblioPage: A dataset of scanned title pages for bibliographic metadata extraction. arXiv.","DOI":"10.1007\/978-3-032-04624-6_17"},{"key":"ref_19","unstructured":"Lopez, P. (2, January September). GROBID: Combining automatic bibliographic data recognition and term extraction for scholarship publications [Conference paper]. 13th European Conference on Research and Advanced Technology for Digital Libraries, Corfu, Greece."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Luo, P., and Liu, Z. (2025). Automatic subject indexing with large language model fine-tuning and controlled vocabulary retrieval. ASLIB Journal of Information Management, Advance online publication.","DOI":"10.1108\/AJIM-05-2025-0253"},{"key":"ref_21","unstructured":"Miller-Nesbitt, A. (2026, February 27). ChatGPT not useful as a tool to streamline library cataloguing processes [Open peer review commentary], Available online: https:\/\/journals.library.ualberta.ca\/eblip\/index.php\/EBLIP\/article\/view\/30524."},{"key":"ref_22","unstructured":"(2026, March 10). PyMuPDF documentation, Available online: https:\/\/pymupdf.readthedocs.io\/en\/latest\/."},{"key":"ref_23","unstructured":"Steinberger, R., Ebrahim, M., and Turchi, M. (, January May). JRC EuroVoc Indexer JEX\u2014A freely available multi-label categorisation tool [Conference paper]. Eighth International Conference on Language Resources and Evaluation (LREC 2012), Istanbul, Turkey."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1080\/01639374.2025.2508956","article-title":"Enhancing cataloging with generative AI: Converting wade-giles to Pinyin","volume":"63","author":"Sun","year":"2025","journal-title":"Cataloging & Classification Quarterly"},{"key":"ref_25","first-page":"265","article-title":"Annif and Finto AI: Developing and implementing automated subject indexing","volume":"13","author":"Suominen","year":"2022","journal-title":"JLIS.it"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"527","DOI":"10.1080\/01639374.2024.2394513","article-title":"Creating and evaluating MARC 21 bibliographic records using ChatGPT","volume":"62","author":"Taniguchi","year":"2024","journal-title":"Cataloging & Classification Quarterly"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"317","DOI":"10.1007\/s10032-015-0249-8","article-title":"CERMINE: Automatic extraction of structured metadata from scientific literature","volume":"18","author":"Tkaczyk","year":"2015","journal-title":"International Journal on Document Analysis and Recognition (IJDAR)"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Turner, M. D. (2025). Large language models can extract metadata for annotation of human neuroimaging publications. Frontiers in Neuroinformatics, 19.","DOI":"10.3389\/fninf.2025.1609077"},{"key":"ref_29","unstructured":"Voskuil, K., and Verberne, S. (2021). Improving reference mining in patents with BERT. arXiv."},{"key":"ref_30","unstructured":"Yang, H., and Hsu, W. (, January November). Automatic metadata information extraction from scientific literature using deep neural networks [Conference paper]. Fourteenth International Conference on Machine Vision (ICMV 2021), Rome, Italy."}],"container-title":["Publications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2304-6775\/14\/1\/19\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T17:27:44Z","timestamp":1773854864000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2304-6775\/14\/1\/19"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,18]]},"references-count":30,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["publications14010019"],"URL":"https:\/\/doi.org\/10.3390\/publications14010019","relation":{},"ISSN":["2304-6775"],"issn-type":[{"value":"2304-6775","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,18]]}}}