{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T08:39:26Z","timestamp":1777365566955,"version":"3.51.4"},"reference-count":38,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2025,4,7]],"date-time":"2025-04-07T00:00:00Z","timestamp":1743984000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Hankuk University of Foreign Studies Research Fund of 2025"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Systems"],"abstract":"<jats:p>Similar patent document retrieval is an essential task that reduces the scope of patent claimants\u2019 searches, and numerous studies have attempted to provide automated patent search services. Recently, Retrieval-Augmented Generation (RAG) based on generative language models has emerged as an excellent method for accessing and utilizing patent knowledge environments. RAG-based patent search services offer enhanced retrieval ranking performance as AI search services by providing document knowledge similar to queries. However, achieving optimal similarity-based document ranking in search services remains a challenging task, as search methods based on document similarity do not adequately address the characteristics of patent documents. Unlike general document retrieval, the similarity of patent documents must take into account prior art relationships. To address this issue, we propose PAI-NET, a deep neural network for computing patent document similarities by incorporating expert knowledge of prior art relationships. We demonstrate that our proposed method outperforms current state-of-the-art models in patent document classification tasks through semantic distance evaluation on the USPD and KPRIS datasets. PAI-NET presents similar document candidates, demonstrating a superior patent search performance improvement of 15% over state-of-the-art methods.<\/jats:p>","DOI":"10.3390\/systems13040259","type":"journal-article","created":{"date-parts":[[2025,4,23]],"date-time":"2025-04-23T20:43:17Z","timestamp":1745440997000},"page":"259","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["PAI-NET: Retrieval-Augmented Generation Patent Network Using Prior Art Information"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9035-8086","authenticated-orcid":false,"given":"Kyung-Yul","family":"Lee","sequence":"first","affiliation":[{"name":"College of Economics and Business, Hankuk University of Foreign Studies, 81, Oedae-ro, Mohyeon-eup, Cheoin-gu, Yongin-si 17035, Gyeonggi-do, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0675-639X","authenticated-orcid":false,"given":"Juho","family":"Bai","sequence":"additional","affiliation":[{"name":"College of Economics and Business, Hankuk University of Foreign Studies, 81, Oedae-ro, Mohyeon-eup, Cheoin-gu, Yongin-si 17035, Gyeonggi-do, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,4,7]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Improving the Domain Adaptation of Retrieval Augmented Generation (RAG) Models for Open Domain Question Answering","volume":"11","author":"Siriwardhana","year":"2022","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_2","unstructured":"Wang, Y., Li, X., Wang, B., Zhou, Y., Ji, H., Chen, H., Zhang, J., Yu, F., Zhao, Z., and Jin, S. (2024). PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods. arXiv."},{"key":"ref_3","unstructured":"Liang, L., Sun, M., Gui, Z., Zhu, Z., Jiang, Z., Zhong, L., Qu, Y., Zhao, P., Bo, Z., and Yang, J. (2024). KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation. arXiv."},{"key":"ref_4","unstructured":"Kang, B., Kim, J., Yun, T.R., and Kim, C.E. (2024). Prompt-RAG: Pioneering Vector Embedding-Free Retrieval-Augmented Generation in Niche Domains, Exemplified by Korean Medicine. arXiv."},{"key":"ref_5","unstructured":"Khanna, S., and Subedi, S. (2024). Tabular Embedding Model (TEM): Finetuning Embedding Models For Tabular RAG Applications. arXiv."},{"key":"ref_6","unstructured":"Nguyen, Z., Annunziata, A., Luong, V., Dinh, S., Le, Q., Ha, A.H., Le, C., Phan, H.A., Raghavan, S., and Nguyen, C. (2024). Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chu, J.M., Lo, H.C., Hsiang, J., and Cho, C.C. (2024, January 16). Patent Response System Optimised for Faithfulness: Procedural Knowledge Embodiment with Knowledge Graph and Retrieval Augmented Generation. Proceedings of the 1st Workshop on Towards Knowledgeable Language Models (KnowLLM 2024), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.knowllm-1.12"},{"key":"ref_8","unstructured":"Wang, S., Yin, X., Wang, M., Guo, R., and Nan, K. (2024). EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent. arXiv."},{"key":"ref_9","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_10","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA. Volume 1 (Long and Short Papers)."},{"key":"ref_11","unstructured":"Dai, Z., Yang, Z., Yang, Y., Carbonell, J.G., Le, Q., and Salakhutdinov, R. (August, January 28). Transformer-XL: Attentive Language Models beyond a Fixed-Length Context. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_12","unstructured":"Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R.R., and Le, Q.V. (2019, January 8\u201314). Xlnet: Generalized autoregressive pretraining for language understanding. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_13","unstructured":"Beltagy, I., Peters, M.E., and Cohan, A. (2020). Longformer: The long-document transformer. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Roudsari, A.H., Afshar, J., Lee, C.C., and Lee, W. (2020, January 19\u201322). Multi-label Patent Classification using Attention-Aware Deep Learning Model. Proceedings of the IEEE International Conference on Big Data and Smart Computing (BigComp), Busan, Republic of Korea.","DOI":"10.1109\/BigComp48618.2020.000-2"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1007\/s11192-018-2905-5","article-title":"DeepPatent: Patent classification with convolutional neural networks and word embedding","volume":"117","author":"Li","year":"2018","journal-title":"Scientometrics"},{"key":"ref_16","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kim, Y. (2014). Convolutional neural networks for sentence classification. arXiv.","DOI":"10.3115\/v1\/D14-1181"},{"key":"ref_18","first-page":"108","article-title":"Domain-specific word embeddings for patent classification","volume":"53","author":"Risch","year":"2019","journal-title":"Data Technol. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2673","DOI":"10.1109\/78.650093","article-title":"Bidirectional recurrent neural networks","volume":"45","author":"Schuster","year":"1997","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lee, J.S., and Hsiang, J. (2019). PatentBERT: Patent Classification with Fine-Tuning a pre-trained BERT Model. arXiv.","DOI":"10.1016\/j.wpi.2020.101965"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hu, J., Li, S., Hu, J., and Yang, G. (2018). A Hierarchical Feature Extraction Model for Multi-Label Mechanical Patent Classification. Sustainability, 10.","DOI":"10.3390\/su10010219"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bai, J., Shim, I., and Park, S. (2020). MEXN: Multi-Stage Extraction Network for Patent Document Classification. Appl. Sci., 10.","DOI":"10.3390\/app10186229"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Song, G., Huang, X., Cao, G., Liu, W., Zhang, J., and Yang, L. (2018, January 12\u201314). Enhanced deep feature representation for patent image classification. Proceedings of the Tenth International Conference on Graphics and Image Processing (ICGIP 2018). International Society for Optics and Photonics, Chengdu, China.","DOI":"10.1117\/12.2524360"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Jiang, S., Luo, J., Pava, G.R., Hu, J., and Magee, C.L. (2020). A CNN-based Patent Image Retrieval Method for Design Ideation. arXiv.","DOI":"10.1115\/1.0001899V"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Csurka, G. (2017). Document image classification, with a specific view on applications of patent images. Current Challenges in Patent Information Retrieval, Springer.","DOI":"10.1007\/978-3-662-53817-3_12"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1080\/01638539809545028","article-title":"An introduction to latent semantic analysis","volume":"25","author":"Landauer","year":"1998","journal-title":"Discourse Processes"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Kim, B.T.S., and Hyun, E.J. (2023). Mapping the Landscape of Blockchain Technology Knowledge: A Patent Co-Citation and Semantic Similarity Approach. Systems, 11.","DOI":"10.3390\/systems11030111"},{"key":"ref_28","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International conference on machine learning. PMLR, Virtual."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Fang, H., and Xie, P. (2020). Cert: Contrastive self-supervised learning for language understanding. arXiv.","DOI":"10.36227\/techrxiv.12308378.v1"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019, January 3\u20137). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_31","unstructured":"Gunel, B., Du, J., Conneau, A., and Stoyanov, V. (2021, January 4). Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning. Proceedings of the International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, Z., Zhou, L., Yang, X., Jia, H., Li, W., and Zhang, J. (2023). User Sentiment Analysis of COVID-19 via Adversarial Training Based on the BERT-FGM-BiGRU Model. Systems, 11.","DOI":"10.3390\/systems11030129"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lin, W., Yu, W., and Xiao, R. (2023). Measuring Patent Similarity Based on Text Mining and Image Recognition. Systems, 11.","DOI":"10.3390\/systems11060294"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MCI.2018.2840738","article-title":"Recent trends in deep learning based natural language processing","volume":"13","author":"Young","year":"2018","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Taigman, Y., Yang, M., Ranzato, M., and Wolf, L. (2014, January 23\u201328). Deepface: Closing the gap to human-level performance in face verification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.220"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1023\/A:1007379606734","article-title":"Multitask learning","volume":"28","author":"Caruana","year":"1997","journal-title":"Mach. Learn."},{"key":"ref_37","unstructured":"Ruder, S. (2017). An overview of multi-task learning in deep neural networks. arXiv."},{"key":"ref_38","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-8954\/13\/4\/259\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:11:43Z","timestamp":1760029903000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-8954\/13\/4\/259"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,7]]},"references-count":38,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,4]]}},"alternative-id":["systems13040259"],"URL":"https:\/\/doi.org\/10.3390\/systems13040259","relation":{},"ISSN":["2079-8954"],"issn-type":[{"value":"2079-8954","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,7]]}}}