{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T20:35:09Z","timestamp":1761165309739,"version":"build-2065373602"},"reference-count":12,"publisher":"Sociedade Brasileira de Computa\u00e7\u00e3o - SBC","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>Entity Matching \u00e9 fundamental para integrar dados de diferentes fontes que se referem \u00e0 mesma entidade. Embora modelos Pr\u00e9-Treinados que adotam t\u00e9cnicas de Entity Blocking sejam amplamente utilizados nessa tarefa, o avan\u00e7o dos Large Language Models (LLMs) sugere novas possibilidades. Este trabalho compara o Ditto, que aplica t\u00e9cnicas de otimiza\u00e7\u00e3o com modelos tradicionais, com o Orca2, um LLM baseado no Llama2 voltado para racioc\u00ednio. Apesar do desempenho inicial inferior, o Orca2 demonstra potencial competitivo, sobretudo com futuras melhorias computacionais. Assim, busca-se avaliar a viabilidade dos LLMs em Entity Matching, analisando precis\u00e3o e custo computacional.<\/jats:p>","DOI":"10.5753\/sbbd.2025.247828","type":"proceedings-article","created":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T19:26:36Z","timestamp":1761074796000},"page":"956-962","source":"Crossref","is-referenced-by-count":0,"title":["Entity Matching com Large Language Models: estudo comparativo com abordagem de Entity Blocking"],"prefix":"10.5753","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-7330-4718","authenticated-orcid":false,"given":"Rodolfo","family":"Bolconte Donato","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tiago","family":"Brasileiro Ara\u00fajo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"3742","published-online":{"date-parts":[[2025,9,29]]},"reference":[{"key":"1","unstructured":"Arvanitis-Kasinikos, I. and Papadakis, G. (2025). Entity matching with 7b llms: A study on prompting strategies and hardware limitations. CEUR Workshop Proceedings."},{"key":"2","doi-asserted-by":"crossref","unstructured":"Barlaug, N. and Gulla, J. A. (2021). Neural networks for entity matching: A survey. ACM Transactions on Knowledge Discovery from Data (TKDD), 15(3):1\u201337.","DOI":"10.1145\/3442200"},{"key":"3","doi-asserted-by":"crossref","unstructured":"Brasileiro Ara\u00fajo, T., Efthymiou, V., Christophides, V., Pitoura, E., and Stefanidis, K. (2025). Treats: Fairness-aware entity resolution over streaming data. Information Systems, 129:102506.","DOI":"10.1016\/j.is.2024.102506"},{"key":"4","doi-asserted-by":"crossref","unstructured":"Christen, P. and Christen, P. (2012). Data matching systems. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection, pages 229\u2013242.","DOI":"10.1007\/978-3-642-31164-2_10"},{"key":"5","doi-asserted-by":"crossref","unstructured":"Kuang, W., Qian, B., Li, Z., Chen, D., Gao, D., Pan, X., Xie, Y., Li, Y., Ding, B., and Zhou, J. (2024). Federatedscope-llm: A comprehensive package for fine-tuning large language models in federated learning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 5260\u20135271.","DOI":"10.1145\/3637528.3671573"},{"key":"6","doi-asserted-by":"crossref","unstructured":"Li, Y., Li, J., Suhara, Y., Doan, A., and Tan, W.-C. (2020). Deep entity matching with pre-trained language models. Proceedings of the VLDB Endowment, 14(1):50\u201360.","DOI":"10.14778\/3421424.3421431"},{"key":"7","unstructured":"Mitra, A., Del Corro, L., Mahajan, S., Codas, A., Simoes, C., Agarwal, S., Chen, X., Razdaibiedina, A., Jones, E., Aggarwal, K., et al. (2023). Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045."},{"key":"8","doi-asserted-by":"crossref","unstructured":"Niven, T. and Kao, H.-Y. (2019). Probing neural network comprehension of natural language arguments. arXiv preprint arXiv:1907.07355.","DOI":"10.18653\/v1\/P19-1459"},{"key":"9","unstructured":"Peeters, R., Der, R. C., and Bizer, C. (2023a). Wdc products: A multi-dimensional entity matching benchmark. arXiv preprint arXiv:2301.09521."},{"key":"10","unstructured":"Peeters, R., Steiner, A., and Bizer, C. (2023b). Entity matching using large language models. arXiv preprint arXiv:2310.11244."},{"key":"11","doi-asserted-by":"crossref","unstructured":"Wang, Y. and Yan, M. (2024). Unsupervised domain adaptation for entity blocking leveraging large language models. In 2024 IEEE International Conference on Big Data (BigData), pages 159\u2013164. IEEE.","DOI":"10.1109\/BigData62323.2024.10825234"},{"key":"12","unstructured":"Zhang, J., Sun, H., and Ho, J. C. (2024). Emba: Entity matching using multi-task learning of bert with attention-over-attention. In EDBT, pages 281\u2013293."}],"event":{"name":"Simp\u00f3sio Brasileiro de Banco de Dados","number":"40","location":"Brasil","acronym":"SBBD 2025"},"container-title":["Anais do XL Simp\u00f3sio Brasileiro de Banco de Dados (SBBD 2025)"],"original-title":[],"link":[{"URL":"https:\/\/sol.sbc.org.br\/index.php\/sbbd\/article\/download\/37309\/37092","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/sol.sbc.org.br\/index.php\/sbbd\/article\/download\/37309\/37092","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T19:26:50Z","timestamp":1761074810000},"score":1,"resource":{"primary":{"URL":"https:\/\/sol.sbc.org.br\/index.php\/sbbd\/article\/view\/37309"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,29]]},"references-count":12,"URL":"https:\/\/doi.org\/10.5753\/sbbd.2025.247828","relation":{},"subject":[],"published":{"date-parts":[[2025,9,29]]}}}