{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T15:44:36Z","timestamp":1782834276181,"version":"3.54.5"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T00:00:00Z","timestamp":1695168000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T00:00:00Z","timestamp":1695168000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Universit\u00e0 degli Studi di Bari Aldo Moro"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Inf Syst"],"published-print":{"date-parts":[[2024,2]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Tenders are powerful means of investment of public funds and represent a strategic development resource. Despite the efforts made so far by governments at national and international levels to digitalise documents related to the Public Administration sector, most of the information is still available in an unstructured format only. With the aim of bridging this gap, we present OIE4PA, our latest study on extracting and classifying relations from tenders of the Public Administration. Our work focuses on the Italian language, where the availability of linguistic resources to perform Natural Language Processing tasks is considerably limited. Nevertheless, OIE4PA adopts a multilingual approach so it can be applied to several languages by providing appropriate training data. Rather than purely training a classifier on a portion of the extracted relations, the backbone idea of our learning strategy is to put a supervised method based on self-training to the proof and to assess whether or not it improves the performance of the classifier. For evaluation purposes, we built a dataset composed of 2,000 triples which have been manually annotated by two human experts. The in-vitro evaluation shows that OIE4PA achieves a MacroF<jats:inline-formula><jats:alternatives><jats:tex-math>$$_1$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:msub>\n                    <mml:mrow\/>\n                    <mml:mn>1<\/mml:mn>\n                  <\/mml:msub>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> equal to <jats:bold>0.89<\/jats:bold> and a <jats:bold>91<\/jats:bold><jats:inline-formula><jats:alternatives><jats:tex-math>$$\\%$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>%<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> accuracy. In addition, OIE4PA was used as the pillar of a prototype search engine, which has been evaluated through an in-vivo experiment with positive feedback from 32 final users, obtaining a SUS score equal to <jats:bold>83.98<\/jats:bold>.<\/jats:p>","DOI":"10.1007\/s10844-023-00814-z","type":"journal-article","created":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T12:02:26Z","timestamp":1695211346000},"page":"273-294","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["OIE4PA: open information extraction for the public administration"],"prefix":"10.1007","volume":"62","author":[{"given":"Lucia","family":"Siciliani","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eleonora","family":"Ghizzota","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pierpaolo","family":"Basile","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pasquale","family":"Lops","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,9,20]]},"reference":[{"key":"814_CR1","unstructured":"Amini, M., Feofanov, V., Pauletto, L., et\u00a0al. (2022). Self-training: A survey. arXiv:2202.12040"},{"key":"#cr-split#-814_CR2.1","unstructured":"Banko, M., Cafarella, M.J., Soderland, S., et\u00a0al. (2007). Open Information Extraction from the Web. In: Veloso MM"},{"key":"#cr-split#-814_CR2.2","unstructured":"(ed) IJCAI 2007, Proceedings of the 20th International Joint Conference on Artificial Intelligence, Hyderabad, India, January 6-12, 2007, pp. 2670-2676, http:\/\/ijcai.org\/Proceedings\/07\/Papers\/429.pdf"},{"key":"814_CR3","volume-title":"Natural language processing for public services","author":"F Barth\u00e9lemy","year":"2022","unstructured":"Barth\u00e9lemy, F., Ghesqui\u00e8re, N., Loozen, N., et al. (2022). Natural language processing for public services. Luxembourg: Publications Office of the European Union."},{"key":"814_CR4","doi-asserted-by":"crossref","unstructured":"Bojanowski, P., Grave, E., Joulin, A., et\u00a0al. (2016). Enriching Word Vectors with Subword Information. arXiv preprint arXiv:1607.04606","DOI":"10.1162\/tacl_a_00051"},{"issue":"194","key":"814_CR5","first-page":"4","volume":"189","author":"J Brooke","year":"1996","unstructured":"Brooke, J., et al. (1996). Sus-a quick and dirty usability scale. Usability evaluation in industry, 189(194), 4\u20137.","journal-title":"Usability evaluation in industry"},{"key":"814_CR6","doi-asserted-by":"crossref","unstructured":"Buchholz, S., Marsi, E. (2006). CoNLL-X Shared Task on Multilingual Dependency Parsing. In: M\u00e0rquez, L., Klein, D. (eds). Proceedings of the Tenth Conference on Computational Natural Language Learning, CoNLL 2006, New York City, USA, June 8-9, 2006, pp. 149\u2013164. https:\/\/aclanthology.org\/W06-2920\/","DOI":"10.3115\/1596276.1596305"},{"key":"814_CR7","doi-asserted-by":"publisher","unstructured":"Cabot, P. H., & Navigli, R., et al. (2021). REBEL: relation extraction by end-to-end language generation. In M. Moens, X. Huang, & L. Specia (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event\/Punta Cana, Dominican Republic (pp. 2370\u20132381). Dominican Republic. https:\/\/doi.org\/10.18653\/v1\/2021.findings-emnlp.204","DOI":"10.18653\/v1\/2021.findings-emnlp.204"},{"key":"814_CR8","unstructured":"Cassotti, P., Siciliani, L., Basile, P., et\u00a0al. (2021). Extracting Relations from Italian Wikipedia using Unsupervised Information Extraction. In: V.W. Anelli, T.D. Noia, N. Ferro, et\u00a0al. (eds). Proceedings of the 11th Italian Information Retrieval Workshop 2021, Bari, Italy, September 13-15, 2021. http:\/\/ceur-ws.org\/Vol-2947\/paper2.pdf"},{"key":"814_CR9","doi-asserted-by":"crossref","unstructured":"Chen, T., Guestrin, C. (2016). Xgboost: A scalable tree boosting system. arXiv:1603.02754","DOI":"10.1145\/2939672.2939785"},{"key":"814_CR10","unstructured":"Christensen, J., Soderland, S., Etzioni, O., et\u00a0al. (2010). Semantic role labeling for open information extraction. In: Proceedings of the NAACL HLT 2010 first international workshop on formalisms and methodology for learning by reading, pp 52\u201360"},{"key":"814_CR11","doi-asserted-by":"publisher","unstructured":"Christensen, J., Mausam, Soderland, S., et\u00a0al. (2011). An analysis of open information extraction based on semantic role labeling. In: M.A. Musen,\u00d3. Corcho (eds) Proceedings of the 6th International Conference on Knowledge Capture (K-CAP 2011), June 26-29, 2011, Banff, Alberta, Canada, pp 113\u2013120, https:\/\/doi.org\/10.1145\/1999676.1999697","DOI":"10.1145\/1999676.1999697"},{"key":"814_CR12","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1145\/2488388.2488420","volume-title":"22nd International World Wide Web Conference, WWW \u201913","author":"L.D Corro","year":"2013","unstructured":"Corro, L. .D., & Gemulla, R., et al. (2013). ClausIE: clause-based open information extraction. In D. Schwabe, V. .A. .F. Almeida, & H. Glaser (Eds.), 22nd International World Wide Web Conference, WWW \u201913 (pp. 355\u2013366). https:\/\/doi.org\/10.1145\/2488388.2488420"},{"key":"814_CR13","doi-asserted-by":"publisher","first-page":"668","DOI":"10.1109\/WAINA.2018.00165","volume-title":"32nd International Conference on Advanced Information Networking and Applications Workshops","author":"E Damiano","year":"2018","unstructured":"Damiano, E., Minutolo, A., & Esposito, M., et al. (2018). Open Information Extraction for Italian Sentences. In L. Barolli, M. Takizawa, & T. Enokido (Eds.), 32nd International Conference on Advanced Information Networking and Applications Workshops (pp. 668\u2013673). AINA 2018 workshops. https:\/\/doi.org\/10.1109\/WAINA.2018.00165"},{"key":"814_CR14","doi-asserted-by":"publisher","unstructured":"Dognin, P., Padhi, I., Melnyk, I., et\u00a0al. (2021). ReGen: Reinforcement learning for text and knowledge base generation using pretrained language models. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, (pp. 1084\u20131099), https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.83","DOI":"10.18653\/v1\/2021.emnlp-main.83"},{"key":"814_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijinfomgt.2019.08.002","volume":"57","author":"YK Dwivedi","year":"2021","unstructured":"Dwivedi, Y. K., Hughes, L., Ismagilova, E., et al. (2021). Artificial intelligence (ai): Multidisciplinary perspectives on emerging challenges, opportunities, and agenda for research, practice and policy. International Journal of Information Management, 57, 101994.","journal-title":"International Journal of Information Management"},{"key":"814_CR16","doi-asserted-by":"publisher","first-page":"3","DOI":"10.5591\/978-1-57735-516-8\/IJCAI11-012","volume-title":"IJCAI 2011, Proceedings of the 22nd International Joint Conference on Artificial Intelligence","author":"O Etzioni","year":"2011","unstructured":"Etzioni, O., Fader, A., Christensen, J., et al. (2011). Open Information Extraction: The Second Generation. In T. Walsh (Ed.), IJCAI 2011, Proceedings of the 22nd International Joint Conference on Artificial Intelligence (pp. 3\u201310). https:\/\/doi.org\/10.5591\/978-1-57735-516-8\/IJCAI11-012"},{"key":"814_CR17","unstructured":"Fan, R., Chang, K., Hsieh, C., et\u00a0al. (2008). LIBLINEAR: A library for large linear classification. J Mach Learn Res 9:1871\u20131874. https:\/\/dl.acm.org\/citation.cfm?id=1442794"},{"key":"814_CR18","doi-asserted-by":"publisher","unstructured":"Gashteovski, K., Gemulla, R., Kotnis, B., et\u00a0al. (2020). On aligning OpenIE extractions with knowledge bases: A case study. In: Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems, Online, (pp. 143\u2013154), https:\/\/doi.org\/10.18653\/v1\/2020.eval4nlp-1.14","DOI":"10.18653\/v1\/2020.eval4nlp-1.14"},{"key":"814_CR19","doi-asserted-by":"publisher","unstructured":"Guarasci, R., Damiano, E., Minutolo, A., et al. (2020). Lexicon-Grammar based open information extraction from natural language sentences in Italian. Expert Syst Appl, 143,. https:\/\/doi.org\/10.1016\/j.eswa.2019.112954","DOI":"10.1016\/j.eswa.2019.112954"},{"key":"814_CR20","doi-asserted-by":"crossref","unstructured":"Josifoski, M., De\u00a0Cao, N., Peyrard, M., et\u00a0al. (2021). GenIE: generative information extraction. arXiv:2112.08340","DOI":"10.18653\/v1\/2022.naacl-main.342"},{"key":"814_CR21","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2019.112954","volume-title":"Understanding the use of emerging technologies in the public sector: A review of horizon 2020 projects","author":"E Kalampokis","year":"2023","unstructured":"Kalampokis, E., Karacapilidis, N., Tsakalidis, D., et al. (2023). Understanding the use of emerging technologies in the public sector: A review of horizon 2020 projects. Digital Government: Research and Practice. https:\/\/doi.org\/10.1016\/j.eswa.2019.112954"},{"key":"814_CR22","unstructured":"Kenton, J.D.M.W.C., Toutanova, L.K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT, (pp. 4171\u20134186)"},{"key":"814_CR23","doi-asserted-by":"crossref","unstructured":"Landis, J.R., Koch, G.G. (1977). The measurement of observer agreement for categorical data. Biometrics (pp. 159\u2013174)","DOI":"10.2307\/2529310"},{"issue":"7","key":"814_CR24","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1080\/10447318.2018.1455307","volume":"34","author":"JR Lewis","year":"2018","unstructured":"Lewis, J. R. (2018). The system usability scale: past, present, and future. International Journal of Human-Computer Interaction, 34(7), 577\u2013590.","journal-title":"International Journal of Human-Computer Interaction"},{"issue":"1","key":"814_CR25","doi-asserted-by":"publisher","DOI":"10.1016\/j.giq.2022.101774","volume":"40","author":"R Madan","year":"2023","unstructured":"Madan, R., & Ashok, M. (2023). AI adoption and diffusion in public administration: A systematic literature review and future research agenda. Gov Inf Q, 40(1), 101774. https:\/\/doi.org\/10.1016\/j.giq.2022.101774","journal-title":"Gov Inf Q"},{"key":"814_CR26","unstructured":"Mausam, Schmitz, M., Soderland, S., et\u00a0al. (2012). Open Language Learning for Information Extraction. In: J. Tsujii, J. Henderson, M. Pasca (eds) Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, EMNLP-CoNLL 2012, July 12-14, 2012, Jeju Island, Korea, (pp. 523\u2013534). https:\/\/aclanthology.org\/D12-1048\/"},{"key":"814_CR27","unstructured":"Niklaus, C., Cetto, M., Freitas, A., et\u00a0al. (2018). A Survey on Open Information Extraction. arXiv:1806.05599"},{"issue":"1","key":"814_CR28","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1145\/2413038.2413056","volume":"38","author":"A Sampaio","year":"2013","unstructured":"Sampaio, A. (2013). Quantifying the user experience: Practical statistics for user research by jeff sauro and james r. lewis. SIGSOFT Softw Eng Notes, 38(1), 57\u201358. https:\/\/doi.org\/10.1145\/2413038.2413056","journal-title":"SIGSOFT Softw Eng Notes"},{"key":"814_CR29","doi-asserted-by":"crossref","unstructured":"Siciliani, L., Cassotti, P., Basile, P., et\u00a0al. (2021). Extracting Relations from Italian Wikipedia using Self-Training. In: E. Fersini, M. Passarotti, V. Patti (eds) Proceedings of the Eighth Italian Conference on Computational Linguistics, CLiC-it 2021, http:\/\/ceur-ws.org\/Vol-3033\/paper28.pdf","DOI":"10.4000\/books.aaccademia.10849"},{"key":"814_CR30","doi-asserted-by":"publisher","first-page":"88","DOI":"10.18653\/v1\/K17-3009","volume-title":"Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies","author":"M Straka","year":"2017","unstructured":"Straka, M., & Strakov\u00e1, J. (2017). Tokenizing, POS tagging, lemmatizing and parsing UD 2.0 with udpipe. In J. Hajic & D. Zeman (Eds.), Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies (pp. 88\u201389). https:\/\/doi.org\/10.18653\/v1\/K17-3009"},{"key":"814_CR31","doi-asserted-by":"crossref","unstructured":"Vo, D., Bagheri, E. (2016). Open Information Extraction. arXiv:1607.02784","DOI":"10.1142\/9789813227927_0001"},{"key":"814_CR32","unstructured":"Wu, F., Weld, D.S. (2010). Open Information Extraction Using Wikipedia. In: J. Hajic, S. Carberry, S. Clark (eds) ACL 2010, Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, July 11-16, 2010, Uppsala, Sweden, (pp. 118\u2013127), https:\/\/aclanthology.org\/P10-1013\/"},{"key":"814_CR33","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526724","volume-title":"Proceedings of the 18th International Conference on World Wide Web, WWW 2009","author":"J Zhu","year":"2009","unstructured":"Zhu, J., Nie, Z., Liu, X., et al. (2009). StatSnowball: a statistical approach to extracting entity relationships. In J. Quemada, G. Le\u00f3n, Y. .S. Maarek, et al. (Eds.), Proceedings of the 18th International Conference on World Wide Web, WWW 2009. . https:\/\/doi.org\/10.1145\/1526709.1526724"}],"container-title":["Journal of Intelligent Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10844-023-00814-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10844-023-00814-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10844-023-00814-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,3,7]],"date-time":"2024-03-07T18:15:20Z","timestamp":1709835320000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10844-023-00814-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,20]]},"references-count":34,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,2]]}},"alternative-id":["814"],"URL":"https:\/\/doi.org\/10.1007\/s10844-023-00814-z","relation":{},"ISSN":["0925-9902","1573-7675"],"issn-type":[{"value":"0925-9902","type":"print"},{"value":"1573-7675","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,20]]},"assertion":[{"value":"22 May 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 August 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 August 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 September 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}},{"value":"The authors have no financial or proprietary interests in any material discussed in this article.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"We declare that this submission follows the policies as outlined in the Guide for Authors.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"The evaluation process does not store sensitive information about participants, and survey answers cannot be attributed to any specific individual.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"All authors agree with the content and give explicit consent to submit.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}