{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:04:14Z","timestamp":1781103854290,"version":"3.54.1"},"reference-count":26,"publisher":"IGI Global Scientific Publishing","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,1,1]]},"abstract":"<p>The objective of this paper is to present a methodology to extract and rank automatically biomedical terms from free text. The authors present new extraction methods taking into account linguistic patterns specialized for the biomedical domain, statistic term extraction measures such as C-value and statistic keyword extraction measures such as Okapi BM25, and TFIDF. These measures are combined in order to improve the extraction process and the authors investigate which combinations are the more relevant associated to different contexts. Experimental results show that an appropriate harmonic mean of C-value associated to keyword extraction measures offers better precision, both for single-word and multi-words term extraction. Experiments describe the extraction of English and French biomedical terms from a corpus of laboratory tests available online. The results are validated by using UMLS (in English) and only MeSH (in French) as reference dictionary.<\/p>","DOI":"10.4018\/ijkdb.2014010101","type":"journal-article","created":{"date-parts":[[2014,4,9]],"date-time":"2014-04-09T08:54:33Z","timestamp":1397033673000},"page":"1-15","source":"Crossref","is-referenced-by-count":9,"title":["Towards a Mixed Approach to Extract Biomedical Terms from Text Corpus"],"prefix":"10.4018","volume":"4","author":[{"given":"Juan Antonio Lossio","family":"Ventura","sequence":"first","affiliation":[{"name":"LIRMM, University Montpellier 2, Montpellier, France & CNRS, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Clement","family":"Jonquet","sequence":"additional","affiliation":[{"name":"LIRMM, University Montpellier 2, Montpellier, France & CNRS, Paris, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mathieu","family":"Roche","sequence":"additional","affiliation":[{"name":"UMR TETIS, Cirad, Irstea, AgroParisTech, Montpellier, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Maguelonne","family":"Teisseire","sequence":"additional","affiliation":[{"name":"UMR TETIS, Cirad, Irstea, AgroParisTech, Montpellier, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"ijkdb.2014010101-0","doi-asserted-by":"crossref","unstructured":"Al Khatib, K., & Amer Badarneh. (2010). Automatic extraction of Arabic multi-word terms. In Proc. of Computer Science and Information Technology (IMCSIT) (pp. 411-418).","DOI":"10.1109\/IMCSIT.2010.5679929"},{"key":"ijkdb.2014010101-1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-00382-0_10"},{"key":"ijkdb.2014010101-2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.48"},{"key":"ijkdb.2014010101-3","doi-asserted-by":"crossref","unstructured":"Eck, N. J., Waltman, L., Noyons, E. C. M., & Buter, R. K. (2010). Automatic term identification for bibliometric mapping. SpringerLink Scientometrics, 82(3).","DOI":"10.1007\/s11192-010-0173-0"},{"key":"ijkdb.2014010101-4","doi-asserted-by":"publisher","DOI":"10.1007\/s007999900023"},{"key":"ijkdb.2014010101-5","doi-asserted-by":"crossref","unstructured":"Gracia, J., Trillo, R., Espinoza, M., & Mena, E. (2006). Querying the web: A multiontology disambiguation method. In Proceedings of the 6th International Conference on Web Engineering (pp. 241\u2013248).","DOI":"10.1145\/1145581.1145630"},{"key":"ijkdb.2014010101-6","doi-asserted-by":"crossref","unstructured":"Gracia, J., Trillo, R., Espinoza, M., & Mena, E. (2008). Web-based measure of semantic relatedness. In Proceedings of the 9th International Conference on Web Information Systems Engineering (pp. 136\u2013150).","DOI":"10.1007\/978-3-540-85481-4_12"},{"key":"ijkdb.2014010101-7","doi-asserted-by":"publisher","DOI":"10.1016\/j.datak.2008.11.002"},{"key":"ijkdb.2014010101-8","unstructured":"Hussey, R., Williams, S., & Mitchell, R. (2012). Automatic keyphrase extraction: A comparison of methods. In Proc. of the International Conference on Information Process, and Knowledge Management (pp. 18-23)."},{"key":"ijkdb.2014010101-9","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-70939-8_6"},{"key":"ijkdb.2014010101-10","unstructured":"Knoth, P., Schmidt, M., Smrz, P., & Zdrahal, Z. (2009). Towards a framework for comparing automatic term recognition methods. In Proceedings of the Conference Znalosti."},{"key":"ijkdb.2014010101-11","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2004.08.004"},{"key":"ijkdb.2014010101-12","unstructured":"Kupsc, A. (2006). Extraction automatique de termes a` partir de textes polonais. Journal Linguistique de Corpus. LabTestOnLine. (n.d.). Retrieved from http:\/\/labtestsonline.org\/"},{"key":"ijkdb.2014010101-13","unstructured":"Lossio Ventura, J. A., Jonquet, C., Roche, M., & Teisseire, M. (2013). Combining C-value and keyword extraction methods for biomedical terms extraction. In Proceedings of the 5th International Symposium on Languages in Biology and Medicine."},{"key":"ijkdb.2014010101-14","doi-asserted-by":"crossref","unstructured":"Lv, Y., & Zhai, C. X. (2011). When documents are very long, BM25 fails! In Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 1103\u20131104).","DOI":"10.1145\/2009916.2010070"},{"key":"ijkdb.2014010101-15","doi-asserted-by":"crossref","unstructured":"Medelyan, O., Eibe, F., & Witten, I. H. (2009). Human-competitive tagging using automatic keyphrase extraction. In Proceedings of the International Conference of Empirical Methods in Natural Language Processing (EMNLP), Singapore.","DOI":"10.3115\/1699648.1699678"},{"key":"ijkdb.2014010101-16","unstructured":"MeSH (Medical Subject Headings) is the NLM controlled vocabulary thesaurus used for indexing articles for PubMed. (2012). Retrieved from http:\/\/www.ncbi.nlm.nih.gov\/mesh"},{"key":"ijkdb.2014010101-17","doi-asserted-by":"crossref","unstructured":"Nenadic, G., Spasic, I., & Ananiadou, S. (2003). Morpho-syntactic clues for terminological processing in Serbian. In Proceedings of the EACL Workshop on Morphological Processing of Slavic Languages (pp. 79\u201386).","DOI":"10.3115\/1613200.1613211"},{"key":"ijkdb.2014010101-18","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gkp440"},{"key":"ijkdb.2014010101-19","unstructured":"Robertson, S. E., Walker, S., & Beaulieu, M. (1999). Okapi at TREC-7: Automatic ad hoc, filtering, VLC and interactive track. IN, 21, 253\u2013264."},{"key":"ijkdb.2014010101-20","doi-asserted-by":"crossref","unstructured":"Sclano, F., & Velardi, P. (2007). TermExtractor: A web application to learn the common terminology of interest groups and research communities. In Enterprise Interoperability II (pp. 287-290).","DOI":"10.1007\/978-1-84628-858-6_32"},{"key":"ijkdb.2014010101-21","unstructured":"Spela Vintar. (2004). Comparative evaluation of c-value in the treatment of nested terms. In Proceedings of the Workshop (Methodologies and Evaluation of Multiword Units in Realworld Applications), LREC (pp. 54\u201357)."},{"key":"ijkdb.2014010101-22","unstructured":"TreeTagger. (n.d.). Retrieved from www.cis.uni- muenchen.de\/~schmid\/tools\/TreeTagger"},{"key":"ijkdb.2014010101-23","unstructured":"Unified Medical Language System (UMLS). (2013). Retrieved from http:\/\/www.nlm.nih.gov\/research\/umls"},{"key":"ijkdb.2014010101-24","unstructured":"Zhang, Y., Milios, E., & Zincirheywood, N. (2004). A comparison of keyword and keyterm-based methods for automatic web site summarization. In Proceedings of the AAAI04 Workshop on Adaptive Text Extraction and Mining (pp. 15\u201320)."},{"key":"ijkdb.2014010101-25","unstructured":"Zhang, Z., Iria, J., Brewster, C., & Ciravegna, F. (2008). A comparative evaluation of term recognition algorithms. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC08)."}],"container-title":["International Journal of Knowledge Discovery in Bioinformatics"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=105097","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,1]],"date-time":"2022-06-01T20:43:36Z","timestamp":1654116216000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/ijkdb.2014010101"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2014,1,1]]},"references-count":26,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,1]]}},"URL":"https:\/\/doi.org\/10.4018\/ijkdb.2014010101","relation":{},"ISSN":["1947-9115","1947-9123"],"issn-type":[{"value":"1947-9115","type":"print"},{"value":"1947-9123","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,1,1]]}}}