{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:39:51Z","timestamp":1750307991002,"version":"3.41.0"},"reference-count":13,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2006,3,1]],"date-time":"2006-03-01T00:00:00Z","timestamp":1141171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2006,3]]},"abstract":"<jats:p>Many previous biological event-extraction systems were based on hand-crafted rules which were specifically tuned to a specific biological application domain. But manually constructing and tuning the rules are time-consuming processes and make the systems less portable. So supervised machine-learning methods were developed to generate the extraction rules automatically, but accepting the trade-off between precision and recall (high recall with low precision, and vice versa) is a barrier to improving performance. To make matters worse, a text in the biological domain is more complex because it often contains more than two biological events in a sentence, and one event in a noun chunk can be an entity for the other event. As a result, there are as yet no systems that give a good performance in extracting events in biological domains by using supervised machine learning.To overcome the limitations of previous systems and the complexity of biological texts, we present the following new ideas. First, we adopted a supervised machine-learning method to reduce the human effort in making extraction rules in order to obtain a highly domain-portable system. Second, we overcame the classical trade-off between precision and recall by using an event component verification method. Thus, machine learning occurs in two phases in our architecture. In the first phase, the system focuses on improving recall in extracting events between biological entities during a supervised machine-learning period. After extracting the biological events with automatically learned rules, in the second phase the system removes incorrect biological events by verifying the extracted event components with a maximum entropy (ME) classification method. In other words, the system targets for high recall in the first phase and tries to achieve high precision with a classifier in the second phase. Finally, we improved a supervised machine-learning algorithm so that it could learn a rule in a noun chunk and a rule extending throughout a sentence at two different levels, separately, for nested biological events.<\/jats:p>","DOI":"10.1145\/1131348.1131353","type":"journal-article","created":{"date-parts":[[2006,7,25]],"date-time":"2006-07-25T14:14:26Z","timestamp":1153836866000},"page":"61-73","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Two-phase learning for biological event extraction and verification"],"prefix":"10.1145","volume":"5","author":[{"given":"Eunju","family":"Kim","sequence":"first","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Song","sequence":"additional","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cheongjae","family":"Lee","sequence":"additional","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kyoungduk","family":"Kim","sequence":"additional","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gary Geunbae","family":"Lee","sequence":"additional","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Byoung-Kee","family":"Yi","sequence":"additional","affiliation":[{"name":"Pohang University of Science and Technology, Pohang, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeongwon","family":"Cha","sequence":"additional","affiliation":[{"name":"Changwon University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2006,3]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Blaschke C. Andrade M. A. Ouzous C. and Valencia A. 1999. Automatic extraction of biological information from scientific text: Protein-protein interactions. Intelligent Systems for Molecular Biology 60--67. Blaschke C. Andrade M. A. Ouzous C. and Valencia A. 1999. Automatic extraction of biological information from scientific text: Protein-protein interactions. Intelligent Systems for Molecular Biology 60--67."},{"key":"e_1_2_1_2_1","unstructured":"Bunescu R. Ge R. Kate R. J. Marcotte E. M. Mooney R. J. Ramani. A. K. and Wong Y. W. 2004. Comparative experiments on learning information extractors for proteins and their interactions. J. Artif. Intell. Medicine (Dec. 2004). Available online: http:\/\/www.sciencedirect.com\/. Bunescu R. Ge R. Kate R. J. Marcotte E. M. Mooney R. J. Ramani. A. K. and Wong Y. W. 2004. Comparative experiments on learning information extractors for proteins and their interactions. J. Artif. Intell. Medicine (Dec. 2004). Available online: http:\/\/www.sciencedirect.com\/."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btg452"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing.","author":"Gildea D.","year":"2001","unstructured":"Gildea , D. 2001 . Corpus variation and parser performance . In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Gildea, D. 2001. Corpus variation and parser performance. In Proceedings of the Conference on Empirical Methods in Natural Language Processing."},{"volume-title":"Pac. Symp. Biocomput.","author":"Park J. C.","key":"e_1_2_1_5_1","unstructured":"Park , J. C. , Kim , H. S. , and Kim , J. J . 2001. Bidirectional incremental parsing for automatic pathway identification with combinatory categorical grammar . Pac. Symp. Biocomput. Park, J. C., Kim, H. S., and Kim, J. J. 2001. Bidirectional incremental parsing for automatic pathway identification with combinatory categorical grammar. Pac. Symp. Biocomput."},{"volume-title":"Pac. Symp. Biocomput. 362--373","author":"Pustejovsky J.","key":"e_1_2_1_6_1","unstructured":"Pustejovsky , J. , Castano , J. , Kotechi , M. , and Cochran , B . 2002. Robust relational parsing over biomedical literature: Extracting inhibit relations . Pac. Symp. Biocomput. 362--373 . Pustejovsky, J., Castano, J., Kotechi, M., and Cochran, B. 2002. Robust relational parsing over biomedical literature: Extracting inhibit relations. Pac. Symp. Biocomput. 362--373."},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Rindflesch T. C. Rayan J. V. and Hunter L. 2000. Extracting molecular binding relationships from biomedical text. In Applied Natural Language Processing. North American Chapter of the Association for Computational Linguistics 188--195. 10.3115\/974147.974173 Rindflesch T. C. Rayan J. V. and Hunter L. 2000. Extracting molecular binding relationships from biomedical text. In Applied Natural Language Processing. North American Chapter of the Association for Computational Linguistics 188--195. 10.3115\/974147.974173","DOI":"10.3115\/974147.974173"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1006\/csla.1996.0011","article-title":"A maximum entropy approach to adaptive statistical language modeling","volume":"10","author":"Rosenfeld R.","year":"1996","unstructured":"Rosenfeld , R. 1996 . A maximum entropy approach to adaptive statistical language modeling . Computer, Speech and Language 10 , 187 -- 228 . Rosenfeld, R. 1996. A maximum entropy approach to adaptive statistical language modeling. Computer, Speech and Language 10, 187--228.","journal-title":"Computer, Speech and Language"},{"volume-title":"Proceedings of the Genome Informatics Workshop, 62--71","author":"Sekimizu T.","key":"e_1_2_1_9_1","unstructured":"Sekimizu , T. , Park , H. S. , and Tsuijii , J . 1998. Identifying the interaction between genes and gene products based on frequently seen verbs in Medline abstracts . In Proceedings of the Genome Informatics Workshop, 62--71 . Sekimizu, T., Park, H. S., and Tsuijii, J. 1998. Identifying the interaction between genes and gene products based on frequently seen verbs in Medline abstracts. In Proceedings of the Genome Informatics Workshop, 62--71."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007562322031"},{"key":"e_1_2_1_11_1","first-page":"541","article-title":"Automatic extraction of protein interactions from scientific abstracts","volume":"5","author":"Thomas J.","year":"2000","unstructured":"Thomas , J. , Milward , D. , Ousounis , C. , Pulman , S. , and Carroll , M. 2000 . Automatic extraction of protein interactions from scientific abstracts . Pac. Symp. Biocomput. 5 , 541 -- 552 . Thomas, J., Milward, D., Ousounis, C., Pulman, S., and Carroll, M. 2000. Automatic extraction of protein interactions from scientific abstracts. Pac. Symp. Biocomput. 5, 541--552.","journal-title":"Pac. Symp. Biocomput."},{"volume-title":"Pac. Symp. Biocomput.","author":"Yakushiji A.","key":"e_1_2_1_13_1","unstructured":"Yakushiji , A. , Tateisi , Y. , and Miyao , Y . 2001. Event extraction from biomedical papers using a full parser . Pac. Symp. Biocomput. Yakushiji, A., Tateisi, Y., and Miyao, Y. 2001. Event extraction from biomedical papers using a full parser. Pac. Symp. Biocomput."},{"volume-title":"Proceedings of the ACL Conference. 10","author":"Zhang T.","key":"e_1_2_1_14_1","unstructured":"Zhang , T. , Damerau , F. , and Johnson , D . 2001. Text chunking using regularized Winnow . In Proceedings of the ACL Conference. 10 .3115\/1073012.1073081 Zhang, T., Damerau, F., and Johnson, D. 2001. Text chunking using regularized Winnow. In Proceedings of the ACL Conference. 10.3115\/1073012.1073081"}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1131348.1131353","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1131348.1131353","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T15:06:16Z","timestamp":1750259176000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1131348.1131353"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,3]]},"references-count":13,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,3]]}},"alternative-id":["10.1145\/1131348.1131353"],"URL":"https:\/\/doi.org\/10.1145\/1131348.1131353","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"type":"print","value":"1530-0226"},{"type":"electronic","value":"1558-3430"}],"subject":[],"published":{"date-parts":[[2006,3]]},"assertion":[{"value":"2006-03-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}