{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T21:17:09Z","timestamp":1777929429410,"version":"3.51.4"},"reference-count":29,"publisher":"Oxford University Press (OUP)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Life stories of diseased and healthy individuals are abundantly available on the Internet. Collecting and mining such online content can offer many valuable insights into patients\u2019 physical and emotional states throughout the pre-diagnosis, diagnosis, treatment and post-treatment stages of the disease compared with those of healthy subjects. However, such content is widely dispersed across the web. Using traditional query-based search engines to manually collect relevant materials is rather labor intensive and often incomplete due to resource constraints in terms of human query composition and result parsing efforts. The alternative option, blindly crawling the whole web, has proven inefficient and unaffordable for e-health researchers.<\/jats:p>\n               <jats:p>Results: We propose a user-oriented web crawler that adaptively acquires user-desired content on the Internet to meet the specific online data source acquisition needs of e-health researchers. Experimental results on two cancer-related case studies show that the new crawler can substantially accelerate the acquisition of highly relevant online content compared with the existing state-of-the-art adaptive web crawling technology. For the breast cancer case study using the full training set, the new method achieves a cumulative precision between 74.7 and 79.4% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 32.8 and 37.0% using the peer method for the same time period. For the lung cancer case study using the full training set, the new method achieves a cumulative precision between 56.7 and 61.2% after 5 h of execution till the end of the 20-h long crawling session as compared with the cumulative precision between 29.3 and 32.4% using the peer method. Using the reduced training set in the breast cancer case study, the cumulative precision of our method is between 44.6 and 54.9%, whereas the cumulative precision of the peer method is between 24.3 and 26.3%; for the lung cancer case study using the reduced training set, the cumulative precisions of our method and the peer method are, respectively, between 35.7 and 46.7% versus between 24.1 and 29.6%. These numbers clearly show a consistently superior accuracy of our method in discovering and acquiring user-desired online content for e-health research.<\/jats:p>\n               <jats:p>Availability and implementation: The implementation of our user-oriented web crawler is freely available to non-commercial users via the following Web site: http:\/\/bsec.ornl.gov\/AdaptiveCrawler.shtml. The Web site provides a step-by-step guide on how to execute the web crawler implementation. In addition, the Web site provides the two study datasets including manually labeled ground truth, initial seeds and the crawling results reported in this article.<\/jats:p>\n               <jats:p>Contact: \u00a0xus1@ornl.gov<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt571","type":"journal-article","created":{"date-parts":[[2013,9,30]],"date-time":"2013-09-30T00:08:14Z","timestamp":1380499694000},"page":"104-114","source":"Crossref","is-referenced-by-count":28,"title":["A user-oriented web crawler for selectively acquiring online content in e-health research"],"prefix":"10.1093","volume":"30","author":[{"given":"Songhua","family":"Xu","sequence":"first","affiliation":[{"name":"Biomedical Science and Engineering Center, Health Data Sciences Institute, Oak Ridge National Laboratory, One Bethel Valley Road, Oak Ridge, TN 37830, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong-Jun","family":"Yoon","sequence":"additional","affiliation":[{"name":"Biomedical Science and Engineering Center, Health Data Sciences Institute, Oak Ridge National Laboratory, One Bethel Valley Road, Oak Ridge, TN 37830, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Georgia","family":"Tourassi","sequence":"additional","affiliation":[{"name":"Biomedical Science and Engineering Center, Health Data Sciences Institute, Oak Ridge National Laboratory, One Bethel Valley Road, Oak Ridge, TN 37830, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2013,9,29]]},"reference":[{"key":"2023012710380807400_btt571-B1","author":"ACS","year":"2013"},{"key":"2023012710380807400_btt571-B2","doi-asserted-by":"crossref","DOI":"10.1145\/775047.775108","article-title":"Collaborative crawling: mining user experiences for topical resource discovery","volume-title":"Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Aggarwal","year":"2002"},{"key":"2023012710380807400_btt571-B3","doi-asserted-by":"crossref","DOI":"10.1145\/371920.371955","article-title":"Intelligent crawling on the World Wide Web with arbitrary predicates","volume-title":"Proceedings of the 10th International Conference on World Wide Web","author":"Aggarwal","year":"2001"},{"key":"2023012710380807400_btt571-B4","doi-asserted-by":"crossref","DOI":"10.1145\/1645953.1646011","article-title":"Adaptive geospatially focused crawling","volume-title":"Proceedings of the 18th ACM Conference on Information and Knowledge Management","author":"Ahlers","year":"2009"},{"key":"2023012710380807400_btt571-B5","doi-asserted-by":"crossref","DOI":"10.1007\/11551188_30","article-title":"Combining text and link analysis for focused crawling","volume-title":"Proceedings of the Third International Conference on Advances in Pattern Recognition - Volume Part I","author":"Almpanidis","year":"2005"},{"key":"2023012710380807400_btt571-B6","doi-asserted-by":"crossref","DOI":"10.1007\/11551362_36","article-title":"Focused crawling using latent semantic indexing: an application for vertical search engines","volume-title":"Proceedings of the 9th European Conference on Research and Advanced Technology for Digital Libraries","author":"Almpanidis","year":"2005"},{"key":"2023012710380807400_btt571-B7","doi-asserted-by":"crossref","DOI":"10.1145\/1273496.1273504","article-title":"Focused crawling with scalable ordinal regression solvers","volume-title":"Proceedings of the 24th international conference on Machine learning","author":"Babaria","year":"2007"},{"key":"2023012710380807400_btt571-B8","doi-asserted-by":"crossref","DOI":"10.1145\/1135777.1136006","article-title":"Focused crawling: experiences in a real world project","volume-title":"Proceedings of the 15th International Conference on World Wide Web","author":"Badia","year":"2006"},{"key":"2023012710380807400_btt571-B9","doi-asserted-by":"crossref","DOI":"10.1145\/1242572.1242632","article-title":"An adaptive crawler for locating hidden web entry points","volume-title":"Proceedings of the 16th International Conference on World Wide Web","author":"Barbosa","year":"2007"},{"key":"2023012710380807400_btt571-B10","doi-asserted-by":"crossref","first-page":"1001","DOI":"10.1016\/j.datak.2009.04.002","article-title":"Improving the performance of focused web crawlers","volume":"68","author":"Batsakis","year":"2009","journal-title":"Data Knowl. Eng."},{"key":"2023012710380807400_btt571-B11","doi-asserted-by":"crossref","DOI":"10.1145\/511446.511466","article-title":"Accelerated focused crawling through online relevance feedback","volume-title":"Proceedings of the 11th international conference on World Wide Web","author":"Chakrabarti","year":"2002"},{"key":"2023012710380807400_btt571-B12","doi-asserted-by":"crossref","first-page":"1057","DOI":"10.1016\/j.camwa.2008.09.021","article-title":"A cross-language focused crawling algorithm based on multiple relevance prediction strategies","volume":"57","author":"Chen","year":"2009","journal-title":"Comput. Math. Appl."},{"key":"2023012710380807400_btt571-B13","doi-asserted-by":"crossref","DOI":"10.1145\/584792.584802","article-title":"Topic-oriented collaborative crawling","volume-title":"Proceedings of the Eleventh International Conference on Information and Knowledge Management","author":"Chung","year":"2002"},{"key":"2023012710380807400_btt571-B14","doi-asserted-by":"crossref","DOI":"10.1145\/1363686.1363953","article-title":"The impact of term selection in genre-aware focused crawling","volume-title":"Proceedings of the 2008 ACM symposium on Applied Computing","author":"de Assis","year":"2008"},{"key":"2023012710380807400_btt571-B15","first-page":"409","article-title":"Focused web crawling: a framework for crawling of country based financial data","volume-title":"Proc. IEEE International Conference on Information and Financial Engineering (ICIFE)","author":"Dey","year":"2010"},{"key":"2023012710380807400_btt571-B16","doi-asserted-by":"crossref","first-page":"24:1","DOI":"10.1145\/2382438.2382443","article-title":"Sentimental spidering: leveraging opinion information in focused crawlers","volume":"30","author":"Fu","year":"2012","journal-title":"ACM Trans. Inf. Syst."},{"key":"2023012710380807400_btt571-B17","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-23863-5_3","article-title":"An extended method for finding related web pages with focused crawling techniques","volume-title":"Proceedings of the 15th International Conference on Knowledge-Based and Intelligent Information and Engineering Systems - Volume Part II","author":"Furuse","year":"2011"},{"key":"2023012710380807400_btt571-B18","doi-asserted-by":"crossref","DOI":"10.1145\/1135777.1135822","article-title":"Geographically focused collaborative crawling","volume-title":"Proceedings of the 15th International Conference on World Wide Web","author":"Gao","year":"2006"},{"key":"2023012710380807400_btt571-B19","doi-asserted-by":"crossref","DOI":"10.1145\/1390334.1390488","article-title":"Guide focused crawler efficiently and effectively using on-line topical importance estimation","volume-title":"Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Guan","year":"2008"},{"key":"2023012710380807400_btt571-B20","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1145\/1656274.1656278","article-title":"The weka data mining software: an update","volume":"11","author":"Hall","year":"2009","journal-title":"ACM SIGKDD Exp. Newslett."},{"key":"2023012710380807400_btt571-B21","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1145\/324133.324140","article-title":"Authoritative sources in a hyperlinked environment","volume":"46","author":"Kleinberg","year":"1999","journal-title":"J. ACM"},{"key":"2023012710380807400_btt571-B22","article-title":"The Boilerpipe library: boilerplate removal and fulltext extraction from html pages","volume-title":"Google Code Base","author":"Kohlschutter","year":"2011"},{"key":"2023012710380807400_btt571-B23","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.1467-8640.2012.00411.x","article-title":"Probabilistic models for focused web crawling","volume":"28","author":"Liu","year":"2012","journal-title":"Comput. Intell."},{"key":"2023012710380807400_btt571-B24","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1145\/1031114.1031117","article-title":"Topical web crawlers: evaluating adaptive algorithms","volume":"4","author":"Menczer","year":"2004","journal-title":"ACM Trans. Internet Technol."},{"key":"2023012710380807400_btt571-B25","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/978-3-540-72079-9_7","volume-title":"The Adaptive Web: Adaptive Focused Crawling","author":"Micarelli","year":"2007"},{"key":"2023012710380807400_btt571-B26","doi-asserted-by":"crossref","first-page":"430","DOI":"10.1145\/1095872.1095875","article-title":"Learning to crawl: comparing classification schemes","volume":"23","author":"Pant","year":"2005","journal-title":"ACM Trans. Inf. Syst."},{"key":"2023012710380807400_btt571-B27","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1109\/TKDE.2006.12","article-title":"Link contexts in classifier-guided topical crawlers","volume":"18","author":"Pant","year":"2006","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"2023012710380807400_btt571-B28","author":"Rose","year":"2012"},{"key":"2023012710380807400_btt571-B29","doi-asserted-by":"crossref","DOI":"10.1145\/1065385.1065455","article-title":"What\u2019s there and what\u2019s not?: focused crawling for missing documents in digital libraries","volume-title":"Proceedings of the 5th ACM\/IEEE-CS Joint Conference on Digital libraries","author":"Zhuang","year":"2005"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/1\/104\/48913998\/bioinformatics_30_1_104.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/1\/104\/48913998\/bioinformatics_30_1_104.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T10:41:42Z","timestamp":1674816102000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/1\/104\/236067"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,9,29]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt571","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,1,1]]},"published":{"date-parts":[[2013,9,29]]}}}