{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T18:09:26Z","timestamp":1754158166563,"version":"3.41.2"},"reference-count":32,"publisher":"Emerald","issue":"3","license":[{"start":{"date-parts":[[2006,5,1]],"date-time":"2006-05-01T00:00:00Z","timestamp":1146441600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,5,1]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>The purpose of this research is to automatically separate and extract meta\u2010data and instance information from various link pages in the web, by utilizing presentation and linkage regularities on the web.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>Research objectives have been achieved through an information extraction system called semantic partitioner that automatically organizes the content in each web page into a hierarchical structure, and an algorithm that interprets and translates these hierarchical structures into logical statements by distinguishing and representing the meta\u2010data and their individual data instances.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>Experimental results for the university domain with 12 computer science department web sites, comprising 361 individual faculty and course home pages indicate that the performance of the meta\u2010data and instance extraction averages 85, 88 percent <jats:italic>F<\/jats:italic>\u2010measure, respectively. Our METEOR system achieves this performance without any domain specific engineering requirement.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>Important contributions of the METEOR system presented in this paper are: it performs extraction without the assumption that the object instance pages are template\u2010driven; it is domain independent and does not require any previously engineered domain ontology; and by interpreting the link pages, it can extract both meta\u2010data, such as concept and attribute names and their relationships, as well as their instances with high accuracy.<\/jats:p><\/jats:sec>","DOI":"10.1108\/14684520610675807","type":"journal-article","created":{"date-parts":[[2006,7,4]],"date-time":"2006-07-04T07:52:03Z","timestamp":1151999523000},"page":"278-296","source":"Crossref","is-referenced-by-count":1,"title":["Gathering meta\u2010data and instances from object referral lists on the web"],"prefix":"10.1108","volume":"30","author":[{"given":"Srinivas","family":"Vadrevu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fatih","family":"Gelgi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saravanakumar","family":"Nagarajan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hasan","family":"Davulcu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022020220471192300_b1","doi-asserted-by":"crossref","unstructured":"Arasu, A. and Garcia\u2010Molina, H. (2003), \u201cExtracting structured data from web pages\u201d, paper presented at ACM SIGMOD Conference on Management of Data, San Diego, CA.","DOI":"10.1145\/872757.872799"},{"key":"key2022020220471192300_b2","doi-asserted-by":"crossref","unstructured":"Arocena, G.O. and Mendelzon, A.O. (1998), \u201cWeboql: restructuring documents, databases, and webs\u201d, paper presented at International Conference on Data Engineering.","DOI":"10.1002\/(SICI)1096-9942(1999)5:3<127::AID-TAPO2>3.0.CO;2-X"},{"key":"key2022020220471192300_b3","unstructured":"Baeza\u2010Yates, R. and Ribeiro\u2010Neto, B. (1999), Modern Information Retrieval, Addison\u2010Wesley, Reading, MA."},{"key":"key2022020220471192300_b4","doi-asserted-by":"crossref","unstructured":"Berwick, R.C. and Pilato, S. (1987), \u201cLearning syntax by automata induction\u201d, Machine Learning, Vol. 2, pp. 9\u201038.","DOI":"10.1007\/BF00058753"},{"key":"key2022020220471192300_b5","unstructured":"Crescenzi, V., Mecca, G. and Merialdo, P. (2001), \u201cRoadrunner: towards automatic data extraction from large web sites\u201d, Proceedings of 27th International Conference on Very Large Data Bases, pp. 109\u201018."},{"key":"key2022020220471192300_b6","unstructured":"Chung, C.Y., Gertz, M. and Sundaresan, N. (2002), \u201cReverse engineering for web data: from visual to semantic structures\u201d, ICDE."},{"key":"key2022020220471192300_b7","doi-asserted-by":"crossref","unstructured":"Davulcu, H., Vadrevu, S. and Nagarajan, S. (2003a), \u201cOntoMiner: bootstrapping and populating ontologies from domain specific web sites\u201d, paper presented at First International Workshop on Semantic Web and Databases, Berlin, September.","DOI":"10.1145\/1013367.1013545"},{"key":"key2022020220471192300_b8","doi-asserted-by":"crossref","unstructured":"Davulcu, H., Vadrevu, S. and Nagarajan, S. (2003b), \u201cOntoMiner: bootstrapping and populating ontologies from domain specific web sites\u201d, IEEE Intelligent Systems, Vol. 18 No. 5.","DOI":"10.1109\/MIS.2003.1234766"},{"key":"key2022020220471192300_b9","doi-asserted-by":"crossref","unstructured":"Davulcu, H., Vadrevu, S. and Nagarajan, S. (2005), \u201cOntoMiner: automated metadata and instance extraction from news websites\u201d, International Journal of Web and Grid Services, Vol. 1 No. 2, pp. 196\u2010221.","DOI":"10.1504\/IJWGS.2005.008320"},{"key":"key2022020220471192300_b10","doi-asserted-by":"crossref","unstructured":"Dill, S., Eiron, N., Gibson, D., Gruhl, D., Guha, R., Jhingran, T.A., Rajagopalan, S., Tomkins, A. and Tomlin, J.A. (2003), \u201cSemtag and seeker: bootstrapping the semantic web via automated semantic annotation\u201d, paper presented at WWW Conference.","DOI":"10.1145\/775152.775178"},{"key":"key2022020220471192300_b11","doi-asserted-by":"crossref","unstructured":"Doorenbos, R.B., Etzioni, O. and Weld, D.S. (1997), \u201cA scalable comparison\u2010shopping agent for the world\u2010wide web\u201d, in Johnson, W.L. and Hayes\u2010Roth, B. (Eds), Proceedings of the First International Conference on Autonomous Agents (Agents'97), ACM Press, Marina del Rey, CA, USA.","DOI":"10.1145\/267658.267666"},{"key":"key2022020220471192300_b12","doi-asserted-by":"crossref","unstructured":"Embley, D.W., Campbell, D.M., Smith, R.D. and Liddle, S.W. (1998), \u201cOntology\u2010based extraction and structuring of information from data\u2010rich unstructured documents\u201d, paper presented at Intl. Conference on Knowledge Management.","DOI":"10.1145\/288627.288641"},{"key":"key2022020220471192300_b13","doi-asserted-by":"crossref","unstructured":"Etzioni, O., Cafarella, M., Downey, D., Kok, S., Popescu, A., Shaked, S.T., Weld, D.S. and Yates, A. (2004), \u201cWeb\u2010scale information extraction in knowitall (preliminary results)\u201d, paper presented at WWW Conference.","DOI":"10.1145\/988672.988687"},{"key":"key2022020220471192300_b14","doi-asserted-by":"crossref","unstructured":"Gold, E.M. (1978), \u201cComplexity of automaton identification from given sets\u201d, Information and Control, Vol. 37, pp. 302\u201020.","DOI":"10.1016\/S0019-9958(78)90562-4"},{"key":"key2022020220471192300_b15","unstructured":"Gus\ufb01eld, D. (1997), Algorithms on Strings, Tree, and Sequences, Cambridge University Press, Cambridge."},{"key":"key2022020220471192300_b16","doi-asserted-by":"crossref","unstructured":"Hammer, J., Garcia\u2010Molina, H., Nestorov, S., Yerneni, R., Breunig, M.M. and Vassalos, V. (1997), \u201cTemplate\u2010based wrappers in the tsimmis system\u201d, paper presented at ACM SIGMOD Conference on Management of Data.","DOI":"10.1145\/253260.253395"},{"key":"key2022020220471192300_b17","doi-asserted-by":"crossref","unstructured":"Hong, T.W. and Clark, K.L. (2001), \u201cUsing grammatical inference to automate information extraction from the web\u201d, Principles of Knowledge Discovery and Data Mining.","DOI":"10.1007\/3-540-44794-6_18"},{"key":"key2022020220471192300_b18","doi-asserted-by":"crossref","unstructured":"Kifer, M., Lausen, G. and Wu, J. (1995), \u201cLogical foundations of object\u2010oriented and frame\u2010based languages\u201d, Journal of the ACM, Vol. 42 No. 4, pp. 741\u2010843.","DOI":"10.1145\/210332.210335"},{"key":"key2022020220471192300_b19","unstructured":"Kushmerick, N., Weld, D.S. and Doorenbos, R.B. (1997), \u201cWrapper induction for information extraction\u201d, paper presented at International Joint Conference on Artificial Intelligence, pp. 729\u201037."},{"key":"key2022020220471192300_b20","doi-asserted-by":"crossref","unstructured":"Lerman, K., Getoor, L., Minton, S. and Knoblock, C. (2004), \u201cUsing the structure of web sites for automatic segmentation of tables\u201d, paper presented at ACM SIGMOD Conference on Management of Data.","DOI":"10.1145\/1007568.1007584"},{"key":"key2022020220471192300_b21","unstructured":"Maedche, A. and Staab, S. (2000), \u201cDiscovering conceptual relations from text\u201d, Proceedings of ECAI\u201000, IOS Press, Amsterdam, pp. 321\u20105."},{"key":"key2022020220471192300_b22","doi-asserted-by":"crossref","unstructured":"Muslea, I., Minton, S. and Knoblock, C. (1999), \u201cA hierarchical approach to wrapper induction\u201d, Proceedings of the Third International Conference on Autonomous Agents (Agents'99), pp. 190\u20107.","DOI":"10.1145\/301136.301191"},{"key":"key2022020220471192300_b23","doi-asserted-by":"crossref","unstructured":"Navigli, R., Velardi, P. and Gangemi, A. (2003), \u201cOntology learning and its application to automated terminology translation\u201d, IEEE Intelligent Systems, Vol. 18 No. 1.","DOI":"10.1109\/MIS.2003.1179190"},{"key":"key2022020220471192300_b25","unstructured":"Rilo, E. and Shepherd, J. (1997), \u201cA corpus\u2010based approach for building semantic lexicons\u201d, in Cardie, C. and Weischedel, R. (Eds), Proceedings of the 2nd Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Somerset, NJ."},{"key":"key2022020220471192300_b26","doi-asserted-by":"crossref","unstructured":"Vadrevu, S., Gelgi, F. and Davulcu, H. (2005a), Semantic partitioning web pages, paper presented at the 6th International Conference on Web Information Systems Engineering (WISE).","DOI":"10.1007\/11581062_9"},{"key":"key2022020220471192300_b27","unstructured":"Vadrevu, S., Nagarajan, S., Gelgi, F. and Davulcu, H. (2005b), \u201cAutomated metadata and instance extraction from news web sites\u201d, paper presented at The IEEE\/WIC\/ACM International Conference on Web Intelligence, September."},{"key":"key2022020220471192300_b28","doi-asserted-by":"crossref","unstructured":"Yang, G. and Kifer, M. (2000), \u201cFlora: implementing an efficient dood system using a tabling logic engine\u201d, paper presented at International Conference on Deductive and Object\u2010Oriented Databases, London.","DOI":"10.1007\/3-540-44957-4_72"},{"key":"key2022020220471192300_b29","doi-asserted-by":"crossref","unstructured":"Yang, G., Kifer, M. and Zhao, C. (2003a), Flora\u20102: A Rule\u2010Based Knowledge Representation and Inference Infrastructure for the Semantic Web, Catania, Sicily.","DOI":"10.1007\/978-3-540-39964-3_43"},{"key":"key2022020220471192300_b30","unstructured":"Yang, G., Tan, W., Mukherjee, S., Ramakrishnan, I.V. and Davulcu, H. (2003b), \u201cOn the power of semantic partitioning of web documents\u201d, paper presented at the Workshop on Information Integration on the Web, Acapulco, Mexico."},{"key":"key2022020220471192300_b31","doi-asserted-by":"crossref","unstructured":"Yang, J., Seo, H. and Choi, J. (2001), \u201cMORPHEUS: a customized comparison\u2010shopping agent\u201d, paper presented at 5th International Conference on Autonomous Agents (Agents\u20102001), Montreal, Canada, pp. 63\u20104.","DOI":"10.1145\/375735.375870"},{"key":"key2022020220471192300_b32","unstructured":"Yang, Y. and Zhang, H. (2001), \u201cHtml page analysis based on visual cues\u201d, ICDAR."},{"key":"key2022020220471192300_frd1","unstructured":"Manola, F. and Miller, E. (2004), Rdf Primer, W3C Recommendation, available at: www.w3.org\/TR\/rdf\u2010primer\/."}],"container-title":["Online Information Review"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/14684520610675807","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/14684520610675807\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/14684520610675807\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T00:40:35Z","timestamp":1753404035000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/oir\/article\/30\/3\/278-296\/314851"}},"subtitle":[],"editor":[{"given":"Miguel\u2010Angel","family":"Sicilia","sequence":"first","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2006,5,1]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2006,5,1]]}},"alternative-id":["10.1108\/14684520610675807"],"URL":"https:\/\/doi.org\/10.1108\/14684520610675807","relation":{},"ISSN":["1468-4527"],"issn-type":[{"type":"print","value":"1468-4527"}],"subject":[],"published":{"date-parts":[[2006,5,1]]}}}