{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T19:04:03Z","timestamp":1754161443626,"version":"3.41.2"},"reference-count":18,"publisher":"Emerald","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2010,11,23]]},"abstract":"<jats:sec>\n                  <jats:title>Purpose<\/jats:title>\n                  <jats:p>The aim of this paper is to propose a strategy for extracting information from web tables.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Design\/methodology\/approach<\/jats:title>\n                  <jats:p>The paper presents a strategy for extracting information from web tables of semi-structured web pages (WPs) by handling the issue of synonym which emerges as these WPs have been designed and created without referring to any standards or guidelines.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Findings<\/jats:title>\n                  <jats:p>The paper finds that this strategy extracts information with high precision, and extracts the attributes besides the sub-attributes that describe the extracted attributes and values of the sub-attributes.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Practical implications<\/jats:title>\n                  <jats:p>Experiment conducted on the Nokia products domain demonstrated that the proposed strategy extracts information from web tables with high precision which is 98.98 percent.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Originality\/value<\/jats:title>\n                  <jats:p>This paper contributes to the research on extracting information.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1108\/17440081011090239","type":"journal-article","created":{"date-parts":[[2010,11,20]],"date-time":"2010-11-20T07:11:57Z","timestamp":1290237117000},"page":"304-318","source":"Crossref","is-referenced-by-count":0,"title":["A strategy for extracting information from semi-structured web pages"],"prefix":"10.1108","volume":"6","author":[{"given":"Mahmoud","family":"Shaker","sequence":"first","affiliation":[{"name":"Department of Computer Science, Faculty of Computer Science and Information Technology, Universiti Putra Malaysia, Serdang, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hamidah","family":"Ibrahim","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computer Science and Information Technology, Universiti Putra Malaysia, Serdang, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aida","family":"Mustapha","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computer Science and Information Technology, Universiti Putra Malaysia, Serdang, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lili","family":"Nurliyana Abdullah","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computer Science and Information Technology, Universiti Putra Malaysia, Serdang, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"2025072819304092000_b1","doi-asserted-by":"crossref","unstructured":"Abels, S.\n           and Hahn, A. (2006), \u201cReclassification of electronic product catalogs: the apricot approach and its evaluation results\u201d, Journal of Informing Science, Vol. 9, pp. 31-47.","DOI":"10.28945\/470"},{"key":"2025072819304092000_b2","doi-asserted-by":"crossref","unstructured":"Arnicans, G.\n           and Karnitis, G. (2006), \u201cIntelligent integration of information from semi-structured web data sources on the base of ontology and meta-models\u201d, Proceedings of the 7th International Baltic Conference, Vilnius, Lithuania, pp. 177-86.","DOI":"10.1109\/DBIS.2006.1678494"},{"key":"2025072819304092000_b3","doi-asserted-by":"crossref","unstructured":"Ashraf, F.\n          , Ozyer, T. and Alhajj, R. (2008), \u201cEmploying clustering techniques for automatic information extraction from HTML documents\u201d, Journal of IEEE Transactions on Systems, Man, and Cybernetics, Part C: Applications and Reviews, Vol. 38, pp. 660-73.","DOI":"10.1109\/TSMCC.2008.923882"},{"key":"2025072819304092000_b4","doi-asserted-by":"crossref","unstructured":"Gatterbauer, W.\n          , Bohunsky, P., Herzog, M., Krupl, B. and Pollak, B. (2007), \u201cTowards domain-independent information extraction from web tables\u201d, Proceedings of the 16th International Conference on World Wide Web, ACM Press, Canada, pp. 71-80.","DOI":"10.1145\/1242572.1242583"},{"key":"2025072819304092000_b5","unstructured":"Hong, F.\n           and Zhao, Z. (2006), \u201cInformation extraction system in large-scale web\u201d, Journal of Communications and Information Technology, ISCIT, Vol. 2, pp. 809-12."},{"key":"2025072819304092000_b6","doi-asserted-by":"crossref","unstructured":"Jung, S.W.\n          , Sung, K.H., Park, T.W. and Kwon, H.C. (2001), \u201cIntelligent integration of information on the internet for travelers on demand\u201d, Proceedings of ISIE, IEEE International Symposium, Korea, Vol. 1, pp. 338-42.","DOI":"10.1109\/ISIE.2001.931810"},{"key":"2025072819304092000_b7","unstructured":"Kaiser, K.\n           (2005), \u201cModeling treatment processes using information extraction\u201d, PhD thesis, Vienna University of Technology, Vienna."},{"key":"2025072819304092000_b8","unstructured":"Kaiser, K.\n           and Miksch, S. (2005), \u201cInformation extraction: a survey\u201d, Institute of Software Technology and Interactive Systems, Vienna University of Technology, Vienna."},{"key":"2025072819304092000_b9","doi-asserted-by":"crossref","unstructured":"Lam, M.I.\n          , Gong, Z. and Muyeba, M. (2008), \u201cA method for web information extraction\u201d, Proceedings of the 10th Asia-Pacific Web Conference, APWeb, Shenyang, Vol. 4976, pp. 383-94.","DOI":"10.1007\/978-3-540-78849-2_39"},{"key":"2025072819304092000_b10","doi-asserted-by":"crossref","unstructured":"Nachouki, G.\n           (2006), \u201cA method for information extraction from the web\u201d, Journal of Information and Communication Technologies, ICTTA, Vol. 1, pp. 517-21.","DOI":"10.1109\/ICTTA.2006.1684424"},{"key":"2025072819304092000_b11","doi-asserted-by":"crossref","unstructured":"Nanda, J.\n          , Thevenot, H.J. and Simpson, T.W. (2005), \u201cProduct family design knowledge representation, integration, and reuse\u201d, Proceedings of Information Reuse and Integration, Las Vegas, NV, pp. 32-7.","DOI":"10.1109\/IRI-05.2005.1506445"},{"issue":"2","key":"2025072819304092000_b12","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1115\/1.2190237","article-title":"A methodology for product family ontology development using formal concept analysis and web ontology language","volume":"6","author":"Nanda","year":"2006","journal-title":"Journal of Computing and Information Science in Engineering"},{"issue":"2","key":"2025072819304092000_b13","first-page":"173","article-title":"Product family design knowledge representation, aggregation, reuse, and analysis","volume":"21","author":"Nanda","year":"2007","journal-title":"Journal of Artificial Intelligence for Engineering Design, Analysis and Manufacturing"},{"key":"2025072819304092000_b14","doi-asserted-by":"crossref","unstructured":"Saggion, H.\n          , Funk, A., Maynard, D. and Bontcheva, K. (2007), \u201cOntology-based information extraction for business intelligence\u201d, Proceedings of the 6th International Semantic Web Conference and the 2nd Asian Semantic Web Conference, Durham, UK, pp. 843-56.","DOI":"10.1007\/978-3-540-76298-0_61"},{"key":"2025072819304092000_b15","doi-asserted-by":"crossref","unstructured":"Son, J.W.\n          , Lee, J.A., Park, S.B., Song, H.J., Lee, S.J. and Park, S.Y. (2008), \u201cDiscriminating meaningful web tables from decorative tables using composite kernel\u201d, Proceedings of the ACM International Conference on Web Intelligence and Intelligent Agent Technology, Vol. 1, pp. 368-71.","DOI":"10.1109\/WIIAT.2008.241"},{"issue":"2","key":"2025072819304092000_b16","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1177\/1063293X07079318","article-title":"An index-based method to manage the tradeoff between diversity and commonality during product family design","volume":"15","author":"Thevenot","year":"2007","journal-title":"Journal of Concurrent Engineering: Research & Applications"},{"issue":"2","key":"2025072819304092000_b17","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/s11280-007-0021-1","article-title":"Information extraction from web pages using presentation regularities and domain knowledge","volume":"10","author":"Vadrevu","year":"2007","journal-title":"Journal of World Wide Web"},{"key":"2025072819304092000_b18","doi-asserted-by":"crossref","unstructured":"Zhao, H.\n          , Meng, W. and Yu, C. (2007), \u201cMining templates from search result records of search engines\u201d, Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Jose, CA, USA, pp. 884-93.","DOI":"10.1145\/1281192.1281286"}],"container-title":["International Journal of Web Information Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/17440081011090239","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/ijwis\/article-pdf\/6\/4\/304\/1115078\/17440081011090239.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ijwis\/article-pdf\/6\/4\/304\/1115078\/17440081011090239.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,28]],"date-time":"2025-07-28T23:30:53Z","timestamp":1753745453000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ijwis\/article\/6\/4\/304\/163780\/A-strategy-for-extracting-information-from-semi"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,11,23]]},"references-count":18,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2010,11,23]]}},"URL":"https:\/\/doi.org\/10.1108\/17440081011090239","relation":{},"ISSN":["1744-0084","1744-0092"],"issn-type":[{"type":"print","value":"1744-0084"},{"type":"electronic","value":"1744-0092"}],"subject":[],"published":{"date-parts":[[2010,11,23]]}}}