{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,4]],"date-time":"2026-04-04T06:11:21Z","timestamp":1775283081165,"version":"3.50.1"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2004,9,1]],"date-time":"2004-09-01T00:00:00Z","timestamp":1093996800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMOD Rec."],"published-print":{"date-parts":[[2004,9]]},"abstract":"<jats:p>The Web has been rapidly \"deepened\" by the prevalence of databases online. With the potentially unlimited information hidden behind their query interfaces, this \"deep Web\" of searchable databses is clearly an important frontier for data access. This paper surveys this relatively unexplored frontier, measuring characteristics pertinent to both exploring and integrating structured Web sources. On one hand, our \"macro\" study surveys the deep Web at large, in April 2004, adopting the random IP-sampling approach, with one million samples. (How large is the deep Web? How is it covered by current directory services?) On the other hand, our \"micro\" study surveys source-specific characteristics over 441 sources in eight representative domains, in December 2002. (How \"hidden\" are deep-Web sources? How do search engines cover their data? How complex and expressive are query forms?) We report our observations and publish the resulting datasets to the research community. We conclude with several implications (of our own) which, while necessarily subjective, might help shape research directions and solutions.<\/jats:p>","DOI":"10.1145\/1031570.1031584","type":"journal-article","created":{"date-parts":[[2005,11,9]],"date-time":"2005-11-09T22:23:27Z","timestamp":1131575007000},"page":"61-70","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":215,"title":["Structured databases on the web"],"prefix":"10.1145","volume":"33","author":[{"given":"Kevin Chen-Chuan","family":"Chang","sequence":"first","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"He","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengkai","family":"Li","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mitesh","family":"Patel","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhen","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2004,9]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"July","year":"2000","unstructured":"BrightPlanet.com. The deep web: Surfacing hidden value. Accessible at http:\/\/brightplanet.com , July 2000 .]] BrightPlanet.com. The deep web: Surfacing hidden value. Accessible at http:\/\/brightplanet.com, July 2000.]]"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1038\/21987"},{"key":"e_1_2_1_3_1","unstructured":"Ed O'Neill Brian Lavoie and Rick Bennett. Web characterization. Accessible at \"http:\/\/wcp.oclc.org\".]]  Ed O'Neill Brian Lavoie and Rick Bennett. Web characterization. Accessible at \"http:\/\/wcp.oclc.org\".]]"},{"key":"e_1_2_1_4_1","unstructured":"GNU. wget. Accessible at \"http:\/\/www.gnu.org\/software\/wget\/wget.html\".]]  GNU. wget. Accessible at \"http:\/\/www.gnu.org\/software\/wget\/wget.html\".]]"},{"key":"e_1_2_1_5_1","volume-title":"The UIUC web integration repository. Computer Science Department","author":"Chen-Chuan Chang Kevin","year":"2003","unstructured":"Kevin Chen-Chuan Chang , Bin He , Chengkai Li , and Zhen Zhang . The UIUC web integration repository. Computer Science Department , University of Illinois at Urbana-Champaign. http :\/\/metaquerier.cs.uiuc.edu\/repository, 2003 .]] Kevin Chen-Chuan Chang, Bin He, Chengkai Li, and Zhen Zhang. The UIUC web integration repository. Computer Science Department, University of Illinois at Urbana-Champaign. http:\/\/metaquerier.cs.uiuc.edu\/repository, 2003.]]"},{"key":"e_1_2_1_6_1","volume-title":"Human Behavior and the Principle of Least Effort","author":"Zipf G. K.","year":"1949","unstructured":"G. K. Zipf . Human Behavior and the Principle of Least Effort . Addison-Wesley , Cambridge, Massachusetts , 1949 .]] G. K. Zipf. Human Behavior and the Principle of Least Effort. Addison-Wesley, Cambridge, Massachusetts, 1949.]]"},{"key":"e_1_2_1_7_1","first-page":"55","volume-title":"WebDB (Informal Proceedings)","author":"Cohen William W.","year":"1999","unstructured":"William W. Cohen . Some practical observations on integration of web information . In WebDB (Informal Proceedings) , pages 55 -- 60 , 1999 .]] William W. Cohen. Some practical observations on integration of web information. In WebDB (Informal Proceedings), pages 55--60, 1999.]]"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/5254.722342"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/290593.290605"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/375663.375671"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304224"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/297117.297123"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/347319.346329"},{"key":"e_1_2_1_14_1","first-page":"14","volume-title":"Proceedings of 24th International Conference on Very Large Data Bases","author":"Meng Weiyi","year":"1998","unstructured":"Weiyi Meng , King-Lup Liu , Clement T. Yu , Xiaodong Wang , Yuhsi Chang , and Naphtali Rishe . Determining text databases to search in the internet . In Proceedings of 24th International Conference on Very Large Data Bases , pages 14 -- 25 , New York City, New York, USA , August 1998 . Morgan Kaufmann.]] Weiyi Meng, King-Lup Liu, Clement T. Yu, Xiaodong Wang, Yuhsi Chang, and Naphtali Rishe. Determining text databases to search in the internet. In Proceedings of 24th International Conference on Very Large Data Bases, pages 14--25, New York City, New York, USA, August 1998. Morgan Kaufmann.]]"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/645502.656100"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007583"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/872757.872784"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1014052.1014071"},{"key":"e_1_2_1_19_1","first-page":"251","volume-title":"Proceedings of the 22nd VLDB Conference","author":"Levy Alon Y.","year":"1996","unstructured":"Alon Y. Levy , Anand Rajaraman , and Joann J. Ordille . Querying heterogeneous information sources using source descriptions . In Proceedings of the 22nd VLDB Conference , pages 251 -- 262 , Bombay, India , 1996 . VLDB Endowment, Saratoga, Calif.]] Alon Y. Levy, Anand Rajaraman, and Joann J. Ordille. Querying heterogeneous information sources using source descriptions. In Proceedings of the 22nd VLDB Conference, pages 251--262, Bombay, India, 1996. VLDB Endowment, Saratoga, Calif.]]"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/645481.655600"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/373626.373713"},{"key":"e_1_2_1_22_1","volume-title":"August","author":"Gravano Luis","year":"1996","unstructured":"Luis Gravano , Chen-Chuan K. Chang , H\u00e9ctor Garc\u00eda-Molina , and Andreas Paepcke . STARTS: Stanford protocol proposal for internet retrieval and search. Accessible at http:\/\/www-db.stanford.edu\/~gravano\/starts.html , August 1996 .]] Luis Gravano, Chen-Chuan K. Chang, H\u00e9ctor Garc\u00eda-Molina, and Andreas Paepcke. STARTS: Stanford protocol proposal for internet retrieval and search. Accessible at http:\/\/www-db.stanford.edu\/~gravano\/starts.html, August 1996.]]"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/645923.670994"},{"key":"e_1_2_1_24_1","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1007\/3-540-48054-4_19","volume-title":"Advances in Conceptual Modeling: ER '99 Workshops on Evolution and Change in Data Management, Reverse Engineering in Information Systems, and the World Wide Web and Conceptual Modeling","author":"Lud\u00e4scher Bertram","year":"1999","unstructured":"Bertram Lud\u00e4scher and Amarnath Gupta . Modeling interactive web sources for information mediation . In Advances in Conceptual Modeling: ER '99 Workshops on Evolution and Change in Data Management, Reverse Engineering in Information Systems, and the World Wide Web and Conceptual Modeling , Paris, France , November 15--18, 1999 , Proceedings, volume 1727 of Lecture Notes in Computer Science , pages 225 -- 238 . Springer , 1999.]] Bertram Lud\u00e4scher and Amarnath Gupta. Modeling interactive web sources for information mediation. In Advances in Conceptual Modeling: ER '99 Workshops on Evolution and Change in Data Management, Reverse Engineering in Information Systems, and the World Wide Web and Conceptual Modeling, Paris, France, November 15--18, 1999, Proceedings, volume 1727 of Lecture Notes in Computer Science, pages 225--238. Springer, 1999.]]"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/645927.672025"},{"key":"e_1_2_1_26_1","volume-title":"ICDE Conference","author":"Caverlee James","year":"2004","unstructured":"James Caverlee , Ling Liu , and David Buttler . Probe, cluster, and discover : Focused extraction of qa-pagelets from the deep web . In ICDE Conference , 2004 .]] James Caverlee, Ling Liu, and David Buttler. Probe, cluster, and discover: Focused extraction of qa-pagelets from the deep web. In ICDE Conference, 2004.]]"},{"key":"e_1_2_1_27_1","first-page":"109","volume-title":"The VLDB Journal 2001","author":"Crescenzi Valter","year":"2001","unstructured":"Valter Crescenzi , Giansalvatore Mecca , and Paolo Merialdo . Roadrunner : Towards automatic data extraction from large web sites . In The VLDB Journal 2001 , pages 109 -- 118 , 2001 .]] Valter Crescenzi, Giansalvatore Mecca, and Paolo Merialdo. Roadrunner: Towards automatic data extraction from large web sites. In The VLDB Journal 2001, pages 109--118, 2001.]]"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/304182.304223"}],"container-title":["ACM SIGMOD Record"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1031570.1031584","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1031570.1031584","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T16:25:03Z","timestamp":1750263903000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1031570.1031584"}},"subtitle":["observations and implications"],"short-title":[],"issued":{"date-parts":[[2004,9]]},"references-count":28,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2004,9]]}},"alternative-id":["10.1145\/1031570.1031584"],"URL":"https:\/\/doi.org\/10.1145\/1031570.1031584","relation":{},"ISSN":["0163-5808"],"issn-type":[{"value":"0163-5808","type":"print"}],"subject":[],"published":{"date-parts":[[2004,9]]},"assertion":[{"value":"2004-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}