{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T10:41:29Z","timestamp":1697107289685},"reference-count":22,"publisher":"Wiley","issue":"12","license":[{"start":{"date-parts":[[2005,8,4]],"date-time":"2005-08-04T00:00:00Z","timestamp":1123113600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Am. Soc. Inf. Sci."],"published-print":{"date-parts":[[2005,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Some collections cover many topics, while others are narrowly focused on a limited number of topics. We introduce the concept of the \u201cscope\u201d of a collection of documents and we compare two ways of measuring it. These measures are based on the distances between documents. The first uses the overlap of words between pairs of documents. The second measure uses a novel method that calculates the semantic relatedness to pairs of words from the documents. Those values are combined to obtain an overall distance between the documents. The main validation for the measures compared Web pages categorized by Yahoo. Sets of pages sampled from broad categories were determined to have a higher scope than sets derived from subcategories. The measure was significant and confirmed the expected difference in scope. Finally, we discuss other measures related to scope.<\/jats:p>","DOI":"10.1002\/asi.20202","type":"journal-article","created":{"date-parts":[[2005,8,4]],"date-time":"2005-08-04T20:49:47Z","timestamp":1123188587000},"page":"1243-1249","source":"Crossref","is-referenced-by-count":1,"title":["Metrics for the scope of a collection"],"prefix":"10.1002","volume":"56","author":[{"given":"Robert B.","family":"Allen","sequence":"first","affiliation":[]},{"given":"Yejun","family":"Wu","sequence":"additional","affiliation":[]}],"member":"311","published-online":{"date-parts":[[2005,8,4]]},"reference":[{"key":"e_1_2_9_2_1","doi-asserted-by":"crossref","unstructured":"Allen R.B. &Wu Y.(2002).Generality of texts. International Conference on Asian Digital Libraries Lecture Notes in Computer Science Vol. 2555 (pp. 111\u2013116).","DOI":"10.1007\/3-540-36227-4_11"},{"issue":"7","key":"e_1_2_9_3_1","article-title":"The University of Michigan Digital Library Project: The testbed","volume":"2","author":"Atkins D.E.","year":"1996","journal-title":"D\u2010Lib Magazine"},{"key":"e_1_2_9_4_1","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/7287.001.0001"},{"key":"e_1_2_9_5_1","volume-title":"Frequency analysis of English usage: Lexicon and grammar","author":"Francis W.N.","year":"1982"},{"key":"e_1_2_9_6_1","first-page":"229","article-title":"GlOSS: Text\u2010source discovery over the Internet","volume":"17","author":"Gravano L.","year":"1999","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_2_9_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/379437.379449"},{"key":"e_1_2_9_8_1","doi-asserted-by":"publisher","DOI":"10.1037\/0033-295X.104.2.211"},{"key":"e_1_2_9_9_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.280.5360.98"},{"key":"e_1_2_9_10_1","unstructured":"Lin D.(1998 July).An information\u2010theoretic definition of similarity. Paper presented at the 15th International Conference on Machine Learning (ICML\u201098) Madison WI."},{"key":"e_1_2_9_11_1","doi-asserted-by":"publisher","DOI":"10.1080\/01690969108406936"},{"key":"e_1_2_9_12_1","volume-title":"Proceedings of the ACM Digital Libraries Conference","author":"Phanouriou C.","year":"1999"},{"key":"e_1_2_9_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/21.24528"},{"key":"e_1_2_9_14_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.514"},{"key":"e_1_2_9_15_1","first-page":"399","volume-title":"Proceedings of the Cognitive Science Society","author":"Resnik P.","year":"2000"},{"key":"e_1_2_9_16_1","first-page":"627","article-title":"Contextual correlates of synonymy","volume":"8","author":"Rubinstein H.","year":"1965","journal-title":"Computational Linguistics"},{"key":"e_1_2_9_17_1","volume-title":"Statistical methods. Ames, IA: Iowa State University","author":"Snedcor G.W..","year":"1967"},{"key":"e_1_2_9_18_1","unstructured":"Steinberg S.(1996).Seek and ye shall find maybe. WIRED Archive 4.05 5."},{"key":"e_1_2_9_18_2","unstructured":"Retrieved June 12 2005 fromhttp:\/www.wired.com\/wired\/archive\/4.05\/indexweb.html"},{"key":"e_1_2_9_19_1","doi-asserted-by":"crossref","unstructured":"Turney P.D.(2001 Sept).Mining the web for synonyms: PMI\u2010IR versus LSA on TOEFL. In European Conference on Machine Learning.","DOI":"10.1007\/3-540-44795-4_42"},{"key":"e_1_2_9_20_1","volume-title":"Information retrieval","author":"van Rijsbergen C.J.","year":"1979"},{"key":"e_1_2_9_21_1","first-page":"365","article-title":"Computer techniques for studying coverage, overlap, and gaps in collections","volume":"12","author":"White H.","year":"1987","journal-title":"Journal of Academic Librarianship"},{"key":"e_1_2_9_22_1","first-page":"133","article-title":"Verb semantics and lexical selection","author":"Wu Z.","year":"1994","journal-title":"Association for Computational Linguistics"}],"container-title":["Journal of the American Society for Information Science and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fasi.20202","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/asi.20202","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,11]],"date-time":"2023-10-11T20:47:21Z","timestamp":1697057241000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/asi.20202"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,8,4]]},"references-count":22,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2005,10]]}},"alternative-id":["10.1002\/asi.20202"],"URL":"https:\/\/doi.org\/10.1002\/asi.20202","archive":["Portico"],"relation":{},"ISSN":["1532-2882","1532-2890"],"issn-type":[{"value":"1532-2882","type":"print"},{"value":"1532-2890","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,8,4]]}}}