{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T20:34:57Z","timestamp":1761597297098,"version":"3.41.2"},"reference-count":42,"publisher":"Emerald","issue":"4","license":[{"start":{"date-parts":[[2011,7,26]],"date-time":"2011-07-26T00:00:00Z","timestamp":1311638400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2011,7,26]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>The purpose of this paper is to propose a framework for describing and evaluating the representativeness of a small set of search results extracted from the original results: this is deemed desirable in information retrieval in enterprise information systems.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>The paper proposes a combined measure, namely <jats:italic>RF<\/jats:italic><jats:sub><jats:italic>\u03b2<\/jats:italic><\/jats:sub>, to evaluate the extracted small set in terms of the notions of coverage and redundancy. Data experiments were conducted on three different extraction strategies to evaluate the representativeness, i.e. coverage and redundancy.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>Both from intuitive and experimental perspectives, the proposed coverage measure, redundancy measure and <jats:italic>RF<\/jats:italic><jats:sub><jats:italic>\u03b2<\/jats:italic><\/jats:sub> measure could effectively evaluate the representativeness.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Research limitations\/implications<\/jats:title><jats:p>The search results, e.g. in the form of documents and texts, are modeled using a vector space model and cosine similarity. Semantic models and linguistic models could be further introduced into this research to improve the proposed measures.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Practical implications<\/jats:title><jats:p>With the rapidly growing need for information retrieval in enterprise information systems, the representativeness of search results become more desirable and important for search engine users. The well\u2010designed representativeness measures will help them achieve satisfactory results.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>The originality of the paper lies in the definition of representativeness of a small set of search results extracted from the original results. This focuses on the two aspects of coverage rate and redundancy rate both from intuitive and experimental perspectives.<\/jats:p><\/jats:sec>","DOI":"10.1108\/17410391111148567","type":"journal-article","created":{"date-parts":[[2011,7,25]],"date-time":"2011-07-25T11:10:11Z","timestamp":1311592211000},"page":"310-321","source":"Crossref","is-referenced-by-count":13,"title":["A combined measure for representative information retrieval in enterprise information systems"],"prefix":"10.1108","volume":"24","author":[{"given":"Baojun","family":"Ma","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Wei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guoqing","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022031520112261800_b1","doi-asserted-by":"crossref","unstructured":"Aliguliyev, R.M. (2009), \u201cClustering of document collection \u2013 a weighting approach\u201d, Expert Systems with Applications, Vol. 36 No. 4, pp. 7904\u201016.","DOI":"10.1016\/j.eswa.2008.11.017"},{"key":"key2022031520112261800_b2","unstructured":"Baeza\u2010Yates, R. and Ribeiro\u2010Neto, B. (1999), Modern Information Retrieval, ACM Press, New York, NY."},{"key":"key2022031520112261800_b3","doi-asserted-by":"crossref","unstructured":"Balog, K. (2007), \u201cPeople search in the enterprise\u201d, Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Amsterdam, The Netherlands, p. 916.","DOI":"10.1145\/1277741.1277985"},{"key":"key2022031520112261800_b4","doi-asserted-by":"crossref","unstructured":"Broder, A.Z. and Ciccolo, A.C. (2004), \u201cTowards the next generation of enterprise search technology\u201d, IBM Systems Journal, Vol. 43 No. 3, pp. 451\u20104.","DOI":"10.1147\/sj.433.0451"},{"key":"key2022031520112261800_b5","doi-asserted-by":"crossref","unstructured":"Bruno, N., Chaudhri, S. and Gravand, L. (2002), \u201cTop\u2010k selection queries over relational databases\u201d, ACM Transactions on Database Systems, Vol. 27 No. 2, pp. 153\u201087.","DOI":"10.1145\/568518.568519"},{"key":"key2022031520112261800_b6","doi-asserted-by":"crossref","unstructured":"Buckley, C. and Voorhees, E.M. (2000), \u201cEvaluating evaluation measure stability\u201d, Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and development in Information Retrieval, ACM Press, New York, NY, pp. 33\u201040.","DOI":"10.1145\/345508.345543"},{"key":"key2022031520112261800_b7","doi-asserted-by":"crossref","unstructured":"Carbonell, J. and Goldstein, J. (1998), \u201cThe use of MMR, diversity\u2010based reranking for reordering documents and producing summaries\u201d, Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM Press, New York, NY, pp. 335\u20106.","DOI":"10.1145\/290941.291025"},{"key":"key2022031520112261800_b8","doi-asserted-by":"crossref","unstructured":"Carpineto, C., Osinski, S., Romano, G. and Weiss, D. (2009), \u201cA survey of web clustering engines\u201d, ACM Computing Surveys, Vol. 41 No. 3, p. 17.","DOI":"10.1145\/1541880.1541884"},{"key":"key2022031520112261800_b9","doi-asserted-by":"crossref","unstructured":"Fagin, R., Lotem, A. and Naor, M. (2003), \u201cOptimal aggregation algorithms for middleware\u201d, Journal of Computer and System Sciences, Vol. 66 No. 4, pp. 614\u201056.","DOI":"10.1016\/S0022-0000(03)00026-6"},{"key":"key2022031520112261800_b10","doi-asserted-by":"crossref","unstructured":"Grabmeier, J. and Rudolph, A. (2002), \u201cTechniques of cluster algorithms in data mining\u201d, Data Mining and Knowledge Discovery, Vol. 6 No. 4, pp. 303\u201060.","DOI":"10.1023\/A:1016308404627"},{"key":"key2022031520112261800_b12","unstructured":"Guntzer, U., Balke, W. and Kie\u00dfling, W. (2000), \u201cOptimizing multi\u2010feature queries for image databases\u201d, Proceedings of the 26th International Conference on Very Large Data Bases, Morgan Kaufmann Publishers, San Francisco, CA, pp. 419\u201028."},{"key":"key2022031520112261800_b13","unstructured":"Han, J.W. and Kamber, M. (2006), Data Mining: Concepts and Techniques, 2nd ed., Morgan Kaufman Publishers, San Francisco, CA."},{"key":"key2022031520112261800_b14","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R. and Friedman, J. (2001), The Elements of Statistical Learning \u2013 Data Mining, Inference, and Prediction, Springer\u2010Verlag, New York, NY.","DOI":"10.1007\/978-0-387-21606-5"},{"key":"key2022031520112261800_b15","unstructured":"Hawking, D. (2004), Challenges in Enterprise Search, ACM International Conference Proceedings Series, Dunedin, New Zealand, Vol. 52, pp. 15\u201024."},{"key":"key2022031520112261800_b18","doi-asserted-by":"crossref","unstructured":"Ilyas, I.F., Beskales, G. and Soliman, M.A. (2008), \u201cSurvey of top\u2010k query processing techniques in relational database systems\u201d, ACM Computing Surveys, Vol. 40 No. 4.","DOI":"10.1145\/1391729.1391730"},{"key":"key2022031520112261800_b19","doi-asserted-by":"crossref","unstructured":"Jain, A.K., Murty, M.N. and Flynn, P.J. (1999), \u201cData clustering: a review\u201d, ACM Computing Surveys, Vol. 31 No. 3, pp. 264\u2010323.","DOI":"10.1145\/331499.331504"},{"key":"key2022031520112261800_b21","doi-asserted-by":"crossref","unstructured":"Kraft, D.E. and Bookstein, A. (1978), \u201cEvaluation of information retrieval system: a decision theory approach\u201d, Journal of the American Society for Information Science, Vol. 29 No. 1, pp. 31\u201040.","DOI":"10.1002\/asi.4630290106"},{"key":"key2022031520112261800_b22","doi-asserted-by":"crossref","unstructured":"Li, Y., Zheng, Z. and Dai, H. (2005), \u201cKDD CUP\u20102005 report: facing a great challenge\u201d, SIGKDD Explorations, Vol. 7 No. 2, pp. 91\u20109.","DOI":"10.1145\/1117454.1117466"},{"key":"key2022031520112261800_b23","doi-asserted-by":"crossref","unstructured":"Lian, X. and Chen, L. (2009), \u201cTop\u2010k dominating queries in uncertain databases\u201d, Proceedings of the 12th International Conference on Extending Database Technology, ACM Press, New York, NY, pp. 660\u201071.","DOI":"10.1145\/1516360.1516437"},{"key":"key2022031520112261800_b24","unstructured":"Liu, B. (2007), Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data, Springer, Berlin, Heidelberg, New York, NY."},{"key":"key2022031520112261800_b25","doi-asserted-by":"crossref","unstructured":"Mamoulis, N., Cheng, K.H., Yiu, M.L. and Cheung, D.W. (2007), \u201cEfficient top\u2010k aggregation of ranked inputs\u201d, ACM Transactions on Database Systems, Vol. 32 No. 3, Article 19.","DOI":"10.1145\/1272743.1272749"},{"key":"key2022031520112261800_b26","doi-asserted-by":"crossref","unstructured":"Manning, C.D., Raghavan, P. and Sch\u00fctze, H. (2009), Introduction to Information Retrieval, Cambridge University Press, Cambridge.","DOI":"10.1017\/CBO9780511809071"},{"key":"key2022031520112261800_b27","doi-asserted-by":"crossref","unstructured":"Marian, A., Bruno, N. and Gravano, L. (2004), \u201cEvaluating top\u2010k queries over web\u2010accessible databases\u201d, ACM Transactions on Database Systems, Vol. 29 No. 2, pp. 319\u201062.","DOI":"10.1145\/1005566.1005569"},{"key":"key2022031520112261800_b28","unstructured":"Page, L., Brin, S., Motwani, R. and Winograd, T. (1998), The Pagerank Citation Ranking: Bringing Order to the Web, Technical Report, Standford InfoLab."},{"key":"key2022031520112261800_b29","unstructured":"Pan, F., Wang, W., Tung, A.K.H. and Yang, J. (2005), \u201cFinding representative set from massive data\u201d, Proceedings of the Fifth IEEE International Conference on Data Mining (ICDM'05), IEEE Computer Society, NW Washington, DC, pp. 338\u201045."},{"key":"key2022031520112261800_b30","doi-asserted-by":"crossref","unstructured":"Papadias, D., Tao, Y., Fu, G. and Seeger, B. (2005), \u201cProgressive skyline computation in database systems\u201d, ACM Transactions on Database Systems, Vol. 30 No. 1, pp. 41\u201082.","DOI":"10.1145\/1061318.1061320"},{"key":"key2022031520112261800_b32","doi-asserted-by":"crossref","unstructured":"Sakai, T. (2007), \u201cOn the reliability of information retrieval metrics based on graded relevance\u201d, Information Processing & Management, Vol. 43 No. 2, pp. 531\u201048.","DOI":"10.1016\/j.ipm.2006.07.020"},{"key":"key2022031520112261800_b33","unstructured":"Salton, G. (1971), The SMART Retrieval System: Experiments in Automatic Document Processing, Prentice Hall, Englewood Cliffs, NJ."},{"key":"key2022031520112261800_b34","unstructured":"Spink, A. and Jansen, B.J. (2004), Web Search: Public Searching of the Web, Kluwer Academic Publishers, New York, NY, Boston, MA, Dordrecht, London, Moscow."},{"key":"key2022031520112261800_b36","unstructured":"van Rijsbergen, C.J. (1979), Information Retrieval, Butterworths, London."},{"key":"key2022031520112261800_b37","doi-asserted-by":"crossref","unstructured":"Yiu, M.L. and Mamoulis, N. (2009), \u201cMulti\u2010dimensional top\u2010k dominating queries\u201d, The VLDB Journal, Vol. 18 No. 3, pp. 695\u2010718.","DOI":"10.1007\/s00778-008-0117-y"},{"key":"key2022031520112261800_b38","doi-asserted-by":"crossref","unstructured":"Zhai, C.X., Cohen, W.W. and Lafferty, J. (2003), \u201cBeyond independent relevance: methods and evaluation metrics for subtopic retrieval\u201d, Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM Press, New York, NY, pp. 10\u201017.","DOI":"10.1145\/860435.860440"},{"key":"key2022031520112261800_b39","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Callan, J. and Minka, T. (2002), \u201cNovelty and redundancy detection in adaptive filtering\u201d, Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'02), Vol. 02, ACM Press, New York, NY, pp. 81\u20108.","DOI":"10.1145\/564376.564393"},{"key":"key2022031520112261800_b41","doi-asserted-by":"crossref","unstructured":"Zhao, Y. and Karypis, G. (2004), \u201cEmpirical and theoretical comparisons of selected criterion functions for document clustering\u201d, Machine Learning, Vol. 55 No. 3, pp. 311\u201031.","DOI":"10.1023\/B:MACH.0000027785.44527.d6"},{"key":"key2022031520112261800_b42","doi-asserted-by":"crossref","unstructured":"Zhao, Y. and Karypis, G. (2005), \u201cHierarchical clustering algorithms for document datasets\u201d, Data Mining and Knowledge Discovery, Vol. 10 No. 2, pp. 141\u201068.","DOI":"10.1007\/s10618-005-0361-3"},{"key":"key2022031520112261800_b40","doi-asserted-by":"crossref","unstructured":"Zhu, M.J., Shi, S.M., Li, M.J. and Wen, J.R. (2009), \u201cEffective top\u2010k computation with term\u2010proximity support\u201d, Information Processing & Management, Vol. 45 No. 4, pp. 401\u201012.","DOI":"10.1016\/j.ipm.2009.04.002"},{"key":"key2022031520112261800_frg11","doi-asserted-by":"crossref","unstructured":"Gordon, M.D. and Lenk, P. (1991), \u201cA utility theoretic examination of the probability ranking principle in information retrieval\u201d, Journal of the American Society for Information Science, Vol. 42 No. 10, pp. 703\u201014.","DOI":"10.1002\/(SICI)1097-4571(199112)42:10<703::AID-ASI3>3.0.CO;2-1"},{"key":"key2022031520112261800_frg16","doi-asserted-by":"crossref","unstructured":"Hua, M., Pei, J., Fu, A.W.C., Lin, X. and Leung, H. (2009), \u201cTop\u2010k typicality queries and efficient query answering methods on large databases\u201d, The VLDB Journal, Vol. 18 No. 3, pp. 809\u201035.","DOI":"10.1007\/s00778-008-0128-8"},{"key":"key2022031520112261800_frg17","doi-asserted-by":"crossref","unstructured":"Huang, A. (2008), \u201cSimilarity measures for text document clustering\u201d, Proceedings of the Sixth New Zealand Computer Science Research Student Conference (NZCSRSC2008), Christchurch, New Zealand, 2008, Vol. 2008, pp. 49\u201056.","DOI":"10.1080\/00480169.2008.36806"},{"key":"key2022031520112261800_frg20","doi-asserted-by":"crossref","unstructured":"Korenius, T., Laurikkala, J. and Juhola, M. (2007), \u201cOn principal component analysis, cosine and Euclidean measures in information retrieval\u201d, Information Sciences, Vol. 177 No. 22, pp. 4893\u2010905.","DOI":"10.1016\/j.ins.2007.05.027"},{"key":"key2022031520112261800_frg31","doi-asserted-by":"crossref","unstructured":"Robertson, S.E. (1977), \u201cThe probability ranking principle in IR\u201d, Journal of Documentation, Vol. 33 No. 4, pp. 294\u2010304.","DOI":"10.1108\/eb026647"},{"key":"key2022031520112261800_frg35","unstructured":"Tang, X.H., Chen, G.Q. and Wei, Q. (2009), \u201cIntroducing relation compactness for generating a flexible size of search results in fuzzy queries\u201d, Proceedings of the Joint 2009 International Fuzzy Systems Association World Congress and 2009 European Society of Fuzzy Logic and Technology Conference, Lisbon, Portugal, Vol. 2009, pp. 1462\u20107."}],"container-title":["Journal of Enterprise Information Management"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/17410391111148567","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/17410391111148567\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/17410391111148567\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T00:19:30Z","timestamp":1753402770000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/jeim\/article\/24\/4\/310-321\/201550"}},"subtitle":[],"editor":[{"given":"Cengiz","family":"Kahraman","sequence":"first","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2011,7,26]]},"references-count":42,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,7,26]]}},"alternative-id":["10.1108\/17410391111148567"],"URL":"https:\/\/doi.org\/10.1108\/17410391111148567","relation":{},"ISSN":["1741-0398"],"issn-type":[{"type":"print","value":"1741-0398"}],"subject":[],"published":{"date-parts":[[2011,7,26]]}}}