{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,13]],"date-time":"2025-12-13T06:41:12Z","timestamp":1765608072120,"version":"3.41.2"},"reference-count":27,"publisher":"Emerald","issue":"5","license":[{"start":{"date-parts":[[2006,9,1]],"date-time":"2006-09-01T00:00:00Z","timestamp":1157068800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,9,1]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>Aims to measure syllable aggregation consistency of Romanized Chinese data in the title fields of bibliographic records. Also aims to verify if the term frequency distributions satisfy conventional bibliometric laws.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>Uses Cooper's interindexer formula to evaluate aggregation consistency within and between two sets of Chinese bibliographic data. Compares the term frequency distributions of polysyllabic words and monosyllabic characters (for vernacular and Romanized data) with the Lotka and the generalised Zipf theoretical distributions. The fits are tested with the Kolmogorov\u2010Smirnov test.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>Finds high internal aggregation consistency within each data set but some aggregation discrepancy between sets. Shows that word (polysyllabic) distributions satisfy Lotka's law but that character (monosyllabic) distributions do not abide by the law.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Research limitations\/implications<\/jats:title><jats:p>The findings are limited to only two sets of bibliographic data (for aggregation consistency analysis) and to one set of data for the frequency distribution analysis. Only two bibliometric distributions are tested. Internal consistency within each database remains fairly high. Therefore the main argument against syllable aggregation does not appear to hold true. The analysis revealed that Chinese words and characters behave differently in terms of frequency distribution but that there is no noticeable difference between vernacular and Romanized data. The distribution of Romanized characters exhibits the worst case in terms of fit to either Lotka's or Zipf's laws, which indicates that Romanized data in aggregated form appear to be a preferable option.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>Provides empirical data on consistency and distribution of Romanized Chinese titles in bibliographic records.<\/jats:p><\/jats:sec>","DOI":"10.1108\/00220410610688750","type":"journal-article","created":{"date-parts":[[2006,9,26]],"date-time":"2006-09-26T07:29:49Z","timestamp":1159255789000},"page":"606-633","source":"Crossref","is-referenced-by-count":1,"title":["Aggregation consistency and frequency of Chinese words and characters"],"prefix":"10.1108","volume":"62","author":[{"given":"Cl\u00e9ment","family":"Arsenault","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022012619520406600_b1","doi-asserted-by":"crossref","unstructured":"Arsenault, C. (2001), \u201cWord division in the transcription of Chinese script in the title fields of bibliographic records\u201d, Cataloging and Classification Quarterly, Vol. 32 No. 3, pp. 109\u201037.","DOI":"10.1300\/J104v32n03_08"},{"key":"key2022012619520406600_b2","unstructured":"Arsenault, C. (2002a), \u201cAnalyse de la consistance dans l'agr\u00e9gation des transcriptions pinyin polysyllabiques dans les bases bibliographiques\u201d, CJILS\/RCSIB, Vol. 26 Nos 2\/3, pp. 91\u2010106."},{"key":"key2022012619520406600_b3","unstructured":"Arsenault, C. (2002b), \u201cPinyin Romanization for OPAC retrieval: is everyone being served?\u201d, Information Technology and Libraries, Vol. 21 No. 2, pp. 45\u201050."},{"key":"key2022012619520406600_b4","unstructured":"Arsenault, C. (2004), \u201cMeasuring and comparing aggregation inconsistency for Chinese titles in two library catalogues\u201d, Proceedings of the CAIS\/ACSI Annual Conference, Winnipeg, Manitoba, Canada, 3\u20105 June, available at: www.cais\u2010acsi.ca\/proceedings\/2004\/arsenault_2004.pdf."},{"key":"key2022012619520406600_b5","doi-asserted-by":"crossref","unstructured":"Chen, Z. and Lee, K.F. (2000), \u201cA new statistical approach to Chinese pinyin input\u201d, Proceedings of the 38th Annual Meeting of the Association for Computational Linguistics (ACL '00), ACL, Hong Kong, pp. 241\u20107, available at: research.microsoft.com\/china\/papers\/Statistical_Chinese_Pinyin_Input.pdf.","DOI":"10.3115\/1075218.1075249"},{"key":"key2022012619520406600_b6","doi-asserted-by":"crossref","unstructured":"Cooper, W.S. (1969), \u201cIs interindexer consistency a hobgoblin?\u201d, American Documentation, Vol. 20 No. 3, pp. 268\u201078.","DOI":"10.1002\/asi.4630200314"},{"key":"key2022012619520406600_b8","doi-asserted-by":"crossref","unstructured":"DeFrancis, J. (1984), The Chinese Language: Fact and Fantasy, University of Hawaii Press, Honolulu, HI.","DOI":"10.1515\/9780824840303"},{"key":"key2022012619520406600_b7","unstructured":"Duanmu, S. (1998), \u201cWordhood in Chinese\u201d, in Packard, J.L. (Ed.), New Approaches to Chinese Word Formation: Morphology, Phonology and the Lexicon in Modern and Ancient Chinese, Mouton de Gruyter, Berlin, pp. 135\u201096."},{"key":"key2022012619520406600_b9","doi-asserted-by":"crossref","unstructured":"Egghe, L. (2005), Power Laws in the Information Production Process: Lotkaian Informetrics, Elsevier, Amsterdam.","DOI":"10.1108\/S1876-0562(2005)05"},{"key":"key2022012619520406600_b10","unstructured":"Egghe, L. and Rousseau, R. (1990), An Introduction to Informetrics, Elsevier, Amsterdam."},{"key":"key2022012619520406600_b11","doi-asserted-by":"crossref","unstructured":"Goldstein, M.L., Morris, S.A. and Yen, G.G. (2004), \u201cProblems with fitting to the power\u2010law distribution\u201d, The European Physical Journal B, Vol. 41 No. 2, pp. 255\u20108.","DOI":"10.1140\/epjb\/e2004-00316-5"},{"key":"key2022012619520406600_b12","doi-asserted-by":"crossref","unstructured":"Ha, L.Q., Sicilia\u2010Garcia, E.I., Ming, J. and Smith, J. (2002), \u201cExtension of Zipf's law to words and phrases\u201d, in Tseng, S.C., Chen, T.E. and Liu, Y.F. (Eds), Coling 2002: Proceedings of the 19th International Conference on Computational Linguistics, Morgan Kaufmann, San Francisco, CA, pp. 315\u201020.","DOI":"10.3115\/1072228.1072345"},{"key":"key2022012619520406600_b13","unstructured":"King, P.L. (1983), \u201cContextual factors in Chinese Pinyin writing\u201d, PhD dissertation, Cornell University, Ithaca, NY."},{"key":"key2022012619520406600_b14","unstructured":"L\u00fc, S. (1979), H\u00e0ny\u00fb Y\u00fbf\u00e2 F\u0113nx\u012b W\u00e8nt\u00ed, (in Chinese), Shangwu Yinshuguan, Beijing."},{"key":"key2022012619520406600_b15","unstructured":"Mair, V.H. (1991), \u201cBuilding the future of information processing in East Asia: demands facing linguistic and technological reality\u201d, in Mair, V.H. and Liu, Y. (Eds), Characters and Computers, IOS Press, Amsterdam, pp. 1\u20108."},{"key":"key2022012619520406600_b16","unstructured":"Mair, V.H. (2001), \u201cPinyin orthographical rules for libraries: a follow\u2010up\u201d, Chinese Librarianship: An International Electronic Journal, Vol. 11, available at: www.iclc.us\/cliej\/cl11.htm."},{"key":"key2022012619520406600_b17","doi-asserted-by":"crossref","unstructured":"Nicholls, P.T. (1989), \u201cBibliometric modeling processes and the empirical validity of Lotka's law\u201d, Journal of the American Society for Information Science, Vol. 40 No. 6, pp. 379\u201085.","DOI":"10.1002\/(SICI)1097-4571(198911)40:6<379::AID-ASI1>3.0.CO;2-Q"},{"key":"key2022012619520406600_b18","unstructured":"OCLC (n.d.), \u201cWorldCat facts and statistics\u201d, available at: www.oclc.org\/worldcat\/statistics\/language.htm (accessed 10 November 2005)."},{"key":"key2022012619520406600_b19","unstructured":"Rousseau, R. (1993), \u201cA table for estimating the exponent in Lotka's law\u201d, Journal of Documentation, Vol. 49 No. 4, pp. 409\u201012."},{"key":"key2022012619520406600_b20","doi-asserted-by":"crossref","unstructured":"Rousseau, R. and Zhang, Q. (1992), \u201cZipf's data on the frequency of Chinese words revisited\u201d, Scientometrics, Vol. 24 No. 2, pp. 201\u201020.","DOI":"10.1007\/BF02017909"},{"key":"key2022012619520406600_b21","doi-asserted-by":"crossref","unstructured":"Suen, C.Y. (1986), Computational Studies of the Most Frequent Chinese Words and Sounds, Word Scientific, Singapore.","DOI":"10.1142\/0219"},{"key":"key2022012619520406600_b22","doi-asserted-by":"crossref","unstructured":"Walther, A. (1926), \u201cAnschauliches zur Riemannschen Zetafunktion\u201d, Acta Mathematica, Vol. 48, pp. 393\u2010400.","DOI":"10.1007\/BF02565343"},{"key":"key2022012619520406600_b23","unstructured":"Wellisch, H.H. (1978), The Conversion of Scripts: Its Nature, History, and Utilization, Wiley, New York, NY."},{"key":"key2022012619520406600_b24","unstructured":"Wilson, C.S. (1999), \u201cInformetrics\u201d, in Williams, M.E. (Ed.), Annual Review of Information Science and Technology, Vol. 34, Information Today, Medford, NJ, pp. 107\u2010247."},{"key":"key2022012619520406600_b25","doi-asserted-by":"crossref","unstructured":"Wolfram, D. and Zhang, J. (2002), \u201cAn investigation of the influence of indexing exhaustivity and term distributions on a document space\u201d, Journal of the American Society for Information Science and Technology, Vol. 53 No. 11, pp. 943\u201052.","DOI":"10.1002\/asi.10121"},{"key":"key2022012619520406600_b27","unstructured":"Zhou, Y. (1993), H\u00e0ny\u00fb P\u012bny\u012bn F\u00e0ng'\u0101n J\u012bch\u00fb Zh\u012bsh\u00ec, (in Chinese), Yuwen Chubanshe, Beijing."},{"key":"key2022012619520406600_b26","unstructured":"Zipf, G.K. (1932), Selected Studies of the Principle of Relative Frequency in Language, Harvard University Press, Cambridge, MA."}],"container-title":["Journal of Documentation"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/00220410610688750","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/00220410610688750\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/00220410610688750\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T23:37:42Z","timestamp":1753400262000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/jd\/article\/62\/5\/606-633\/204566"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,9,1]]},"references-count":27,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2006,9,1]]}},"alternative-id":["10.1108\/00220410610688750"],"URL":"https:\/\/doi.org\/10.1108\/00220410610688750","relation":{},"ISSN":["0022-0418"],"issn-type":[{"type":"print","value":"0022-0418"}],"subject":[],"published":{"date-parts":[[2006,9,1]]}}}