{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T17:53:15Z","timestamp":1754157195790,"version":"3.41.2"},"reference-count":7,"publisher":"Emerald","issue":"4","license":[{"start":{"date-parts":[[2007,11,6]],"date-time":"2007-11-06T00:00:00Z","timestamp":1194307200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,11,6]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>The purpose of this paper is to develop automated methods for creating metadata for documents in an institutional repository.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>Two methods are examined for automatically building metadata in an institutional repository context. Text mining techniques are employed to discover relationships among documents with similar content, from which are inferred possible values for missing or incomplete metadata elements. Machine learning techniques are used to identify and extract specific metadata element values from document content.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>Text mining techniques can be used to cluster documents with similar content. This allows values for metadata elements, like keyword, to be projected from documents with established metadata to related documents. Machine learning techniques are found to be reasonably accurate for extracting from documents values for metadata elements, such as, title, author, and abstract. Results show sufficient promise to support the next phase of the project: the development of assistive tools for use by metadata specialists to create or edit document metadata.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>This paper focuses on the use of automated metadata extraction techniques to assist metadata creation, lessening the time and effort required to add documents to institutional repositories.<\/jats:p><\/jats:sec>","DOI":"10.1108\/10650750710831547","type":"journal-article","created":{"date-parts":[[2007,10,20]],"date-time":"2007-10-20T05:27:46Z","timestamp":1192858066000},"page":"403-410","source":"Crossref","is-referenced-by-count":1,"title":["New possibilities for metadata creation in an institutional repository context"],"prefix":"10.1108","volume":"23","author":[{"given":"Alan","family":"Burk","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Muhammad","family":"Al\u2010Digeil","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dominic","family":"Forest","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jennifer","family":"Whitney","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022031020054598200_b1","unstructured":"Al\u2010Digeil, M., Burk, A., Forest, D. and Whitney, J. (2005), \u201cAutomatic metadata extraction and creation in an institutional repository context\u201d, Proceedings 2005 CaSTA Symposium, 5 October 2005, University of Alberta, Edmonton, Alberta."},{"key":"key2022031020054598200_b2","unstructured":"Barton, J., Currier, S. and Hey, J.M.N. (2003), \u201cBuilding quality assurance into metadata creation: an analysis based on the learning objects and e\u2010prints communities of practice\u201d, in Sutton, S., Greenberg, J., Tennis, J. (eds), Proceedings 2003 Dublin Core Conference: Supporting Communities of Discourse and Practice \u2013 Metadata Research and Applications, Seattle, Washington, available at: http:\/\/eprints.soton.ac.uk\/20\/02\/201_paper60.pdf."},{"key":"key2022031020054598200_b3","doi-asserted-by":"crossref","unstructured":"Foster, N.F. and Gibbons, S. (2005), \u201cUnderstanding faculty to improve content recruitment for institutional repositories\u201d, D\u2010Lib Magazine, Vol. 11 No. 1, available at: www.dlib.org\/dlib\/january05\/foster\/01foster.html.","DOI":"10.1045\/january2005-foster"},{"key":"key2022031020054598200_b4","unstructured":"Han, H., Giles, C.L., Manavoglu, E., Zha, H., Zhang, Z. and Fox, E.A. (2003), \u201cAutomatic document metadata extraction using support vector machines\u201d, Proceedings of the 3rd ACM\/IEEE\u2010CS Joint Conference on Digital Libraries, Houston, Texas."},{"key":"key2022031020054598200_b5","unstructured":"Harnard, S. (2003), \u201cOpen access to peer\u2010reviewed research through author\/institution self\u2010archiving: maximizing research impact by maximizing online access\u201d, in Law, D. and Andrews, J. (Eds), Digital Libraries: Policy Planning and Practice, Ashgate Publishing, Aldershot."},{"key":"key2022031020054598200_b6","unstructured":"Hearst, M. (2003), \u201cWhat is text mining?\u201d, unpublished essay, available at: www.ischool.berkeley.edu\/ \u223c\u2009hearst\/text\u2010mining.html."},{"key":"key2022031020054598200_b7","unstructured":"Seymore, K., McCallum, A. and Rosenfeld, R. (1999), \u201cLearning hidden Markov model structure for information extraction\u201d, in AAAI Workshop on Machine Learning for Information Extraction, available at: www.cs.cmu.edu\/ \u223c\u2009kseymore\/papers\/ie_aaai99.ps.gz."}],"container-title":["OCLC Systems &amp; Services: International digital library perspectives"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/10650750710831547","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/10650750710831547\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/10650750710831547\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T23:39:02Z","timestamp":1753400342000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/dlp\/article\/23\/4\/403-410\/305081"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,11,6]]},"references-count":7,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,11,6]]}},"alternative-id":["10.1108\/10650750710831547"],"URL":"https:\/\/doi.org\/10.1108\/10650750710831547","relation":{},"ISSN":["1065-075X"],"issn-type":[{"type":"print","value":"1065-075X"}],"subject":[],"published":{"date-parts":[[2007,11,6]]}}}