{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T17:54:40Z","timestamp":1754157280691,"version":"3.41.2"},"reference-count":10,"publisher":"Emerald","issue":"1","license":[{"start":{"date-parts":[[2008,2,15]],"date-time":"2008-02-15T00:00:00Z","timestamp":1203033600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2008,2,15]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>The purpose of this paper is to develop a system that can convert PDF files to XML files.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>The system works with XML as an information display model and XSLT as an information extraction rule. The process is illustrated by converting a scientific and technological paper in PDF to a valid XML file.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>Because the PDF file adopts the self\u2010descriptive definition, its content information and the display information exists in different objects; therefore, it is not easy to directly extract information from the PDF source file. The undirected way to solve this problem in the system design was to convert the PDF source file to a relatively easy processing intermediate format, which can then be automatically converted to the target file in accordance with relevant rules.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>It is important to be able to easily and conveniently extract information from PDF files and this paper shows how it can be done. The design ideas contained in the paper can also be applied to information extraction from other types of files.<\/jats:p><\/jats:sec>","DOI":"10.1108\/02640470810851743","type":"journal-article","created":{"date-parts":[[2008,2,9]],"date-time":"2008-02-09T07:05:23Z","timestamp":1202540723000},"page":"68-74","source":"Crossref","is-referenced-by-count":6,"title":["Converting PDF files to XML files"],"prefix":"10.1108","volume":"26","author":[{"given":"Wende","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022021319495572700_b1","unstructured":"Acrobat 7.0 SDK (2005), \u201cGuide to SDK Samples\u201d, available at: http:\/\/partners.adobe.com\/asn\/developer\/acrosdk\/main.html."},{"key":"key2022021319495572700_b2","unstructured":"Adobe Systems Incorporated (2001), \u201cAdobe System Inc., Reference, Adobe Portable Document Format Version .4_[M].3rd\u201d, available at: www.adobe.com\/support\/downloads\/product.jsp?product=44&platform=Windows."},{"key":"key2022021319495572700_b3","unstructured":"(The) Apache Software Foundation (2005), \u201cXalan\u2010Java Version 2.7.0\u201d. available at: http:\/\/xml.apache.org\/xalan\u2010j\/."},{"key":"key2022021319495572700_b4","doi-asserted-by":"crossref","unstructured":"Chang, C.\u2010H. and Lui, S.\u2010C. (2001), \u201cIEPAD: information extraction based on pattern discovery\u201d, Proceedings of 2001 International World Wide Web Conference, pp. 681\u20108.","DOI":"10.1145\/371920.372182"},{"key":"key2022021319495572700_b5","unstructured":"Crescenzi, V. and Mecca, G. (2003), On Automatic Information Extraction from Large Web Sites, Technical Report DIA\u201076\u20102003, Road Runner Project, pp. 19\u201023."},{"key":"key2022021319495572700_b6","unstructured":"Sourceforge.net (2005), \u201cAbout SAX\u201d, available at: www.saxproject.org\/."},{"key":"key2022021319495572700_b7","unstructured":"W3C (2000), \u201cExtensible Markup Language 1.0: second edition\u201d, available at: www.w3.org\/TR\/REC\u2010xml."},{"key":"key2022021319495572700_b8","unstructured":"W3C (2005a), \u201cDocument Object Model (DOM)\u201d, available at: www.w3.org\/DOM\/."},{"key":"key2022021319495572700_b9","unstructured":"W3C (2005b), \u201cXSL Transformations (XSLT) Version 1.0\u201d, available at: www.w3.org\/TR\/xslt."},{"key":"key2022021319495572700_b10","unstructured":"Yang, D. (1999), \u201cDesign and implementation of the Chinese character PDF reader for user\u201d, Computer Application, Vol. 19 No. 6, pp. 1\u20104."}],"container-title":["The Electronic Library"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/02640470810851743","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/02640470810851743\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/02640470810851743\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T23:42:06Z","timestamp":1753400526000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/el\/article\/26\/1\/68-74\/42252"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,2,15]]},"references-count":10,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2008,2,15]]}},"alternative-id":["10.1108\/02640470810851743"],"URL":"https:\/\/doi.org\/10.1108\/02640470810851743","relation":{},"ISSN":["0264-0473"],"issn-type":[{"type":"print","value":"0264-0473"}],"subject":[],"published":{"date-parts":[[2008,2,15]]}}}