{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T04:00:23Z","timestamp":1754107223247,"version":"3.41.0"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,12,7]],"date-time":"2024-12-07T00:00:00Z","timestamp":1733529600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nd\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001659","name":"German Research Foundation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Comput. Cult. Herit."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>The automatic analysis of images in the historical sciences often requires the identification of objects. Object identification is a well-researched problem for modern photographs; however, for historical material, annotations are often necessary. We present a solution for finding objects without manual work. The method consists of a style transfer of images from the COCO dataset into the domain using CycleGAN and training with items obtained through pseudo-labelling on the original and the additional transferred COCO images. Different strategies to assemble the dataset are compared. The best method obtains an F1 score of 0.58 for 15 object types without any labelling.<\/jats:p>","DOI":"10.1145\/3699963","type":"journal-article","created":{"date-parts":[[2024,10,15]],"date-time":"2024-10-15T12:19:09Z","timestamp":1728994749000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Object Detection in Historical Images: Transfer Learning and Pseudo Labelling"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4181-7968","authenticated-orcid":false,"given":"Yongho","family":"Kim","sequence":"first","affiliation":[{"name":"Faculty of Mathematics, Otto von Guericke University Magdeburg, Magdeburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6559-0510","authenticated-orcid":false,"given":"Chanjong","family":"Im","sequence":"additional","affiliation":[{"name":"Social Science, Otto von Guericke University Magdeburg, Magdeburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8398-9699","authenticated-orcid":false,"given":"Thomas","family":"Mandl","sequence":"additional","affiliation":[{"name":"Information Science, University of Hildesheim, Hildesheim, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,7]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_3_2_2_2","DOI":"10.1353\/uni.2021.0002"},{"issue":"9394","key":"e_1_3_2_3_2","first-page":"397","article-title":"Examples of Challenges and Opportunities in Visual Analysis in the Digital Humanities","author":"Rushmeier Holly","year":"2015","unstructured":"Holly Rushmeier, Ruggero Pintus, Ying Yang, Christiana Wong, and David Li. 2015. Examples of Challenges and Opportunities in Visual Analysis in the Digital Humanities. Human Vision and Electronic Imaging XX 9394 (2015), 397\u2013405.","journal-title":"Human Vision and Electronic Imaging"},{"doi-asserted-by":"publisher","key":"e_1_3_2_4_2","DOI":"10.1632\/pmla.2020.135.1.130"},{"key":"e_1_3_2_5_2","volume-title":"Digitale Bildwissenschaft","author":"Kohle Hubertus","year":"2013","unstructured":"Hubertus Kohle. 2013. Digitale Bildwissenschaft. H\u00fclsbusch."},{"unstructured":"Leonardo Impett and Fabian Offert. 2023. There Is a Digital Art History. arXiv:2308.07464. Retrieved from https:\/\/arxiv.org\/abs\/2308.07464","key":"e_1_3_2_6_2"},{"key":"e_1_3_2_7_2","first-page":"255","volume-title":"Post-Proceedings of the 5th Conference Digital Humanities in the Nordic Countries (DHN \u201920)","volume":"2865","author":"Kim Yongho","year":"2020","unstructured":"Yongho Kim, Thomas Mandl, Chanjong Im, Sebastian Schmideler, and Wiebke Helm. 2020. Applying Computer Vision Systems to Historical Book Illustrations: Challenges and First Results. In Post-Proceedings of the 5th Conference Digital Humanities in the Nordic Countries (DHN \u201920), CEUR Workshop Proceedings, Vol. 2865, CEUR-WS.org, 255\u2013260. Retrieved from https:\/\/ceur-ws.org\/Vol-2865\/poster7.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_2_8_2","DOI":"10.1145\/3474085.3478564"},{"doi-asserted-by":"publisher","key":"e_1_3_2_9_2","DOI":"10.5220\/0010390606220629"},{"doi-asserted-by":"publisher","key":"e_1_3_2_10_2","DOI":"10.1145\/3458885"},{"key":"e_1_3_2_11_2","volume-title":"International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations. Yoshua Bengio and Yann LeCun (Eds.), Retrieved from http:\/\/arxiv.org\/abs\/1409.1556"},{"doi-asserted-by":"publisher","key":"e_1_3_2_12_2","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_13_2","first-page":"6105","volume-title":"Proceedings of the 36th International Conference on Machine Learning","volume":"97","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 97, PMLR, 6105\u20136114. Retrieved from http:\/\/proceedings.mlr.press\/v97\/tan19a.html"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR \u201921)","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the 9th International Conference on Learning Representations (ICLR \u201921). OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"doi-asserted-by":"publisher","key":"e_1_3_2_15_2","DOI":"10.1109\/CVPR.2009.5206848"},{"doi-asserted-by":"publisher","key":"e_1_3_2_16_2","DOI":"10.1007\/978-3-319-10602-1_48"},{"doi-asserted-by":"publisher","key":"e_1_3_2_17_2","DOI":"10.5281\/ZENODO.4622007"},{"unstructured":"Dong-Hyun Lee. 2013. Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:18507866","key":"e_1_3_2_18_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_19_2","DOI":"10.1007\/s11831-019-09388-y"},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems 2020 (NeurIPS \u201920)","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. In Proceedings of the Annual Conference on Neural Information Processing Systems 2020 (NeurIPS \u201920). Hugo Larochelle, Marc\u2019Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.), Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html"},{"doi-asserted-by":"publisher","key":"e_1_3_2_21_2","DOI":"10.1109\/ICCV.2017.244"},{"unstructured":"Leon A. Gatys Alexander S. Ecker and Matthias Bethge. 2015. A neural algorithm of artistic style. arXiv:1508.06576. Retrieved from https:\/\/arxiv.org\/abs\/1508.06576","key":"e_1_3_2_22_2"},{"key":"e_1_3_2_23_2","first-page":"91","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Proceedings of the Annual Conference on Neural Information Processing Systems. 91\u201399. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2015\/hash\/14bfa6bb14875e45bba028a21ed38046-Abstract.html"},{"unstructured":"Glenn Jocher. 2020. YOLOv5. GitHub. Retrieved from https:\/\/github.com\/ultralytics\/yolov5","key":"e_1_3_2_24_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_25_2","DOI":"10.1007\/s11263-009-0275-4"},{"unstructured":"Tianhe Ren Jianwei Yang Shilong Liu Ailing Zeng Feng Li Hao Zhang Hongyang Li Zhaoyang Zeng and Lei Zhang. 2023. A strong and reproducible object detector with only public datasets. arXiv:2304.13027. Retrieved from https:\/\/arxiv.org\/abs\/2304.13027","key":"e_1_3_2_26_2"},{"unstructured":"Toru Ogawa Atsushi Otsubo Rei Narita Yusuke Matsui Toshihiko Yamasaki and Kiyoharu Aizawa. 2018. Object detection for comics using Manga109 annotations. arXiv:1803.08670. Retrieved from https:\/\/arxiv.org\/abs\/1803.08670","key":"e_1_3_2_27_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_28_2","DOI":"10.1007\/978-3-319-46448-0_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_29_2","DOI":"10.1109\/CVPR.2017.690"},{"unstructured":"Alexey Bochkovskiy Chien-Yao Wang and Hong-Yuan Mark Liao. 2020. YOLOv4: Optimal speed and accuracy of object detection. arXiv:2004.10934. Retrieved from https:\/\/arxiv.org\/abs\/2004.10934","key":"e_1_3_2_30_2"},{"unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An incremental improvement. arXiv:1804.02767. Retrieved from https:\/\/arxiv.org\/abs\/1804.02767","key":"e_1_3_2_31_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_32_2","DOI":"10.1109\/CVPR.2016.91"},{"doi-asserted-by":"publisher","key":"e_1_3_2_33_2","DOI":"10.1109\/ICIEM51511.2021.9445365"},{"doi-asserted-by":"publisher","key":"e_1_3_2_34_2","DOI":"10.1007\/s00607-020-00869-8"},{"doi-asserted-by":"publisher","key":"e_1_3_2_35_2","DOI":"10.1007\/s11042-021-11754-7"},{"doi-asserted-by":"publisher","key":"e_1_3_2_36_2","DOI":"10.1145\/3633454"},{"doi-asserted-by":"publisher","key":"e_1_3_2_37_2","DOI":"10.11588\/arthistoricum.413.c5767"},{"doi-asserted-by":"publisher","key":"e_1_3_2_38_2","DOI":"10.1093\/llc\/fqz022"},{"doi-asserted-by":"publisher","key":"e_1_3_2_39_2","DOI":"10.1145\/3476887.3476893"},{"doi-asserted-by":"publisher","key":"e_1_3_2_40_2","DOI":"10.1371\/journal.pone.0248414"},{"unstructured":"Babak Saleh and Ahmed M. Elgammal. 2015. Large-scale classification of fine-art paintings: Learning the right metric on the right feature. arXiv:1505.00855. Retrieved from https:\/\/arxiv.org\/abs\/1505.00855","key":"e_1_3_2_41_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_42_2","DOI":"10.1007\/s00138-014-0621-6"},{"doi-asserted-by":"publisher","key":"e_1_3_2_43_2","DOI":"10.1016\/j.eswa.2019.05.036"},{"unstructured":"Alexander Dunst and Rita Hartel. 2018. Hin zu einer Visuellen Stilometrie: Automatische Genre-und Autorunterscheidung in graphischen Narrativen. In Kritik der digitalen Vernunft. 5. Tagung \u201cDigital Humanities im deutschsprachigen Raum\u201d. Retrieved from http:\/\/dhd2018.uni-koeln.de\/wp-content\/uploads\/boa-DHd2018-web-ISBN.pdf","key":"e_1_3_2_44_2"},{"doi-asserted-by":"publisher","key":"e_1_3_2_45_2","DOI":"10.30687\/978-88-6969-332-8\/030"},{"doi-asserted-by":"publisher","key":"e_1_3_2_46_2","DOI":"10.1109\/MetroArchaeo43810.2018.9089828"},{"doi-asserted-by":"publisher","key":"e_1_3_2_47_2","DOI":"10.1007\/s42803-023-00070-1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_48_2","DOI":"10.11588\/arthistoricum.413.c5769"},{"key":"e_1_3_2_49_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML)","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 139, PMLR, 8748\u20138763. Retrieved from http:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"key":"e_1_3_2_50_2","first-page":"192","volume-title":"Lernen, Wissen, Daten, Analysen (LWDA) Conference Proceedings","volume":"3630","author":"Diem Sebastian","year":"2023","unstructured":"Sebastian Diem and Thomas Mandl. 2023. Automatic Classification of Portraits: Application of Transformer and CNN Based Models for an Art Historic Dataset. In Lernen, Wissen, Daten, Analysen (LWDA) Conference Proceedings Michael Leyer and Johannes Wichmann (Eds.), CEUR Workshop Proceedings, Vol. 3630, CEUR-WS.org, 192\u2013206. Retrieved from https:\/\/ceur-ws.org\/Vol-3630\/LWDA2023-paper18.pdf"},{"doi-asserted-by":"publisher","key":"e_1_3_2_51_2","DOI":"10.13173\/WIF.4.00I"},{"key":"e_1_3_2_52_2","first-page":"255","volume-title":"Post-Proceedings of the 5th Conference Digital Humanities in the Nordic Countries (DHN \u201920)","volume":"2865","author":"Kim Yongho","year":"2020","unstructured":"Yongho Kim, Thomas Mandl, Chanjong Im, Sebastian Schmideler, and Wiebke Helm. 2020. Applying Computer Vision Systems to Historical Book Illustrations: Challenges and First Results. In Post-Proceedings of the 5th Conference Digital Humanities in the Nordic Countries (DHN \u201920), CEUR Workshop Proceedings, Vol. 2865, CEUR-WS.org, 255\u2013260. Retrieved from https:\/\/ceur-ws.org\/Vol-2865\/poster7.pdf"}],"container-title":["Journal on Computing and Cultural Heritage"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3699963","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3699963","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:05:57Z","timestamp":1750291557000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3699963"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,7]]},"references-count":51,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3699963"],"URL":"https:\/\/doi.org\/10.1145\/3699963","relation":{},"ISSN":["1556-4673","1556-4711"],"issn-type":[{"type":"print","value":"1556-4673"},{"type":"electronic","value":"1556-4711"}],"subject":[],"published":{"date-parts":[[2024,12,7]]},"assertion":[{"value":"2024-02-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}