{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T19:49:50Z","timestamp":1774986590178,"version":"3.50.1"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001659","name":"German Research Foundation","doi-asserted-by":"crossref","award":["ref. 414984028 and ref. 556566056"],"award-info":[{"award-number":["ref. 414984028 and ref. 556566056"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]},{"name":"European Union Horizon 2020","award":["957407"],"award-info":[{"award-number":["957407"]}]},{"DOI":"10.13039\/501100006374","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIS-2420577 and IIS-2420691"],"award-info":[{"award-number":["IIS-2420577 and IIS-2420691"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,6,17]]},"abstract":"<jats:p>Transfer learning is an effective technique for tuning a deep learning model when training data or computational resources are limited. Instead of training a new model from scratch, the parameters of an existing base model are adjusted for the new task. The accuracy of such a fine-tuned model depends on the suitability of the base model chosen. Model search automates the selection of such a base model by evaluating the suitability of candidate models for a specific task. This entails inference with each candidate model on task-specific data. With thousands of models available through model stores, the computational cost of model search is a major bottleneck for efficient transfer learning.<\/jats:p>\n                  <jats:p>\n                    In this work, we present\n                    <jats:italic toggle=\"yes\">Alsatian<\/jats:italic>\n                    , a novel model search system. Based on the observation that many candidate models overlap to a significant extent and following a careful bottleneck analysis, we propose optimization techniques that are applicable to many model search frameworks. These optimizations include: (i) splitting models into individual blocks that can be shared across models, (ii) caching of intermediate inference results and model blocks, and (iii) selecting a beneficial search order for models to maximize sharing of cached results. In our evaluation on state-of-the-art deep learning models from computer vision and natural language processing, we show that\n                    <jats:italic toggle=\"yes\">Alsatian<\/jats:italic>\n                    outperforms baselines by up to 14x.\n                  <\/jats:p>","DOI":"10.1145\/3725264","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:23:29Z","timestamp":1750281809000},"page":"1-27","source":"Crossref","is-referenced-by-count":1,"title":["Alsatian: Optimizing Model Search for Deep Transfer Learning"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-2569-2549","authenticated-orcid":false,"given":"Nils","family":"Strassenburg","sequence":"first","affiliation":[{"name":"Hasso Plattner Institute, University of Potsdam, Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2887-2452","authenticated-orcid":false,"given":"Boris","family":"Glavic","sequence":"additional","affiliation":[{"name":"University of Illinois Chicago, Chicago, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-3335-8045","authenticated-orcid":false,"given":"Tilmann","family":"Rabl","sequence":"additional","affiliation":[{"name":"Hasso Plattner Institute, University of Potsdam, Potsdam, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2024. CS231n Convolutional Neural Networks for Visual Recognition. https:\/\/cs231n.github.io\/transfer-learning\/"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00653"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824035"},{"key":"e_1_2_2_4_1","first-page":"19301","article-title":"Scalable diverse model selection for accessible transfer learning","volume":"34","author":"Bolya Daniel","year":"2021","unstructured":"Daniel Bolya, Rohit Mittapalli, and Judy Hoffman. 2021. Scalable diverse model selection for accessible transfer learning. Advances in Neural Information Processing Systems, Vol. 34 (2021), 19301-19312.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_5_1","first-page":"446","volume-title":"proceedings, part VI 13","author":"Bossard Lukas","year":"2014","unstructured":"Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014. Food-101--mining discriminative components with random forests. In Computer vision--ECCV 2014: 13th European conference, zurich, Switzerland, September 6--12, 2014, proceedings, part VI 13. Springer, 446-461."},{"key":"e_1_2_2_6_1","unstructured":"Sasank Chilamkurthy. 2024. Transfer Learning for Computer Vision Tutorial - PyTorch Tutorials 2.2.2cu121 documentation. https:\/\/pytorch.org\/tutorials\/beginner\/transfer_learning_tutorial.html"},{"key":"e_1_2_2_7_1","volume-title":"Deep learning with Python","author":"Chollet Francois","unstructured":"Francois Chollet. 2021. Deep learning with Python,. Simon and Schuster."},{"key":"e_1_2_2_8_1","volume-title":"Jonah B Gelbach","author":"Dai Timothy","year":"2024","unstructured":"Timothy Dai, Austin Peters, Jonah B Gelbach, David Freeman Engstrom, and Daniel Kang. 2024. tailwiz: Empowering Domain Experts with Easy-to-Use, Task-Specific Natural Language Processing Models. In DEEM. 12-22."},{"key":"e_1_2_2_9_1","unstructured":"Cl\u00e9ment Delangue. 2023. Hugging Face just crossed 1 000 000 free public models. https:\/\/x.com\/clementdelangue\/status\/1839375655688884305's=43"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389715"},{"key":"e_1_2_2_11_1","unstructured":"Aditya Deshpande Alessandro Achille Avinash Ravichandran Hao Li Luca Zancato Charless Fowlkes Rahul Bhotika Stefano Soatto and Pietro Perona. 2021. A linearized framework and a new benchmark for model selection for fine-tuning. arXiv preprint arXiv:2102.00084 (2021)."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19--1423"},{"key":"e_1_2_2_13_1","volume-title":"International Conference on Learning Representations,. https:\/\/openreview.net\/forum?id=YicbFdNTTy","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations,. https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_2_2_14_1","volume-title":"Imagenette: A smaller subset of 10 classes from Imagenet. https:\/\/github.com\/fastai\/imagenette","year":"2024","unstructured":"Fastai. 2024. Imagenette: A smaller subset of 10 classes from Imagenet. https:\/\/github.com\/fastai\/imagenette"},{"key":"e_1_2_2_15_1","volume-title":"Choose Your Transformer: Improved Transferability Estimation of Transformer Models on Classification Tasks. In Findings of the Association for Computational Linguistics: ACL","author":"Garbaciauskas Lukas","year":"2024","unstructured":"Lukas Garbaciauskas, Max Ploner, and Alan Akbik. 2024. Choose Your Transformer: Improved Transferability Estimation of Transformer Models on Classification Tasks. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar, (Eds.). 12752-12768."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526173"},{"key":"e_1_2_2_17_1","volume-title":"Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow,. \"O'Reilly Media","author":"G\u00e9ron Aur\u00e9lien","unstructured":"Aur\u00e9lien G\u00e9ron. 2022. Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow,. \"O'Reilly Media, Inc.\"."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_2_19_1","volume-title":"International conference on machine learning. PMLR, 2790-2799","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International conference on machine learning. PMLR, 2790-2799."},{"key":"e_1_2_2_20_1","volume-title":"Proceedings of the 39th International Conference on Machine Learning, (Proceedings of Machine Learning Research","volume":"9225","author":"Huang Long-Kai","year":"2022","unstructured":"Long-Kai Huang, Junzhou Huang, Yu Rong, Qiang Yang, and Ying Wei. 2022. Frustratingly Easy Transferability Estimation. In Proceedings of the 39th International Conference on Machine Learning, (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, (Eds.). PMLR, 9201-9225. https:\/\/proceedings.mlr.press\/v162\/huang22d.html"},{"key":"e_1_2_2_21_1","volume-title":"Hugging Face: Machine Learning Platform. https:\/\/huggingface.co\/","author":"Face Hugging","year":"2024","unstructured":"Hugging Face. 2024. Hugging Face: Machine Learning Platform. https:\/\/huggingface.co\/"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.blackboxnlp-1.19"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.3390\/app10103359"},{"key":"e_1_2_2_24_1","unstructured":"Aditya Khosla Nityananda Jayadevaprakash Bangpeng Yao and Li Fei-Fei. 2011. ImageNet Dogs Dataset. http:\/\/vision.stanford.edu\/aditya86\/ImageNetDogs\/"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00277"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.77"},{"key":"e_1_2_2_27_1","unstructured":"Jaejun Lee Raphael Tang and Jimmy Lin. 2019. What would elsa do? freezing layers during transformer fine-tuning. arXiv preprint arXiv:1911.03090 (2019)."},{"key":"e_1_2_2_28_1","unstructured":"Yoonho Lee Annie S Chen Fahim Tajwar Ananya Kumar Huaxiu Yao Percy Liang and Chelsea Finn. 2022. Surgical fine-tuning improves adaptation to distribution shifts. arXiv preprint arXiv:2210.11466 (2022)."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00354"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00269"},{"key":"e_1_2_2_31_1","first-page":"278","article-title":"SmartLite","volume":"17","author":"Lin Qiuru","year":"2023","unstructured":"Qiuru Lin, Sai Wu, Junbo Zhao, Jian Dai, Meng Shi, Gang Chen, and Feifei Li. 2023. SmartLite: A DBMS-Based Serving System for DNN Inference in Resource-Constrained Environments. PVLDB, Vol. 17, 3 (2023), 278-291.","journal-title":"PVLDB"},{"key":"e_1_2_2_32_1","doi-asserted-by":"crossref","unstructured":"Abhilasha Lodha Gayatri Belapurkar Saloni Chalkapurkar Yuanming Tao Reshmi Ghosh Samyadeep Basu Dmitrii Petrov and Soundararajan Srinivasan. 2023. On surgical fine-tuning for language encoders. arXiv preprint arXiv:2310.17041 (2023).","DOI":"10.18653\/v1\/2023.findings-emnlp.204"},{"key":"e_1_2_2_33_1","first-page":"142","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies,. Association for Computational Linguistics","author":"Maas Andrew L.","year":"2011","unstructured":"Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning Word Vectors for Sentiment Analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies,. Association for Computational Linguistics, Portland, Oregon, USA, 142-150. http:\/\/www.aclweb.org\/anthology\/P11--1015"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.112"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517846"},{"key":"e_1_2_2_36_1","volume-title":"Proceedings of the 37th International Conference on Machine Learning,. PMLR, 7294-7305","author":"Nguyen Cuong","year":"2020","unstructured":"Cuong Nguyen, Tal Hassner, Matthias Seeger, and Cedric Archambeau. 2020. LEEP: A New Measure to Evaluate Transferability of Learned Representations. In Proceedings of the 37th International Conference on Machine Learning,. PMLR, 7294-7305. https:\/\/proceedings.mlr.press\/v119\/nguyen20b.html"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452788"},{"key":"e_1_2_2_38_1","volume-title":"Scalable Transfer Learning with Expert Models. In International Conference on Learning Representations,. https:\/\/openreview.net\/forum?id=23ZjUGpjcc","author":"Puigcerver Joan","year":"2021","unstructured":"Joan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cedric Renggli, Andr\u00e9 Susano Pinto, Sylvain Gelly, Daniel Keysers, and Neil Houlsby. 2021. Scalable Transfer Learning with Expert Models. In International Conference on Learning Representations,. https:\/\/openreview.net\/forum?id=23ZjUGpjcc"},{"key":"e_1_2_2_39_1","unstructured":"PyTorch Team. 2024. Models and Pre-trained Weights. https:\/\/pytorch.org\/vision\/stable\/models.html"},{"key":"e_1_2_2_40_1","unstructured":"Sebastian Raschka. 2024. Build a Large Language Model (From Scratch). Manning."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00899"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/3565816.3565831"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.48786\/EDBT.2022.12"},{"key":"e_1_2_2_45_1","volume-title":"International conference on machine learning. PMLR, 10096-10106","author":"Tan Mingxing","year":"2021","unstructured":"Mingxing Tan and Quoc Le. 2021. Efficientnetv2: Smaller models and faster training. In International conference on machine learning. PMLR, 10096-10106."},{"key":"e_1_2_2_46_1","unstructured":"Hugging Face team. 2025 a. Conditional DETR model with ResNet-50 backbone. https:\/\/huggingface.co\/microsoft\/conditional-detr-resnet-50"},{"key":"e_1_2_2_47_1","unstructured":"Hugging Face team. 2025 b. DETR (End-to-End Object Detection) model with ResNet-50 backbone. https:\/\/huggingface.co\/facebook\/detr-resnet-50"},{"key":"e_1_2_2_48_1","unstructured":"Hugging Face team. 2025 c. DETR (End-to-End Object Detection) model with ResNet-50 backbone (dilated C5 stage). https:\/\/huggingface.co\/facebook\/detr-resnet-50-dc5"},{"key":"e_1_2_2_49_1","unstructured":"Hugging Face team. 2025 d. ResNet. https:\/\/huggingface.co\/microsoft\/resnet-18"},{"key":"e_1_2_2_50_1","unstructured":"Hugging Face team. 2025 e. ResNet-152 v1.5. https:\/\/huggingface.co\/microsoft\/resnet-152"},{"key":"e_1_2_2_51_1","unstructured":"Hugging Face team. 2025 f. TrOCR (base-sized model fine-tuned on SROIE). https:\/\/huggingface.co\/microsoft\/trocr-base-printed"},{"key":"e_1_2_2_52_1","unstructured":"Hugging Face team. 2025 g. Vision Transformer (base-sized model). https:\/\/huggingface.co\/google\/vit-base-patch16--224-in21k"},{"key":"e_1_2_2_53_1","unstructured":"Hugging Face team. 2025 h. Vision Transformer (large-sized model) trained using DINOv2. https:\/\/huggingface.co\/facebook\/dinov2-large"},{"key":"e_1_2_2_54_1","unstructured":"Hugging Face team. 2025 i. YOLOS (small-sized) model. https:\/\/huggingface.co\/hustvl\/yolos-small"},{"key":"e_1_2_2_55_1","unstructured":"TensorFlow. 2024. Transfer Learning and Fine-Tuning. https:\/\/www.tensorflow.org\/tutorials\/images\/transfer_learning"},{"key":"e_1_2_2_56_1","unstructured":"TensorFlow Team. 2024. TensorFlow Hub: A Repository of Trained Machine Learning Models. https:\/\/www.tensorflow.org\/hub"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939502.2939516"},{"key":"e_1_2_2_58_1","volume-title":"Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","unstructured":"Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2022. CUB-200--2011. doi:10.22002\/D1.20098","DOI":"10.22002\/D1.20098"},{"key":"e_1_2_2_60_1","volume-title":"Helix: Holistic optimization for accelerating iterative machine learning. arXiv preprint arXiv:1812.05762","author":"Xin Doris","year":"2018","unstructured":"Doris Xin, Stephen Macke, Litian Ma, Jialin Liu, Shuchen Song, and Aditya Parameswaran. 2018. Helix: Holistic optimization for accelerating iterative machine learning. arXiv preprint arXiv:1812.05762, (2018)."},{"key":"e_1_2_2_61_1","volume-title":"How transferable are features in deep neural networks? Advances in neural information processing systems","author":"Yosinski Jason","year":"2014","unstructured":"Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014. How transferable are features in deep neural networks? Advances in neural information processing systems, Vol. 27 (2014)."},{"key":"e_1_2_2_62_1","volume-title":"International Conference on Machine Learning,. PMLR, 12133-12143","author":"You Kaichao","year":"2021","unstructured":"Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. 2021. Logme: Practical assessment of pre-trained models for transfer learning. In International Conference on Machine Learning,. PMLR, 12133-12143."},{"key":"e_1_2_2_63_1","article-title":"Ranking and Tuning Pre-Trained Models: A New Paradigm for Exploiting Model Hubs","volume":"23","author":"You Kaichao","year":"2022","unstructured":"Kaichao You, Yong Liu, Ziyang Zhang, Jianmin Wang, Michael I. Jordan, and Mingsheng Long. 2022. Ranking and Tuning Pre-Trained Models: A New Paradigm for Exploiting Model Hubs. J. Mach. Learn. Res., Vol. 23 (2022), 209:1-209:47.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_2_64_1","first-page":"39","article-title":"Accelerating the machine learning lifecycle with MLflow","volume":"41","author":"Zaharia Matei","year":"2018","unstructured":"Matei Zaharia, Andrew Chen, Aaron Davidson, Ali Ghodsi, Sue Ann Hong, Andy Konwinski, Siddharth Murching, Tomas Nykodym, Paul Ogilvie, Mani Parkhe, and others. 2018. Accelerating the machine learning lifecycle with MLflow. IEEE Data Eng. Bull., Vol. 41, 4 (2018), 39-45.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_2_65_1","first-page":"818","volume-title":"Proceedings, Part I 13","author":"Zeiler Matthew D","year":"2014","unstructured":"Matthew D Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6--12, 2014, Proceedings, Part I 13. Springer, 818-833."},{"key":"e_1_2_2_66_1","volume-title":"Smola","author":"Zhang Aston","year":"2023","unstructured":"Aston Zhang, Zachary C. Lipton, Mu Li, and Alexander J. Smola. 2023. Dive into Deep Learning,. Cambridge University Press."},{"key":"e_1_2_2_67_1","first-page":"28566","article-title":"Gradient-based Parameter Selection for Efficient Fine-Tuning","author":"Zhang Zhi","year":"2024","unstructured":"Zhi Zhang, Qizhe Zhang, Zijun Gao, Renrui Zhang, Ekaterina Shutova, Shiji Zhou, and Shanghang Zhang. 2024. Gradient-based Parameter Selection for Efficient Fine-Tuning. In CVPR. 28566-28577.","journal-title":"CVPR."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725264","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:56:42Z","timestamp":1774983402000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725264"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,17]]},"references-count":67,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,17]]}},"alternative-id":["10.1145\/3725264"],"URL":"https:\/\/doi.org\/10.1145\/3725264","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,17]]}}}