{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T10:25:27Z","timestamp":1785147927434,"version":"3.55.0"},"publisher-location":"Cham","reference-count":44,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783031726422","type":"print"},{"value":"9783031726439","type":"electronic"}],"license":[{"start":{"date-parts":[[2024,11,22]],"date-time":"2024-11-22T00:00:00Z","timestamp":1732233600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,11,22]],"date-time":"2024-11-22T00:00:00Z","timestamp":1732233600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Power-law scaling indicates that large-scale training with uniform sampling is prohibitively slow. Active learning methods aim to increase data efficiency by prioritizing learning on the most relevant examples. Despite their appeal, these methods have yet to be widely adopted since no one algorithm has been shown to a) generalize across models and tasks b) scale to large datasets and c) yield overall FLOP savings when accounting for the overhead of data selection. In this work we propose a method which satisfies these three properties, leveraging small, cheap proxy models to estimate \u201clearnability\u201d scores for datapoints, which are used to prioritize data for training much larger models. As a result, models trained using our methods \u2013 <jats:italic>ClassAct<\/jats:italic> and <jats:italic>ActiveCLIP<\/jats:italic> \u2013 require 46% and 51% fewer training updates and up to 25% less total computation to reach the same performance as uniformly-trained visual classifiers on JFT and multimodal models on ALIGN, respectively. Finally, we find our data-prioritization scheme to be complementary with recent data-curation and learning objectives, yielding a new state-of-the-art in several multimodal transfer tasks.<\/jats:p>","DOI":"10.1007\/978-3-031-72643-9_16","type":"book-chapter","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T20:46:50Z","timestamp":1732222010000},"page":"264-280","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding"],"prefix":"10.1007","author":[{"given":"Talfan","family":"Evans","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shreya","family":"Pathak","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hamza","family":"Merzic","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonathan","family":"Schwarz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryutaro","family":"Tanno","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Olivier J.","family":"H\u00e9naff","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,11,22]]},"reference":[{"key":"16_CR1","unstructured":"Abbas, A., Tirumala, K., Simig, D., Ganguli, S., Morcos, A.S.: SemDeDup: data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540 (2023)"},{"key":"16_CR2","unstructured":"Alayrac, J.B., et\u00a0al.: Flamingo: a visual language model for few-shot learning. In: Advances in Neural Information Processing Systems (2022)"},{"key":"16_CR3","unstructured":"Brown, T., et\u00a0al.: Language models are few-shot learners. In: Advances in Neural Information Processing Systems (2020)"},{"key":"16_CR4","unstructured":"Campbell, T., Broderick, T.: Bayesian coreset construction via greedy iterative geodesic ascent. In: International Conference on Machine Learning, pp. 698\u2013706. PMLR (2018)"},{"key":"16_CR5","unstructured":"Cassirer, A., et al.: Reverb: a framework for experience replay (2021)"},{"key":"16_CR6","unstructured":"Chen, X., et\u00a0al.: PaLI: a jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794 (2022)"},{"issue":"240","key":"16_CR7","first-page":"1","volume":"24","author":"A Chowdhery","year":"2023","unstructured":"Chowdhery, A., et al.: PaLM: scaling language modeling with pathways. J. Mach. Learn. Res. 24(240), 1\u2013113 (2023)","journal-title":"J. Mach. Learn. Res."},{"key":"16_CR8","unstructured":"Coleman, C., et al.: Selection via proxy: efficient data selection for deep learning. arXiv preprint arXiv:1906.11829 (2019)"},{"key":"16_CR9","unstructured":"Dosovitskiy, A., et\u00a0al.: An image is worth 16\u00a0$$\\times $$\u00a016 words: transformers for image recognition at scale. In: International Conference on Learning Representations (2021)"},{"key":"16_CR10","unstructured":"Espeholt, L., et\u00a0al.: IMPALA: scalable distributed deep-RL with importance weighted actor-learner architectures. In: International Conference on Machine Learning, pp. 1407\u20131416. PMLR (2018)"},{"key":"16_CR11","doi-asserted-by":"crossref","unstructured":"Feldman, V.: Does learning require memorization? A short tale about a long tail. In: Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, pp. 954\u2013959 (2020)","DOI":"10.1145\/3357713.3384290"},{"key":"16_CR12","unstructured":"Gadre, S.Y., et\u00a0al.: DataComp: in search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108 (2023)"},{"key":"16_CR13","unstructured":"Gunasekar, S., et\u00a0al.: Textbooks are all you need. arXiv preprint arXiv:2306.11644 (2023)"},{"key":"16_CR14","doi-asserted-by":"crossref","unstructured":"Har-Peled, S., Kushal, A.: Smaller coresets for k-median and k-means clustering. In: Proceedings of the Twenty-First Annual Symposium on Computational Geometry, pp. 126\u2013134 (2005)","DOI":"10.1145\/1064092.1064114"},{"key":"16_CR15","doi-asserted-by":"crossref","unstructured":"Hessel, J., Holtzman, A., Forbes, M., Bras, R.L., Choi, Y.: ClipScore: a reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718 (2021)","DOI":"10.18653\/v1\/2021.emnlp-main.595"},{"key":"16_CR16","unstructured":"Hoffmann, J., et\u00a0al.: Training compute-optimal large language models. In: Advances in Neural Information Processing Systems (2022)"},{"key":"16_CR17","doi-asserted-by":"publisher","unstructured":"Ilharco, G., et al.: Openclip, July 2021. https:\/\/doi.org\/10.5281\/zenodo.5143773, if you use this software, please cite it as below","DOI":"10.5281\/zenodo.5143773"},{"key":"16_CR18","unstructured":"Jouppi, N.P., et\u00a0al.: In-datacenter performance analysis of a tensor processing unit. In: Proceedings of the 44th Annual International Symposium on Computer Architecture, pp. 1\u201312 (2017)"},{"key":"16_CR19","unstructured":"Kaplan, J., et al.: Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)"},{"key":"16_CR20","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","volume-title":"Computer Vision \u2013 ECCV 2014","author":"T-Y Lin","year":"2014","unstructured":"Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740\u2013755. Springer, Cham (2014). https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48"},{"issue":"4","key":"16_CR21","doi-asserted-by":"publisher","first-page":"986","DOI":"10.1214\/aoms\/1177728069","volume":"27","author":"DV Lindley","year":"1956","unstructured":"Lindley, D.V.: On a measure of the information provided by an experiment. Ann. Math. Stat. 27(4), 986\u20131005 (1956)","journal-title":"Ann. Math. Stat."},{"key":"16_CR22","unstructured":"Loshchilov, I., Hutter, F.: Online batch selection for faster training of neural networks. arXiv preprint arXiv:1511.06343 (2015)"},{"key":"16_CR23","unstructured":"Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: International Conference on Learning Representations (2019)"},{"issue":"4","key":"16_CR24","doi-asserted-by":"publisher","first-page":"590","DOI":"10.1162\/neco.1992.4.4.590","volume":"4","author":"DJ MacKay","year":"1992","unstructured":"MacKay, D.J.: Information-based objective functions for active data selection. Neural Comput. 4(4), 590\u2013604 (1992)","journal-title":"Neural Comput."},{"key":"16_CR25","doi-asserted-by":"crossref","unstructured":"Mahmoud, A., et al.: SIEVE: multimodal dataset pruning using image captioning models. arXiv preprint arXiv:2310.02110 (2023)","DOI":"10.1109\/CVPR52733.2024.02116"},{"key":"16_CR26","unstructured":"Marion, M., \u00dcst\u00fcn, A., Pozzobon, L., Wang, A., Fadaee, M., Hooker, S.: When less is more: investigating data pruning for pretraining LLMs at scale. arXiv preprint arXiv:2309.04564 (2023)"},{"key":"16_CR27","unstructured":"Mindermann, S., et\u00a0al.: Prioritized training on points that are learnable, worth learning, and not yet learnt. In: International Conference on Machine Learning, pp. 15630\u201315649. PMLR (2022)"},{"key":"16_CR28","first-page":"20596","volume":"34","author":"M Paul","year":"2021","unstructured":"Paul, M., Ganguli, S., Dziugaite, G.K.: Deep learning on a data diet: finding important examples early in training. Adv. Neural. Inf. Process. Syst. 34, 20596\u201320607 (2021)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"16_CR29","doi-asserted-by":"crossref","unstructured":"Prabhu, A., Dognin, C., Singh, M.: Sampling bias in deep active classification: an empirical study. arXiv preprint arXiv:1909.09389 (2019)","DOI":"10.18653\/v1\/D19-1417"},{"key":"16_CR30","unstructured":"Radford, A., et\u00a0al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning (2021)"},{"key":"16_CR31","unstructured":"Schaul, T., Quan, J., Antonoglou, I., Silver, D.: Prioritized experience replay. arXiv preprint arXiv:1511.05952 (2015)"},{"key":"16_CR32","first-page":"25278","volume":"35","author":"C Schuhmann","year":"2022","unstructured":"Schuhmann, C., et al.: LAION-5B: an open large-scale dataset for training next generation image-text models. Adv. Neural. Inf. Process. Syst. 35, 25278\u201325294 (2022)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"16_CR33","unstructured":"Schuhmann, C., et al.: LAION-400M: open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114 (2021)"},{"key":"16_CR34","unstructured":"Settles, B.: Active learning literature survey (2009)"},{"key":"16_CR35","first-page":"19523","volume":"35","author":"B Sorscher","year":"2022","unstructured":"Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., Morcos, A.: Beyond neural scaling laws: beating power law scaling via data pruning. Adv. Neural. Inf. Process. Syst. 35, 19523\u201319536 (2022)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"16_CR36","doi-asserted-by":"crossref","unstructured":"Sun, C., Shrivastava, A., Singh, S., Gupta, A.: Revisiting unreasonable effectiveness of data in deep learning era. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 843\u2013852 (2017)","DOI":"10.1109\/ICCV.2017.97"},{"key":"16_CR37","unstructured":"Sun, Q., Fang, Y., Wu, L., Wang, X., Cao, Y.: EVA-CLIP: improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389 (2023)"},{"key":"16_CR38","unstructured":"Toneva, M., Sordoni, A., Combes, R.T.D., Trischler, A., Bengio, Y., Gordon, G.J.: An empirical study of example forgetting during deep neural network learning. arXiv preprint arXiv:1812.05159 (2018)"},{"key":"16_CR39","unstructured":"Xie, S.M., et al.: DoReMi: optimizing data mixtures speeds up language model pretraining (2023)"},{"key":"16_CR40","unstructured":"Xie, S.M., Santurkar, S., Ma, T., Liang, P.: Data selection for language models via importance resampling. arXiv preprint arXiv:2302.03169 (2023)"},{"key":"16_CR41","unstructured":"Yang, F., et al.: Launchpad: a programming model for distributed machine learning research. arXiv preprint arXiv:2106.04516 (2021). https:\/\/arxiv.org\/abs\/2106.04516"},{"key":"16_CR42","unstructured":"Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., Wu, Y.: CoCa: contrastive captioners are image-text foundation models. In: Transactions on Machine Learning Research (2022)"},{"key":"16_CR43","doi-asserted-by":"crossref","unstructured":"Zhai, X., Kolesnikov, A., Houlsby, N., Beyer, L.: Scaling vision transformers. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 12104\u201312113 (2022)","DOI":"10.1109\/CVPR52688.2022.01179"},{"key":"16_CR44","doi-asserted-by":"crossref","unstructured":"Zhai, X., Mustafa, B., Kolesnikov, A., Beyer, L.: Sigmoid loss for language image pre-training. arXiv preprint arXiv:2303.15343 (2023)","DOI":"10.1109\/ICCV51070.2023.01100"}],"container-title":["Lecture Notes in Computer Science","Computer Vision \u2013 ECCV 2024"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-031-72643-9_16","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T21:29:15Z","timestamp":1732224555000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-031-72643-9_16"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,22]]},"ISBN":["9783031726422","9783031726439"],"references-count":44,"URL":"https:\/\/doi.org\/10.1007\/978-3-031-72643-9_16","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,22]]},"assertion":[{"value":"22 November 2024","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"ECCV","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"European Conference on Computer Vision","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Milan","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Italy","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2024","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"29 September 2024","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"4 October 2024","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"18","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"eccv2024","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/eccv2024.ecva.net\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}