{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:19:44Z","timestamp":1778080784505,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,29]],"date-time":"2023-10-29T00:00:00Z","timestamp":1698537600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,11,2]]},"DOI":"10.1145\/3607827.3616843","type":"proceedings-article","created":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T22:09:13Z","timestamp":1698358153000},"page":"61-67","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Subsampling of Frequent Words in Text for Pre-training a Vision-Language Model"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7129-6799","authenticated-orcid":false,"given":"Mingliang","family":"Liang","sequence":"first","affiliation":[{"name":"Radboud University, Nijmegen, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1430-4721","authenticated-orcid":false,"given":"Martha","family":"Larson","sequence":"additional","affiliation":[{"name":"Radboud University, Nijmegen, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,29]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"e_1_3_2_1_2_1","volume-title":"International conference on machine learning.","author":"Chen Mark","year":"2020","unstructured":"Mark Chen , Alec Radford , Rewon Child , Jeffrey Wu , Heewoo Jun , David Luan , and Ilya Sutskever . 2020 b. Generative pretraining from pixels . In International conference on machine learning. Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020b. Generative pretraining from pixels. In International conference on machine learning."},{"key":"e_1_3_2_1_3_1","volume-title":"International Conference on Machine Learning.","author":"Chen Ting","year":"2020","unstructured":"Ting Chen , Simon Kornblith , Mohammad Norouzi , and Geoffrey Hinton . 2020 a. A Simple Framework for Contrastive Learning of Visual Representations . In International Conference on Machine Learning. Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020a. A Simple Framework for Contrastive Learning of Visual Representations. In International Conference on Machine Learning."},{"key":"e_1_3_2_1_4_1","volume-title":"Describing Textures in the Wild. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition.","author":"Cimpoi M.","unstructured":"M. Cimpoi , S. Maji , I. Kokkinos , S. Mohamed , , and A. Vedaldi . 2014 . Describing Textures in the Wild. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi. 2014. Describing Textures in the Wild. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_1_5_1","volume-title":"An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In International Conference on Artificial Intelligence and Statistics.","author":"Coates Adam","year":"2011","unstructured":"Adam Coates , Andrew Ng , and Honglak Lee . 2011 . An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In International Conference on Artificial Intelligence and Statistics. Adam Coates, Andrew Ng, and Honglak Lee. 2011. An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In International Conference on Artificial Intelligence and Statistics."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_7_1","volume-title":"International Conference on Learning Representations.","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy , Lucas Beyer , Alexander Kolesnikov , Dirk Weissenborn , Xiaohua Zhai , Thomas Unterthiner , Mostafa Dehghani , Matthias Minderer , Georg Heigold , Sylvain Gelly , Jakob Uszkoreit , and Neil Houlsby . 2021 . An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale . In International Conference on Learning Representations. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_8_1","unstructured":"Andrea Frome Gregory S. Corrado Jonathon Shlens Samy Bengio Jeffrey Dean Marc'Aurelio Ranzato and Tom\u00e1 s Mikolov. 2013. DeViSE: A Deep Visual-Semantic Embedding Model. In Advances in Neural Information Processing Systems.  Andrea Frome Gregory S. Corrado Jonathon Shlens Samy Bengio Jeffrey Dean Marc'Aurelio Ranzato and Tom\u00e1 s Mikolov. 2013. DeViSE: A Deep Visual-Semantic Embedding Model. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913491297"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_11_1","unstructured":"Gabriel Ilharco Mitchell Wortsman Ross Wightman Cade Gordon Nicholas Carlini Rohan Taori Achal Dave Vaishaal Shankar Hongseok Namkoong John Miller Hannaneh Hajishirzi Ali Farhadi and Ludwig Schmidt. 2021. OpenCLIP.  Gabriel Ilharco Mitchell Wortsman Ross Wightman Cade Gordon Nicholas Carlini Rohan Taori Achal Dave Vaishaal Shankar Hongseok Namkoong John Miller Hannaneh Hajishirzi Ali Farhadi and Ludwig Schmidt. 2021. OpenCLIP."},{"key":"e_1_3_2_1_12_1","volume-title":"International Conference on Machine Learning.","author":"Jia Chao","year":"2021","unstructured":"Chao Jia , Yinfei Yang , Ye Xia , Yi-Ting Chen , Zarana Parekh , Hieu Pham , Quoc Le , Yun-Hsuan Sung , Zhen Li , and Tom Duerig . 2021 . Scaling up visual and vision-language representation learning with noisy text supervision . In International Conference on Machine Learning. Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02240"},{"key":"e_1_3_2_1_14_1","volume-title":"SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations.","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter . 2017 . SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations. Ilya Loshchilov and Frank Hutter. 2017. SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_15_1","volume-title":"Decoupled Weight Decay Regularization. In International Conference on Learning Representations.","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter . 2019 . Decoupled Weight Decay Regularization. In International Conference on Learning Representations. Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_16_1","volume-title":"Efficient Estimation of Word Representations in Vector Space. In International Conference on Learning Representations.","author":"Mikolov Tom\u00e1","year":"2013","unstructured":"Tom\u00e1 s Mikolov , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013 a. Efficient Estimation of Word Representations in Vector Space. In International Conference on Learning Representations. Tom\u00e1 s Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a. Efficient Estimation of Word Representations in Vector Space. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_17_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013b. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems.  Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013b. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_1_18_1","volume-title":"SLIP: Self-supervision Meets Language-Image Pre-training. In 1European conference on computer vision.","author":"Mu Norman","year":"2022","unstructured":"Norman Mu , Alexander Kirillov , David A. Wagner , and Saining Xie . 2022 . SLIP: Self-supervision Meets Language-Image Pre-training. In 1European conference on computer vision. Norman Mu, Alexander Kirillov, David A. Wagner, and Saining Xie. 2022. SLIP: Self-supervision Meets Language-Image Pre-training. In 1European conference on computer vision."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"e_1_3_2_1_20_1","volume-title":"Cats and Dogs. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"Parkhi Omkar M.","unstructured":"Omkar M. Parkhi , Andrea Vedaldi , Andrew Zisserman , and C. V. Jawahar . 2012 . Cats and Dogs. In IEEE Conference on Computer Vision and Pattern Recognition. Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. 2012. Cats and Dogs. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_1_21_1","volume-title":"Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning.","author":"Radford Alec","year":"2021","unstructured":"Alec Radford , Jong Wook Kim , Chris Hallacy , Aditya Ramesh , Gabriel Goh , Sandhini Agarwal , Girish Sastry , Amanda Askell , Pamela Mishkin , Jack Clark , Gretchen Krueger , and Ilya Sutskever . 2021 . Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In International Conference on Machine Learning."},{"key":"e_1_3_2_1_22_1","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. arxiv: 2204.06125  Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. arxiv: 2204.06125"},{"key":"e_1_3_2_1_23_1","volume-title":"High-Resolution Image Synthesis With Latent Diffusion Models. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition.","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach , Andreas Blattmann , Dominik Lorenz , Patrick Esser , and Bj\u00f6rn Ommer . 2022 . High-Resolution Image Synthesis With Latent Diffusion Models. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_1_24_1","volume-title":"Conference on Neural Information Processing Systems, Datasets and Benchmarks Track.","author":"Schuhmann Christoph","year":"2022","unstructured":"Christoph Schuhmann , Romain Beaumont , Richard Vencu , Cade W Gordon , Ross Wightman , Mehdi Cherti , Theo Coombes , Aarush Katta , Clayton Mullis , Mitchell Wortsman , Patrick Schramowski , Srivatsa R Kundurthy , Katherine Crowson , Ludwig Schmidt , Robert Kaczmarczyk , and Jenia Jitsev . 2022 . LAION-5B: An open large-scale dataset for training next generation image-text models . In Conference on Neural Information Processing Systems, Datasets and Benchmarks Track. Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022. LAION-5B: An open large-scale dataset for training next generation image-text models. In Conference on Neural Information Processing Systems, Datasets and Benchmarks Track."},{"key":"e_1_3_2_1_25_1","unstructured":"Christoph Schuhmann Richard Vencu Romain Beaumont Robert Kaczmarczyk Clayton Mullis Aarush Katta Theo Coombes Jenia Jitsev and Aran Komatsuzaki. 2021. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. arxiv: 2111.02114  Christoph Schuhmann Richard Vencu Romain Beaumont Robert Kaczmarczyk Clayton Mullis Aarush Katta Theo Coombes Jenia Jitsev and Aran Komatsuzaki. 2021. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. arxiv: 2111.02114"},{"key":"e_1_3_2_1_26_1","volume-title":"Image Alt-text Dataset For Automatic Image Captioning. In Annual Meeting of the Association for Computational Linguistics.","author":"Sharma Piyush","year":"2018","unstructured":"Piyush Sharma , Nan Ding , Sebastian Goodman , and Radu Soricut . 2018 . Conceptual Captions: A Cleaned, Hypernymed , Image Alt-text Dataset For Automatic Image Captioning. In Annual Meeting of the Association for Computational Linguistics. Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_2_1_27_1","volume-title":"Representation Learning with Contrastive Predictive Coding. arxiv","author":"van den Oord Aaron","year":"1807","unstructured":"Aaron van den Oord , Yazhe Li , and Oriol Vinyals . 2019. Representation Learning with Contrastive Predictive Coding. arxiv : 1807 .03748 Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019. Representation Learning with Contrastive Predictive Coding. arxiv: 1807.03748"},{"key":"e_1_3_2_1_28_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Bastiaan S. Veeling Jasper Linmans Jim Winkens Taco Cohen and Max Welling. 2018. Rotation Equivariant CNNs for Digital Pathology. In Medical Image Computing and Computer Assisted Intervention.  Bastiaan S. Veeling Jasper Linmans Jim Winkens Taco Cohen and Max Welling. 2018. Rotation Equivariant CNNs for Digital Pathology. In Medical Image Computing and Computer Assisted Intervention.","DOI":"10.1007\/978-3-030-00934-2_24"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01759"}],"event":{"name":"MM '23: The 31st ACM International Conference on Multimedia","location":"Ottawa ON Canada","acronym":"MM '23","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 1st Workshop on Large Generative Models Meet Multimodal Applications"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607827.3616843","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3607827.3616843","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:05Z","timestamp":1750178765000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3607827.3616843"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,29]]},"references-count":31,"alternative-id":["10.1145\/3607827.3616843","10.1145\/3607827"],"URL":"https:\/\/doi.org\/10.1145\/3607827.3616843","relation":{},"subject":[],"published":{"date-parts":[[2023,10,29]]},"assertion":[{"value":"2023-10-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}