{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,8]],"date-time":"2025-07-08T04:09:00Z","timestamp":1751947740504,"version":"3.41.2"},"reference-count":73,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T00:00:00Z","timestamp":1751846400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T00:00:00Z","timestamp":1751846400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004687","name":"Universidad de Murcia","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004687","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Transformer-based multimodal models often require expensive, full-model training on task-specific all-modality datasets to achieve high accuracy on targeted downstream tasks. To reduce this significant cost, we introduce SAFFE, a methodology for building accurate, task-specific multimodal models with minimal training, using only standard GPU hardware. SAFFE leverages off-the-shelf, pre-trained, frozen unimodal encoders for each input modality (e.g., text, image, or audio) and connects them through a lightweight, trainable component called the FusionAlign Module (FAM). FAM is a bottleneck mid-fusion neural network, trained on the target dataset to align the outputs of the independently pre-trained unimodal encoders. This approach eliminates the need for end-to-end training while maintaining strong accuracy for the downstream task. As a proof of concept, we validate SAFFE on image retrieval and language understanding tasks. SAFFE-derived models outperform state-of-the-art multimodal systems on datasets such as CIFAR-10, ImageNet-100, and COCO, achieving competitive results with significantly fewer trainable parameters and training time.<\/jats:p>","DOI":"10.1007\/s11227-025-07473-7","type":"journal-article","created":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T12:08:36Z","timestamp":1751890116000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Saffe: Multimodal Model Composition with Semantic-Alignment Fusion of Frozen Encoders"],"prefix":"10.1007","volume":"81","author":[{"given":"Maithri","family":"Kulasekara","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan F.","family":"Ingl\u00e9s-Romero","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baldomero","family":"Imbern\u00f3n","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jos\u00e9 L.","family":"Abell\u00e1n","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,7]]},"reference":[{"key":"7473_CR1","first-page":"24206","volume":"34","author":"H Akbari","year":"2021","unstructured":"Akbari H, Yuan L, Qian R, Chuang W-H, Chang S-F, Cui Y, Gong B (2021) Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems 34:24206\u201324221","journal-title":"Advances in Neural Information Processing Systems"},{"key":"7473_CR2","doi-asserted-by":"crossref","unstructured":"Gu Z, Lang B, Yue T, Huang L (2017) Learning joint multimodal representation based on multi-fusion deep neural networks. In: Neural Information Processing: 24th International Conference, ICONIP 2017, Guangzhou, China, November 14-18, 2017, Proceedings, Part II 24, pp. 276\u2013285 Springer","DOI":"10.1007\/978-3-319-70096-0_29"},{"key":"7473_CR3","first-page":"14200","volume":"34","author":"A Nagrani","year":"2021","unstructured":"Nagrani A, Yang S, Arnab A, Jansen A, Schmid C, Sun C (2021) Attention bottlenecks for multimodal fusion. Advances in neural information processing systems 34:14200\u201314213","journal-title":"Advances in neural information processing systems"},{"key":"7473_CR4","doi-asserted-by":"crossref","unstructured":"Girdhar R, Singh M, Ravi N, Van Der\u00a0Maaten L, Joulin A, Misra I (2022) Omnivore: A single model for many visual modalities. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16102\u201316112","DOI":"10.1109\/CVPR52688.2022.01563"},{"key":"7473_CR5","doi-asserted-by":"crossref","unstructured":"Guzhov A, Raue F, Hees J, Dengel A (2022) Audioclip: Extending clip to image, text and audio. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 976\u2013980 IEEE","DOI":"10.1109\/ICASSP43922.2022.9747631"},{"key":"7473_CR6","doi-asserted-by":"crossref","unstructured":"Girdhar R, El-Nouby A, Liu Z, Singh M, Alwala KV, Joulin A, Misra I (2023) Imagebind: One embedding space to bind them all. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 15180\u201315190","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"7473_CR7","unstructured":"Zhang Y, Gong K, Zhang K, Li H, Qiao Y, Ouyang W, Yue X (2023) Meta-transformer: A unified framework for multimodal learning. arXiv preprint arXiv:2307.10802"},{"key":"7473_CR8","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740\u2013755 Springer","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"7473_CR9","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J et al (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763 PMLR"},{"key":"7473_CR10","unstructured":"Jia C, Yang Y, Xia Y, Chen Y-T, Parekh Z, Pham H, Le QV, Sung Y, Li Z, Duerig T (2021) Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision arxiv: 2102.05918"},{"key":"7473_CR11","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.111037","volume":"280","author":"J Liu","year":"2023","unstructured":"Liu J, Mao Y, Huang Z, Ye Y (2023) A bottleneck network with light attention for multimodal clustering. Knowledge-Based Systems 280:111037","journal-title":"Knowledge-Based Systems"},{"key":"7473_CR12","doi-asserted-by":"crossref","unstructured":"Shvetsova N, Chen B, Rouditchenko A, Thomas S, Kingsbury B, Feris RS, Harwath D, Glass J, Kuehne H (2022) Everything at once-multi-modal fusion transformer for video retrieval. In: Proceedings of the Ieee\/cvf Conference on Computer Vision and Pattern Recognition, pp. 20020\u201320029","DOI":"10.1109\/CVPR52688.2022.01939"},{"key":"7473_CR13","doi-asserted-by":"crossref","unstructured":"Li Y, Jiang S, Hu B, Wang L, Zhong W, Luo W, Ma L, Zhang M (2024) Uni-moe: Scaling unified multimodal llms with mixture of experts. arXiv preprint arXiv:2405.11273","DOI":"10.1109\/TPAMI.2025.3532688"},{"key":"7473_CR14","doi-asserted-by":"crossref","unstructured":"Piergiovanni A, Noble I, Kim D, Ryoo MS, Gomes V, Angelova A (2024) Mirasol3b: A multimodal autoregressive model for time-aligned and contextual modalities. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 26804\u201326814","DOI":"10.1109\/CVPR52733.2024.02531"},{"key":"7473_CR15","doi-asserted-by":"crossref","unstructured":"Xie S, Sun C, Huang J, Tu Z, Murphy K (2018) Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 305\u2013321","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"7473_CR16","doi-asserted-by":"crossref","unstructured":"Wang X, Girshick R, Gupta A, He K (2018) Non-local neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7794\u20137803","DOI":"10.1109\/CVPR.2018.00813"},{"key":"7473_CR17","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114939","volume":"177","author":"Q Chen","year":"2021","unstructured":"Chen Q, Wang W, Huang K, De S, Coenen F (2021) Multi-modal generative adversarial networks for traffic event detection in smart cities. Expert Systems with Applications 177:114939","journal-title":"Expert Systems with Applications"},{"key":"7473_CR18","first-page":"23716","volume":"35","author":"J-B Alayrac","year":"2022","unstructured":"Alayrac J-B, Donahue J, Luc P, Miech A, Barr I, Hasson Y, Lenc K, Mensch A, Millican K, Reynolds M et al (2022) Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35:23716\u201323736","journal-title":"Advances in neural information processing systems"},{"key":"7473_CR19","unstructured":"Koh JY, Salakhutdinov R, Fried D (2023) Grounding language models to images for multimodal inputs and outputs. In: International Conference on Machine Learning, pp. 17283\u201317300 PMLR"},{"key":"7473_CR20","unstructured":"Li J, Li D, Xiong C, Hoi S (2022) Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning, pp. 12888\u201312900 PMLR"},{"key":"7473_CR21","unstructured":"Merullo J, Castricato L, Eickhoff C, Pavlick E (2023) Linearly mapping from image to text space. In: The Eleventh International Conference on Learning Representations"},{"key":"7473_CR22","unstructured":"Vaswani A (2017) Attention is all you need. Advances in Neural Information Processing Systems"},{"issue":"6","key":"7473_CR23","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1007\/s00138-021-01249-8","volume":"32","author":"SY Boulahia","year":"2021","unstructured":"Boulahia SY, Amamra A, Madi MR, Daikh S (2021) Early, intermediate and late fusion strategies for robust deep learning-based multimodal action recognition. Machine Vision and Applications 32(6):121","journal-title":"Machine Vision and Applications"},{"key":"7473_CR24","doi-asserted-by":"crossref","unstructured":"Yang Z, Wang J, Tang Y, Chen K, Zhao H, Torr PH (2022) Lavt: Language-aware vision transformer for referring image segmentation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 18155\u201318165","DOI":"10.1109\/CVPR52688.2022.01762"},{"key":"7473_CR25","doi-asserted-by":"crossref","unstructured":"Zheng G, Zhou X, Li X, Qi Z, Shan Y, Li X (2023) Layoutdiffusion: Controllable diffusion model for layout-to-image generation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 22490\u201322499","DOI":"10.1109\/CVPR52729.2023.02154"},{"key":"7473_CR26","doi-asserted-by":"publisher","first-page":"6126","DOI":"10.1609\/aaai.v38i6.28429","volume":"38","author":"T Wu","year":"2024","unstructured":"Wu T, Li X, Qi Z, Hu D, Wang X, Shan Y, Li X (2024) Spherediffusion: Spherical geometry-aware distortion resilient diffusion model. Proceedings of the AAAI Conference on Artificial Intelligence 38:6126\u20136134","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"issue":"7","key":"7473_CR27","doi-asserted-by":"publisher","first-page":"5961","DOI":"10.1109\/TCYB.2021.3052522","volume":"52","author":"Y Guo","year":"2021","unstructured":"Guo Y, Gao L, Song J, Wang P, Sebe N, Shen HT, Li X (2021) Relation regularized scene graph generation. IEEE Transactions on Cybernetics 52(7):5961\u20135972","journal-title":"IEEE Transactions on Cybernetics"},{"key":"7473_CR28","doi-asserted-by":"crossref","unstructured":"Li X, Zheng G, Yu Y, Ji N, Li X (2024) Relationship-incremental scene graph generation by a divide-and-conquer pipeline with feature adapter. IEEE Transactions on Image Processing","DOI":"10.1109\/TIP.2024.3384096"},{"key":"7473_CR29","doi-asserted-by":"crossref","unstructured":"Morency L-P, Baltru\u0161aitis T (2017) Multimodal machine learning: integrating language, vision and speech. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts, pp. 3\u20135","DOI":"10.18653\/v1\/P17-5002"},{"key":"7473_CR30","first-page":"25278","volume":"35","author":"C Schuhmann","year":"2022","unstructured":"Schuhmann C, Beaumont R, Vencu R, Gordon C, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M et al (2022) Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems 35:25278\u201325294","journal-title":"Advances in Neural Information Processing Systems"},{"key":"7473_CR31","unstructured":"Wang B, Komatsuzaki A (2021) GPT-J-6B: A 6 billion parameter autoregressive language model"},{"key":"7473_CR32","doi-asserted-by":"crossref","unstructured":"Morency L-P, Baltru\u0161aitis T (2017) Multimodal machine learning: integrating language, vision and speech. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts, pp. 3\u20135","DOI":"10.18653\/v1\/P17-5002"},{"key":"7473_CR33","doi-asserted-by":"crossref","unstructured":"Wang W, Tran D, Feiszli M (2020) What makes training multi-modal classification networks hard? In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 12695\u201312705","DOI":"10.1109\/CVPR42600.2020.01271"},{"key":"7473_CR34","doi-asserted-by":"crossref","unstructured":"Cherti M, Beaumont R, Wightman R, Wortsman M, Ilharco G, Gordon C, Schuhmann C, Schmidt L, Jitsev J (2023) Reproducible scaling laws for contrastive language-image learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 2818\u20132829","DOI":"10.1109\/CVPR52729.2023.00276"},{"key":"7473_CR35","unstructured":"Chen S, He X, Guo L, Zhu X, Wang W, Tang J, Liu J (2023) Valor: Vision-audio-language omni-perception pretraining model and dataset. arXiv preprint arXiv:2304.08345"},{"key":"7473_CR36","doi-asserted-by":"crossref","unstructured":"Hu R, Singh A (2021) Unit: Multimodal multitask learning with a unified transformer. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 1439\u20131449","DOI":"10.1109\/ICCV48922.2021.00147"},{"key":"7473_CR37","doi-asserted-by":"crossref","unstructured":"Zhai X, Mustafa B, Kolesnikov A, Beyer L (2023) Sigmoid loss for language image pre-training. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), pp. 11975\u201311986","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"7473_CR38","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248\u2013255 Ieee","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"7473_CR39","unstructured":"Krizhevsky A (2009) Learning multiple layers of features from tiny images. Technical report"},{"key":"7473_CR40","doi-asserted-by":"crossref","unstructured":"Kalantidis Y, Tolias G et al (2024) Label propagation for zero-shot classification with vision-language models. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 23209\u201323218","DOI":"10.1109\/CVPR52733.2024.02190"},{"key":"7473_CR41","unstructured":"Najdenkoska I, Zhen X, Worring M (2022) Meta-learning makes a better multimodal few-shot learner. In: Sixth Workshop on Meta-Learning at the Conference on Neural Information Processing Systems"},{"key":"7473_CR42","unstructured":"Devlin J (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805"},{"key":"7473_CR43","doi-asserted-by":"crossref","unstructured":"Reimers N (2019) Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084","DOI":"10.18653\/v1\/D19-1410"},{"key":"7473_CR44","unstructured":"Liu Y (2019) Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692364"},{"key":"7473_CR45","unstructured":"Dosovitskiy A (2020) An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"7473_CR46","doi-asserted-by":"crossref","unstructured":"Gong Y, Chung Y-A, Glass J (2021) Ast: Audio spectrogram transformer. arXiv preprint arXiv:2104.01778","DOI":"10.21437\/Interspeech.2021-698"},{"key":"7473_CR47","doi-asserted-by":"crossref","unstructured":"Yao M, Tao D, Gao R, Qi P (2024) Anomaly detection for mec enabled hierarchical industrial iot with transformer enhanced variational auto encoder. IEEE Transactions on Industrial Informatics","DOI":"10.1109\/TII.2024.3421600"},{"key":"7473_CR48","doi-asserted-by":"crossref","unstructured":"Arnab A, Dehghani M, Heigold G, Sun C, Lu\u010di\u0107 M, Schmid C (2021) Vivit: A video vision transformer. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 6836\u20136846","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"7473_CR49","doi-asserted-by":"crossref","unstructured":"Liao J, Shi Y, Gong M, Shou L, Qu H, Zeng M (2021) Improving zero-shot neural machine translation on language-specific encoders-decoders. In: 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1\u20138 IEEE","DOI":"10.1109\/IJCNN52387.2021.9534401"},{"key":"7473_CR50","doi-asserted-by":"crossref","unstructured":"Cho Y, Yu H, Kang S-J (2023) Cross-aware early fusion with stage-divided vision and language transformer encoders for referring image segmentation. IEEE Transactions on Multimedia","DOI":"10.1109\/TMM.2023.3340062"},{"key":"7473_CR51","doi-asserted-by":"crossref","unstructured":"Ding H, Liu C, Wang S, Jiang X (2021) Vision-language transformer and query generation for referring segmentation. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 16321\u201316330","DOI":"10.1109\/ICCV48922.2021.01601"},{"key":"7473_CR52","first-page":"200","volume":"34","author":"M Tsimpoukelli","year":"2021","unstructured":"Tsimpoukelli M, Menick JL, Cabi S, Eslami S, Vinyals O, Hill F (2021) Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems 34:200\u2013212","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"10s","key":"7473_CR53","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3505244","volume":"54","author":"S Khan","year":"2022","unstructured":"Khan S, Naseer M, Hayat M, Zamir SW, Khan FS, Shah M (2022) Transformers in vision: A survey. ACM computing surveys (CSUR) 54(10s):1\u201341","journal-title":"ACM computing surveys (CSUR)"},{"key":"7473_CR54","unstructured":"Pang Z, Xie Z, Man Y, Wang Y-X (2023) Frozen transformers in language models are effective visual encoder layers. arXiv preprint arXiv:2310.12973"},{"key":"7473_CR55","unstructured":"Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M-A, Lacroix T, Rozi\u00e8re B, Goyal N, Hambro E, Azhar F et al (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971"},{"key":"7473_CR56","doi-asserted-by":"crossref","unstructured":"Liu Z (2024) Improving the inference efficiency of transformer models in machine translation tasks by low-rank decomposition methods. In: 2024 International Conference on Electronics and Devices, Computational Science (ICEDCS), pp. 930\u2013936 IEEE","DOI":"10.1109\/ICEDCS64328.2024.00173"},{"key":"7473_CR57","doi-asserted-by":"crossref","unstructured":"Qin B, Li J, Tang S, Zhuang Y (2025) Dba: Efficient transformer with dynamic bilinear low-rank attention. IEEE Transactions on Neural Networks and Learning Systems","DOI":"10.1109\/TNNLS.2025.3527046"},{"key":"7473_CR58","unstructured":"Chen Y, Shang J, Zhang Z, Sheng J, Liu T, Wang S, Sun Y, Wu H, Wang H (2024) Mixture of hidden-dimensions transformer. arXiv preprint arXiv:2412.05644"},{"key":"7473_CR59","unstructured":"Cukierski W (2013) Dogs vs. Cats. https:\/\/kaggle.com\/competitions\/dogs-vs-cats. Kaggle"},{"key":"7473_CR60","unstructured":"CIFAR 10. (2021) https:\/\/huggingface.co\/datasets\/uoft-cs\/cifar10. Hugging Face"},{"key":"7473_CR61","unstructured":"Shekhar A (2021) ImageNet100. https:\/\/www.kaggle.com\/datasets\/ambityga\/imagenet100. Kaggle"},{"key":"7473_CR62","unstructured":"CIFAR 100. (2021) https:\/\/huggingface.co\/datasets\/uoft-cs\/cifar100. Hugging Face"},{"key":"7473_CR63","unstructured":"COCO. (2021) https:\/\/huggingface.co\/datasets\/detection-datasets\/coco. huggingface"},{"key":"7473_CR64","unstructured":"Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L et al (2019) Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32"},{"key":"7473_CR65","unstructured":"Wu B, Xu C, Dai X, Wan A, Zhang P, Yan Z, Tomizuka M, Gonzalez J, Keutzer K, Vajda P (2020) Visual Transformers: Token-based Image Representation and Processing for Computer Vision"},{"key":"7473_CR66","unstructured":"Wolf T (2019) Huggingface\u2019s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771"},{"key":"7473_CR67","unstructured":"Kingma DP (2014) Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980"},{"key":"7473_CR68","first-page":"26418","volume":"35","author":"J Gu","year":"2022","unstructured":"Gu J, Meng X, Lu G, Hou L, Minzhe N, Liang X, Yao L, Huang R, Zhang W, Jiang X et al (2022) Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. Advances in Neural Information Processing Systems 35:26418\u201326431","journal-title":"Advances in Neural Information Processing Systems"},{"key":"7473_CR69","doi-asserted-by":"crossref","unstructured":"He K, Chen X, Xie S, Li Y, Doll\u00e1r P, Girshick R (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000\u201316009","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"7473_CR70","unstructured":"Yang A, Pan J, Lin J, Men R, Zhang Y, Zhou J, Zhou C (2022) Chinese clip: Contrastive vision-language pretraining in chinese. arXiv preprint arXiv:2211.01335"},{"key":"7473_CR71","doi-asserted-by":"crossref","unstructured":"Jian Y, Liu T, Tao Y, Zhang C, Vosoughi S, Yang H (2023) Expedited training of visual conditioned language generation via redundancy reduction. arXiv preprint arXiv:2310.03291","DOI":"10.18653\/v1\/2024.acl-long.19"},{"key":"7473_CR72","first-page":"9694","volume":"34","author":"J Li","year":"2021","unstructured":"Li J, Selvaraju R, Gotmare A, Joty S, Xiong C, Hoi SCH (2021) Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems 34:9694\u20139705","journal-title":"Advances in neural information processing systems"},{"key":"7473_CR73","doi-asserted-by":"crossref","unstructured":"Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers. In: European Conference on Computer Vision, pp. 213\u2013229 Springer","DOI":"10.1007\/978-3-030-58452-8_13"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07473-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-025-07473-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07473-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,7]],"date-time":"2025-07-07T12:08:49Z","timestamp":1751890129000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-025-07473-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,7]]},"references-count":73,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["7473"],"URL":"https:\/\/doi.org\/10.1007\/s11227-025-07473-7","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,7]]},"assertion":[{"value":"16 May 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 July 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"1114"}}