{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T01:42:13Z","timestamp":1784338933297,"version":"3.55.0"},"reference-count":194,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,12,27]],"date-time":"2024-12-27T00:00:00Z","timestamp":1735257600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,12,27]],"date-time":"2024-12-27T00:00:00Z","timestamp":1735257600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Key Laboratory of Flight Techniques and Flight Safety","award":["FZ2022KF06"],"award-info":[{"award-number":["FZ2022KF06"]}]},{"name":"Sichuan Province Engineering Technology Research Center of General Aircraft Maintenance","award":["GAMRC2023YB06"],"award-info":[{"award-number":["GAMRC2023YB06"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In recent years, large-scale artificial intelligence (AI) models have become a focal point in technology, attracting widespread attention and acclaim. Notable examples include Google\u2019s BERT and OpenAI\u2019s GPT, which have scaled their parameter sizes to hundreds of billions or even tens of trillions. This growth has been accompanied by a significant increase in the amount of training data, significantly improving the capabilities and performance of these models. Unlike previous reviews, this paper provides a comprehensive discussion of the algorithmic principles of large-scale AI models and their industrial applications from multiple perspectives. We first outline the evolutionary history of these models, highlighting milestone algorithms while exploring their underlying principles and core technologies. We then evaluate the challenges and limitations of large-scale AI models, including computational resource requirements, model parameter inflation, data privacy concerns, and specific issues related to multi-modal AI models, such as reliance on text-image pairs, inconsistencies in understanding and generation capabilities, and the lack of true \u201cmulti-modality\u201d. Various industrial applications of these models are also presented. Finally, we discuss future trends, predicting further expansion of model scale and the development of cross-modal fusion. This study provides valuable insights to inform and inspire future future research and practice.<\/jats:p>","DOI":"10.1007\/s44267-024-00065-8","type":"journal-article","created":{"date-parts":[[2024,12,27]],"date-time":"2024-12-27T09:39:59Z","timestamp":1735292399000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":72,"title":["An overview of large AI models and their applications"],"prefix":"10.1007","volume":"2","author":[{"given":"Xiaoguang","family":"Tu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhi","family":"He","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhi-Hao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3508-756X","authenticated-orcid":false,"given":"Jian","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,12,27]]},"reference":[{"issue":"3","key":"65_CR1","doi-asserted-by":"publisher","first-page":"685","DOI":"10.1007\/s12525-021-00475-2","volume":"31","author":"C. Janiesch","year":"2021","unstructured":"Janiesch, C., Zschech, P., & Heinrich, K. (2021). Machine learning and deep learning. Electronic Markets, 31(3), 685\u2013695.","journal-title":"Electronic Markets"},{"key":"65_CR2","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1016\/j.neucom.2016.12.038","volume":"234","author":"W. Liu","year":"2017","unstructured":"Liu, W., Wang, Z., Liu, X., Zeng, N., Liu, Y., & Alsaadi, F. E. (2017). A survey of deep neural network architectures and their applications. Neurocomputing, 234, 11\u201326.","journal-title":"Neurocomputing"},{"key":"65_CR3","first-page":"163","volume-title":"Proceedings of the 55th international scientific conference on information, communication and energy systems and technologies","author":"L. Sandjakoska","year":"2020","unstructured":"Sandjakoska, L., & Stojanovska, F. (2020). How initialization is related to deep neural networks generalization capability: experimental study. In Proceedings of the 55th international scientific conference on information, communication and energy systems and technologies (pp. 163\u2013166). Piscataway: IEEE."},{"issue":"48","key":"65_CR4","first-page":"1","volume":"13","author":"E. H. Ettaouil","year":"2021","unstructured":"Ettaouil, E. H. (2021). Generalization ability augmentation and regularization of deep convolutional neural networks using l1\/2 pooling. International Journal on Technical and Physical Problems of Engineering, 13(48), 1\u20136.","journal-title":"International Journal on Technical and Physical Problems of Engineering"},{"key":"65_CR5","first-page":"1","volume-title":"Proceedings of the 7th international conference on learning representations","author":"C. Singh","year":"2019","unstructured":"Singh, C., Murdoch, W. J., & Yu, B. (2019). Hierarchical interpretations for neural network predictions. In Proceedings of the 7th international conference on learning representations (pp. 1\u201326). Retrieved November 11, 2024, from https:\/\/openreview.net\/forum?id=SkEqro0ctQ."},{"key":"65_CR6","unstructured":"Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., et\u00a0al. (2022). Emergent abilities of large language models. Retrieved November 11, 2024, from https:\/\/openreview.net\/forum?id=yzkSU5zdwD."},{"key":"65_CR7","first-page":"28785","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"S. Swaminathan","year":"2023","unstructured":"Swaminathan, S., Dedieu, A., Vasudeva Raju, R., Shanahan, M., Lazaro-Gredilla, M., & George, D. (2023). Schema-learning and rebinding as mechanisms of in-context learning and emergence. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 28785\u201328804). Red Hook: Curran Associates."},{"key":"65_CR8","unstructured":"Boiko, D. A., MacKnight, R., & Gomes, G. (2023). Emergent autonomous scientific research capabilities of large language models. arXiv preprint. arXiv:2304.05332."},{"key":"65_CR9","first-page":"5998","volume-title":"Proceedings of the 31st international conference on neural information processing systems","author":"A. Vaswani","year":"2017","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017). Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, et al. (Eds.), Proceedings of the 31st international conference on neural information processing systems (pp. 5998\u20136008). Red Hook: Curran Associates."},{"key":"65_CR10","first-page":"4171","volume-title":"Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: human language technologies","author":"J. Devlin","year":"2019","unstructured":"Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 conference of the North American chapter of the Association for Computational Linguistics: human language technologies (pp. 4171\u20134186). Stroudsburg: ACL."},{"key":"65_CR11","unstructured":"Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. Retrieved November 11, 2024, from https:\/\/cdn.openai.com\/research-covers\/language-unsupervised\/language_understanding_paper.pdf."},{"key":"65_CR12","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. Retrieved November 11, 2024, from https:\/\/cdn.openai.com\/better-language-models\/language_models_are_unsupervised_multitask_learners.pdf."},{"key":"65_CR13","unstructured":"Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., et\u00a0al. (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint. arXiv:2204.05862."},{"key":"65_CR14","unstructured":"Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., et\u00a0al. (2022). OPT: open pre-trained transformer language models. arXiv preprint. arXiv:2205.01068."},{"key":"65_CR15","first-page":"1","volume":"24","author":"A. Chowdhery","year":"2023","unstructured":"Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., et al. (2023). PaLM: scaling language modeling with pathways. Journal of Machine Learning Research, 24, 1\u2013113.","journal-title":"Journal of Machine Learning Research"},{"key":"65_CR16","unstructured":"Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., et\u00a0al. (2021). ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint. arXiv:2107.02137."},{"key":"65_CR17","unstructured":"Wang, S., Sun, Y., Xiang, Y., Wu, Z., Ding, S., Gong, W., et\u00a0al. (2021). ERNIE 3.0 Titan: exploring larger-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint. arXiv:2112.12731."},{"key":"65_CR18","unstructured":"Lin, J., Men, R., Yang, A., Zhou, C., Ding, M., Zhang, Y., et\u00a0al. (2021). M6: a Chinese multimodal pretrainer. arXiv preprint. arXiv:2103.00823."},{"key":"65_CR19","unstructured":"Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., et\u00a0al. (2021). PanGu-\u03b1: large-scale autoregressive pretrained Chinese language models with auto-parallel computation. arXiv preprint. arXiv:2104.12369."},{"key":"65_CR20","unstructured":"Ren, X., Zhou, P., Meng, X., Huang, X., Wang, Y., Wang, W., et\u00a0al. (2023). Pangu-\u03a3: Towards trillion parameter language model with sparse heterogeneous computing. arXiv preprint. arXiv:2303.10845."},{"key":"65_CR21","first-page":"1932","volume-title":"Proceedings of the 38th AAAI conference on artificial intelligence","author":"Z. Gu","year":"2024","unstructured":"Gu, Z., Zhu, B., Zhu, G., Chen, Y., Tang, M., & Wang, J. (2024). AnomalyGPT: detecting industrial anomalies using large vision-language models. In M. J. Wooldridge, J. G. Dy, & S. Natarajan (Eds.), Proceedings of the 38th AAAI conference on artificial intelligence (pp. 1932\u20131940). Palo Alto: AAAI Press."},{"issue":"2","key":"65_CR22","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1145\/3624724","volume":"67","author":"M. Shanahan","year":"2024","unstructured":"Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68\u201379.","journal-title":"Communications of the ACM"},{"key":"65_CR23","unstructured":"Bahng, H., Jahanian, A., Sankaranarayanan, S., & Isola, P. (2022). Exploring visual prompts for adapting large-scale models. arXiv preprint. arXiv:2203.17274."},{"key":"65_CR24","unstructured":"Liu, W., & Zuo, Y. (2023). Stone needle: a general multimodal large-scale model framework towards healthcare. arXiv preprint. arXiv:2306.16034."},{"issue":"2","key":"65_CR25","doi-asserted-by":"publisher","DOI":"10.1002\/mef2.43","volume":"2","author":"D. Q. Wang","year":"2023","unstructured":"Wang, D. Q., Feng, L. Y., Ye, J. G., Zou, J. G., & Zheng, Y. F. (2023). Accelerating the integration of ChatGPT and other large-scale AI models into biomedical research and healthcare. MedComm\u2013Future Medicine, 2(2), e43.","journal-title":"MedComm\u2013Future Medicine"},{"key":"65_CR26","unstructured":"Yang, Y., Uy, M. C. S., & Huang, A. (2020). FinBERT: a pretrained language model for financial communications. arXiv preprint. arXiv:2006.08097."},{"key":"65_CR27","unstructured":"Zhang, H., Dereck, S. S., Wang, Z., Lv, X., Xu, K., Wu, L., et\u00a0al. (2023). Large scale foundation models for intelligent manufacturing applications: a survey. arXiv preprint. arXiv:2312.06718."},{"issue":"9","key":"65_CR28","doi-asserted-by":"publisher","first-page":"3718","DOI":"10.1109\/TITS.2019.2932038","volume":"21","author":"L. Zhou","year":"2019","unstructured":"Zhou, L., Zhang, S., Yu, J., & Chen, X. (2019). Spatial-temporal deep tensor neural networks for large-scale urban network speed prediction. IEEE Transactions on Intelligent Transportation Systems, 21(9), 3718\u20133729.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"issue":"9","key":"65_CR29","doi-asserted-by":"publisher","first-page":"13489","DOI":"10.1002\/er.6679","volume":"45","author":"J. Devaraj","year":"2021","unstructured":"Devaraj, J., Madurai Elavarasan, R., Shafiullah, G. M., Jamal, T., & Khan, I. (2021). A holistic review on energy forecasting using big data and deep learning models. International Journal of Energy Research, 45(9), 13489\u201313530.","journal-title":"International Journal of Energy Research"},{"key":"65_CR30","doi-asserted-by":"publisher","first-page":"3540","DOI":"10.18653\/v1\/D18-1390","volume-title":"Proceedings of the 2018 conference on empirical methods in natural language processing","author":"H. Zhong","year":"2018","unstructured":"Zhong, H., Guo, Z., Tu, C., Xiao, C., Liu, Z., & Sun, M. (2018). Legal judgment prediction via topological learning. In E. Riloff, D. Chiang, J. Hockenmaier, et al. (Eds.), Proceedings of the 2018 conference on empirical methods in natural language processing (pp. 3540\u20133549). Stroudsburg: ACL."},{"key":"65_CR31","doi-asserted-by":"publisher","DOI":"10.1155\/2022\/1549842","volume":"2022","author":"A. Abozeid","year":"2022","unstructured":"Abozeid, A., Alanazi, R., Elhadad, A., Taloba, A. I., & Abd El-Aziz, R. M. (2022). A large-scale dataset and deep learning model for detecting and counting olive trees in satellite imagery. Computational Intelligence and Neuroscience, 2022, 1549842.","journal-title":"Computational Intelligence and Neuroscience"},{"issue":"2","key":"65_CR32","doi-asserted-by":"publisher","first-page":"48","DOI":"10.1109\/MCI.2014.2307227","volume":"9","author":"E. Cambria","year":"2014","unstructured":"Cambria, E., & White, B. (2014). Jumping NLP curves: a review of natural language processing research. IEEE Computational Intelligence Magazine, 9(2), 48\u201357.","journal-title":"IEEE Computational Intelligence Magazine"},{"issue":"9","key":"65_CR33","first-page":"20","volume":"12","author":"K. Sharifani","year":"2022","unstructured":"Sharifani, K., Amini, M., Akbari, Y., & Godarzi, A. J. (2022). Operating machine learning across natural language processing techniques for improvement of fabricated news model. International Journal of Science and Information System Research, 12(9), 20\u201344.","journal-title":"International Journal of Science and Information System Research"},{"key":"65_CR34","first-page":"1529","volume-title":"Proceedings of the third international conference on intelligent communication technologies and virtual mobile networks","author":"T. P. Nagarhalli","year":"2021","unstructured":"Nagarhalli, T. P., Vaze, V., & Rana, N. K. (2021). Impact of machine learning in natural language processing: a review. In Proceedings of the third international conference on intelligent communication technologies and virtual mobile networks (pp. 1529\u20131534). Piscataway: IEEE."},{"key":"65_CR35","doi-asserted-by":"publisher","DOI":"10.1186\/s12874-021-01347-1","volume":"21","author":"C. J. Harrison","year":"2021","unstructured":"Harrison, C. J., & Sidey-Gibbons, C. J. (2021). Machine learning in medicine: a practical introduction to natural language processing. BMC Medical Research Methodology, 21, 158.","journal-title":"BMC Medical Research Methodology"},{"issue":"2","key":"65_CR36","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1561\/2200000096","volume":"16","author":"L. Wu","year":"2023","unstructured":"Wu, L., Chen, Y., Shen, K., Guo, X., Gao, H., Li, S., et al. (2023). Graph neural networks for natural language processing: a survey. Foundations and Trends in Machine Learning, 16(2), 119\u2013328.","journal-title":"Foundations and Trends in Machine Learning"},{"key":"65_CR37","doi-asserted-by":"publisher","first-page":"345","DOI":"10.1613\/jair.4992","volume":"57","author":"Y. Goldberg","year":"2016","unstructured":"Goldberg, Y. (2016). A primer on neural network models for natural language processing. Journal of Artificial Intelligence Research, 57, 345\u2013420.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"65_CR38","first-page":"64","volume-title":"Proceedings of the international conference on information systems and computer aided education","author":"W. Wang","year":"2018","unstructured":"Wang, W., & Gang, J. (2018). Application of convolutional neural network in natural language processing. In Proceedings of the international conference on information systems and computer aided education (pp. 64\u201370). Piscataway: IEEE."},{"key":"65_CR39","volume-title":"Introduction to natural language processing","author":"Q. Zhang","year":"2023","unstructured":"Zhang, Q., Gui, T., & Huang, X. (2023). Introduction to natural language processing. Beijing: Electronic Industry Press."},{"issue":"5","key":"65_CR40","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.48.291","volume":"48","author":"A. Fine","year":"1982","unstructured":"Fine, A. (1982). Hidden variables, joint probability, and the Bell inequalities. Physical Review Letters, 48(5), 291.","journal-title":"Physical Review Letters"},{"issue":"4","key":"65_CR41","doi-asserted-by":"publisher","first-page":"532","DOI":"10.1109\/PROC.1976.10159","volume":"64","author":"F. Jelinek","year":"1976","unstructured":"Jelinek, F. (1976). Continuous speech recognition by statistical methods. Proceedings of the IEEE, 64(4), 532\u2013556.","journal-title":"Proceedings of the IEEE"},{"key":"65_CR42","unstructured":"Yu, D., Seltzer, M. L., Li, J., Huang, J. T., & Seide, F. (2013). Feature learning in deep neural networks-studies on speech recognition tasks. arXiv preprint. arXiv:1301.3605."},{"issue":"4","key":"65_CR43","first-page":"467","volume":"18","author":"P. F. Brown","year":"1992","unstructured":"Brown, P. F., Della Pietra, V. J., Desouza, P. V., Lai, J. C., & Mercer, R. L. (1992). Class-based n-gram models of natural language. Computational Linguistics, 18(4), 467\u2013480.","journal-title":"Computational Linguistics"},{"key":"65_CR44","first-page":"258","volume-title":"Proceedings of the 49th annual meeting of the Association for Computational Linguistics: human language technologies","author":"A. Pauls","year":"2011","unstructured":"Pauls, A., & Klein, D. (2011). Faster and smaller n-gram language models. In D. Lin, Y. Matsumoto, & R. Mihalcea (Eds.), Proceedings of the 49th annual meeting of the Association for Computational Linguistics: human language technologies (pp. 258\u2013267). Stroudsburg: ACL."},{"key":"65_CR45","first-page":"758","volume-title":"Proceddings of the international work-conference on artificial neural networks","author":"M. Verleysen","year":"2005","unstructured":"Verleysen, M., & Fran\u00e7ois, D. (2005). The curse of dimensionality in data mining and time series prediction. In J. Cabestany, A. Prieto, & F. S. Hern\u00e1ndez (Eds.), Proceddings of the international work-conference on artificial neural networks (pp. 758\u2013770). Berlin: Springer."},{"issue":"4","key":"65_CR46","doi-asserted-by":"publisher","first-page":"724","DOI":"10.1109\/TASL.2008.2012323","volume":"17","author":"T. Hirsimaki","year":"2009","unstructured":"Hirsimaki, T., Pylkkonen, J., & Kurimo, M. (2009). Importance of high-order n-gram models in morph-based speech recognition. IEEE Transactions on Audio, Speech, and Language Processing, 17(4), 724\u2013732.","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"issue":"352","key":"65_CR47","doi-asserted-by":"publisher","DOI":"10.1136\/bmj.i1981","volume":"353","author":"S. Greenland","year":"2016","unstructured":"Greenland, S., Mansournia, M. A., & Altman, D. G. (2016). Sparse data bias: a problem hiding in plain sight. The BMJ, 353(352), i1981.","journal-title":"The BMJ"},{"issue":"1","key":"65_CR48","doi-asserted-by":"publisher","first-page":"93","DOI":"10.1016\/j.specom.2003.08.002","volume":"42","author":"J. R. Bellegarda","year":"2004","unstructured":"Bellegarda, J. R. (2004). Statistical language model adaptation: review and perspectives. Speech Communication, 42(1), 93\u2013108.","journal-title":"Speech Communication"},{"issue":"8","key":"65_CR49","doi-asserted-by":"publisher","first-page":"1270","DOI":"10.1109\/5.880083","volume":"88","author":"R. Rosenfeld","year":"2000","unstructured":"Rosenfeld, R. (2000). Two decades of statistical language modeling: where do we go from here? Proceedings of the IEEE, 88(8), 1270\u20131278.","journal-title":"Proceedings of the IEEE"},{"key":"65_CR50","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1007\/978-3-642-24797-2_4","volume-title":"Supervised sequence labelling with recurrent neural networks","author":"A. Graves","year":"2012","unstructured":"Graves, A., & Graves, A. (2012). Long short-term memory. In A. Graves (Ed.), Supervised sequence labelling with recurrent neural networks (pp. 37\u201345). Berlin: Springer."},{"key":"65_CR51","doi-asserted-by":"crossref","unstructured":"Pelevina, M., Arefyev, N., Biemann, C., & Panchenko, A. (2017). Making sense of word embeddings. arXiv preprint. arXiv:1708.03390.","DOI":"10.18653\/v1\/W16-1620"},{"issue":"8","key":"65_CR52","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S. Hochreiter","year":"1997","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"65_CR53","unstructured":"Yin, W., Kann, K., Yu, M., & Sch\u00fctze, H. (2017). Comparative study of CNN and RNN for natural language processing. arXiv preprint. arXiv:1702.01923."},{"key":"65_CR54","first-page":"248","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"J. Deng","year":"2009","unstructured":"Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Li, F.-F. (2009). ImageNet: a large-scale hierarchical image database. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 248\u2013255). Piscataway: IEEE."},{"key":"65_CR55","unstructured":"Hong, T., Kim, D., Ji, M., Hwang, W., Nam, D., & Park, S. (2020). BROS: a pre-trained language model for understanding texts in document. Retrieved Nomber 11, 2024, from https:\/\/openreview.net\/pdf?id=punMXQEsPr0."},{"key":"65_CR56","first-page":"2227","volume-title":"Proceedings of the 2018 conference of the North American chapter of the Association for Computational Linguistics: human language technologies","author":"M. E. Peters","year":"2018","unstructured":"Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., et al. (2018). Deep contextualized word representations. In M. A. Walker, H. Ji, & A. Stent (Eds.), Proceedings of the 2018 conference of the North American chapter of the Association for Computational Linguistics: human language technologies (pp. 2227\u20132237). Stroudsburg: ACL."},{"issue":"2","key":"65_CR57","doi-asserted-by":"publisher","first-page":"329","DOI":"10.1037\/0033-2909.88.2.329","volume":"88","author":"G. Felsten","year":"1980","unstructured":"Felsten, G., & Wasserman, G. S. (1980). Visual masking: mechanisms and theories. Psychological Bulletin, 88(2), 329.","journal-title":"Psychological Bulletin"},{"issue":"2","key":"65_CR58","doi-asserted-by":"publisher","DOI":"10.1002\/gamm.202100008","volume":"44","author":"L. Ruthotto","year":"2021","unstructured":"Ruthotto, L., & Haber, E. (2021). An introduction to deep generative modeling. GAMM-Mitteilungen, 44(2), e202100008.","journal-title":"GAMM-Mitteilungen"},{"key":"65_CR59","first-page":"1191","volume-title":"Proceedings of the 17th international conference on machine learning","author":"T. Zhang","year":"2000","unstructured":"Zhang, T., & Oles, F. (2000). The value of unlabeled data for classification problems. In P. Langley (Ed.), Proceedings of the 17th international conference on machine learning (pp. 1191\u20131198). San Francisco: Morgan Kaufmann Publishers."},{"key":"65_CR60","first-page":"5900","volume-title":"Proceedings of the IEEE international conference on acoustics, speech and signal processing","author":"Z. Tang","year":"2016","unstructured":"Tang, Z., Wang, D., & Zhang, Z. (2016). Recurrent neural network training with dark knowledge transfer. In Proceedings of the IEEE international conference on acoustics, speech and signal processing (pp. 5900\u20135904). Piscataway: IEEE."},{"key":"65_CR61","first-page":"1","volume":"21","author":"C. Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"65_CR62","first-page":"1","volume":"25","author":"H. W. Chung","year":"2024","unstructured":"Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., et al. (2024). Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25, 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"65_CR63","first-page":"320","volume-title":"Proceedings of the 60th annual meeting of the association for computational linguistics","author":"Z. Du","year":"2021","unstructured":"Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., et al. (2021). GLM: general language model pretraining with autoregressive blank infilling. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th annual meeting of the association for computational linguistics (pp. 320\u2013325). Stroudsburg: ACL."},{"key":"65_CR64","doi-asserted-by":"publisher","first-page":"4487","DOI":"10.18653\/v1\/P19-1441","volume-title":"Proceedings of the 57th annual meeting of the association for computational linguistics","author":"X. Liu","year":"2019","unstructured":"Liu, X., He, P., Chen, W., & Gao, J. (2019). Multi-task deep neural networks for natural language understanding. In A. Korhonen, D. R. Traum, & L. M\u00e0rquez (Eds.), Proceedings of the 57th annual meeting of the association for computational linguistics (pp. 4487\u20134496). Stroudsburg: ACL."},{"key":"65_CR65","unstructured":"Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., & Socher, R. (2019). CTRL: a conditional transformer language model for controllable generation. arXiv preprint. arXiv:1909.05858."},{"key":"65_CR66","unstructured":"Sun, F. K., & Lai, C. I. (2020). Conditioned natural language generation using only unconditioned language model: an exploration. arXiv preprint. arXiv:2011.07347."},{"key":"65_CR67","unstructured":"Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., et\u00a0al. (2023). PaLM 2 technical report. arXiv preprint. arXiv:2305.10403."},{"key":"65_CR68","first-page":"430","volume":"4","author":"P. Barham","year":"2022","unstructured":"Barham, P., Chowdhery, A., Dean, J., Ghemawat, S., Hand, S., Hurt, D., et al. (2022). Pathways: asynchronous distributed dataflow for ML. Proceedings of Machine Learning and Systems, 4, 430\u2013449.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"65_CR69","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., et\u00a0al. (2023). LLaMA: open and efficient foundation language models. arXiv preprint. arXiv:2302.13971."},{"key":"65_CR70","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.127063","volume":"568","author":"J. Su","year":"2024","unstructured":"Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., & Liu, Y. (2024). Roformer: enhanced transformer with rotary position embedding. Neurocomputing, 568, 127063.","journal-title":"Neurocomputing"},{"key":"65_CR71","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., et\u00a0al. (2023). LLaMA 2: open foundation and fine-tuned chat models. arXiv preprint. arXiv:2307.09288."},{"key":"65_CR72","first-page":"699","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Chen","year":"2020","unstructured":"Chen, T., Liu, S., Chang, S., Cheng, Y., Amini, L., & Wang, Z. (2020). Adversarial robustness: from self-supervised pre-training to fine-tuning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 699\u2013708). Piscataway: IEEE."},{"issue":"1","key":"65_CR73","doi-asserted-by":"publisher","first-page":"116","DOI":"10.4156\/jcit.vol5.issue1.13","volume":"5","author":"Y. Tao","year":"2010","unstructured":"Tao, Y., Xia, Y., Xu, T., & Chi, X. (2010). Research progress of the scale invariant feature transform (SIFT) descriptors. Journal of Convergence Information Technology, 5(1), 116\u2013121.","journal-title":"Journal of Convergence Information Technology"},{"key":"65_CR74","first-page":"172","volume-title":"Proceedings of the the 5th international conference on business and industrial research","author":"T. Surasak","year":"2018","unstructured":"Surasak, T., Takahiro, I., Cheng, C., Wang, C., & Sheng, P. (2018). Histogram of oriented gradients for human detection in video. In Proceedings of the the 5th international conference on business and industrial research (pp. 172\u2013176). Piscataway: IEEE."},{"issue":"5","key":"65_CR75","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s41870-017-0080-1","volume":"13","author":"M. A. Chandra","year":"2021","unstructured":"Chandra, M. A., & Bedi, S. S. (2021). Survey on SVM and their application in image classification. International Journal of Information Technology, 13(5), 1\u201311.","journal-title":"International Journal of Information Technology"},{"issue":"2","key":"65_CR76","first-page":"130","volume":"27","author":"Y. Y. Song","year":"2015","unstructured":"Song, Y. Y., & Ying, L. U. (2015). Decision tree methods: applications for classification and prediction. Shanghai Archives of Psychiatry, 27(2), 130.","journal-title":"Shanghai Archives of Psychiatry"},{"issue":"6","key":"65_CR77","doi-asserted-by":"publisher","first-page":"745","DOI":"10.1016\/S1874-1029(13)60052-X","volume":"39","author":"Y. Cao","year":"2013","unstructured":"Cao, Y., Miao, Q., Liu, J., & Gao, L. (2013). Advance and prospects of AdaBoost algorithm. Acta Automatica Sinica, 39(6), 745\u2013758.","journal-title":"Acta Automatica Sinica"},{"key":"65_CR78","first-page":"1106","volume-title":"Proceedings of the 26th international conference on neural information processing systems","author":"A. Krizhevsky","year":"2012","unstructured":"Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. In P. L. Bartlett, F. C. N. Pereira, C. J. C. Burges, et al. (Eds.), Proceedings of the 26th international conference on neural information processing systems (pp. 1106\u20131114). Red Hook: Curran Associates."},{"key":"65_CR79","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint. arXiv:1409.1556."},{"key":"65_CR80","first-page":"1","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"C. Szegedy","year":"2015","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., et al. (2015). Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1\u20139). Piscataway: IEEE."},{"key":"65_CR81","first-page":"770","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"K. He","year":"2016","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770\u2013778). Piscataway: IEEE."},{"key":"65_CR82","unstructured":"Chen, X., Fang, H., Lin, T. Y., Vedantam, R., Gupta, S., Doll\u00e1r, P., et\u00a0al. (2015). Microsoft COCO captions: data collection and evaluation server. arXiv preprint. arXiv:1504.00325."},{"key":"65_CR83","first-page":"1","volume-title":"Proeedings of the 9th international conference on learning representations","author":"A. Dosovitskiy","year":"2020","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2020). An image is worth 16 \u00d7 16 words: transformers for image recognition at scale. In Proeedings of the 9th international conference on learning representations (pp. 1\u201321). Retrieved November 14, 2024, from https:\/\/openreview.net\/forum?id=YicbFdNTTy."},{"key":"65_CR84","first-page":"1691","volume-title":"Proceedings of the 37th international conference on machine learning","author":"M. Chen","year":"2020","unstructured":"Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., et al. (2020). Generative pretraining from pixels. In Proceedings of the 37th international conference on machine learning (pp. 1691\u20131703). Retrieved November 14, 2024, from http:\/\/proceedings.mlr.press\/v119\/chen20s.html."},{"key":"65_CR85","first-page":"491","volume-title":"Proceedings of the 16th European conference on computer vision","author":"A. Kolesnikov","year":"2020","unstructured":"Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., et al. (2020). Big transfer (bit): general visual representation learning. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 491\u2013507). Cham: Springer."},{"key":"65_CR86","first-page":"15908","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"K. Han","year":"2021","unstructured":"Han, K., Xiao, A., Wu, E., Guo, J., Xu, C., & Wang, Y. (2021). Transformer in transformer. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, et al. (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 15908\u201315919). Red Hook: Curran Associates."},{"key":"65_CR87","first-page":"213","volume-title":"Proceedings of the 16th European conference on computer vision","author":"N. Carion","year":"2020","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 213\u2013229). Cham: Springer."},{"key":"65_CR88","first-page":"1601","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Dai","year":"2021","unstructured":"Dai, Z., Cai, B., Lin, Y., & Chen, J. (2021). UP-DETR: unsupervised pre-training for object detection with transformers. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 1601\u20131610). Piscataway: IEEE."},{"key":"65_CR89","first-page":"1","volume-title":"Proceedings of the 32nd British machine vision conference","author":"M. Zheng","year":"2021","unstructured":"Zheng, M., Gao, P., Zhang, R., Li, K., Wang, X., Li, H., et al. (2021). End-to-end object detection with adaptive clustering transformer. In Proceedings of the 32nd British machine vision conference (pp. 1\u201314). Swansea: BMVA Press."},{"key":"65_CR90","first-page":"10337","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Chen","year":"2020","unstructured":"Chen, Y., Cao, Y., Hu, H., & Wang, L. (2020). Memory enhanced global-local aggregation for video object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10337\u201310346). Piscataway: IEEE."},{"key":"65_CR91","first-page":"528","volume-title":"Proceedings of the 16th European conference on computer vision","author":"Y. Zeng","year":"2020","unstructured":"Zeng, Y., Fu, J., & Chao, H. (2020). Learning joint spatial-temporal transformations for video inpainting. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 528\u2013543). Cham: Springer."},{"key":"65_CR92","first-page":"813","volume-title":"Proceedings of the 38th international conference on machine learning","author":"G. Bertasius","year":"2021","unstructured":"Bertasius, G., Wang, H., & Torresani, L. (2021). Is space-time attention all you need for video understanding? In Proceedings of the 38th international conference on machine learning (pp. 813\u2013824). Retrieved November 13, 2024, from http:\/\/proceedings.mlr.press\/v139\/bertasius21a.html."},{"key":"65_CR93","volume-title":"Parallel computing works!","author":"G. C. Fox","year":"1994","unstructured":"Fox, G. C., Williams, R. D., & Messina, G. C. (1994). Parallel computing works! Amsterdam: Elsevier."},{"key":"65_CR94","first-page":"1","volume-title":"Proceedings of the NAACL-HLT 2012 workshop: will we ever really replace the n-gram model? On the future of language modeling for HLT","author":"H. S. Le","year":"2012","unstructured":"Le, H. S., Allauzen, A., & Yvon, F. (2012). Measuring the influence of long range dependencies with neural network language models. In Proceedings of the NAACL-HLT 2012 workshop: will we ever really replace the n-gram model? On the future of language modeling for HLT (pp. 1\u201310). Stroudsburg: ACL."},{"key":"65_CR95","first-page":"16000","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"K. He","year":"2022","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., & Girshick, R. (2022). Masked autoencoders are scalable vision learners. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 16000\u201316009). Piscataway: IEEE."},{"key":"65_CR96","first-page":"9640","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"X. Chen","year":"2021","unstructured":"Chen, X., Xie, S., & He, K. (2021). An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 9640\u20139649). Piscataway: IEEE."},{"key":"65_CR97","first-page":"15637","volume-title":"Proceedings of the 33th international conference on neural information processing systems","author":"D. Hendrycks","year":"2019","unstructured":"Hendrycks, D., Mazeika, M., Kadavath, S., & Song, D. (2019). Using self-supervised learning can improve model robustness and uncertainty. In H. Wallach, H. Larochelle, A. Beygelzimer, et al. (Eds.), Proceedings of the 33th international conference on neural information processing systems (pp. 15637\u201315648). Red Hook: Curran Associates."},{"key":"65_CR98","first-page":"22243","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"T. Chen","year":"2020","unstructured":"Chen, T., Kornblith, S., Swersky, K., Norouzi, M., & Hinton, G. E. (2020). Big self-supervised models are strong semi-supervised learners. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 22243\u201322255). Red Hook: Curran Associates."},{"key":"65_CR99","first-page":"65743","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"N. Parthasarathy","year":"2023","unstructured":"Parthasarathy, N., Eslami, S. M., Carreira, J., & Henaff, O. (2023). Self-supervised video pretraining yields robust and more human-aligned visual representations. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 65743\u201365765). Red Hook: Curran Associates."},{"key":"65_CR100","first-page":"24206","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"H. Akbari","year":"2021","unstructured":"Akbari, H., Yuan, L., Qian, R., Chuang, W. H., Chang, S. F., Cui, Y., et al. (2021). VATT: transformers for multimodal self-supervised learning from raw video, audio and text. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, et al. (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 24206\u201324221). Red Hook: Curran Associates."},{"key":"65_CR101","doi-asserted-by":"publisher","first-page":"47646","DOI":"10.1109\/ACCESS.2024.3383047","volume":"12","author":"R. B. Figueiredo","year":"2024","unstructured":"Figueiredo, R. B., & Mendes, H. A. (2024). Analyzing information leakage on video object detection datasets by splitting images into clusters with high spatiotemporal correlation. IEEE Access, 12, 47646\u201347655.","journal-title":"IEEE Access"},{"key":"65_CR102","first-page":"14549","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"L. Wang","year":"2023","unstructured":"Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., et al. (2023). Videomae v2: scaling video masked autoencoders with dual masking. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 14549\u201314560). Piscataway: IEEE."},{"key":"65_CR103","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"65_CR104","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"X. Zou","year":"2023","unstructured":"Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J., et al. (2023). Segment everything everywhere all at once. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201313). Red Hook: Curran Associates."},{"key":"65_CR105","first-page":"4077","volume-title":"Proceedings of the 31st international conference on neural information processing systems","author":"J. Snell","year":"2017","unstructured":"Snell, J., Swersky, K., & Zemel, R. (2017). Prototypical networks for few-shot learning. In I. Guyon, U. von Luxburg, S. Bengio, et al. (Eds.), Proceedings of the 31st international conference on neural information processing systems (pp. 4077\u20134087). Red Hook: Curran Associates."},{"key":"65_CR106","first-page":"472","volume-title":"Proceedings of the 17th European conference on computer vision","author":"T. Xi","year":"2022","unstructured":"Xi, T., Sun, Y., Yu, D., Li, B., Peng, N., Zhang, G., et al. (2022). UFO: unified feature optimization. In S. Avidan, G. J. Brostow, M. Ciss\u00e9, et al. (Eds.), Proceedings of the 17th European conference on computer vision (pp. 472\u2013488). Cham: Springer."},{"key":"65_CR107","unstructured":"Shao, J., Chen, S., Li, Y., Wang, K., Yin, Z., He, Y., et\u00a0al. (2021). Intern: a new learning paradigm towards general vision. arXiv preprint. arXiv:2111.08687."},{"key":"65_CR108","first-page":"5198","volume-title":"Proceedings of the 32nd AAAI conference on artificial intelligence","author":"D. Kiela","year":"2018","unstructured":"Kiela, D., Grave, E., Joulin, A., & Mikolov, T. (2018). Efficient large-scale multi-modal classification. In S. A. McIlraith & K. Q. Weinberger (Eds.), Proceedings of the 32nd AAAI conference on artificial intelligence (pp. 5198\u20135204). Palo Alto: AAAI Press."},{"key":"65_CR109","doi-asserted-by":"publisher","first-page":"376","DOI":"10.1016\/j.patcog.2018.08.007","volume":"86","author":"H. Chen","year":"2019","unstructured":"Chen, H., Li, Y., & Su, D. (2019). Multi-modal fusion network with multi-scale multi-path and cross-modal interactions for RGB-D salient object detection. Pattern Recognition, 86, 376\u2013385.","journal-title":"Pattern Recognition"},{"key":"65_CR110","first-page":"2425","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"S. Antol","year":"2015","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., et al. (2015). VQA: visual question answering. In Proceedings of the IEEE international conference on computer vision (pp. 2425\u20132433). Piscataway: IEEE."},{"issue":"9","key":"65_CR111","doi-asserted-by":"publisher","first-page":"10850","DOI":"10.1109\/TPAMI.2023.3261988","volume":"45","author":"F. A. Croitoru","year":"2023","unstructured":"Croitoru, F. A., Hondru, V., Ionescu, R. T., & Shah, M. (2023). Diffusion models in vision: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 10850\u201310869.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"1","key":"65_CR112","doi-asserted-by":"publisher","first-page":"53","DOI":"10.1109\/MSP.2017.2765202","volume":"35","author":"A. Creswell","year":"2018","unstructured":"Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., & Bharath, A. A. (2018). Generative adversarial networks: an overview. IEEE Signal Processing Magazine, 35(1), 53\u201365.","journal-title":"IEEE Signal Processing Magazine"},{"key":"65_CR113","first-page":"16784","volume-title":"Proceedings of the international conference on machine learning","author":"A. Nichol","year":"2021","unstructured":"Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., et al. (2021). GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In K. Chaudhuri, S. Jegelka, L. Song, et al. (Eds.), Proceedings of the international conference on machine learning (pp. 16784\u201316804). Retrieved November 14, 2024, from https:\/\/proceedings.mlr.press\/v162\/nichol22a.html."},{"key":"65_CR114","first-page":"8748","volume-title":"Proceedings of the 38th international conference on machine learning","author":"A. Radford","year":"2021","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al. (2021). Learning transferable visual models from natural language supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th international conference on machine learning (pp. 8748\u20138763). Retrieved November 14, 2024, from http:\/\/proceedings.mlr.press\/v139\/radford21a.html."},{"key":"65_CR115","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv preprint. arXiv:2204.06125."},{"key":"65_CR116","first-page":"10684","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Rombach","year":"2022","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10684\u201310695). Piscataway: IEEE."},{"key":"65_CR117","first-page":"6840","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"J. Ho","year":"2020","unstructured":"Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 6840\u20136851). Red Hook: Curran Associates."},{"key":"65_CR118","first-page":"36479","volume-title":"Proceedings of the 36th international conference on neural information processing systems","author":"C. Saharia","year":"2022","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., et al. (2022). Photorealistic text-to-image diffusion models with deep language understanding. In S. Koyejo, S. Mohamed, A. Agarwal, et al. (Eds.), Proceedings of the 36th international conference on neural information processing systems (pp. 36479\u201336494). Red Hook: Curran Associates."},{"issue":"6","key":"65_CR119","first-page":"3666","volume":"6","author":"M. A. Awel","year":"2019","unstructured":"Awel, M. A., & Abidi, A. I. (2019). Review on optical character recognition. International Journal of Research in Engineering and Technology, 6(6), 3666\u20133669.","journal-title":"International Journal of Research in Engineering and Technology"},{"key":"65_CR120","unstructured":"Wang, J., Yang, Z., Hu, X., Li, L., Lin, K., Gan, Z., et\u00a0al. (2022). GIT: a generative image-to-text transformer for vision and language. arXiv preprint. arXiv:2205.14100."},{"key":"65_CR121","unstructured":"Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., et\u00a0al. (2023). CogVLM: visual expert for large language models. arXiv preprint. arXiv:2311.03079."},{"key":"65_CR122","doi-asserted-by":"publisher","first-page":"30","DOI":"10.18653\/v1\/2022.emnlp-main.3","volume-title":"Proceedings of the 2022 conference on empirical methods in natural language processing","author":"M. Geva","year":"2022","unstructured":"Geva, M., Caciularu, A., Wang, K. R., & Goldberg, Y. (2022). Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Y. Goldberg, Z. Kozareva, & Y. Zhang (Eds.), Proceedings of the 2022 conference on empirical methods in natural language processing (pp. 30\u201345). Stroudsburg: ACL."},{"key":"65_CR123","unstructured":"He, X., Wei, L., Xie, L., & Tian, Q. (2024). Incorporating visual experts to resolve the information loss in multimodal large language models. arXiv preprint. arXiv:2401.03105."},{"key":"65_CR124","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., et\u00a0al. (2023). GPT-4 technical report. arXiv preprint. arXiv:2303.08774."},{"key":"65_CR125","first-page":"1","volume-title":"Proceedings of the 12th international conference on learning representations","author":"D. Zhu","year":"2024","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2024). MiniGPT-4: enhancing vision-language understanding with advanced large language models. In Proceedings of the 12th international conference on learning representations (pp. 1\u201317). Retrieved November 14, 2024, from https:\/\/openreview.net\/forum?id=1tZbq88f27."},{"key":"65_CR126","unstructured":"Wang, M., Xing, J., & Liu, Y. (2021). ActionCLIP: a new paradigm for video action recognition. arXiv preprint. arXiv:2109.08472."},{"key":"65_CR127","first-page":"4904","volume-title":"Proceedings of the 38th international conference on machine learning","author":"C. Jia","year":"2021","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y. T., Parekh, Z., Pham, H., et al. (2021). Scaling up visual and vision-language representation learning with noisy text supervision. In M. Meila & T. Zhang (Eds.), Proceedings of the 38th international conference on machine learning (pp. 4904\u20134916). Retrieved November 14, 2024, from http:\/\/proceedings.mlr.press\/v139\/jia21b.html."},{"key":"65_CR128","first-page":"12888","volume-title":"International conference on machine learning","author":"J. Li","year":"2022","unstructured":"Li, J., Li, D., Xiong, C., & Hoi, S. (2022). BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In K. Chaudhuri, S. Jegelka, L. Song, et al. (Eds.). International conference on machine learning (pp. 12888\u201312900). Retrieved November 14, 2024, from https:\/\/proceedings.mlr.press\/v162\/li22n.html."},{"key":"65_CR129","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., et\u00a0al. (2022). Image as a foreign language: BEiT pretraining for all vision and vision-language tasks. arXiv preprint. arXiv:2208.10442.","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"65_CR130","first-page":"1","volume-title":"Proceedings of the 12th international conference on learning representations","author":"X. Chen","year":"2022","unstructured":"Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A. J., Padlewski, P., Salz, D., et al. (2022). PaLI: a jointly-scaled multilingual language-image model. In Proceedings of the 12th international conference on learning representations. (pp. 1\u201333). Retrieved November 12, 2024, from https:\/\/openreview.net\/forum?id=mWVoBz4W0u."},{"key":"65_CR131","first-page":"7480","volume-title":"Proceedings of the international conference on machine learning","author":"M. Dehghani","year":"2023","unstructured":"Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., et al. (2023). Scaling vision transformers to 22 billion parameters. In A. Krause, E. Brunskill, K. Cho, et al. (Eds.), Proceedings of the international conference on machine learning (pp. 7480\u20137512). Retrieved November 14, 2024, from https:\/\/proceedings.mlr.press\/v202\/dehghani23a.html."},{"key":"65_CR132","first-page":"8469","volume-title":"Proceedings of the international conference on machine learning","author":"D. Driess","year":"2023","unstructured":"Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., et al. (2023). PaLM-E: an embodied multimodal language model. In A. Krause, E. Brunskill, K. Cho, et al. (Eds.), Proceedings of the international conference on machine learning (pp. 8469\u20138488). Retrieved November 14, 2024, from https:\/\/proceedings.mlr.press\/v202\/driess23a.html."},{"key":"65_CR133","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"H. Liu","year":"2023","unstructured":"Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201325). Red Hook: Curran Associates."},{"key":"65_CR134","first-page":"26296","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"H. Liu","year":"2024","unstructured":"Liu, H., Li, C., Li, Y., & Lee, Y. J. (2024). Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 26296\u201326306). Piscataway: IEEE."},{"key":"65_CR135","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"C. Li","year":"2023","unstructured":"Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., et al. (2023). LLaVA-Med: training a large language-and-vision assistant for biomedicine in one day. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201317). Red Hook: Curran Associates."},{"key":"65_CR136","unstructured":"Dai, W., Li, J., Li, D., Tiong, A. M., Zhao, J., Wang, W., et\u00a0al. (2023). InstructBLIP: towards general-purpose vision-language models with instruction tuning (pp. 1\u201317). Red Hook: Curran Associates."},{"key":"65_CR137","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., et\u00a0al. (2023). Qwen-VL: a frontier large vision-language model with versatile abilities. arXiv preprint. arXiv:2308.12966."},{"key":"65_CR138","first-page":"24185","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Chen","year":"2023","unstructured":"Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., et al. (2023). Intern VL: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 24185\u201324198). Piscataway: IEEE."},{"key":"65_CR139","first-page":"1","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"Y. Xia","year":"2023","unstructured":"Xia, Y., Huang, H., Zhu, J., & Zhao, Z. (2023). Achieving cross modal generalization with multimodal unified representation. In A. Oh, T. Neumann, A. Globerson, et al. (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 1\u201315). Red Hook: Curran Associates."},{"key":"65_CR140","doi-asserted-by":"publisher","first-page":"59800","DOI":"10.1109\/ACCESS.2021.3070212","volume":"9","author":"G. Joshi","year":"2021","unstructured":"Joshi, G., Walambe, R., & Kotecha, K. (2021). A review on explainability in multimodal deep neural nets. IEEE Access, 9, 59800\u201359821.","journal-title":"IEEE Access"},{"key":"65_CR141","first-page":"223","volume-title":"Proceedings of the IEEE Asia Pacific conference on circuits and systems","author":"M. Wang","year":"2018","unstructured":"Wang, M., Lu, S., Zhu, D., Lin, J., & Wang, Z. (2018). A high-speed and low-complexity architecture for softmax function in deep learning. In Proceedings of the IEEE Asia Pacific conference on circuits and systems (pp. 223\u2013226). Piscataway: IEEE."},{"key":"65_CR142","first-page":"580","volume-title":"Proceedings of the 32nd international conference on neural information processing systems","author":"B. Hanin","year":"2018","unstructured":"Hanin, B. (2018). Which neural net architectures give rise to exploding and vanishing gradients? In S. Bengio, H. Wallach, H. Larochelle, et al. (Eds.), Proceedings of the 32nd international conference on neural information processing systems (pp. 580\u2013589). Red Hook: Curran Associates."},{"key":"65_CR143","unstructured":"Han, Z., Gao, C., Liu, J., & Zhang, S. Q. (2024). Parameter-efficient fine-tuning for large AI models: a comprehensive survey. arXiv preprint. arXiv:2403.14608."},{"issue":"2","key":"65_CR144","doi-asserted-by":"publisher","first-page":"3153","DOI":"10.1109\/LRA.2020.2974682","volume":"5","author":"A. Loquercio","year":"2020","unstructured":"Loquercio, A., Segu, M., & Scaramuzza, D. (2020). A general framework for uncertainty estimation in deep learning. IEEE Robotics and Automation Letters, 5(2), 3153\u20133160.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"65_CR145","first-page":"4582","volume-title":"Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing","author":"X. L. Li","year":"2021","unstructured":"Li, X. L., & Liang, P. (2021). Prefix-tuning: optimizing continuous prompts for generation. In C. Zong, F. Xia, W. Li, et al. (Eds.), Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing (pp. 4582\u20134597). Stroudsburg: ACL."},{"key":"65_CR146","doi-asserted-by":"crossref","unstructured":"Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., et\u00a0al. (2021). P-tuning v2: prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint. arXiv:2110.07602.","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"65_CR147","doi-asserted-by":"publisher","first-page":"3505","DOI":"10.1145\/3394486.3406703","volume-title":"Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining","author":"J. Rasley","year":"2020","unstructured":"Rasley, J., Rajbhandari, S., Ruwase, O., & He, Y. (2020). DeepSpeed: system optimizations enable training deep learning models with over 100 billion parameters. In R. Gupta, Y. Liu, J. Tang, et al. (Eds.), Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining (pp. 3505\u20133506). New York: ACM."},{"issue":"5\u20136","key":"65_CR148","doi-asserted-by":"publisher","first-page":"781","DOI":"10.1016\/j.neunet.2005.06.003","volume":"18","author":"T. Taskaya-Temizel","year":"2005","unstructured":"Taskaya-Temizel, T., & Casey, M. C. (2005). A comparative study of autoregressive neural network hybrids. Neural Networks, 18(5\u20136), 781\u2013789.","journal-title":"Neural Networks"},{"issue":"12","key":"65_CR149","doi-asserted-by":"publisher","first-page":"5586","DOI":"10.1109\/TKDE.2021.3070203","volume":"34","author":"Y. Zhang","year":"2021","unstructured":"Zhang, Y., & Yang, Q. (2021). A survey on multi-task learning. IEEE Transactions on Knowledge and Data Engineering, 34(12), 5586\u20135609.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"65_CR150","first-page":"11272","volume-title":"Proceedings of the 61st annual meeting of the Association for Computational Linguistics","author":"H. Ivison","year":"2023","unstructured":"Ivison, H., Bhagia, A., Wang, Y., Hajishirzi, H., & Peters, M. (2023). HINT: hypernetwork instruction tuning for efficient zero-& few-shot generalisation. In A. Rogers, J. L. Boyd-Graber, & N. Okazaki (Eds.), Proceedings of the 61st annual meeting of the Association for Computational Linguistics (pp. 11272\u201311288). Stroudsburg: ACL."},{"key":"65_CR151","first-page":"8187","volume-title":"Proceedings of the 62nd annual meeting of the Association for Computational Linguistics","author":"K. Lv","year":"2024","unstructured":"Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., & Qiu, X. (2024). Full parameter fine-tuning for large language models with limited resources. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd annual meeting of the Association for Computational Linguistics (pp. 8187\u20138198). Stroudsburg: ACL."},{"key":"65_CR152","unstructured":"Shen, L., Sun, Y., Yu, Z., Ding, L., Tian, X., & Tao, D. (2023). On efficient training of large-scale deep learning models: a literature review. arXiv preprint. arXiv:2304.03589."},{"key":"65_CR153","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1145\/3357223.3362707","volume-title":"Proceedings of the ACM symposium on cloud computing","author":"J. J. Dai","year":"2019","unstructured":"Dai, J. J., Wang, Y., Qiu, X., Ding, D., Zhang, Y., Wang, Y., et al. (2019). BigDL: a distributed deep learning framework for big data. In Proceedings of the ACM symposium on cloud computing (pp. 50\u201360). New York: ACM."},{"key":"65_CR154","doi-asserted-by":"publisher","DOI":"10.1016\/j.physd.2019.132306","volume":"404","author":"A. Sherstinsky","year":"2020","unstructured":"Sherstinsky, A. (2020). Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network. Physica D. Nonlinear Phenomena, 404, 132306.","journal-title":"Physica D. Nonlinear Phenomena"},{"issue":"2","key":"65_CR155","doi-asserted-by":"publisher","first-page":"227","DOI":"10.3102\/1076998619872761","volume":"45","author":"B. Pang","year":"2020","unstructured":"Pang, B., Nijkamp, E., & Wu, Y. N. (2020). Deep learning with tensorflow: a review. Journal of Educational and Behavioral Statistics, 45(2), 227\u2013248.","journal-title":"Journal of Educational and Behavioral Statistics"},{"key":"65_CR156","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-5364-9","volume-title":"Deep learning with Python: learn best practices of deep learning models with PyTorch","author":"N. Ketkar","year":"2021","unstructured":"Ketkar, N., & Moolayil, J. (2021). Deep learning with Python: learn best practices of deep learning models with PyTorch. Berkeley: Apress."},{"key":"65_CR157","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-2766-4","volume-title":"Deep learning with Python: a hands-on introduction","author":"N. Ketkar","year":"2017","unstructured":"Ketkar, N. (2017). Deep learning with Python: a hands-on introduction. Berkeley: Apress."},{"issue":"1","key":"65_CR158","first-page":"4","volume":"9","author":"J. Richardson","year":"2013","unstructured":"Richardson, J., McLeod, S., Flora, K., Sauers, N., Kannan, S., & Sincar, M. (2013). Large-scale 1: 1 computing initiatives: an open access database. International Journal of Education and Development Using ICT, 9(1), 4\u201318.","journal-title":"International Journal of Education and Development Using ICT"},{"key":"65_CR159","first-page":"1","volume-title":"Proceedings of the IEEE 6th international conference on biometrics: theory, applications and systems","author":"A. Sapkota","year":"2013","unstructured":"Sapkota, A., & Boult, T. E. (2013). Large scale unconstrained open set face database. In Proceedings of the IEEE 6th international conference on biometrics: theory, applications and systems (pp. 1\u20138). Piscataway: IEEE."},{"key":"65_CR160","doi-asserted-by":"publisher","first-page":"921","DOI":"10.1093\/nar\/gku955","volume":"43","author":"Y. Igarashi","year":"2015","unstructured":"Igarashi, Y., Nakatsu, N., Yamashita, T., Ono, A., Ohno, Y., Urushidani, T., et al. (2015). Open TG-GATEs: a large-scale toxicogenomics database. Nucleic Acids Research, 43, 921\u2013927.","journal-title":"Nucleic Acids Research"},{"key":"65_CR161","unstructured":"Bandy, J., & Vincent, N. (2021). Addressing\u201d documentation debt\u201d in machine learning research: a retrospective datasheet for bookcorpus. arXiv preprint. arXiv:2105.05241."},{"issue":"1","key":"65_CR162","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1109\/MSP.2020.2984428","volume":"38","author":"D. Deter","year":"2020","unstructured":"Deter, D., Wang, C., Cook, A., & Perry, N. K. (2020). Simulating the autonomous future: a look at virtual vehicle environments and how to validate simulation using public data sets. IEEE Signal Processing Magazine, 38(1), 111\u2013121.","journal-title":"IEEE Signal Processing Magazine"},{"key":"65_CR163","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1145\/3548785.3548793","volume-title":"Proceedings of the 26th international database engineered applications symposium","author":"M. Endres","year":"2022","unstructured":"Endres, M., Mannarapotta Venugopal, A., & Tran, T. S. (2022). Synthetic data generation: a comparative study. In B. C. Desai & P. Z. Revesz (Eds.), Proceedings of the 26th international database engineered applications symposium (pp. 94\u2013102). New York: ACM."},{"key":"65_CR164","doi-asserted-by":"publisher","first-page":"387","DOI":"10.1007\/978-3-319-98812-2_35","volume-title":"Proceedings of the 29th international conference on database and expert systems applications","author":"A. Dandekar","year":"2018","unstructured":"Dandekar, A., Zen, R. A., & Bressan, S. (2018). A comparative study of synthetic dataset generation techniques. In S. Hartmann, H. Ma, A. Hameurlain, et al. (Eds.), Proceedings of the 29th international conference on database and expert systems applications (pp. 387\u2013395). Cham: Springer."},{"issue":"6","key":"65_CR165","doi-asserted-by":"publisher","first-page":"5553","DOI":"10.1109\/TDSC.2024.3379434","volume":"21","author":"C. Gao","year":"2024","unstructured":"Gao, C., Jajodia, S., Pugliese, A., & Subrahmanian, V. S. (2024). FakeDB: generating fake synthetic databases. IEEE Transactions on Dependable and Secure Computing, 21(6), 5553\u20135564.","journal-title":"IEEE Transactions on Dependable and Secure Computing"},{"key":"65_CR166","first-page":"16784","volume-title":"Proceedings of the international conference on machine learning","author":"A. Q. Nichol","year":"2024","unstructured":"Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., et al. (2024). GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In K. Chaudhuri, S. Jegelka, L. Song, et al. (Eds.), Proceedings of the international conference on machine learning (pp. 16784\u201316804). Retrieved November 12, 2024, from https:\/\/proceedings.mlr.press\/v162\/nichol22a.html."},{"key":"65_CR167","unstructured":"Sharir, O., Peleg, B., & Shoham, Y. (2020). The cost of training NLP models: a concise overview. arXiv preprint. arXiv:2004.08900."},{"issue":"6","key":"65_CR168","first-page":"1","volume":"30","author":"X. Chen","year":"2021","unstructured":"Chen, X., Wang, G., Zhou, Y., & Li, R. (2021). The high-quality development of China\u2019s publication printing industry from an environmental perspective. Journal of Global Information Management, 30(6), 1\u201318.","journal-title":"Journal of Global Information Management"},{"key":"65_CR169","unstructured":"Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B. (2019). Megatron-LM: training multi-billion parameter language models using model parallelism. arXiv preprint. arXiv:1909.08053."},{"key":"65_CR170","unstructured":"Wang, G., Qin, H., Jacobs, S. A., Holmes, C., Rajbhandari, S., Ruwase, O., et\u00a0al. (2023). Zero++: extremely efficient collective communication for giant model training. arXiv preprint. arXiv:2306.10209."},{"key":"65_CR171","first-page":"551","volume-title":"Proceedings of the 2021 USENIX annual technical conference","author":"J. Ren","year":"2021","unstructured":"Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., et al. (2021). Zero-offload: democratizing billion-scale model training. In I. Calciu & G. Kuenning (Eds.), Proceedings of the 2021 USENIX annual technical conference (pp. 551\u2013564). Berkeley: USENIX Association."},{"key":"65_CR172","unstructured":"Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., et\u00a0al. (2022). Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint. arXiv:2201.11990."},{"key":"65_CR173","unstructured":"Le Scao, T., Fan, A., Akiki, C., Pavlick, E., Ili\u0107, S., Hesslow, D., et\u00a0al. (2023). Bloom: a 176b-parameter open-access multilingual language model. arXiv preprint. arXiv:2211.05100."},{"key":"65_CR174","first-page":"1","volume-title":"Proceedings of the 13th EuroSys conference","author":"Y. Peng","year":"2018","unstructured":"Peng, Y., Bao, Y., Chen, Y., Wu, C., & Guo, C. (2018). Optimus: an efficient dynamic resource scheduler for deep learning clusters. In R. Oliveira, P. Felber, & Y. C. Hu (Eds.), Proceedings of the 13th EuroSys conference (pp. 1\u201314). New York: ACM."},{"issue":"9","key":"65_CR175","doi-asserted-by":"publisher","first-page":"514","DOI":"10.1038\/s41928-020-00476-7","volume":"3","author":"Y. Ding","year":"2020","unstructured":"Ding, Y., Jiang, W., Lou, Q., Liu, J., Xiong, J., Hu, X. S., et al. (2020). Hardware design and the competency awareness of a neural network. Nature Electronics, 3(9), 514\u2013523.","journal-title":"Nature Electronics"},{"key":"65_CR176","doi-asserted-by":"publisher","first-page":"1374","DOI":"10.1123\/ijspp.2023-0154","volume":"12","author":"D. Boullosa","year":"2023","unstructured":"Boullosa, D., Claudino, J. G., Fernandez-Fernandez, J., Bok, D., Loturco, I., Stults-Kolehmainen, M., et al. (2023). The fine-tuning approach for training monitoring. International Journal of Sports Physiology and Performance, 12, 1374\u20131379.","journal-title":"International Journal of Sports Physiology and Performance"},{"key":"65_CR177","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-89010-0","volume-title":"Multivariate statistical machine learning methods for genomic prediction","author":"O. A. Montesinos L\u00f3pez","year":"2022","unstructured":"Montesinos L\u00f3pez, O. A., Montesinos L\u00f3pez, A., & Crossa, J. (2022). Multivariate statistical machine learning methods for genomic prediction. Cham: Springer."},{"key":"65_CR178","first-page":"353","volume-title":"The 16th European conference on computer vision","author":"S. Stoll","year":"2020","unstructured":"Stoll, S., Hadfield, S., & Bowden, R. (2020). Signsynth: data-driven sign language video generation. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), The 16th European conference on computer vision (pp. 353\u2013370). Cham: Springer."},{"issue":"1","key":"65_CR179","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1007\/s10462-022-10189-2","volume":"56","author":"Y. Luo","year":"2023","unstructured":"Luo, Y., Yu, X., Yang, D., & Zhou, B. (2023). A survey of intelligent transmission line inspection based on unmanned aerial vehicle. Artificial Intelligence Review, 56(1), 173\u2013201.","journal-title":"Artificial Intelligence Review"},{"issue":"3","key":"65_CR180","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1609\/aimag.v36i3.2601","volume":"36","author":"J. Wu","year":"2015","unstructured":"Wu, J., Williams, K. M., Chen, H. H., Khabsa, M., Caragea, C., Tuarob, S., et al. (2015). Citeseerx: AI in a digital library search engine. AI Magazine, 36(3), 35\u201348.","journal-title":"AI Magazine"},{"key":"65_CR181","unstructured":"Kim, S. Y., Fan, Z., Noller, Y., & Roychoudhury, A. (2024). Codexity: secure AI-assisted code generation. arXiv preprint. arXiv:2405.03927."},{"key":"65_CR182","first-page":"1","volume-title":"Proceedings of the international conference for high performance computing, networking, storage and analysis","author":"R. Y. Aminabadi","year":"2022","unstructured":"Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., et al. (2022). Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale. In Proceedings of the international conference for high performance computing, networking, storage and analysis (pp. 1\u201315). Piscataway: IEEE."},{"key":"65_CR183","unstructured":"Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: accurate post-training quantization for generative pre-trained transformers. arXiv preprint. arXiv:2210.17323."},{"key":"65_CR184","first-page":"1","volume-title":"Proceedings of the 41st international conference on machine learning","author":"H. Qin","year":"2024","unstructured":"Qin, H., Ma, X., Zheng, X., Li, X., Zhang, Y., Liu, S., et al. (2024). Accurate LoRA-finetuning quantization of LLMs via information retention. In Proceedings of the 41st international conference on machine learning (pp. 1\u201319). Retrieved November 14, 2024, from https:\/\/openreview.net\/forum?id=jQ92egz5Ym."},{"key":"65_CR185","first-page":"1","volume-title":"Proceedings of the 41st international conference on machine learning","author":"W. Huang","year":"2024","unstructured":"Huang, W., Liu, Y., Qin, H., Li, Y., Zhang, S., Liu, X., et al. (2024). BiLLM: pushing the limit of post-training quantization for LLMs. In Proceedings of the 41st international conference on machine learning (pp. 1\u201320). Retrieved November 14, 2024, from https:\/\/openreview.net\/forum?id=qOl2WWOqFg."},{"key":"65_CR186","first-page":"1314","volume-title":"Proceedings of the IEEE symposium on security and privacy","author":"X. Pan","year":"2020","unstructured":"Pan, X., Zhang, M., Ji, S., & Yang, M. (2020). Privacy risks of general-purpose language models. In Proceedings of the IEEE symposium on security and privacy (pp. 1314\u20131331). Piscataway: IEEE."},{"key":"65_CR187","first-page":"14200","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"A. Nagrani","year":"2021","unstructured":"Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., & Sun, C. (2021). Attention bottlenecks for multimodal fusion. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, et al. (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 14200\u201314213). Red Hook: Curran Associates."},{"issue":"7","key":"65_CR188","first-page":"3366","volume":"44","author":"M. De Lange","year":"2021","unstructured":"De Lange, M., Aljundi, R., Masana, M., Parisot, S., Jia, X., Leonardis, A., et al. (2021). A continual learning survey: defying forgetting in classification tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(7), 3366\u20133385.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"12","key":"65_CR189","doi-asserted-by":"publisher","first-page":"1185","DOI":"10.1038\/s42256-022-00568-3","volume":"4","author":"G. M. van de Ven","year":"2022","unstructured":"van de Ven, G. M., Tuytelaars, T., & Tolias, A. S. (2022). Three types of incremental learning. Nature Machine Intelligence, 4(12), 1185\u20131197.","journal-title":"Nature Machine Intelligence"},{"issue":"1","key":"65_CR190","first-page":"857","volume":"35","author":"X. Liu","year":"2021","unstructured":"Liu, X., Zhang, F., Hou, Z., Mian, L., Wang, Z., Zhang, J., et al. (2021). Self-supervised learning: generative or contrastive. IEEE Transactions on Knowledge and Data Engineering, 35(1), 857\u2013876.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"issue":"1","key":"65_CR191","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1093\/nsr\/nwx106","volume":"5","author":"Z. H. Zhou","year":"2018","unstructured":"Zhou, Z. H. (2018). A brief introduction to weakly supervised learning. National Science Review, 5(1), 44\u201353.","journal-title":"National Science Review"},{"issue":"2","key":"65_CR192","first-page":"163","volume":"5","author":"V. Shah","year":"2021","unstructured":"Shah, V., & Konda, S. R. (2021). Neural networks and explainable AI: bridging the gap between models and interpretability. International Journal of Computer Science and Technology, 5(2), 163\u2013176.","journal-title":"International Journal of Computer Science and Technology"},{"key":"65_CR193","unstructured":"Li, Y., Du, M., Song, R., Wang, X., & Wang, Y. (2023). A survey on fairness in large language models. arXiv preprint. arXiv:2308.10149."},{"issue":"8","key":"65_CR194","doi-asserted-by":"publisher","first-page":"1655","DOI":"10.1109\/JPROC.2019.2921977","volume":"107","author":"J. Chen","year":"2019","unstructured":"Chen, J., & Ran, X. (2019). Deep learning with edge computing: a review. Proceedings of the IEEE, 107(8), 1655\u20131674.","journal-title":"Proceedings of the IEEE"}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00065-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-024-00065-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00065-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,27]],"date-time":"2024-12-27T22:09:20Z","timestamp":1735337360000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-024-00065-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,27]]},"references-count":194,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["65"],"URL":"https:\/\/doi.org\/10.1007\/s44267-024-00065-8","relation":{},"ISSN":["2731-9008"],"issn-type":[{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,27]]},"assertion":[{"value":"14 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 November 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 November 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 December 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"34"}}