{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T23:11:23Z","timestamp":1783552283263,"version":"3.55.0"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2024,7,19]],"date-time":"2024-07-19T00:00:00Z","timestamp":1721347200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,19]],"date-time":"2024-07-19T00:00:00Z","timestamp":1721347200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["952215"],"award-info":[{"award-number":["952215"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2024,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper introduces an Automated Machine Learning (AutoML) framework specifically designed to efficiently synthesize end-to-end multimodal machine learning pipelines. Traditional reliance on the computationally demanding Neural Architecture Search is minimized through the strategic integration of pre-trained transformer models. This innovative approach enables the effective unification of diverse data modalities into high-dimensional embeddings, streamlining the pipeline development process. We leverage an advanced Bayesian Optimization strategy, informed by meta-learning, to facilitate the warm-starting of the pipeline synthesis, thereby enhancing computational efficiency. Our methodology demonstrates its potential to create advanced and custom multimodal pipelines within limited computational resources. Extensive testing across 23 varied multimodal datasets indicates the promise and utility of our framework in diverse scenarios. The results contribute to the ongoing efforts in the AutoML field, suggesting new possibilities for efficiently handling complex multimodal data. This research represents a step towards developing more efficient and versatile tools in multimodal machine learning pipeline development, acknowledging the collaborative and ever-evolving nature of this field.\n<\/jats:p>","DOI":"10.1007\/s10994-024-06568-1","type":"journal-article","created":{"date-parts":[[2024,7,19]],"date-time":"2024-07-19T14:02:34Z","timestamp":1721397754000},"page":"7011-7053","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Towards efficient AutoML: a pipeline synthesis approach leveraging pre-trained transformers for multimodal data"],"prefix":"10.1007","volume":"113","author":[{"given":"Ambarish","family":"Moharil","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joaquin","family":"Vanschoren","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Prabhant","family":"Singh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Damian","family":"Tamburri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,7,19]]},"reference":[{"key":"6568_CR1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1505.00468","author":"A Agrawal","year":"2015","unstructured":"Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Batra, D., & Parikh, D. (2015). VQA: Visual question answering. ArXiv. https:\/\/doi.org\/10.48550\/ARXIV.1505.00468","journal-title":"ArXiv"},{"key":"6568_CR2","doi-asserted-by":"publisher","unstructured":"Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., & Auli, M. (2022). data2vec: A General framework for self-supervised learning in speech, vision and language. arXivhttps:\/\/doi.org\/10.48550\/ARXIV.2202.03555. https:\/\/arxiv.org\/abs\/2202.03555","DOI":"10.48550\/ARXIV.2202.03555"},{"key":"6568_CR3","unstructured":"Barrett, L. F. (2017). How emotions are made: The secret life of the brain. Houghton Mifflin Harcourt. https:\/\/books.google.nl\/books?id=hN8MBgAAQBAJ"},{"key":"6568_CR4","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1810.04805","author":"J Devlin","year":"2018","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1810.04805","journal-title":"arXiv"},{"key":"6568_CR5","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2010.11929","author":"A Dosovitskiy","year":"2020","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2010.11929","journal-title":"arXiv"},{"key":"6568_CR6","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2202.10936","author":"Y Du","year":"2022","unstructured":"Du, Y., Liu, Z., Li, J., & Zhao, W. X. (2022). A survey of vision-language pre-trained models. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2202.10936","journal-title":"arXiv"},{"key":"6568_CR7","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1808.05377","author":"T Elsken","year":"2018","unstructured":"Elsken, T., Metzen, J. H., & Hutter, F. (2018). Neural architecture search: A survey. Journal of Machine Learning Research. https:\/\/doi.org\/10.48550\/ARXIV.1808.05377","journal-title":"Journal of Machine Learning Research"},{"key":"6568_CR8","doi-asserted-by":"publisher","unstructured":"Erickson, N., Shi, X., Sharpnack, J., & Smola, A. (2022). Multimodal automl for image, text and tabular data. In Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining. KDD \u201922 (pp. 4786\u20134787). Association for Computing Machinery. https:\/\/doi.org\/10.1145\/3534678.3542616","DOI":"10.1145\/3534678.3542616"},{"key":"6568_CR9","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2003.06505","author":"N Erickson","year":"2020","unstructured":"Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., & Smola, A. (2020). AutoGluon-tabular: Robust and Accurate AutoML for structured data. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2003.06505","journal-title":"arXiv"},{"key":"6568_CR10","doi-asserted-by":"publisher","unstructured":"Ferraro, F., Mostafazadeh, N., Huang, T.K., Vanderwende, L., Devlin, J., Galley, M., & Mitchell, M. (2015). A survey of current datasets for vision and language research. In L. M\u00e0rquez, C. Callison-Burch, J. Su, D. Pighin, Y. Marton (Eds) Proceedings of the 2015 conference on empirical methods in natural language processing, EMNLP 2015, Lisbon, Portugal, September 17\u201321, 2015 (pp. 207\u2013213). The Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/v1\/d15-1021","DOI":"10.18653\/v1\/d15-1021"},{"key":"6568_CR11","doi-asserted-by":"publisher","unstructured":"Feurer, M., Eggensperger, K., Falkner, S., Lindauer, M., & Hutter, F. (2020). Auto-sklearn 2.0: Hands-free automl via meta-learning. https:\/\/doi.org\/10.48550\/ARXIV.2007.04074","DOI":"10.48550\/ARXIV.2007.04074"},{"key":"6568_CR12","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2004.05439","author":"T Hospedales","year":"2020","unstructured":"Hospedales, T., Antoniou, A., Micaelli, P., & Storkey, A. (2020). Meta-learning in neural networks: A survey. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2004.05439","journal-title":"arXiv"},{"key":"6568_CR13","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-05318-5","volume-title":"Automated machine learning: Methods, systems, challenges","author":"F Hutter","year":"2019","unstructured":"Hutter, F., Kotthoff, L., & Vanschoren, J. (2019). Automated machine learning: Methods, systems, challenges (1st ed.). Springer.","edition":"1"},{"key":"6568_CR14","unstructured":"Khan, S. H., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., & Shah, M. (2021) Transformers in vision: A survey. CoRR arxiv:2101.01169"},{"key":"6568_CR15","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1909.11942","author":"Z Lan","year":"2019","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2019). ALBERT: A lite BERT for self-supervised learning of language representations. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1909.11942","journal-title":"arXiv"},{"key":"6568_CR16","unstructured":"Liang, P. P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M. A., Zhu, Y., Salakhutdinov, R., & Morency, L. (2021). Multibench: Multiscale benchmarks for multimodal representation learning. In J. Vanschoren, & S. Yeung (Eds.) Proceedings of the neural information processing systems track on datasets and benchmarks 1, NeurIPS datasets and benchmarks 2021, December 2021, Virtualhttps:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper\/2021\/hash\/37693cfc748049e45d87b8c7d8b9aacd-Abstract-round1.html"},{"key":"6568_CR17","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2107.07651","author":"J Li","year":"2021","unstructured":"Li, J., Selvaraju, R. R., Gotmare, A. D., Joty, S., Xiong, C., & Hoi, S. (2021). Align before fuse: Vision and language representation learning with momentum distillation. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2107.07651","journal-title":"arXiv"},{"key":"6568_CR18","doi-asserted-by":"publisher","unstructured":"Liu, H., Simonyan, K., & Yang, Y. (2018). DARTS: Differentiable architecture search. arXivhttps:\/\/doi.org\/10.48550\/ARXIV.1806.09055 . https:\/\/arxiv.org\/abs\/1806.09055","DOI":"10.48550\/ARXIV.1806.09055"},{"key":"6568_CR19","doi-asserted-by":"publisher","first-page":"1","DOI":"10.48550\/ARXIV.1907.11692","volume":"1","author":"Y Liu","year":"2019","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv, 1, 1. https:\/\/doi.org\/10.48550\/ARXIV.1907.11692","journal-title":"arXiv"},{"issue":"9","key":"6568_CR20","doi-asserted-by":"publisher","first-page":"3108","DOI":"10.1109\/TPAMI.2021.3075372","volume":"43","author":"Z Liu","year":"2021","unstructured":"Liu, Z., Pavao, A., Xu, Z., Escalera, S., Ferreira, F., Guyon, I., Hong, S., Hutter, F., Ji, R., J\u00fanior, J. C. S. J., Li, G., Lindauer, M., Luo, Z., Madadi, M., Nierhoff, T., Niu, K., Pan, C., Stoll, D., Treguer, S., \u2026 Zhang, Y. (2021). Winning solutions and post-challenge analyses of the ChaLearn AutoDL challenge 2019. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(9), 3108\u20133125. https:\/\/doi.org\/10.1109\/TPAMI.2021.3075372","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"6568_CR21","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1908.03557","author":"LH Li","year":"2019","unstructured":"Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., & Chang, K.-W. (2019). VisualBERT: A simple and performant baseline for vision and language. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1908.03557","journal-title":"arXiv"},{"key":"6568_CR22","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2004.06165","author":"X Li","year":"2020","unstructured":"Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., & Gao, J. (2020). Oscar: Object-semantics aligned pre-training for vision-language tasks. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2004.06165","journal-title":"arXiv"},{"key":"6568_CR23","doi-asserted-by":"publisher","first-page":"605","DOI":"10.1613\/jair.4377","volume":"51","author":"P Nguyen","year":"2014","unstructured":"Nguyen, P., Hilario, M., & Kalousis, A. (2014). Using meta-mining to support data mining workflow planning and optimization. J. Artif. Intell. Res., 51, 605\u2013644. https:\/\/doi.org\/10.1613\/jair.4377","journal-title":"J. Artif. Intell. Res."},{"key":"6568_CR24","doi-asserted-by":"crossref","unstructured":"Olson, R. S., & Moore, J. H. (2016). Evaluation of a tree-based pipeline optimization tool for automating data science. In Proceedings of the genetic and evolutionary computation conference 2016 (pp. 485\u2013492). ACM.","DOI":"10.1145\/2908812.2908918"},{"key":"6568_CR25","unstructured":"Ordonez, V., Kulkarni, G., & Berg, T.L. (2011). Im2text: Describing images using 1 million captioned photographs. In J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. C. N. Pereira, K. Q. Weinberger (Eds.) Advances in neural information processing systems 24: 25th annual conference on neural information processing systems 2011. Proceedings of a meeting held 12\u201314 December 2011, Granada, Spain (pp. 1143\u20131151). https:\/\/proceedings.neurips.cc\/paper\/2011\/hash\/5dd9db5e033da9c6fb5ba83c7a7ebea9-Abstract.html"},{"key":"6568_CR26","unstructured":"\u00d6zt\u00fcrk, E., Ferreira, F., Jomaa, H., Schmidt-Thieme, L., Grabocka, J., & Hutter, F. (2022). Zero-shot AutoML with pretrained models. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, & S. Sabato (Eds.) Proceedings of the 39th international conference on machine learning. Proceedings of machine learning research (Vol. 162, pp. 17138\u201317155). PMLR. https:\/\/proceedings.mlr.press\/v162\/ozturk22a.html"},{"key":"6568_CR27","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1505.04870","author":"BA Plummer","year":"2015","unstructured":"Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., & Lazebnik, S. (2015). Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1505.04870","journal-title":"arXiv"},{"issue":"10","key":"6568_CR28","doi-asserted-by":"publisher","first-page":"1872","DOI":"10.1007\/s11431-020-1647-3","volume":"63","author":"X Qiu","year":"2020","unstructured":"Qiu, X., Sun, T., Xu, Y., Shao, Y., Dai, N., & Huang, X. (2020). Pre-trained models for natural language processing: A survey. Science China Technological Sciences, 63(10), 1872\u20131897. https:\/\/doi.org\/10.1007\/s11431-020-1647-3","journal-title":"Science China Technological Sciences"},{"key":"6568_CR29","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2103.00020","author":"A Radford","year":"2021","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2103.00020","journal-title":"arXiv"},{"key":"6568_CR30","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2111.02705","author":"X Shi","year":"2021","unstructured":"Shi, X., Mueller, J., Erickson, N., Li, M., & Smola, A. J. (2021). Benchmarking multimodal AutoML for tabular data with text fields. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2111.02705","journal-title":"arXiv"},{"key":"6568_CR31","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2112.04482","author":"A Singh","year":"2021","unstructured":"Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., & Kiela, D. (2021). FLAVA: A foundational language and vision alignment model. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.2112.04482","journal-title":"arXiv"},{"key":"6568_CR32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.48550\/ARXIV.1908.08530","volume":"1","author":"W Su","year":"2019","unstructured":"Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., & Dai, J. (2019). VL-BERT: Pre-training of generic visual-linguistic representations. arXiv, 1, 1. https:\/\/doi.org\/10.48550\/ARXIV.1908.08530","journal-title":"arXiv"},{"key":"6568_CR33","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1908.07490","author":"H Tan","year":"2019","unstructured":"Tan, H., & Bansal, M. (2019). LXMERT: Learning cross-modality encoder representations from transformers. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1908.07490","journal-title":"arXiv"},{"key":"6568_CR34","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1208.3719","author":"C Thornton","year":"2012","unstructured":"Thornton, C., Hutter, F., Hoos, H. H., & Leyton-Brown, K. (2012). Auto-WEKA: Combined selection and hyperparameter optimization of classification algorithms. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1208.3719","journal-title":"arXiv"},{"key":"6568_CR35","doi-asserted-by":"publisher","first-page":"31640","DOI":"10.7554\/eLife.31640","volume":"7","author":"MJ Van Ackeren","year":"2018","unstructured":"Van Ackeren, M. J., Barbero, F. M., Mattioni, S., Bottini, R., & Collignon, O. (2018). Neuronal populations in the occipital cortex of the blind synchronize to the temporal dynamics of speech. eLife, 7, 31640. https:\/\/doi.org\/10.7554\/eLife.31640","journal-title":"eLife"},{"key":"6568_CR36","unstructured":"Wistuba, M., Rawat, A., & Pedapati, T. (2019). A survey on neural architecture search. IBM Research AIarXiv:1905.01392 [cs.LG]."},{"issue":"3","key":"6568_CR37","doi-asserted-by":"publisher","first-page":"619","DOI":"10.1007\/s40745-021-00326-z","volume":"10","author":"A Zehtab-Salmasi","year":"2021","unstructured":"Zehtab-Salmasi, A., Feizi-Derakhshi, A.-R., Nikzad-Khasmakhi, N., Asgari-Chenaghlu, M., & Nabipour, S. (2021). Multimodal price prediction. Annals of Data Science, 10(3), 619\u2013635.","journal-title":"Annals of Data Science"},{"key":"6568_CR38","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.1904.12054","author":"M-A Z\u00f6ller","year":"2019","unstructured":"Z\u00f6ller, M.-A., & Huber, M. F. (2019). Benchmark and survey of automated machine learning frameworks. Journal of Artificial Intelligence Research. https:\/\/doi.org\/10.48550\/ARXIV.1904.12054","journal-title":"Journal of Artificial Intelligence Research"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06568-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-024-06568-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06568-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T17:41:15Z","timestamp":1723052475000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-024-06568-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,19]]},"references-count":38,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2024,9]]}},"alternative-id":["6568"],"URL":"https:\/\/doi.org\/10.1007\/s10994-024-06568-1","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,19]]},"assertion":[{"value":"4 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 May 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 May 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 July 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest. The institutions involved in this research are \u2018@tue.nl\u2019 and \u2018@jads.nl.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not Applicable. This research did not involve human participants, their data, or animals.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not Applicable. This research did not involve human participants.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not Applicable. This manuscript does not contain any individual person\u2019s data in any form.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}]}}