{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T02:45:06Z","timestamp":1784342706810,"version":"3.55.0"},"reference-count":62,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T00:00:00Z","timestamp":1669852800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T00:00:00Z","timestamp":1669852800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100017561","name":"Future of Life Institute","doi-asserted-by":"crossref","award":["RFP2-152"],"award-info":[{"award-number":["RFP2-152"]}],"id":[{"id":"10.13039\/100017561","id-type":"DOI","asserted-by":"crossref"}]},{"name":"MIT-Spain - INDITEX Sustainability Seed Fund","award":["COST-OMIZE"],"award-info":[{"award-number":["COST-OMIZE"]}]},{"DOI":"10.13039\/501100010198","name":"Ministerio de Econom\u00eda, Industria y Competitividad, Gobierno de Espa\u00f1a","doi-asserted-by":"publisher","award":["RTI2018-094403-B-C32"],"award-info":[{"award-number":["RTI2018-094403-B-C32"]}],"id":[{"id":"10.13039\/501100010198","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003359","name":"Generalitat Valenciana","doi-asserted-by":"publisher","award":["PROMETEO\/2019\/098"],"award-info":[{"award-number":["PROMETEO\/2019\/098"]}],"id":[{"id":"10.13039\/501100003359","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003359","name":"Generalitat Valenciana","doi-asserted-by":"publisher","award":["NNEST\/2021\/317"],"award-info":[{"award-number":["NNEST\/2021\/317"]}],"id":[{"id":"10.13039\/501100003359","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["952215 (TAILOR)"],"award-info":[{"award-number":["952215 (TAILOR)"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","award":["HR00112120007 ReCOG-A"],"award-info":[{"award-number":["HR00112120007 ReCOG-A"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Universidad Polit\u00e8cnica de Val\u00e8ncia"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The automation of data science and other data manipulation processes depend on the integration and formatting of \u2018messy\u2019 data. Data wrangling is an umbrella term for these tedious and time-consuming tasks. Tasks such as transforming dates, units or names expressed in different formats have been challenging for machine learning because (1) users expect to solve them with short cues or few examples, and (2) the problems depend heavily on domain knowledge. Interestingly, large language models today (1) can infer from very few examples or even a short clue in natural language, and (2) can integrate vast amounts of domain knowledge. It is then an important research question to analyse whether language models are a promising approach for data wrangling, especially as their capabilities continue growing. In this paper we apply different variants of the language model Generative Pre-trained Transformer (GPT) to five batteries covering a wide range of data wrangling problems. We compare the effect of prompts and few-shot regimes on their results and how they compare with specialised data wrangling systems and other tools. Our major finding is that they appear as a powerful tool for a wide range of data wrangling tasks. We provide some guidelines about how they can be integrated into data processing pipelines, provided the users can take advantage of their flexibility and the diversity of tasks to be addressed. However, reliability is still an important issue to overcome.<\/jats:p>","DOI":"10.1007\/s10994-022-06259-9","type":"journal-article","created":{"date-parts":[[2022,12,1]],"date-time":"2022-12-01T22:13:48Z","timestamp":1669932828000},"page":"2053-2082","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":40,"title":["Can language models automate data wrangling?"],"prefix":"10.1007","volume":"112","author":[{"given":"Gonzalo","family":"Jaimovitch-L\u00f3pez","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"C\u00e8sar","family":"Ferri","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jos\u00e9","family":"Hern\u00e1ndez-Orallo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2902-6477","authenticated-orcid":false,"given":"Fernando","family":"Mart\u00ednez-Plumed","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mar\u00eda Jos\u00e9","family":"Ram\u00edrez-Quintana","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,12,1]]},"reference":[{"issue":"2","key":"6259_CR1","doi-asserted-by":"publisher","first-page":"113","DOI":"10.14429\/dsj.66.9463","volume":"66","author":"P Ashok","year":"2016","unstructured":"Ashok, P., & Nawaz, G. K. (2016). Outlier detection method on uci repository dataset by entropy based rough k-means. Defence Science Journal, 66(2), 113\u2013121.","journal-title":"Defence Science Journal"},{"key":"6259_CR2","doi-asserted-by":"publisher","first-page":"164380","DOI":"10.1109\/ACCESS.2020.3021596","volume":"8","author":"P Bellmann","year":"2020","unstructured":"Bellmann, P., & Schwenker, F. (2020). Ordinal classification: Working definition and detection of ordinal structures. IEEE Access, 8, 164380\u2013164391. https:\/\/doi.org\/10.1109\/ACCESS.2020.3021596","journal-title":"IEEE Access"},{"key":"6259_CR3","doi-asserted-by":"crossref","unstructured":"Ben-Gal, I. (2005). Outlier detection. In Data mining and knowledge discovery handbook (pp. 131\u2013146). Springer.","DOI":"10.1007\/0-387-25465-X_7"},{"key":"6259_CR4","doi-asserted-by":"crossref","unstructured":"Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610\u2013623). FAccT \u201921.","DOI":"10.1145\/3442188.3445922"},{"key":"6259_CR5","first-page":"1137","volume":"3","author":"Y Bengio","year":"2003","unstructured":"Bengio, Y., Ducharme, R., Vincent, P., & Janvin, C. (2003). A neural probabilistic language model. The Journal of Machine Learning Research, 3, 1137\u20131155.","journal-title":"The Journal of Machine Learning Research"},{"key":"6259_CR6","unstructured":"Bhupatiraju, S., Singh, R., Mohamed, A. R., & Kohli, P. (2017). Deep API programmer: Learning to program with APIs. arXiv preprint arXiv:1704.04327."},{"key":"6259_CR7","unstructured":"BIG-bench collaboration. (2022). Beyond the imitation game: Measuring and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615. https:\/\/github.com\/google\/BIG-bench\/"},{"key":"6259_CR8","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., & Agarwal, S. (2020). Language models are few-shot learners. arXiv preprint arXiv:2005.14165."},{"issue":"2","key":"6259_CR9","doi-asserted-by":"publisher","first-page":"288","DOI":"10.1109\/TPAMI.2008.72","volume":"31","author":"Y Chen","year":"2008","unstructured":"Chen, Y., Dang, X., Peng, H., & Bart, H. L. (2008). Outlier detection with the kernelized spatial depth function. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(2), 288\u2013305.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"6259_CR10","doi-asserted-by":"crossref","unstructured":"Contreras-Ochando, L., Ferri, C., & Hern\u00e1ndez-Orallo, J. (2019a). Automating common data science matrix transformations. In ECMLPKDD workshop on Automating Data Science. ECML-PKDD \u201919.","DOI":"10.1007\/978-3-030-43823-4_2"},{"key":"6259_CR11","unstructured":"Contreras-Ochando, L., Ferri, C., Hern\u00e1ndez-Orallo, J., Mart\u00ednez-Plumed, F., Ram\u00edrez-Quintana, M. J., & Katayama, S. (2019b). Automated data transformation with inductive programming and dynamic background knowledge. In Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2019. ECML-PKDD \u201919."},{"key":"6259_CR12","doi-asserted-by":"crossref","unstructured":"Cropper, A., Tamaddoni, A., & Muggleton, S. H. (2015). Meta-interpretive learning of data transformation programs. In Inductive Logic Programming (pp. 46\u201359).","DOI":"10.1007\/978-3-319-40566-7_4"},{"key":"6259_CR13","doi-asserted-by":"crossref","unstructured":"Das, K., & Schneider, J. (2007). Detecting anomalous records in categorical datasets. In Proceedings of the 13th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 220\u2013229).","DOI":"10.1145\/1281192.1281219"},{"key":"6259_CR14","doi-asserted-by":"crossref","unstructured":"Das, K., Schneider, J., & Neill, D. B. (2008). Anomaly pattern detection in categorical datasets. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 169\u2013176).","DOI":"10.1145\/1401890.1401915"},{"issue":"3","key":"6259_CR15","doi-asserted-by":"publisher","first-page":"76","DOI":"10.1145\/3495256","volume":"65","author":"T De Bie","year":"2022","unstructured":"De Bie, T., De Raedt, L., Hern\u00e1ndez-Orallo, J., Hoos, H. H., Smyth, P., & Williams, C. K. I. (2022). Automating data science: Prospects and challenges. Communications of the ACM, 65(3), 76\u201387.","journal-title":"Communications of the ACM"},{"key":"6259_CR16","unstructured":"Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805."},{"key":"6259_CR17","unstructured":"Dua, D., & Graff, C. (2017). UCI machine learning repository. http:\/\/archive.ics.uci.edu\/ml."},{"key":"6259_CR18","doi-asserted-by":"crossref","unstructured":"Ellis, K., & Gulwani, S. (2017). Learning to learn programs from examples: Going beyond program structure. In IJCAI (pp. 1638\u20131645).","DOI":"10.24963\/ijcai.2017\/227"},{"issue":"7","key":"6259_CR19","doi-asserted-by":"publisher","first-page":"3217","DOI":"10.1002\/int.22415","volume":"36","author":"MP Fernando","year":"2021","unstructured":"Fernando, M. P., C\u00e8sar, F., David, N., & Jos\u00e9, H. O. (2021). Missing the missing values: The ugly duckling of fairness in machine learning. International Journal of Intelligent Systems, 36(7), 3217\u20133258.","journal-title":"International Journal of Intelligent Systems"},{"key":"6259_CR20","volume-title":"Introducing Microsoft Power BI","author":"A Ferrari","year":"2016","unstructured":"Ferrari, A., & Russo, M. (2016). Introducing Microsoft Power BI. Microsoft Press."},{"key":"6259_CR21","first-page":"473","volume":"16","author":"T Furche","year":"2016","unstructured":"Furche, T., Gottlob, G., Libkin, L., Orsi, G., & Paton, N. W. (2016). Data wrangling for big data. Challenges and opportunities. EDBT, 16, 473\u2013478.","journal-title":"EDBT"},{"key":"6259_CR22","doi-asserted-by":"crossref","unstructured":"Gao, T., Fisch, A., & Chen, D. (2020). Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.","DOI":"10.18653\/v1\/2021.acl-long.295"},{"issue":"1","key":"6259_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s41044-016-0014-0","volume":"1","author":"S Garc\u00eda","year":"2016","unstructured":"Garc\u00eda, S., Ram\u00edrez-Gallego, S., Luengo, J., Ben\u00edtez, J. M., & Herrera, F. (2016). Big data preprocessing: Methods and prospects. Big Data Analytics, 1(1), 1\u201322.","journal-title":"Big Data Analytics"},{"key":"6259_CR24","doi-asserted-by":"crossref","unstructured":"Gulwani, S. (2011). Automating string processing in spreadsheets using input-output examples. In Procs. 38th Principles of Programming Languages (pp. 317\u2013330).","DOI":"10.1145\/1926385.1926423"},{"issue":"11","key":"6259_CR25","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1145\/2736282","volume":"58","author":"S Gulwani","year":"2015","unstructured":"Gulwani, S., Hern\u00e1ndez-Orallo, J., Kitzelmann, E., Muggleton, S. H., Schmid, U., & Zorn, B. (2015). Inductive programming meets the real world. Communications of the ACM, 58(11), 90\u201399.","journal-title":"Communications of the ACM"},{"key":"6259_CR26","doi-asserted-by":"crossref","unstructured":"Ham, K. (2013). OpenRefine (version 2.5). http:\/\/openrefine.org.free\/ Open-source tool for cleaning and transforming data. Journal of the Medical Library Association: JMLA, 101 (3), 233.","DOI":"10.3163\/1536-5050.101.3.020"},{"issue":"1","key":"6259_CR27","doi-asserted-by":"publisher","first-page":"103","DOI":"10.2298\/CSIS0501103H","volume":"2","author":"Z He","year":"2005","unstructured":"He, Z., Xu, X., Huang, Z. J., & Deng, S. (2005). Fp-outlier: Frequent pattern based outlier detection. Computer Science and Information Systems, 2(1), 103\u2013118.","journal-title":"Computer Science and Information Systems"},{"key":"6259_CR28","unstructured":"Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. In ICLR."},{"key":"6259_CR29","unstructured":"Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021). Measuring mathematical problem solving with the MATH dataset. In CoRR. arxiv:2103.03874."},{"key":"6259_CR30","doi-asserted-by":"crossref","unstructured":"Hulsebos, M., Hu, K., Bakker, M., Zgraggen, E., Satyanarayan, A., Kraska, T., Demiralp, \u00c7., & Hidalgo, C. (2019). Sherlock: A deep learning approach to semantic data type detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (pp. 1500\u20131508).","DOI":"10.1145\/3292500.3330993"},{"key":"6259_CR31","doi-asserted-by":"crossref","unstructured":"Izacard, G., & Grave, E. (2020). Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282.","DOI":"10.18653\/v1\/2021.eacl-main.74"},{"key":"6259_CR32","doi-asserted-by":"crossref","unstructured":"Jaimovitch-Lopez, G., Ferri, C., Hernandez-Orallo, J., Martinez-Plumed, F., & Ramirez-Quintana, M. J. (2021). Can language models automate data wrangling?. In ECML\/PKDD Workshop on Automated Data Science (ADS2021). https:\/\/sites.google.com\/view\/autods.","DOI":"10.1007\/s10994-022-06259-9"},{"key":"6259_CR33","doi-asserted-by":"crossref","unstructured":"Kandel, S., Paepcke, A., Hellerstein, J., & Heer, J. (2011). Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 3363\u20133372). ACM.","DOI":"10.1145\/1978942.1979444"},{"key":"6259_CR34","doi-asserted-by":"crossref","unstructured":"Lazarevic, A., & Kumar, V. (2005). Feature bagging for outlier detection. In Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining (pp. 157\u2013166).","DOI":"10.1145\/1081870.1081891"},{"key":"6259_CR35","doi-asserted-by":"crossref","unstructured":"Lu, Y., Bartolo, M., Moore, A., Riedel, S., & Stenetorp, P. (2021). Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.","DOI":"10.18653\/v1\/2022.acl-long.556"},{"key":"6259_CR36","unstructured":"Nazabal, A., Williams, C. K., Colavizza, G., Smith, C. R., & Williams, A. (2020). Data engineering for data analytics: A classification of the issues, and case studies. arXiv preprint arXiv:2004.12929."},{"issue":"1","key":"6259_CR37","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1007\/s10618-011-0234-x","volume":"25","author":"K Noto","year":"2012","unstructured":"Noto, K., Brodley, C., & Slonim, D. (2012). Frac: A feature-modeling approach for semi-supervised and unsupervised anomaly detection. Data Mining and Knowledge Discovery, 25(1), 109\u2013133.","journal-title":"Data Mining and Knowledge Discovery"},{"key":"6259_CR38","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825\u20132830.","journal-title":"Journal of Machine Learning Research"},{"key":"6259_CR39","doi-asserted-by":"crossref","unstructured":"Petrova-Antonova, D., & Tancheva, R. (2020). Data cleaning: A case study with OpenRefine and Trifacta Wrangler. In International Conference on the Quality of Information and Communications Technology (pp. 32\u201340). Springer.","DOI":"10.1007\/978-3-030-58793-2_3"},{"key":"6259_CR40","unstructured":"Porwal, U., & Mukund, S. (2017). Outlier detection by consistent data selection method. arXiv preprint arXiv:1712.04129."},{"key":"6259_CR41","unstructured":"Puri, R., & Catanzaro, B. (2019). Zero-shot text classification with generative language models. arXiv preprint arXiv:1912.10165."},{"issue":"8","key":"6259_CR42","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9.","journal-title":"OpenAI Blog"},{"key":"6259_CR43","unstructured":"Raman, V., & Hellerstein, J. M. (2001). Potter\u2019s wheel: An interactive data cleaning system. In VLDB (Vol.\u00a01, pp. 381\u2013390)."},{"key":"6259_CR44","unstructured":"Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., & Eccles, T. (2022). A generalist agent. arXiv preprint arXiv:2205.06175."},{"issue":"3","key":"6259_CR45","doi-asserted-by":"publisher","first-page":"581","DOI":"10.1093\/biomet\/63.3.581","volume":"63","author":"DB Rubin","year":"1976","unstructured":"Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581\u2013592.","journal-title":"Biometrika"},{"key":"6259_CR46","doi-asserted-by":"crossref","unstructured":"Schick, T., & Sch\u00fctze, H. (2020). Exploiting cloze questions for few-shot text classification and natural language inference. arXiv preprint arXiv:2001.07676.","DOI":"10.18653\/v1\/2021.eacl-main.20"},{"issue":"4","key":"6259_CR47","doi-asserted-by":"publisher","first-page":"656","DOI":"10.1002\/j.1538-7305.1949.tb00928.x","volume":"28","author":"CE Shannon","year":"1949","unstructured":"Shannon, C. E. (1949). Communication theory of secrecy systems. The Bell System Technical Journal, 28(4), 656\u2013715.","journal-title":"The Bell System Technical Journal"},{"key":"6259_CR48","doi-asserted-by":"crossref","unstructured":"Shi, Y., Li, W., & Sha, F. (2016). Metric learning for ordinal data. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol.\u00a030).","DOI":"10.1609\/aaai.v30i1.10280"},{"key":"6259_CR49","doi-asserted-by":"crossref","unstructured":"Singh, R., & Gulwani, S. (2015). Predicting a correct program in programming by example. In International Conference on Computer Aided Verification (pp. 398\u2013414). Springer.","DOI":"10.1007\/978-3-319-21690-4_23"},{"key":"6259_CR50","doi-asserted-by":"crossref","unstructured":"Singh, R., & Gulwani, S. (2016). Transforming spreadsheet data types using examples. In Proceedings of the 43rd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (pp. 343\u2013356).","DOI":"10.1145\/2837614.2837668"},{"key":"6259_CR51","volume-title":"Tableau Desktop Pocket Reference","author":"R Sleeper","year":"2021","unstructured":"Sleeper, R. (2021). Tableau Desktop Pocket Reference. O\u2019Reilly Media Inc."},{"key":"6259_CR52","unstructured":"Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., & Zhang, E. (2022). Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990."},{"key":"6259_CR53","unstructured":"Tamkin, A., Brundage, M., Clark, J., & Ganguli, D. (2021). Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503."},{"key":"6259_CR54","unstructured":"Terrizzano, I. G., Schwarz, P. M., Roth, M., & Colino, J. E. (2015). Data wrangling: The challenging journey from the wild to the lake. In CIDR."},{"key":"6259_CR55","unstructured":"Trifacta (2022): Trifacta Wrangler. https:\/\/www.trifacta.com"},{"key":"6259_CR56","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. arXiv preprint arXiv:1706.03762."},{"key":"6259_CR57","unstructured":"Wei, J., Bosma, M. P., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., & Le, Q. V. (2022). Finetuned language models are zero-shot learners. https:\/\/openreview.net\/forum?id=gEZrGCozdqR"},{"key":"6259_CR58","doi-asserted-by":"crossref","unstructured":"Wu, B., Szekely, P., & Knoblock, C. A. (2012). Learning data transformation rules through examples: Preliminary results. In Information Integration on the Web (p.\u00a08).","DOI":"10.1145\/2331801.2331809"},{"key":"6259_CR59","doi-asserted-by":"crossref","unstructured":"Xu, S., Semnani, S. J., Campagna, G., & Lam, M. S. (2020). AutoQA: From databases to QA semantic parsers with only synthetic training data. In EMNLP.","DOI":"10.18653\/v1\/2020.emnlp-main.31"},{"key":"6259_CR60","unstructured":"Zeng, W., Ren, X., Su, T., Wang, H., Liao, Y., Wang, Z., Jiang, X., Yang, Z., Wang, K., Zhang, X., & Li, C. (2021). Pangu-$$\\alpha$$: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369."},{"key":"6259_CR61","doi-asserted-by":"crossref","unstructured":"Zhang, D., Suhara, Y., Li, J., Hulsebos, M., Demiralp, \u00c7., & Tan, W. C. (2019). Sato: Contextual semantic type detection in tables. arXiv preprint arXiv:1911.06311.","DOI":"10.14778\/3407790.3407793"},{"key":"6259_CR62","doi-asserted-by":"crossref","unstructured":"Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., & Fedus, W. (2022). Designing effective sparse expert models. arXiv preprint arXiv:2202.08906.","DOI":"10.1109\/IPDPSW55747.2022.00171"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06259-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-022-06259-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-022-06259-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,6]],"date-time":"2023-06-06T19:10:54Z","timestamp":1686078654000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-022-06259-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,1]]},"references-count":62,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["6259"],"URL":"https:\/\/doi.org\/10.1007\/s10994-022-06259-9","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,1]]},"assertion":[{"value":"25 January 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 August 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 September 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 December 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"No conflicts of interest or competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}}]}}