{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T06:51:34Z","timestamp":1777013494532,"version":"3.51.4"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,13]],"date-time":"2023-06-13T00:00:00Z","timestamp":1686614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,6,13]]},"abstract":"<jats:p>Software systems that learn from data with machine learning (ML) are used in critical decision-making processes. Unfortunately, real-world experience shows that the pipelines for data preparation, feature encoding and model training in ML systems are often brittle with respect to their input data. As a consequence, data scientists have to run different kinds of data centric what-if analyses to evaluate the robustness and reliability of such pipelines, e.g., with respect to data errors or preprocessing techniques. These what-if analyses follow a common pattern: they take an existing ML pipeline, create a pipeline variant by introducing a small change, and execute this pipeline variant to see how the change impacts the pipeline's output score. The application of existing analysis techniques to ML pipelines is technically challenging as they are hard to integrate into existing pipeline code and their execution introduces large overheads due to repeated work.<\/jats:p>\n          <jats:p>We propose mlwhatif to address these integration and efficiency challenges for data-centric what-if analyses on ML pipelines. mlwhatif enables data scientists to declaratively specify what-if analyses for an ML pipeline, and to automatically generate, optimize and execute the required pipeline variants. Our approach employs pipeline patches to specify changes to the data, operators and models of a pipeline. Based on these patches, we define a multi-query optimizer for efficiently executing the resulting pipeline variants jointly, with four subsumption-based optimization rules. Subsequently, we detail how to implement the pipeline variant generation and optimizer of mlwhatif. For that, we instrument native ML pipelines written in Python to extract dataflow plans with re-executable operators.<\/jats:p>\n          <jats:p>We experimentally evaluate mlwhatif, and find that its speedup scales linearly with the number of pipeline variants in applicable cases, and is invariant to the input data size. In end-to-end experiments with four analyses on more than 60 pipelines, we show speedups of up to 13x compared to sequential execution, and find that the speedup is invariant to the model and featurization in the pipeline. Furthermore, we confirm the low instrumentation overhead of mlwhatif.<\/jats:p>","DOI":"10.1145\/3589273","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T20:26:45Z","timestamp":1687292805000},"page":"1-26","source":"Crossref","is-referenced-by-count":18,"title":["Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9884-9517","authenticated-orcid":false,"given":"Stefan","family":"Grafberger","sequence":"first","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0183-6910","authenticated-orcid":false,"given":"Paul","family":"Groth","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4722-5840","authenticated-orcid":false,"given":"Sebastian","family":"Schelter","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, et al. Tfx: A tensorflow-based production-scale machine learning platform. KDD","author":"Baylor Denis","year":"2017","unstructured":"Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc, et al. Tfx: A tensorflow-based production-scale machine learning platform. KDD (2017)."},{"key":"e_1_2_2_2_1","volume-title":"Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline. ESEC\/FSE","author":"Biswas Sumon","year":"2021","unstructured":"Sumon Biswas and Hridesh Rajan. Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline. ESEC\/FSE (2021)."},{"key":"e_1_2_2_3_1","volume-title":"et al . SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle. CIDR","author":"Boehm Matthias","year":"2020","unstructured":"Matthias Boehm, Iulian Antonov, Sebastian Baunsgaard, Mark Dokter, et al . SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle. CIDR (2020)."},{"key":"e_1_2_2_4_1","volume-title":"Data Validation for Machine Learning. MLSys","author":"Breck Eric","year":"2019","unstructured":"Eric Breck, Neoklis Polyzotis, Sudip Roy, Steven Whang, and Martin Zinkevich. Data Validation for Machine Learning. MLSys (2019)."},{"key":"e_1_2_2_5_1","first-page":"1","volume":"45","author":"Breiman Leo","year":"2001","unstructured":"Leo Breiman. Random forests. JMLR 45, 1 (2001).","journal-title":"JMLR"},{"key":"e_1_2_2_6_1","volume-title":"Bahareh Sadat Arab, and Boris Glavic. Efficient Answering of Historical What-If Queries. SIGMOD","author":"Campbell Felix S.","year":"2022","unstructured":"Felix S. Campbell, Bahareh Sadat Arab, and Boris Glavic. Efficient Answering of Historical What-If Queries. SIGMOD (2022)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_2_8_1","unstructured":"Fran\u00e7ois Chollet et al. Keras. https:\/\/github.com\/fchollet\/keras."},{"key":"e_1_2_2_9_1","volume-title":"Zoi Kaoudi, Tilmann Rabl, and Volker Markl. Materialization and Reuse Optimizations for Production Data Science Pipelines. SIGMOD","author":"Derakhshan Behrouz","year":"2021","unstructured":"Behrouz Derakhshan, Alireza Rezaei Mahdiraji, Zoi Kaoudi, Tilmann Rabl, and Volker Markl. Materialization and Reuse Optimizations for Production Data Science Pipelines. SIGMOD (2021)."},{"key":"e_1_2_2_10_1","volume-title":"Caravan: Provisioning for What-If Analysis. CIDR","author":"Deutch Daniel","year":"2013","unstructured":"Daniel Deutch, Zachary G Ives, Tova Milo, and Val Tannen. Caravan: Provisioning for What-If Analysis. CIDR (2013)."},{"key":"e_1_2_2_11_1","unstructured":"DS3Lab ETH Zuerich. DSPipes. https:\/\/github.com\/DS3Lab\/datascope-pipelines\/tree\/main\/dspipes."},{"key":"e_1_2_2_12_1","volume-title":"Complaint-Driven Training Data Debugging at Interactive Speeds. SIGMOD","author":"Flokas Lampros","year":"2022","unstructured":"Lampros Flokas, Weiyuan Wu, Yejia Liu, Jiannan Wang, Nakul Verma, and Eugene Wu. Complaint-Driven Training Data Debugging at Interactive Speeds. SIGMOD (2022)."},{"key":"e_1_2_2_13_1","unstructured":"Fran\u00e7ois Chollet. Simple MNIST convnet. https:\/\/keras.io\/examples\/vision\/mnist_convnet\/."},{"key":"e_1_2_2_14_1","volume-title":"HypeR: Hypothetical Reasoning With What-If and How-To Queries Using a Probabilistic Causal Approach. SIGMOD","author":"Galhotra Sainyam","year":"2022","unstructured":"Sainyam Galhotra, Amir Gilad, Sudeepa Roy, and Babak Salimi. HypeR: Hypothetical Reasoning With What-If and How-To Queries Using a Probabilistic Causal Approach. SIGMOD (2022)."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3533028.3533303"},{"key":"e_1_2_2_16_1","volume-title":"Data distribution debugging in machine learning pipelines. VLDBJ","author":"Grafberger Stefan","year":"2022","unstructured":"Stefan Grafberger, Paul Groth, Julia Stoyanovich, and Sebastian Schelter. Data distribution debugging in machine learning pipelines. VLDBJ (2022)."},{"key":"e_1_2_2_17_1","volume-title":"MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines. SIGMOD","author":"Grafberger Stefan","year":"2021","unstructured":"Stefan Grafberger, Shubha Guha, Julia Stoyanovich, and Sebastian Schelter. MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines. SIGMOD (2021)."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319888"},{"key":"e_1_2_2_19_1","volume-title":"Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song. Efficient task-specific data valuation for nearest neighbor algorithms. VLDB","author":"Jia Ruoxi","year":"2019","unstructured":"Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song. Efficient task-specific data valuation for nearest neighbor algorithms. VLDB (2019)."},{"key":"e_1_2_2_20_1","unstructured":"Sayash Kapoor and Arvind Narayanan. Leakage and the Reproducibility Crisis in ML-based Science. https:\/\/arxiv.org\/abs\/2207.07048"},{"key":"e_1_2_2_21_1","unstructured":"Bojan Karla? David Dao Matteo Interlandi Bo Li Sebastian Schelter Wentao Wu and Ce Zhang. 2022. Data Debugging with Shapley Importance over End-to-End Machine Learning Pipelines. https:\/\/arxiv.org\/abs\/2204.11131"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994514"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2926534.2926540"},{"key":"e_1_2_2_24_1","volume-title":"An Intermediate Representation for Optimizing Machine Learning Pipelines. PVLDB","author":"Kunft Andreas","year":"2019","unstructured":"Andreas Kunft, Asterios Katsifodimos, Sebastian Schelter, Sebastian Bre\u00df, Tilmann Rabl, and Volker Markl. An Intermediate Representation for Optimizing Machine Learning Pipelines. PVLDB (2019)."},{"key":"e_1_2_2_25_1","volume-title":"Cleanml: A benchmark for joint data cleaning and machine learning [experiments and analysis]. ICDE","author":"Li Peng","year":"2019","unstructured":"Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. Cleanml: A benchmark for joint data cleaning and machine learning [experiments and analysis]. ICDE (2019)."},{"key":"e_1_2_2_26_1","first-page":"12","volume":"13","author":"Mahdavi Mohammad","year":"2020","unstructured":"Mohammad Mahdavi and Ziawasch Abedjan. Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning. PVLDB 13, 12 (2020).","journal-title":"Transfer Learning. PVLDB"},{"key":"e_1_2_2_27_1","volume-title":"Semi-Supervised Data Cleaning with Raha and Baran. CIDR","author":"Mahdavi Mohammad","year":"2021","unstructured":"Mohammad Mahdavi and Ziawasch Abedjan. Semi-Supervised Data Cleaning with Raha and Baran. CIDR (2021)."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/869655"},{"key":"e_1_2_2_29_1","volume-title":"pandas: a foundational Python library for data analysis and statistics. Python for high performance and scientific computing 14, 9","author":"McKinney Wes","year":"2011","unstructured":"Wes McKinney et al . pandas: a foundational Python library for data analysis and statistics. Python for high performance and scientific computing 14, 9 (2011)."},{"key":"e_1_2_2_30_1","first-page":"1","volume":"17","author":"Meng Xiangrui","year":"2016","unstructured":"Xiangrui Meng, Joseph Bradley, Burak Yavuz, Evan Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, DB Tsai, Manish Amde, Sean Owen, et al. Mllib: Machine learning in apache spark. JMLR 17, 1 (2016).","journal-title":"Mllib: Machine learning in apache spark. JMLR"},{"key":"e_1_2_2_31_1","unstructured":"Yuval Moskovitch Jinyang Li and H. V. Jagadish. Bias Analysis and Mitigation in Data-Driven Tools Using Provenance (TaPP '22)."},{"key":"e_1_2_2_32_1","volume-title":"Confident learning: Estimating uncertainty in dataset labels. JAIR 70","author":"Northcutt Curtis","year":"2021","unstructured":"Curtis Northcutt, Lu Jiang, and Isaac Chuang. Confident learning: Estimating uncertainty in dataset labels. JAIR 70 (2021)."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3213880.3213890"},{"key":"e_1_2_2_34_1","volume-title":"End-to-End Optimization of Machine Learning Prediction Queries. SIGMOD","author":"Park Kwanghyun","year":"2022","unstructured":"Kwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen, Matteo Interlandi, and Konstantinos Karanasos. End-to-End Optimization of Machine Learning Prediction Queries. SIGMOD (2022)."},{"key":"e_1_2_2_35_1","volume-title":"et al. Scikit-learn: Machine learning in Python. JMLR 12","author":"Pedregosa Fabian","year":"2011","unstructured":"Fabian Pedregosa, Ga\u00ebl Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. Scikit-learn: Machine learning in Python. JMLR 12 (2011)."},{"key":"e_1_2_2_36_1","volume-title":"Eduardo C de Almeida, and Felix Naumann. Efficient detection of data dependency violations. CIKM","author":"Pena Eduardo HM","year":"2020","unstructured":"Eduardo HM Pena, Edson R Lucas Filho, Eduardo C de Almeida, and Felix Naumann. Efficient detection of data dependency violations. CIKM (2020)."},{"key":"e_1_2_2_37_1","first-page":"12","volume":"13","author":"Petersohn Devin","year":"2020","unstructured":"Devin Petersohn, Stephen Macke, Doris Xin, William Ma, Doris Lee, Xiangxi Mo, Joseph E. Gonzalez, Joseph M. Hellerstein, Anthony D. Joseph, and Aditya Parameswaran. Towards Scalable Dataframe Systems. PVLDB 13, 12 (2020).","journal-title":"Aditya Parameswaran. Towards Scalable Dataframe Systems. PVLDB"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551842"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299887.3299891"},{"key":"e_1_2_2_40_1","volume-title":"Interpretable data-based explanations for fairness debugging. SIGMOD","author":"Pradhan Romila","year":"2021","unstructured":"Romila Pradhan, Jiongli Zhu, Boris Glavic, and Babak Salimi. Interpretable data-based explanations for fairness debugging. SIGMOD (2021)."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/319702.319729"},{"key":"e_1_2_2_43_1","volume-title":"Efficient and Extensible Algorithms for Multi Query Optimization. SIGMOD","author":"Roy Prasan","year":"2000","unstructured":"Prasan Roy, S. Seshadri, S. Sudarshan, and Siddhesh Bhobe. Efficient and Extensible Algorithms for Multi Query Optimization. SIGMOD (2000)."},{"key":"e_1_2_2_44_1","volume-title":"On challenges in machine learning model management","author":"Schelter Sebastian","year":"2018","unstructured":"Sebastian Schelter, Felix Biessmann, Tim Januschowski, David Salinas, Stephan Seufert, and Gyuri Szarvas. On challenges in machine learning model management. IEEE Data Engineering Bulletin (2018)."},{"key":"e_1_2_2_45_1","volume-title":"Screening Native ML Pipelines with \"ArgusEyes\". CIDR","author":"Schelter Sebastian","year":"2022","unstructured":"Sebastian Schelter, Stefan Grafberger, Shubha Guha, Olivier Sprangers, Bojan Karla?, and Ce Zhang. Screening Native ML Pipelines with \"ArgusEyes\". CIDR (2022)."},{"key":"e_1_2_2_46_1","volume-title":"Fairprep: Promoting data to a first-class citizen in studies on fairness-enhancing interventions. EDBT","author":"Schelter Sebastian","year":"2019","unstructured":"Sebastian Schelter, Yuxuan He, Jatin Khilnani, and Julia Stoyanovich. Fairprep: Promoting data to a first-class citizen in studies on fairness-enhancing interventions. EDBT (2019)."},{"key":"e_1_2_2_47_1","volume-title":"JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. EDBT","author":"Schelter Sebastian","year":"2021","unstructured":"Sebastian Schelter, Tammo Rukat, and Felix Biessmann. JENGA - A Framework to Study the Impact of Data Errors on the Predictions of Machine Learning Models. EDBT (2021)."},{"key":"e_1_2_2_48_1","unstructured":"Maximilian E. Sch\u00fcle Luca Scalerandi Alfons Kemper and Thomas Neumann. Blue Elephants Inspecting Pandas: Inspection and Execution of Machine Learning Pipelines in SQL. (2023)."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/42201.42202"},{"key":"e_1_2_2_50_1","volume-title":"Keystoneml: Optimizing pipelines for large-scale advanced analytics. ICDE","author":"Sparks Evan R","year":"2017","unstructured":"Evan R Sparks, Shivaram Venkataraman, Tomer Kaftan, Michael J Franklin, and Benjamin Recht. Keystoneml: Optimizing pipelines for large-scale advanced analytics. ICDE (2017)."},{"key":"e_1_2_2_51_1","volume-title":"Responsible Data Management. Commun. ACM","author":"Stoyanovich Julia","year":"2022","unstructured":"Julia Stoyanovich, Bill Howe, Serge Abiteboul, H.V. Jagadish, and Sebastian Schelter. Responsible Data Management. Commun. ACM (2022)."},{"key":"e_1_2_2_52_1","first-page":"2","volume":"27","author":"Subramanian Subbu N.","year":"1998","unstructured":"Subbu N. Subramanian and Shivakumar Venkataraman. Cost-Based Optimization of Decision Support Queries Using Transient-Views. SIGMOD Record 27, 2 (1998).","journal-title":"Decision Support Queries Using Transient-Views. SIGMOD Record"},{"key":"e_1_2_2_53_1","volume-title":"Bernd Bischl, and Luis Torgo. OpenML: networked science in machine learning. KDD","author":"Vanschoren Joaquin","year":"2014","unstructured":"Joaquin Vanschoren, Jan N Van Rijn, Bernd Bischl, and Luis Torgo. OpenML: networked science in machine learning. KDD (2014)."},{"key":"e_1_2_2_54_1","volume-title":"The What-If Tool: Interactive Probing of Machine Learning Models","author":"Wexler James","year":"2019","unstructured":"James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viegas, and Jimbo Wilson. The What-If Tool: Interactive Probing of Machine Learning Models. IEEE Transactions on Visualization and Computer Graphics (2019)."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/3297753.3297763"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457566"},{"key":"e_1_2_2_57_1","unstructured":"CleanML benchmark. https:\/\/github.com\/chu-data-lab\/CleanML."},{"key":"e_1_2_2_58_1","unstructured":"FairPreprocessing benchmark. https:\/\/github.com\/sumonbis\/FairPreprocessing\/tree\/c644dd38615f34dba39320397fb00d5509602864\/benchmark."},{"key":"e_1_2_2_59_1","unstructured":"Permutation Feature Importance. https:\/\/scikit-learn.org\/stable\/modules\/permutation_importance.html."},{"key":"e_1_2_2_60_1","volume-title":"https:\/\/github.com\/schelterlabs\/jenga\/blob\/c219c645c664d2e81b7dfab2c51262e64e20f4ab\/src\/jenga\/basis.py#L25","author":"Jenga Task","unstructured":"Task abstraction in Jenga. https:\/\/github.com\/schelterlabs\/jenga\/blob\/c219c645c664d2e81b7dfab2c51262e64e20f4ab\/src\/jenga\/basis.py#L25."},{"key":"e_1_2_2_61_1","unstructured":"Anaconda.com. 2020. The State of Data Science. https:\/\/www.anaconda.com\/state-of-data-science-2020."},{"key":"e_1_2_2_62_1","volume-title":"Ray: A distributed framework for emerging {AI} applications. OSDI","author":"Moritz Philipp","year":"2018","unstructured":"Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Pual, and Michael Jordan. Ray: A distributed framework for emerging {AI} applications. OSDI (2018)."},{"key":"e_1_2_2_63_1","unstructured":"Linea.py. 2022. Move fast from data science prototype to pipeline. https:\/\/lineapy.org."},{"key":"e_1_2_2_64_1","unstructured":"Databricks. 2022. Mlflow recipes. https:\/\/www.mlflow.org\/docs\/latest\/recipes.html."},{"key":"e_1_2_2_65_1","unstructured":"Ray. 2022. Ray Dataset API. https:\/\/docs.ray.io\/en\/latest\/data\/api\/dataset.html."},{"key":"e_1_2_2_66_1","volume-title":"Retiring Adult: New Datasets for Fair Machine Learning. NeurIPS","author":"Ding Frances","year":"2018","unstructured":"Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt. Retiring Adult: New Datasets for Fair Machine Learning. NeurIPS (2018)."},{"key":"e_1_2_2_67_1","unstructured":"Cardio Data Set. 2022. Cardio Kaggle Dataset. https:\/\/www.kaggle.com\/datasets\/mdshamimrahman\/cardio-data-set."},{"key":"e_1_2_2_68_1","volume-title":"Can Foundation Models Wrangle Your Data?. VLDB","author":"Narayan Avanika","year":"2022","unstructured":"Avanika Narayan, Ines Chami, Laurel Orr, and Christopher R\u00e9. Can Foundation Models Wrangle Your Data?. VLDB (2022)."},{"key":"e_1_2_2_69_1","volume-title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers, and Iryna Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP (2019)."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589273","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589273","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:54Z","timestamp":1750182534000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589273"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,13]]},"references-count":69,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,13]]}},"alternative-id":["10.1145\/3589273"],"URL":"https:\/\/doi.org\/10.1145\/3589273","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,13]]}}}