{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,2]],"date-time":"2025-11-02T07:06:04Z","timestamp":1762067164824,"version":"build-2065373602"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,8]]},"abstract":"<jats:p>We live in the gilded age of data-driven computing. With public clouds offering virtually unlimited amounts of compute and storage, enterprises collecting data about every aspect of their businesses, and advances in analytics and machine learning technologies, data driven decision making is now timely, cost-effective, and therefore, pervasive. Alas, only a handful of power users can wield today's powerful data engineering tools. For one thing, most solutions require knowledge of specific programming interfaces or libraries. Furthermore, running them requires complex configurations and knowledge of the underlying cloud for cost-effectiveness.<\/jats:p>\n          <jats:p>We decided that a fundamental redesign is in order to democratize data engineering for the masses at cloud scale. The result is Informatica Cloud Data Integration - Elastic (CDI-E). Since the early 1990s, Informatica has been a pioneer and industry leader in building no-code data engineering tools. Non-experts can express complex data engineering tasks using a graphical user interface (GUI). Informatica CDI-E is built to incorporate the simplicity of GUI in the design layer with an elastic and highly scalable run time to handle data in any format without little to no user input using automated optimizations. Users upload their data to the cloud in any format and can immediately use them in conjunction with their data management and analytic tools of choice using CDI-E GUI. Implementation began in the Spring of 2017, and Informatica CDI-E has been generally available since the Summer of 2019. Today, CDI-E is used in production by a growing number of small and large enterprises to make sense of data in arbitrary formats.<\/jats:p>\n          <jats:p>In this paper, we describe the architecture of Informatica CDI-E and its novel no-code data engineering interface. The paper highlights some of the key features of CDI-E: simplicity without loss in productivity and extreme elasticity. It concludes with lessons we learned and an outlook of the future.<\/jats:p>","DOI":"10.14778\/3554821.3554825","type":"journal-article","created":{"date-parts":[[2022,9,29]],"date-time":"2022-09-29T22:28:39Z","timestamp":1664490519000},"page":"3319-3331","source":"Crossref","is-referenced-by-count":2,"title":["CDI-E"],"prefix":"10.14778","volume":"15","author":[{"given":"Prakash","family":"Das","sequence":"first","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shivangi","family":"Srivastava","sequence":"additional","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Valentin","family":"Moskovich","sequence":"additional","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anmol","family":"Chaturvedi","sequence":"additional","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anant","family":"Mittal","sequence":"additional","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongqin","family":"Xiao","sequence":"additional","affiliation":[{"name":"Informatica"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mosharaf","family":"Chowdhury","sequence":"additional","affiliation":[{"name":"University of Michigan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,9,29]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n.d.]. Amazon Redshift. https:\/\/aws.amazon.com\/redshift\/.  [n.d.]. Amazon Redshift. https:\/\/aws.amazon.com\/redshift\/."},{"key":"e_1_2_1_2_1","unstructured":"[n.d.]. Amazon Simple Storage Service (S3). https:\/\/aws.amazon.com\/s3.  [n.d.]. Amazon Simple Storage Service (S3). https:\/\/aws.amazon.com\/s3."},{"key":"e_1_2_1_3_1","unstructured":"[n.d.]. Apache Avro. https:\/\/avro.apache.org.  [n.d.]. Apache Avro. https:\/\/avro.apache.org."},{"key":"e_1_2_1_4_1","unstructured":"[n.d.]. Apache Cassandra. https:\/\/cassandra.apache.org.  [n.d.]. Apache Cassandra. https:\/\/cassandra.apache.org."},{"key":"e_1_2_1_5_1","unstructured":"[n.d.]. Apache Hadoop. https:\/\/www.hadoop.apache.org.  [n.d.]. Apache Hadoop. https:\/\/www.hadoop.apache.org."},{"key":"e_1_2_1_6_1","unstructured":"[n.d.]. Apache Hive. https:\/\/www.hive.apache.org.  [n.d.]. Apache Hive. https:\/\/www.hive.apache.org."},{"key":"e_1_2_1_7_1","unstructured":"[n.d.]. Apache ORC. https:\/\/orc.apache.org.  [n.d.]. Apache ORC. https:\/\/orc.apache.org."},{"key":"e_1_2_1_8_1","unstructured":"[n.d.]. Apache Parquet. https:\/\/parquet.apache.org.  [n.d.]. Apache Parquet. https:\/\/parquet.apache.org."},{"key":"e_1_2_1_9_1","unstructured":"[n.d.]. Apache Spark. https:\/\/spark.apache.org.  [n.d.]. Apache Spark. https:\/\/spark.apache.org."},{"key":"e_1_2_1_10_1","unstructured":"[n.d.]. AWS Glue. https:\/\/aws.amazon.com\/glue\/.  [n.d.]. AWS Glue. https:\/\/aws.amazon.com\/glue\/."},{"key":"e_1_2_1_11_1","unstructured":"[n.d.]. AWS Lambda. https:\/\/docs.aws.amazon.com\/lambda\/.  [n.d.]. AWS Lambda. https:\/\/docs.aws.amazon.com\/lambda\/."},{"key":"e_1_2_1_12_1","unstructured":"[n.d.]. Azure Blob Storage. https:\/\/docs.microsoft.com\/en-us\/azure\/storage\/blobs\/storage-blobs-overview.  [n.d.]. Azure Blob Storage. https:\/\/docs.microsoft.com\/en-us\/azure\/storage\/blobs\/storage-blobs-overview."},{"key":"e_1_2_1_13_1","unstructured":"[n.d.]. Azure Functions. https:\/\/azure.microsoft.com\/en-us\/services\/functions\/.  [n.d.]. Azure Functions. https:\/\/azure.microsoft.com\/en-us\/services\/functions\/."},{"key":"e_1_2_1_14_1","unstructured":"[n.d.]. Azure Logic Apps. https:\/\/azure.microsoft.com\/en-us\/services\/logic-apps\/#overview.  [n.d.]. Azure Logic Apps. https:\/\/azure.microsoft.com\/en-us\/services\/logic-apps\/#overview."},{"key":"e_1_2_1_15_1","unstructured":"[n.d.]. Azure Synapse Analytics. https:\/\/azure.microsoft.com\/en-us\/services\/synapse-analytics\/.  [n.d.]. Azure Synapse Analytics. https:\/\/azure.microsoft.com\/en-us\/services\/synapse-analytics\/."},{"key":"e_1_2_1_16_1","unstructured":"[n.d.]. BigQuery. https:\/\/cloud.google.com\/bigquery.  [n.d.]. BigQuery. https:\/\/cloud.google.com\/bigquery."},{"key":"e_1_2_1_17_1","unstructured":"[n.d.]. Cloud Functions. https:\/\/cloud.google.com\/functions.  [n.d.]. Cloud Functions. https:\/\/cloud.google.com\/functions."},{"key":"e_1_2_1_18_1","unstructured":"[n.d.]. Cloud Storage. https:\/\/cloud.google.com\/storage.  [n.d.]. Cloud Storage. https:\/\/cloud.google.com\/storage."},{"key":"e_1_2_1_19_1","unstructured":"[n.d.]. Data Factory. https:\/\/azure.microsoft.com\/en-us\/services\/data-factory\/.  [n.d.]. Data Factory. https:\/\/azure.microsoft.com\/en-us\/services\/data-factory\/."},{"key":"e_1_2_1_20_1","unstructured":"[n.d.]. Databricks. https:\/\/databricks.com.  [n.d.]. Databricks. https:\/\/databricks.com."},{"key":"e_1_2_1_21_1","unstructured":"[n.d.]. Extensible Markup Language (XML) - W3C. https:\/\/www.w3.org\/xml.  [n.d.]. Extensible Markup Language (XML) - W3C. https:\/\/www.w3.org\/xml."},{"key":"e_1_2_1_22_1","unstructured":"[n.d.]. Fast Healthcare Interoperability Resources. https:\/\/en.wikipedia.org\/wiki\/Fast_Healthcare_Interoperability_Resources.  [n.d.]. Fast Healthcare Interoperability Resources. https:\/\/en.wikipedia.org\/wiki\/Fast_Healthcare_Interoperability_Resources."},{"key":"e_1_2_1_23_1","unstructured":"[n.d.]. Health Insurance Portability and Accountability Act. https:\/\/www.hhs.gov\/hipaa.  [n.d.]. Health Insurance Portability and Accountability Act. https:\/\/www.hhs.gov\/hipaa."},{"key":"e_1_2_1_24_1","unstructured":"[n.d.]. HL7 International. https:\/\/hl7.org.  [n.d.]. HL7 International. https:\/\/hl7.org."},{"key":"e_1_2_1_25_1","unstructured":"[n.d.]. Informatica Cloud Application Integration. https:\/\/www.informatica.com\/products\/cloud-application-integration.html.  [n.d.]. Informatica Cloud Application Integration. https:\/\/www.informatica.com\/products\/cloud-application-integration.html."},{"key":"e_1_2_1_26_1","unstructured":"[n.d.]. Informatica's Cost Optimization Engine. https:\/\/www.informatica.com\/lp\/informaticas-cost-optimization-engine_4257.html.  [n.d.]. Informatica's Cost Optimization Engine. https:\/\/www.informatica.com\/lp\/informaticas-cost-optimization-engine_4257.html."},{"key":"e_1_2_1_27_1","unstructured":"[n.d.]. JavaScript Object Notation. https:\/\/www.json.org.  [n.d.]. JavaScript Object Notation. https:\/\/www.json.org."},{"key":"e_1_2_1_28_1","unstructured":"[n.d.]. Kubernetes Scheduler. https:\/\/www.kubernetes.io\/docs\/concepts\/scheduling-eviction\/kube-scheduler\/.  [n.d.]. Kubernetes Scheduler. https:\/\/www.kubernetes.io\/docs\/concepts\/scheduling-eviction\/kube-scheduler\/."},{"key":"e_1_2_1_29_1","unstructured":"[n.d.]. LVM. https:\/\/www.tecmint.com\/create-lvm-storage-in-linux\/.  [n.d.]. LVM. https:\/\/www.tecmint.com\/create-lvm-storage-in-linux\/."},{"key":"e_1_2_1_30_1","unstructured":"[n.d.]. LVM1. https:\/\/www.tecmint.com\/extend-and-reduce-lvms-in-linux\/.  [n.d.]. LVM1. https:\/\/www.tecmint.com\/extend-and-reduce-lvms-in-linux\/."},{"key":"e_1_2_1_31_1","unstructured":"[n.d.]. MongoDB. https:\/\/www.mongodb.com.  [n.d.]. MongoDB. https:\/\/www.mongodb.com."},{"key":"e_1_2_1_32_1","unstructured":"[n.d.]. Nucleus Research. https:\/\/nucleusresearch.com\/.  [n.d.]. Nucleus Research. https:\/\/nucleusresearch.com\/."},{"key":"e_1_2_1_33_1","unstructured":"[n.d.]. Nucleus Research Informatica. https:\/\/nucleusresearch.com\/research\/single\/roi-guidebook-informatica\/.  [n.d.]. Nucleus Research Informatica. https:\/\/nucleusresearch.com\/research\/single\/roi-guidebook-informatica\/."},{"key":"e_1_2_1_34_1","unstructured":"[n.d.]. presto. https:\/\/prestodb.io.  [n.d.]. presto. https:\/\/prestodb.io."},{"key":"e_1_2_1_35_1","unstructured":"[n.d.]. Running Spark on Kubernetes. https:\/\/www.spark.apache.org\/docs\/latest\/running-on-kubernetes.html.  [n.d.]. Running Spark on Kubernetes. https:\/\/www.spark.apache.org\/docs\/latest\/running-on-kubernetes.html."},{"key":"e_1_2_1_36_1","unstructured":"[n.d.]. SWIFT EDI Document Standard. https:\/\/www.edibasics.com\/edi-resources\/document-standards\/swift\/.  [n.d.]. SWIFT EDI Document Standard. https:\/\/www.edibasics.com\/edi-resources\/document-standards\/swift\/."},{"key":"e_1_2_1_37_1","unstructured":"[n.d.]. Taints and Tolerations. https:\/\/www.kubernetes.io\/docs\/concepts\/scheduling-eviction\/taint-and-toleration\/.  [n.d.]. Taints and Tolerations. https:\/\/www.kubernetes.io\/docs\/concepts\/scheduling-eviction\/taint-and-toleration\/."},{"key":"e_1_2_1_38_1","unstructured":"[n.d.]. TPC. https:\/\/www.tpc.org.  [n.d.]. TPC. https:\/\/www.tpc.org."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2903741"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002974.2002977"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_2_1_42_1","volume-title":"Ives","author":"Doan AnHai","year":"2012","unstructured":"AnHai Doan , Alon Y. Halevy , and Zachary G . Ives . 2012 . Principles of Data Integration. Morgan Kaufmann . https:\/\/research.cs.wisc.edu\/dibook\/ AnHai Doan, Alon Y. Halevy, and Zachary G. Ives. 2012. Principles of Data Integration. Morgan Kaufmann. https:\/\/research.cs.wisc.edu\/dibook\/"},{"volume-title":"Big data integration","author":"Dong Xin Luna","key":"e_1_2_1_43_1","unstructured":"Xin Luna Dong and Divesh Srivastava . 2013. Big data integration . In ICDE. IEEE , 1245--1248. Xin Luna Dong and Divesh Srivastava. 2013. Big data integration. In ICDE. IEEE, 1245--1248."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742795"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749432"},{"key":"e_1_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Jiawei Jiang Shaoduo Gan Yue Liu Fanlin Wang Gustavo Alonso Ana Klimovic Ankit Singla Wentao Wu and Ce Zhang. 2021. Towards demystifying serverless machine learning training. In ACM SIGMOD. 857--871.  Jiawei Jiang Shaoduo Gan Yue Liu Fanlin Wang Gustavo Alonso Ana Klimovic Ankit Singla Wentao Wu and Ce Zhang. 2021. Towards demystifying serverless machine learning training. In ACM SIGMOD. 857--871.","DOI":"10.1145\/3448016.3459240"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732977.2732995"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/543613.543644"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3386126"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3240491"},{"key":"e_1_2_1_51_1","volume-title":"Lambada: Interactive data analytics on cold data using serverless cloud infrastructure. In ACM SIGMOD. 115--130.","author":"M\u00fcller Ingo","year":"2020","unstructured":"Ingo M\u00fcller , Renato Marroqu\u00edn , and Gustavo Alonso . 2020 . Lambada: Interactive data analytics on cold data using serverless cloud infrastructure. In ACM SIGMOD. 115--130. Ingo M\u00fcller, Renato Marroqu\u00edn, and Gustavo Alonso. 2020. Lambada: Interactive data analytics on cold data using serverless cloud infrastructure. In ACM SIGMOD. 115--130."},{"key":"e_1_2_1_52_1","volume-title":"David DeWitt, and Samuel Madden.","author":"Perron Matthew","year":"2020","unstructured":"Matthew Perron , Raul Castro Fernandez , David DeWitt, and Samuel Madden. 2020 . Starling : A scalable query engine on cloud functions. In ACM SIGMOD. 131--141. Matthew Perron, Raul Castro Fernandez, David DeWitt, and Samuel Madden. 2020. Starling: A scalable query engine on cloud functions. In ACM SIGMOD. 131--141."},{"key":"e_1_2_1_53_1","unstructured":"Qifan Pu Shivaram Venkataraman and Ion Stoica. 2019. Shuffling fast and slow: Scalable analytics on serverless infrastructure. In USENIX NSDI. 193--206.  Qifan Pu Shivaram Venkataraman and Ion Stoica. 2019. Shuffling fast and slow: Scalable analytics on serverless infrastructure. In USENIX NSDI. 193--206."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-019-00589-2"},{"key":"e_1_2_1_55_1","volume-title":"numpywren: Serverless linear algebra. arXiv preprint arXiv:1810.09679","author":"Shankar Vaishaal","year":"2018","unstructured":"Vaishaal Shankar , Karl Krauth , Qifan Pu , Eric Jonas , Shivaram Venkataraman , Ion Stoica , Benjamin Recht , and Jonathan Ragan-Kelley . 2018. numpywren: Serverless linear algebra. arXiv preprint arXiv:1810.09679 ( 2018 ). Vaishaal Shankar, Karl Krauth, Qifan Pu, Eric Jonas, Shivaram Venkataraman, Ion Stoica, Benjamin Recht, and Jonathan Ragan-Kelley. 2018. numpywren: Serverless linear algebra. arXiv preprint arXiv:1810.09679 (2018)."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1987.10488205"},{"key":"e_1_2_1_57_1","first-page":"3","article-title":"Data Integration: The Current Status and the Way Forward","volume":"41","author":"Stonebraker Michael","year":"2018","unstructured":"Michael Stonebraker , Ihab F Ilyas , 2018 . Data Integration: The Current Status and the Way Forward . IEEE Data Engineering Bulletin 41 , 2 (2018), 3 -- 9 . Michael Stonebraker, Ihab F Ilyas, et al. 2018. Data Integration: The Current Status and the Way Forward. IEEE Data Engineering Bulletin 41, 2 (2018), 3--9.","journal-title":"IEEE Data Engineering Bulletin"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735508.2735514"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687609"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/988672.988711"},{"key":"e_1_2_1_61_1","volume-title":"Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)","author":"Zaharia Matei","year":"2012","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Tathagata Das , Ankur Dave , Justin Ma , Murphy McCauly , Michael J Franklin , Scott Shenker , and Ion Stoica . 2012 . Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12) . 15--28. Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). 15--28."},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing","author":"Zaharia Matei","year":"2010","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Michael J. Franklin , Scott Shenker , and Ion Stoica . 2010 . Spark: Cluster Computing with Working Sets . In Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing ( Boston, MA) (HotCloud'10). USENIX Association, USA, 10. Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2010. Spark: Cluster Computing with Working Sets. In Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (Boston, MA) (HotCloud'10). USENIX Association, USA, 10."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3554821.3554825","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:22:06Z","timestamp":1672226526000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3554821.3554825"}},"subtitle":["an elastic cloud service for data engineering"],"short-title":[],"issued":{"date-parts":[[2022,8]]},"references-count":62,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2022,8]]}},"alternative-id":["10.14778\/3554821.3554825"],"URL":"https:\/\/doi.org\/10.14778\/3554821.3554825","relation":{},"ISSN":["2150-8097"],"issn-type":[{"type":"print","value":"2150-8097"}],"subject":[],"published":{"date-parts":[[2022,8]]}}}