{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T11:53:16Z","timestamp":1780487596196,"version":"3.54.1"},"reference-count":15,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T00:00:00Z","timestamp":1705363200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Columbus State University","award":["30177"],"award-info":[{"award-number":["30177"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>We have designed a real-world smart building energy fault detection (SBFD) system on a cloud-based Databricks workspace, a high-performance computing (HPC) environment for big-data-intensive applications powered by Apache Spark. By avoiding a Smart Building Diagnostics as a Service approach and keeping a tightly centralized design, the rapid development and deployment of the cloud-based SBFD system was achieved within one calendar year. Thanks to Databricks\u2019 built-in scheduling interface, a continuous pipeline of real-time ingestion, integration, cleaning, and analytics workflows capable of energy consumption prediction and anomaly detection was implemented and deployed in the cloud. The system currently provides fault detection in the form of predictions and anomaly detection for 96 buildings on an active military installation. The system\u2019s various jobs all converge within 14 min on average. It facilitates the seamless interaction between our workspace and a cloud data lake storage provided for secure and automated initial ingestion of raw data provided by a third party via the Secure File Transfer Protocol (SFTP) and BLOB (Binary Large Objects) file system secure protocol drivers. With a powerful Python binding to the Apache Spark distributed computing framework, PySpark, these actions were coded into collaborative notebooks and chained into the aforementioned pipeline. The pipeline was successfully managed and configured throughout the lifetime of the project and is continuing to meet our needs in deployment. In this paper, we outline the general architecture and how it differs from previous smart building diagnostics initiatives, present details surrounding the underlying technology stack of our data pipeline, and enumerate some of the necessary configuration steps required to maintain and develop this big data analytics application in the cloud.<\/jats:p>","DOI":"10.3390\/computers13010023","type":"journal-article","created":{"date-parts":[[2024,1,16]],"date-time":"2024-01-16T04:03:30Z","timestamp":1705377810000},"page":"23","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Cloud-Based Infrastructure and DevOps for Energy Fault Detection in Smart Buildings"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-1161-4777","authenticated-orcid":false,"given":"Kaleb","family":"Horvath","sequence":"first","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed Riduan","family":"Abid","sequence":"additional","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2731-3237","authenticated-orcid":false,"given":"Thomas","family":"Merino","sequence":"additional","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryan","family":"Zimmerman","sequence":"additional","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2801-1962","authenticated-orcid":false,"given":"Yesem","family":"Peker","sequence":"additional","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shamim","family":"Khan","sequence":"additional","affiliation":[{"name":"TSYS School of Computer Science, Turner College of Business, Columbus State University, Columbus, GA 31907, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,1,16]]},"reference":[{"key":"ref_1","first-page":"297","article-title":"Machine learning for Energy Consumption Prediction and Scheduling in Smart Buildings","volume":"2","author":"Bourhnane","year":"2020","journal-title":"Spring Nat. Appl. Sci. J."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Mohamed, N., Lazarova-Molnar, S., and Al-Jaroodi, J. (2016, January 27\u201329). SBDaaS: Smart Building Diagnostics as a Service on the Cloud. Proceedings of the 2016 2nd International Conference on Intelligent Green Building and Smart Grid (IBSG), Prague, Czech Republic.","DOI":"10.1109\/IGBSG.2016.7539417"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Stamatescu, I., Bolboaca, V., and Stamatescu, G. (2018, January 4\u20137). Distributed Monitoring of Smart Buildings with Cloud Backend Infrastructure. Proceedings of the 2018 International Conference on Control, Decision and Information Technologies (CoDIT\u201918), Orlando, FL, USA.","DOI":"10.1109\/CoDIT.2018.8394917"},{"key":"ref_4","first-page":"32","article-title":"Big data processing for smart grids","volume":"10","author":"Benhaddou","year":"2015","journal-title":"IADIS Int. J. Comput. Sci. Inf. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Morisio, M., Torchiano, M., and Jedlitschka, A. (2020). Product-Focused Software Process Improvement, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-030-64148-1"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2852082","article-title":"Supporting the design of machine learning workflows with a recommendation system","volume":"6","author":"Jannach","year":"2016","journal-title":"ACM Trans. Interact. Intell. Syst. (TiiS)"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1145\/98163.98169","article-title":"Distributed file systems: Concepts and examples","volume":"22","author":"Levy","year":"1990","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_8","unstructured":"Vangoor, B.K.R., Tarasov, V., and Zadok, E. (March, January 27). To FUSE or not to FUSE: Performance of User-Space file systems. Proceedings of the 15th USENIX Conference on File and Storage Technologies (FAST\u201917), Santa Clara, CA, USA."},{"key":"ref_9","unstructured":"Ravat, F., and Zhao, Y. (2019, January 26\u201329). Data lakes: Trends and perspectives. Proceedings of the Database and Expert Systems Applications: 30th International Conference, DEXA 2019, Linz, Austria. Proceedings, Part I 30."},{"key":"ref_10","unstructured":"Chaimov, N., Malony, A., Canon, S., Iancu, C., Ibrahim, K.Z., and Srinivasan, J. (June, January 31). Scaling Spark on HPC systems. Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing, Kyoto, Japan."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shvachko, K., Kuang, H., Radia, S., and Chansler, R. (2010, January 3\u20137). The hadoop distributed file system. Proceedings of the 2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST), Incline Village, NV, USA.","DOI":"10.1109\/MSST.2010.5496972"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Camacho-Rodr\u00edguez, J., Chauhan, A., Gates, A., Koifman, E., O\u2019Malley, O., Garg, V., Haindrich, Z., Shelukhin, S., Jayachandran, P., and Seth, S. (July, January 30). Apache hive: From mapreduce to enterprise-grade big data warehousing. Proceedings of the 2019 International Conference on Management of Data, Amsterdam, The Netherlands.","DOI":"10.1145\/3299869.3314045"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1007\/s41060-016-0027-9","article-title":"Big data analytics on Apache Spark","volume":"1","author":"Salloum","year":"2016","journal-title":"Int. J. Data Sci. Anal."},{"key":"ref_14","unstructured":"Erich, F., Amrit, C., and Daneva, M. (2014). Report: DevOps Literature Review, National Institute of Advanced Industrial Science and Technology."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1145\/2934664","article-title":"Apache Spark: A unified engine for big data processing","volume":"59","author":"Zaharia","year":"2016","journal-title":"Commun. ACM"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/1\/23\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:47:33Z","timestamp":1760104053000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/1\/23"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,16]]},"references-count":15,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["computers13010023"],"URL":"https:\/\/doi.org\/10.3390\/computers13010023","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,16]]}}}