{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,14]],"date-time":"2025-06-14T04:09:07Z","timestamp":1749874147647,"version":"3.41.0"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2017,1,4]],"date-time":"2017-01-04T00:00:00Z","timestamp":1483488000000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Computing"],"published-print":{"date-parts":[[2017,8]]},"DOI":"10.1007\/s00607-016-0534-5","type":"journal-article","created":{"date-parts":[[2017,1,4]],"date-time":"2017-01-04T14:15:45Z","timestamp":1483539345000},"page":"765-790","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Energy prediction of CUDA application instances using dynamic regression models"],"prefix":"10.1007","volume":"99","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1718-2635","authenticated-orcid":false,"given":"R. S.","family":"Rejitha","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shajulin","family":"Benedict","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Suja A.","family":"Alex","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shany","family":"Infanto","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2017,1,4]]},"reference":[{"key":"534_CR1","doi-asserted-by":"crossref","unstructured":"Alcides F, Bruno C (2013) miniumGPU: An Intelligent Framework for GPU Programming, Facing the Multicore-Challenge III, Volume 7686 of the series Lecture Notes in Computer Science, pp 96\u2013107","DOI":"10.1007\/978-3-642-35893-7_9"},{"key":"534_CR2","unstructured":"Ashwin MA, Mayank D, Wu CF (2016) CampProf: A Visual Performance Analysis Tool for Memory Bound GPU Kernels, in https:\/\/vtechworks.lib.vt.edu\/bitstream\/handle\/10919\/19729\/CampProf-TechReport.pdf?sequence=3&isAllowed=y . Accessed on 10 July 2016"},{"key":"534_CR3","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1007\/s11227-005-2335-z","volume":"34","author":"DA Bacigalupo","year":"2005","unstructured":"Bacigalupo DA, Jarvis SA, He L, Spooner DP, Dillenberger DN, Nudd GR (2005) An investigation into the application of different performance prediction methods to distributed enterprise applications. J Supercomput 34:2","journal-title":"J Supercomput"},{"key":"534_CR4","doi-asserted-by":"crossref","unstructured":"Barnes BJ, Rountree B, Lowenthal DK, Reeves J, de Supinski B, Schulz M (2008) A regression-based approach to scalability prediction. In: 22nd annual international conference on supercomputing","DOI":"10.1145\/1375527.1375580"},{"key":"534_CR5","doi-asserted-by":"publisher","unstructured":"Benedict S, Rejitha RS, Phillip G, Prodan R, Fahringer T (2015) Energy prediction of OpenMP applications using random forest modeling approach. In: iWAPT2015 @ IPDPS, pp 1251\u20131260. doi: 10.1109\/IPDPSW.2015.12","DOI":"10.1109\/IPDPSW.2015.12"},{"key":"534_CR6","doi-asserted-by":"publisher","unstructured":"Boyer M, Jiayuan M, Kumaran K (2013) Improving GPU performance prediction with data transfer modeling. In: 2013 IEEE 27th international parallel and distributed processing symposium workshops & PhD forum (IPDPSW), pp 1097\u20131106. doi: 10.1109\/IPDPSW.2013.236","DOI":"10.1109\/IPDPSW.2013.236"},{"key":"534_CR7","doi-asserted-by":"crossref","unstructured":"Brehm J, Worley P (1997) Performance prediction for complex parallel applications. In: 11th international symposium on parallel processing","DOI":"10.1109\/IPPS.1997.580884"},{"key":"534_CR8","unstructured":"Cao J, Jarvis SA, Spoon DP, Turner JD, Kerbyson DJ, Nudd GR (2002) Performance prediction technology for agent-based resource management in grid environments. In: 16th international parallel and distributed processing symposium, DC, USA, Washington"},{"key":"534_CR9","doi-asserted-by":"crossref","unstructured":"Carrington L, Snavely A, Wolter N (2006) A performance prediction framework for scientific applications. Future Gener Comput Syst 22(3)","DOI":"10.1016\/j.future.2004.11.019"},{"key":"534_CR10","unstructured":"CUDA application catalog, in http:\/\/www.nvidia.com\/content\/gpu-applications\/PDF\/gpu-applications-catalog.pdf . Accessed in Nov 2016"},{"issue":"10","key":"534_CR11","doi-asserted-by":"publisher","first-page":"3787","DOI":"10.1007\/s11227-015-1467-z","volume":"71","author":"T Elizabeth","year":"2015","unstructured":"Elizabeth T, Nathan C, David AP, John B, Barry IP, Dave H (2015) Parallel cuda implementation of conflict detection for application to airspace deconfliction. J Supercomput 71(10):3787\u20133810. doi: 10.1007\/s11227-015-1467-z","journal-title":"J Supercomput"},{"key":"534_CR12","doi-asserted-by":"publisher","unstructured":"Filipovic J, Benkner S (2015) OpenCL kernel fusion for GPU, Xeon Phi and CPU. In: Proceedings of IEEE international symposium on computer architecture and high performance computing (SBAC-PAD), Florianopolis, Brazil, October 2015, IEEE Computer Society. doi: 10.1109\/SBAC-PAD.2015.29","DOI":"10.1109\/SBAC-PAD.2015.29"},{"key":"534_CR13","doi-asserted-by":"publisher","unstructured":"Haifeng W, Yunpeng C. Predicting power consumption of GPUs with fuzzy wavelet neural networks. Parallel Comput 44:18\u201336. doi: 10.1016\/j.parco.2015.02.002","DOI":"10.1016\/j.parco.2015.02.002"},{"issue":"3","key":"534_CR14","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1145\/1555815.1555775","volume":"37","author":"S Hong","year":"2009","unstructured":"Hong S, Kim H (2009) An analytical model for a GPU architecture with memory-level and thread-level parallelism awareness. SIGARCH Com-put Archit News 37(3):152\u2013163. doi: 10.1145\/1555815.1555775","journal-title":"SIGARCH Com-put Archit News"},{"issue":"12","key":"534_CR15","doi-asserted-by":"crossref","first-page":"1374","DOI":"10.1109\/12.817403","volume":"48","author":"MA Iverson","year":"1999","unstructured":"Iverson MA, Ozguner F, Potter L (1999) Statistical prediction of task execution times through analytic benchmarking for scheduling in a heterogeneous environment. IEEE Trans Comput 48(12):1374\u20131379","journal-title":"IEEE Trans Comput"},{"key":"534_CR16","doi-asserted-by":"publisher","unstructured":"Jaewoong S, Aniruddha D, Hyesoon K, Richard V. A performance analysis framework for identifying potential benefits in GPGPU applications. In: Proceedings of the 17th ACM SIGPLAN symposium on principles and practice of parallel programming (PPoPP \u201912), ACM, New York, NY, USA, 11\u201322. doi: 10.1145\/2145816.2145819","DOI":"10.1145\/2145816.2145819"},{"key":"534_CR17","doi-asserted-by":"crossref","unstructured":"Kapadia N, Fortes J, Brodley C (1999) Predictive application-performance modeling in a computational grid environment. In: 8th international symposium on high performance distributed computing","DOI":"10.1109\/HPDC.1999.805281"},{"key":"534_CR18","first-page":"547","volume":"2011","author":"Y Kubota","year":"2011","unstructured":"Kubota Y, Takahashi D (2011) Optimization of Sparse Matrix-Vector Multiplication by Auto Selecting Storage Schemes on GPU. LNCS, in ICCSA 2011:547\u2013561","journal-title":"LNCS, in ICCSA"},{"key":"534_CR19","unstructured":"Li H, Groep D, Templon J, Wolters L (2004) Predicting job start times on clusters. In: International symposium on cluster computing and the grid"},{"key":"534_CR20","doi-asserted-by":"publisher","unstructured":"Monakov A, Lokhmotov A, Avetisyan A (2010) Automatically tuning sparse matrix-vector multiplication for GPU architectures. LNCS, Springer, pp 111\u2013125. doi: 10.1007\/978-3-642-11515-8_10","DOI":"10.1007\/978-3-642-11515-8_10"},{"key":"534_CR21","doi-asserted-by":"crossref","unstructured":"Nadeem F, Murtaza Y, Prodan R, Fahringer T (2006) International conference on e-science and grid computing. In: Soft benchmarks-based application performance prediction using a minimum training set, Amsterdam, Netherlands","DOI":"10.1109\/E-SCIENCE.2006.261155"},{"key":"534_CR22","unstructured":"Peppher project, http:\/\/www.peppher.edu . Accessed in Nov 2016"},{"key":"534_CR23","doi-asserted-by":"publisher","unstructured":"Shajulin B, Rejitha RS, Alex SA. Energy and performance prediction of CUDA applications using dynamic regression models. In: Proceedings of the 9th India software engineering conference (ISEC \u201916). ACM, New York, NY, USA, 37\u201347. doi: 10.1145\/2856636.2856643","DOI":"10.1145\/2856636.2856643"},{"issue":"10","key":"534_CR24","doi-asserted-by":"publisher","first-page":"1370","DOI":"10.1016\/j.jpdc.2008.05.014","volume":"68","author":"C Shuai","year":"2008","unstructured":"Shuai C, Michael B, Jiayuan M, David T, Sheaffer JW, Kevin S (2008) A performance study of general-purpose applications on graphics processors using CUDA. J Parallel Distrib Comput 68(10):1370\u20131380. doi: 10.1016\/j.jpdc.2008.05.014","journal-title":"J Parallel Distrib Comput"},{"key":"534_CR25","doi-asserted-by":"publisher","unstructured":"Shuaiwen Song, Chunyi Su, Rountree B., Cameron K.W., \u201cA Simplified and Accurate Model of Power-Performance Efficiency on Emergent GPU Architectures,\u201d 2013 IEEE 27th International Symposium on Parallel & Distributed Processing (IPDPS), pp. 673\u2013686, 20\u201324 May 2013. doi: 10.1109\/IPDPS.2013.73","DOI":"10.1109\/IPDPS.2013.73"},{"key":"534_CR26","doi-asserted-by":"publisher","unstructured":"Sierra-Canto X, Madera-Ramirez F, Uc-Cetina V (2010) Parallel Training of a Back-Propagation Neural Network Using CUDA,\u201d 2010 Ninth International Conference on in Machine Learning and Applications (ICMLA), pp. 307\u2013312, 12\u201314 Dec. doi: 10.1109\/ICMLA.2010.52","DOI":"10.1109\/ICMLA.2010.52"},{"issue":"9","key":"534_CR27","doi-asserted-by":"crossref","first-page":"1007","DOI":"10.1016\/j.jpdc.2004.06.008","volume":"64","author":"W Smith","year":"2004","unstructured":"Smith W, Foster I, Taylor V (2004) Predicting application run times with historical information. J Parallel Distributed Comput 64(9):1007\u20131016","journal-title":"J Parallel Distributed Comput"},{"key":"534_CR28","doi-asserted-by":"crossref","unstructured":"Snavely A, Carrington L, Wolter N, Labarta J, Badia R, Purkayastha A (2002) A framework for performance modeling and prediction,\u201d in Supercomputing Conference","DOI":"10.1109\/SC.2002.10004"},{"issue":"2","key":"534_CR29","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1007\/s10586-007-0039-2","volume":"11","author":"S Sodhi","year":"2008","unstructured":"Sodhi S, Subhlok J, Xu Q (2008) Performance prediction with skeletons. Cluster Comput 11(2):151\u2013165","journal-title":"Cluster Comput"},{"key":"534_CR30","unstructured":"R. Susukita, H. Ando, M. Aoyagi, H. Honda, Y. Inadomi, K. Inoue, S. Ishizuki, Y. Kimura, H. Komatsu, M. Kurokawa, K. J. Murakami, H. Shibamura, S. Yamamura and Y. Yu, \u201cPerformance prediction of large-scale parallell system and application using macro-level simulation,\u201d in Supercomputing Conference, 2008"},{"key":"534_CR31","doi-asserted-by":"crossref","unstructured":"Takefusa A, Tatebe O, Matsuoka S, Morita Y (2003) Performance analysis of scheduling and replication algorithms on Grid Datafarm architecture for high-energy physics applications,\u201d in 12th International Symposium on High Performance Distributed Computing","DOI":"10.1109\/HPDC.2003.1210014"},{"key":"534_CR32","doi-asserted-by":"crossref","unstructured":"Taylor V, Wu X, Geisler J, Stevens R (2002) \u201cUsing Kernel Couplings to Predict Parallel Application Performance,\u201d in 11th IEEE International Symposium on High Performance Distributed Computing. DC, USA, Washington","DOI":"10.1109\/HPDC.2002.1029910"},{"key":"534_CR33","doi-asserted-by":"publisher","unstructured":"Tingxing D, Dobrev V, Kolev T, Rieben R, Tomov S, Dongarra J (2014) A Step towards Energy Efficient Computing: Redesigning a Hydrodynamic Application on CPU-GPU. In Parallel and Distributed Processing Symposium, 2014 IEEE 28th International, pp. 972\u2013981, 19\u201323. doi: 10.1109\/IPDPS.2014.103","DOI":"10.1109\/IPDPS.2014.103"},{"key":"534_CR34","doi-asserted-by":"crossref","unstructured":"Tirado-Ramos A, Tsouloupas G, Dikaiakos M, Sloot P (2005) Grid Resource Selection by Application Benchmarking for Computational Haemodynamics Applications. In: International Conference on Computational Science, 2005","DOI":"10.1007\/11428831_66"},{"key":"534_CR35","doi-asserted-by":"publisher","unstructured":"Usman D, Johan E, Christoph WK (2011) Auto-tuning SkePU: a multi-backend skeleton programming framework for multi-GPU systems. In: Proceedings of the 4th international workshop on multicore software engineering (IWMSE\u201911), ACM, New York, NY, USA, pp 25\u201332. doi: 10.1145\/1984693.1984697","DOI":"10.1145\/1984693.1984697"},{"key":"534_CR36","doi-asserted-by":"publisher","unstructured":"Vedran M, Martina HD, Nataa HB (2014) Optimizing ELARS Algorithms using NVIDIA CUDA Heterogeneous Parallel Programming Platform\u201d. ICT Innovations 2014, pp. 135\u2013144. doi: 10.1007\/978-3-319-09879-1_14","DOI":"10.1007\/978-3-319-09879-1_14"},{"key":"534_CR37","doi-asserted-by":"crossref","unstructured":"Vraalsen F, Aydt RA, Mendes CL, Reed DA (2001) Performance Contracts: Predicting and Monitoring Grid Application Behavior,\u201d in 2nd International Workshop on Grid Computing","DOI":"10.1007\/3-540-45644-9_15"},{"key":"534_CR38","doi-asserted-by":"crossref","unstructured":"Wu Y, Liu L, Mao J, Yang G, Zheng W (2007) An analytical model for performance evaluation in a computational grid,\u201d in 3rd workshop on High performance computing in China","DOI":"10.1145\/1375783.1375813"},{"key":"534_CR39","unstructured":"Wu X, Taylor V, Paris J (2006) A Web-based Prophesy Automated Performance Modeling System,\u201d in International Conference on Web Technologies, Applications and Services"},{"key":"534_CR40","doi-asserted-by":"publisher","unstructured":"Xiangzheng S, Yunquan Z, Ting W, Xianyi Z, Liang Y, Li R (2011) Optimizing SpMV for Diagonal Sparse Matrices on GPU,\u201d 2011 International Conference on in Parallel Processing (ICPP), pp 492\u2013501. doi: 10.1109\/ICPP.2011.53","DOI":"10.1109\/ICPP.2011.53"},{"key":"534_CR41","unstructured":"Yang LT, Ma X, Mueller F (2005) Cross-Platform Performance Prediction of Parallel Applications Using Partial Execution,\u201d in Supercomputing"},{"issue":"4","key":"534_CR42","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1016\/j.peva.2005.01.008","volume":"63","author":"EJH Yero","year":"2006","unstructured":"Yero EJH, Henriques MAA (2006) Contention-sensitive static performance prediction for parallel distributed applications. Perform Evaluat 63(4):265\u2013277","journal-title":"Perform Evaluat"},{"key":"534_CR43","doi-asserted-by":"crossref","unstructured":"Yooseong K, Shrivastava A (2011) CuMAPz: a tool to analyze memory access patterns in CUDA\u201d, in Design Automation Conference (DAC), 2011 48th ACM\/EDAC\/IEEE, pp 128\u2013133","DOI":"10.1145\/2024724.2024754"}],"container-title":["Computing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s00607-016-0534-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s00607-016-0534-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s00607-016-0534-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,14]],"date-time":"2025-06-14T00:17:28Z","timestamp":1749860248000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s00607-016-0534-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,1,4]]},"references-count":43,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2017,8]]}},"alternative-id":["534"],"URL":"https:\/\/doi.org\/10.1007\/s00607-016-0534-5","relation":{},"ISSN":["0010-485X","1436-5057"],"issn-type":[{"type":"print","value":"0010-485X"},{"type":"electronic","value":"1436-5057"}],"subject":[],"published":{"date-parts":[[2017,1,4]]}}}