{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,1]],"date-time":"2025-07-01T12:48:21Z","timestamp":1751374101012,"version":"3.41.0"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2013,12,1]],"date-time":"2013-12-01T00:00:00Z","timestamp":1385856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2013,12]]},"abstract":"<jats:p>We present a novel code generation scheme for GPUs. Its key feature is the platform-aware generation of a heterogeneous pool of threads. This exposes more data-sharing opportunities among the concurrent threads and reduces the memory requirements that would otherwise exceed the capacity of the on-chip memory. Instead of the conventional strategy of focusing on exposing as much parallelism as possible, our scheme leverages on the phased nature of memory access patterns found in many applications that exhibit massive parallelism. We demonstrate the effectiveness of our code generation strategy on a computational systems biology application. This application consists of computing a Dynamic Bayesian Network (DBN) approximation of the dynamics of signalling pathways described as a system of Ordinary Differential Equations (ODEs). The approximation algorithm involves (i) sampling many (of the order of a few million) times from the set of initial states, (ii) generating trajectories through numerical integration, and (iii) storing the statistical properties of this set of trajectories in Conditional Probability Tables (CPTs) of a DBN via a prespecified discretization of the time and value domains. The trajectories can be computed in parallel. However, the intermediate data needed for computing them, as well as the entries for the CPTs, are too large to be stored locally. Our experiments show that the proposed code generation scheme scales well, achieving significant performance improvements on three realistic signalling pathways models. These results suggest how our scheme could be extended to deal with other applications involving systems of ODEs.<\/jats:p>","DOI":"10.1145\/2541228.2555311","type":"journal-article","created":{"date-parts":[[2014,1,14]],"date-time":"2014-01-14T13:39:57Z","timestamp":1389706797000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["GPU code generation for ODE-based applications with phased shared-data access patterns"],"prefix":"10.1145","volume":"10","author":[{"given":"Andrei","family":"Hagiescu","sequence":"first","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bing","family":"Liu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"R.","family":"Ramanathan","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sucheendra K.","family":"Palaniappan","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Cui","sequence":"additional","affiliation":[{"name":"Advanced Digital Science Centre, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bipasa","family":"Chattopadhyay","sequence":"additional","affiliation":[{"name":"University of North Carolina"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"P. S.","family":"Thiagarajan","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weng-Fai","family":"Wong","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Compilers: Principles, Techniques and Tools, 2 ed","author":"Aho A. V.","year":"2006","unstructured":"Aho , A. V. , Lam , M. S. , Sethi , R. , and Ullman , J. D . 2006 . Compilers: Principles, Techniques and Tools, 2 ed . Addison Wesley . Aho, A. V., Lam, M. S., Sethi, R., and Ullman, J. D. 2006. Compilers: Principles, Techniques and Tools, 2 ed. Addison Wesley."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1038\/ncb1497"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/1025127.1025992"},{"key":"e_1_2_1_4_1","first-page":"325","article-title":"Heterogeneous multicore parallel programming for graphics processing units. Sci","volume":"17","author":"Bodin F.","year":"2009","unstructured":"Bodin , F. and Bihan , S. 2009 . Heterogeneous multicore parallel programming for graphics processing units. Sci . Program. 17 , 4, 325 -- 336 . Bodin, F. and Bihan, S. 2009. Heterogeneous multicore parallel programming for graphics processing units. Sci. Program. 17, 4, 325--336.","journal-title":"Program."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1088\/1478-3967\/1\/3\/006"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2011.50"},{"key":"e_1_2_1_7_1","volume-title":"NVIDIA's Fermi: The First Complete GPU Computing Architecture. Retrived","author":"Glaskowsky P. N.","year":"2013","unstructured":"Glaskowsky , P. N. 2009. NVIDIA's Fermi: The First Complete GPU Computing Architecture. Retrived December 2, 2013 from http:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/P.Glaskowsky_NVIDIA's_Fermi-The_First_Complete_GPU.pdf. Glaskowsky, P. N. 2009. NVIDIA's Fermi: The First Complete GPU Computing Architecture. Retrived December 2, 2013 from http:\/\/www.nvidia.com\/content\/PDF\/fermi_white_papers\/P.Glaskowsky_NVIDIA's_Fermi-The_First_Complete_GPU.pdf."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jtbi.2008.01.006"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.52"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555775"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950365.1950409"},{"key":"e_1_2_1_12_1","unstructured":"Khronos. 2012. Khronos OpenCL. http:\/\/www.khronos.org\/opencl\/.  Khronos. 2012. Khronos OpenCL. http:\/\/www.khronos.org\/opencl\/."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature01254"},{"key":"e_1_2_1_14_1","unstructured":"Koller D. and Friedman N. 2009. Probabilistic Graphical Models: Principles and Techniques (Adaptive Computation and Machine Learning). MIT Press.   Koller D. and Friedman N. 2009. Probabilistic Graphical Models: Principles and Techniques (Adaptive Computation and Machine Learning). MIT Press."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1582710.1582711"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/bts166"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tcs.2011.01.021"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-03845-7_17"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1001059"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Maedo A. Ozaki Y. Sivakumaran S. Akiyama T. Urakubo H. Usami A. Sato M. Kaibuchi K. and Kuroda S. 2006. Ca&lt;sup;&gt;2&plus;&lt;\/sup;&gt;-independent phospholipase A2-dependent sustained Rho-kinase activation exhibits all-or-none response. Genes to Cells 11 1071--1083.  Maedo A. Ozaki Y. Sivakumaran S. Akiyama T. Urakubo H. Usami A. Sato M. Kaibuchi K. and Kuroda S. 2006. Ca&lt;sup;&gt;2&plus;&lt;\/sup;&gt;-independent phospholipase A2-dependent sustained Rho-kinase activation exhibits all-or-none response. Genes to Cells 11 1071--1083.","DOI":"10.1111\/j.1365-2443.2006.01001.x"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/272991.272995"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2010.41"},{"key":"e_1_2_1_24_1","unstructured":"NVIDIA. 2012. NVIDIA CUDA.  NVIDIA. 2012. NVIDIA CUDA."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2008.917757"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2007.01012.x"},{"volume-title":"Proceedings of the 11th International Conference on Computational Systems Biology (CMSB'13)","author":"Palaniappan S. K.","key":"e_1_2_1_27_1","unstructured":"Palaniappan , S. K. , Gyori , B. M. , Liu , B. , Hsu , D. , and Thiagarajan , P. S . 2013. Statistical model checking based calibration and analysis of bio-pathway models . In Proceedings of the 11th International Conference on Computational Systems Biology (CMSB'13) . Palaniappan, S. K., Gyori, B. M., Liu, B., Hsu, D., and Thiagarajan, P. S. 2013. Statistical model checking based calibration and analysis of bio-pathway models. In Proceedings of the 11th International Conference on Computational Systems Biology (CMSB'13)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1345206.1345220"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1375527.1375572"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1735688.1735697"},{"volume-title":"Proceedings of the 28th IEEE International Parallel & Distributed Processing Symposuim (IPDPS'10)","author":"Ye X.","key":"e_1_2_1_31_1","unstructured":"Ye , X. , Fan , D. , Lin , W. , Yuan , N. , and Ienne , P . 2010. High performance comparison-based sorting algorithm on many-core GPUs . In Proceedings of the 28th IEEE International Parallel & Distributed Processing Symposuim (IPDPS'10) . 1--10. Ye, X., Fan, D., Lin, W., Yuan, N., and Ienne, P. 2010. High performance comparison-based sorting algorithm on many-core GPUs. In Proceedings of the 28th IEEE International Parallel & Distributed Processing Symposuim (IPDPS'10). 1--10."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ic.2006.05.002"},{"volume-title":"Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA'11)","author":"Zhang Y.","key":"e_1_2_1_33_1","unstructured":"Zhang , Y. and Owens , J. D . 2011. A quantitative performance analysis model for GPU architectures . In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA'11) . IEEE Computer Society, Washington, DC, 382--393. Zhang, Y. and Owens, J. D. 2011. A quantitative performance analysis model for GPU architectures. In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA'11). IEEE Computer Society, Washington, DC, 382--393."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2541228.2555311","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2541228.2555311","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:35:01Z","timestamp":1750232101000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2541228.2555311"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,12]]},"references-count":32,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2013,12]]}},"alternative-id":["10.1145\/2541228.2555311"],"URL":"https:\/\/doi.org\/10.1145\/2541228.2555311","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2013,12]]},"assertion":[{"value":"2012-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}