{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T01:17:00Z","timestamp":1782782220705,"version":"3.54.5"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2025,3,24]],"date-time":"2025-03-24T00:00:00Z","timestamp":1742774400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,24]],"date-time":"2025-03-24T00:00:00Z","timestamp":1742774400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"University of Innsbruck and Medical University of Innsbruck"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Time-of-Flight (ToF) camera systems are increasingly capable of analyzing larger 3D spaces and providing more detailed and precise results. To increase the speed-to-solution during development, testing and validation of such systems, light propagation simulation is employed. One such simulation, RSim, was previously performed on single workstations, however, the increase in detail required for newer ToF hardware necessitates cluster-level parallelism in order to maintain an experiment latency which enables productive design work. Celerity is a high-level parallel API and runtime system for clusters of accelerators intended to simplify the development of domain science applications. It automatically manages data and work distribution, while also transparently enabling asynchronous compute and communication overlapping. In this paper, we present a use case study of porting the full RSim application to GPU clusters using the Celerity system. In order to improve scalability, a new parallelization scheme was employed for the core simulation task, and Celerity was extended with a high-level split constraints feature which enables this scheme. We present strong- and weak-scaling experiments for the resulting application on three accelerator clusters and up to 128 GPUs, and also evaluate the relative programming effort required to distribute the application on multiple GPUs using different APIs.<\/jats:p>","DOI":"10.1007\/s10766-025-00787-2","type":"journal-article","created":{"date-parts":[[2025,3,27]],"date-time":"2025-03-27T00:36:51Z","timestamp":1743035811000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Celerity-RSim: Porting Light Propagation Simulation to Accelerator Clusters Using a High-Level API"],"prefix":"10.1007","volume":"53","author":[{"given":"Peter","family":"Thoman","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Philipp","family":"Gschwandtner","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Facundo","family":"Molina Heredina","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Fahringer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,24]]},"reference":[{"key":"787_CR1","doi-asserted-by":"publisher","unstructured":"Afzal, A., Schmitt, C., Alhaddad, S., Grynko, Y., Teich, J., Forstner, J., Hannig, F.: Solving Maxwell\u2019s equations with modern c++ and sycl: A case study. In: 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP), pp.\u00a01\u20138 (2018). https:\/\/doi.org\/10.1109\/ASAP.2018.8445127","DOI":"10.1109\/ASAP.2018.8445127"},{"key":"787_CR2","doi-asserted-by":"publisher","unstructured":"Bauer, M., Treichler, S., Slaughter, E., Aiken, A.: Legion: expressing locality and independence with logical regions. In: SC\u201912: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pp. 1\u201311 (2012). https:\/\/doi.org\/10.1109\/SC.2012.71","DOI":"10.1109\/SC.2012.71"},{"key":"787_CR3","doi-asserted-by":"publisher","unstructured":"Bender, M., Demaine, E., Farach-Colton, M.: Cache-oblivious b-trees. In: Proceedings 41st Annual Symposium on Foundations of Computer Science, pp. 399\u2013409 (2000). https:\/\/doi.org\/10.1109\/SFCS.2000.892128","DOI":"10.1109\/SFCS.2000.892128"},{"key":"787_CR4","doi-asserted-by":"publisher","unstructured":"Breyer, M., Dai\u00df, G., Pfl\u00fcger, D.: Performance-portable distributed k-nearest neighbors using locality-sensitive hashing and sycl. In: Proceedings of the 9th International Workshop on OpenCL. IWOCL\u201921, Association for Computing Machinery, New York, NY, USA (2021). https:\/\/doi.org\/10.1145\/3456669.3456692,","DOI":"10.1145\/3456669.3456692"},{"issue":"9","key":"787_CR5","doi-asserted-by":"publisher","first-page":"9409","DOI":"10.1007\/s11227-022-05040-y","volume":"79","author":"M de Castro","year":"2023","unstructured":"de Castro, M., Santamaria-Valenzuela, I., Torres, Y., Gonzalez-Escribano, A., Llanos, D.R.: Epsilod: efficient parallel skeleton for generic iterative stencil computations in distributed GPUS. J. Supercomput. 79(9), 9409\u20139442 (2023). https:\/\/doi.org\/10.1007\/s11227-022-05040-y","journal-title":"J. Supercomput."},{"key":"787_CR6","doi-asserted-by":"publisher","unstructured":"Deakin, T., McIntosh-Smith, S.: Evaluating the performance of hpc-style sycl applications. In: Proceedings of the International Workshop on OpenCL. IWOCL \u201920, Association for Computing Machinery, New York, NY, USA (2020). https:\/\/doi.org\/10.1145\/3388333.3388643","DOI":"10.1145\/3388333.3388643"},{"key":"787_CR7","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1016\/j.cam.2014.02.011","volume":"270","author":"E D\u2019Azevedo","year":"2014","unstructured":"D\u2019Azevedo, E., Hu, Z., Su, S.Q., Wong, K.: Solving a large scale radiosity problem on gpu-based parallel computers. J. Comput. Appl. Math. 270, 109\u2013120 (2014). https:\/\/doi.org\/10.1016\/j.cam.2014.02.011","journal-title":"J. Comput. Appl. Math."},{"issue":"1","key":"787_CR8","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1007\/s10766-017-0490-5","volume":"46","author":"A Ernstsson","year":"2018","unstructured":"Ernstsson, A., Li, L., Kessler, C.: Skepu 2: flexible and type-safe skeleton programming for heterogeneous parallel systems. Int. J. Parallel Prog. 46(1), 62\u201380 (2018). https:\/\/doi.org\/10.1007\/s10766-017-0490-5","journal-title":"Int. J. Parallel Prog."},{"key":"787_CR9","doi-asserted-by":"publisher","first-page":"5205","DOI":"10.1007\/s11042-015-2943-4","volume":"75","author":"Z Fu","year":"2016","unstructured":"Fu, Z., Li, J.: Gpu-based image method for room impulse response calculation. Multimedia Tools Appl. 75, 5205\u20135221 (2016)","journal-title":"Multimedia Tools Appl."},{"issue":"10","key":"787_CR10","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1109\/MCG.1984.6429331","volume":"4","author":"AS Glassner","year":"1984","unstructured":"Glassner, A.S.: Space subdivision for fast ray tracing. IEEE Comput. Graph. Appl. 4(10), 15\u201324 (1984)","journal-title":"IEEE Comput. Graph. Appl."},{"key":"787_CR11","volume-title":"An Introduction to Ray Tracing","author":"AS Glassner","year":"1989","unstructured":"Glassner, A.S.: An Introduction to Ray Tracing. Elsevier, Amsterdam (1989)"},{"key":"787_CR12","doi-asserted-by":"publisher","unstructured":"Gschwandtner, P., Kissmann, R., Huber, D., Salzmann, P., Knorr, F., Thoman, P., Fahringer, T.: Porting real-world applications to GPU clusters: a celerity and cronos case study. In: 2021 IEEE 17th International Conference on eScience (eScience), pp. 90\u201398 (2021). https:\/\/doi.org\/10.1109\/eScience51609.2021.00019","DOI":"10.1109\/eScience51609.2021.00019"},{"key":"787_CR13","volume-title":"Time-of-Flight Cameras: Principles, Methods and Applications","author":"M Hansard","year":"2012","unstructured":"Hansard, M., Lee, S., Choi, O., Horaud, R.P.: Time-of-Flight Cameras: Principles, Methods and Applications. Springer, Berlin (2012)"},{"issue":"1","key":"787_CR14","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1111\/j.1467-8659.2010.01844.x","volume":"30","author":"M Hapala","year":"2011","unstructured":"Hapala, M., Havran, V.: Kd-tree traversal algorithms for ray tracing. Comput. Graph. Forum 30(1), 199\u2013213 (2011)","journal-title":"Comput. Graph. Forum"},{"key":"787_CR15","doi-asserted-by":"crossref","unstructured":"Hossain, M.M., Tucker, T.M., Kurfess, T.R., Vuduc, R.W.: A GPU-parallel construction of volumetric tree. In: Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms, pp.\u00a01\u20134 (2015)","DOI":"10.1145\/2833179.2833191"},{"key":"787_CR16","unstructured":"Hranitzky, R.: A Scalable multi-DSP System for Room Impulse Response Simulation. Ph.D. thesis, Technical University of Graz (1997)"},{"key":"787_CR17","doi-asserted-by":"crossref","unstructured":"Karras, T., Aila, T.: Fast parallel construction of high-quality bounding volume hierarchies. In: Proceedings of the 5th High-Performance Graphics Conference, pp. 89\u201399. ACM (2013)","DOI":"10.1145\/2492045.2492055"},{"key":"787_CR18","doi-asserted-by":"crossref","unstructured":"Knorr, F., Thoman, P., Fahringer, T.: Declarative data flow in a graph-based distributed memory runtime system. Int. J. Parallel Program. 1\u201322 (2022)","DOI":"10.21203\/rs.3.rs-2045925\/v1"},{"issue":"1","key":"787_CR19","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1121\/1.2936367","volume":"124","author":"EA Lehmann","year":"2008","unstructured":"Lehmann, E.A., Johansson, A.M.: Prediction of energy decay in room impulse responses simulated with an image-source model. J. Acoust. Soc. Am. 124(1), 269\u2013277 (2008). https:\/\/doi.org\/10.1121\/1.2936367","journal-title":"J. Acoust. Soc. Am."},{"key":"787_CR20","unstructured":"Message Passing Interface Forum: MPI: A Message-Passing Interface Standard, Version 3.1 (2015). https:\/\/www.mpi-forum.org\/docs\/mpi-3.1\/mpi31-report.pdf"},{"issue":"7","key":"787_CR21","doi-asserted-by":"publisher","first-page":"5117","DOI":"10.1007\/s11227-019-02829-2","volume":"76","author":"A Rasch","year":"2020","unstructured":"Rasch, A., Bigge, J., Wrodarczyk, M., Schulze, R., Gorlatch, S.: dOCAL: high-level distributed programming with OpenCL and CUDA. J. Supercomput. 76(7), 5117\u20135138 (2020). https:\/\/doi.org\/10.1007\/s11227-019-02829-2","journal-title":"J. Supercomput."},{"key":"787_CR22","doi-asserted-by":"crossref","unstructured":"Salzmann, P., Knorr, F., Thoman, P., Gschwandtner, P., Cosenza, B., Fahringer, T.: An asynchronous dataflow-driven execution model for distributed accelerator computing. In: 2023 IEEE\/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid), pp. 82\u201393. IEEE (2023)","DOI":"10.1109\/CCGrid57682.2023.00018"},{"key":"787_CR23","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.2943779","author":"C Schissler","year":"2017","unstructured":"Schissler, C., Manocha, D.: Interactive sound propagation and rendering for large multi-source scenes. ACM Trans. Graph. (2017). https:\/\/doi.org\/10.1145\/3072959.2943779","journal-title":"ACM Trans. Graph."},{"key":"787_CR24","unstructured":"The Khronos Group: SYCL Specification, Version 2020 Revision 8 (2023) https:\/\/registry.khronos.org\/SYCL\/specs\/sycl-2020\/html\/sycl-2020.html"},{"key":"787_CR25","doi-asserted-by":"crossref","unstructured":"Thoman, P., Salzmann, P., Cosenza, B., Fahringer, T.: Celerity: High-level C++ for accelerator clusters. In: Euro-Par 2019: Parallel Processing: 25th International Conference on Parallel and Distributed Computing, G\u00f6ttingen, Germany, August 26\u201330, 2019, Proceedings 25, pp. 291\u2013303. Springer (2019)","DOI":"10.1007\/978-3-030-29400-7_21"},{"key":"787_CR26","doi-asserted-by":"publisher","unstructured":"Thoman, P., Wippler, M., Hranitzky, R., Fahringer, T.: Rtx-rsim: accelerated Vulkan room response simulation for time-of-flight imaging. In: Proceedings of the International Workshop on OpenCL. IWOCL\u201920, Association for Computing Machinery, New York, NY, USA (2020). https:\/\/doi.org\/10.1145\/3388333.3388662","DOI":"10.1145\/3388333.3388662"},{"key":"787_CR27","doi-asserted-by":"publisher","unstructured":"Thoman, P., Wippler, M., Hranitzky, R., Gschwandtner, P., Fahringer, T.: Multi-GPU room response simulation with hardware raytracing. Concurr. Comput. Pract. Exp. 34(4), e6663 (2022). https:\/\/doi.org\/10.1002\/cpe.6663, (https:\/\/onlinelibrary.wiley.com\/doi\/abs\/10.1002\/cpe.6663).","DOI":"10.1002\/cpe.6663"},{"issue":"4","key":"787_CR28","doi-asserted-by":"publisher","first-page":"805","DOI":"10.1109\/TPDS.2021.3097283","volume":"33","author":"CR Trott","year":"2022","unstructured":"Trott, C.R., Lebrun-Grandi\u00e9, D., Arndt, D., Ciesko, J., Dang, V., Ellingwood, N., Gayatri, R., Harvey, E., Hollman, D.S., Ibanez, D., Liber, N., Madsen, J., Miles, J., Poliakoff, D., Powell, A., Rajamanickam, S., Simberg, M., Sunderland, D., Turcksin, B., Wilke, J.: Kokkos 3: programming model extensions for the exascale era. IEEE Trans. Parallel Distrib. Syst. 33(4), 805\u2013817 (2022). https:\/\/doi.org\/10.1109\/TPDS.2021.3097283","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"issue":"1","key":"787_CR29","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1121\/1.398336","volume":"86","author":"M Vorl\u00e4nder","year":"1989","unstructured":"Vorl\u00e4nder, M.: Simulation of the transient and steady-state sound propagation in rooms using a new combined ray-tracing\/image-source algorithm. J. Acoust. Soc. Am. 86(1), 172\u2013178 (1989)","journal-title":"J. Acoust. Soc. Am."}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-025-00787-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-025-00787-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-025-00787-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,5,16]],"date-time":"2025-05-16T15:23:01Z","timestamp":1747408981000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-025-00787-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,24]]},"references-count":29,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["787"],"URL":"https:\/\/doi.org\/10.1007\/s10766-025-00787-2","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"value":"0885-7458","type":"print"},{"value":"1573-7640","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,24]]},"assertion":[{"value":"28 September 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 March 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"17"}}