{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T05:00:15Z","timestamp":1771650015223,"version":"3.50.1"},"reference-count":18,"publisher":"Information Processing Society of Japan","issue":"0","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Journal of Information Processing"],"published-print":{"date-parts":[[2026]]},"DOI":"10.2197\/ipsjjip.34.132","type":"journal-article","created":{"date-parts":[[2026,2,14]],"date-time":"2026-02-14T22:09:56Z","timestamp":1771106996000},"page":"132-139","source":"Crossref","is-referenced-by-count":0,"title":["Single-precision Matrix Multiplication Performance on Cerebras CS-2: Evaluation and Modelling of Performance, Scalability and Energy Efficiency"],"prefix":"10.2197","volume":"34","author":[{"given":"Takaaki","family":"Miyajima","sequence":"first","affiliation":[{"name":"Meiji University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryunosuke","family":"Matsuzaki","sequence":"additional","affiliation":[{"name":"Meiji University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daichi","family":"Mukunoki","sequence":"additional","affiliation":[{"name":"Information Technology Center, Nagoya University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1012","reference":[{"key":"1","doi-asserted-by":"crossref","unstructured":"[1] Lie, S.: Cerebras Architecture Deep Dive: First Look Inside the HW\/SW Co-Design for Deep Learning: Cerebras Systems, <i>2022 IEEE Hot Chips 34 Symposium<\/i>, <i>HCS 2022<\/i>, pp.1-34, IEEE (online), DOI: 10.1109\/HCS55958.2022.9895479 (2022).","DOI":"10.1109\/HCS55958.2022.9895479"},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] Orenes-Vera, M., Sharapov, I., Schreiber, R., Jacquelin, M., Vandermersch, P. and Chetlur, S.: Wafer-Scale Fast Fourier Transforms, <i>Proc. 37th ACM International Conference on Supercomputing<\/i>, <i>ICS &apos;23<\/i>, pp.180-191, Association for Computing Machinery (online), DOI: 10.1145\/3577193.3593708 (2023).","DOI":"10.1145\/3577193.3593708"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] Jacquelin, M., Araya-Polo, M. and Meng, J.: Scalable distributed high-order stencil computations, <i>Proc. International Conference on High Performance Computing, Networking, Storage and Analysis<\/i>, <i>SC &apos;22<\/i>, IEEE Press (2022).","DOI":"10.1109\/SC41404.2022.00035"},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] Rocki, K., Van Essendelft, D., Sharapov, I., Schreiber, R., Morrison, M., Kibardin, V., Portnoy, A., Dietiker, J.F., Syamlal, M. and James, M.: Fast Stencil-Code Computation on a Wafer-Scale Processor, <i>Proc. International Conference for High Performance Computing, Networking, Storage and Analysis<\/i>, <i>SC &apos;20<\/i>, IEEE Press (2020).","DOI":"10.1109\/SC41405.2020.00062"},{"key":"5","unstructured":"[5] Miyajima, T., Matsuzaki, R. and Fukuoka, L.: STREAM Benchmark on Cerebras WSE-2 (POSTER), <i>Proc. 39th International Conference ISC High Performance 2024 Research Paper<\/i>, No.10 (2024)."},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] Mukunoki, D. and Imamura, T.: Performance Analysis of 2D-compatible 2.5D-PDGEMM on Knights Landing Cluster, <i>Computational Science - ICCS 2018<\/i>, Shi, Y., Fu, H., Tian, Y., Krzhizhanovskaya, V.V., Lees, M.H., Dongarra, J. and Sloot, P.M.A. (Eds.), Cham, Springer International Publishing, pp.853-858 (2018).","DOI":"10.1007\/978-3-319-93713-7_85"},{"key":"7","doi-asserted-by":"crossref","unstructured":"[7] Matsuzaki, R., Mukunoki, D. and Miyajima, T.: Performance evaluation and modelling of single-precision matrix multiplication on Cerebras CS-2, <i>Proc. SC &apos;24 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis<\/i>, <i>SC-W &apos;24<\/i>, pp.727-731, IEEE Press (online), DOI: 10.1109\/SCW63240.2024.00101 (2025).","DOI":"10.1109\/SCW63240.2024.00101"},{"key":"8","unstructured":"[8] Cerebras Systems: Documentation for Developing with CSL - SDK Documentation (1.0.0), available from &lt;https:\/\/sdk.cerebras.net&gt;."},{"key":"9","unstructured":"[9] Petitet, A., Whaley, R., Dongarra, J. and Cleary, A.: HPL - a Portable Implementation of the High-Performance Linpack Benchmark for Distributed-Memory Computers (2008)."},{"key":"10","doi-asserted-by":"publisher","unstructured":"[10] Dongarra, J.J., Luszczek, P. and Petitet, A.: The LINPACK Benchmark: past, present and future, <i>Concurrency and Computation: Practice and Experience<\/i>, Vol.15, No.9, pp.803-820 (online), DOI: 10.1002\/cpe.728 (2003).","DOI":"10.1002\/cpe.728"},{"key":"11","unstructured":"[11] Meuer, H.W., Strohmaier, E., Dongarra, J. and Simon, H.D.: <i>The TOP500: History, Trends, and Future Directions in High Performance Computing<\/i>, Chapman &amp; Hall\/CRC, 1st edition (2014)."},{"key":"12","unstructured":"[12] van de Geijn, R.A. and Watts, J.: SUMMA: Scalable Universal Matrix Multiplication Algorithm, Technical Report, USA (1995)."},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] Solomonik, E. and Demmel, J.: Matrix Multiplication on Multidimensional Torus Networks, <i>High Performance Computing for Computational Science - VECPAR 2012<\/i>, pp.201-215, Springer Berlin Heidelberg (2013).","DOI":"10.1007\/978-3-642-38718-0_21"},{"key":"14","unstructured":"[14] Cerebras Systems: GEMM with Collective Operations - SDK Documentation (1.0.0), available from &lt;https:\/\/sdk.cerebras.net\/csl\/code-examples\/benchmark-gemm-collectives&gt;."},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] Kwasniewski, G., Kabi\u0107, M., Besta, M., VandeVondele, J., Solc\u00e0, R. and Hoefler, T.: Red-blue pebbling revisited: Near optimal parallel matrix-matrix multiplication, <i>Proc. International Conference for High Performance Computing, Networking, Storage and Analysis<\/i>, <i>SC \u201919<\/i>, Association for Computing Machinery (online), DOI: 10.1145\/3295500.3356181 (2019).","DOI":"10.1145\/3295500.3356181"},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] Solomonik, E. and Demmel, J.: Communication-Optimal Parallel 2.5D Matrix Multiplication and LU Factorization Algorithms, <i>Euro-Par 2011 Parallel Processing<\/i>, Jeannot, E., Namyst, R. and Roman, J. (Eds.), pp.90-109, Springer Berlin Heidelberg (2011).","DOI":"10.1007\/978-3-642-23397-5_10"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] Georganas, E., Gonzalez-Dominguez, J., Solomonik, E., Zheng, Y., Tourino, J. and Yelick, K.: Communication avoiding and overlapping for numerical linear algebra, <i>SC &apos;12: Proc. International Conference on High Performance Computing, Networking, Storage and Analysis<\/i>, pp.1-11 (online), DOI: 10.1109\/SC.2012.32 (2012).","DOI":"10.1109\/SC.2012.32"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] Demmel, J., Eliahu, D., Fox, A., Kamil, S., Lipshitz, B., Schwartz, O. and Spillinger, O.: Communication-Optimal Parallel Recursive Rectangular Matrix Multiplication, <i>2013 IEEE 27th International Symposium on Parallel and Distributed Processing<\/i>, pp.261-272 (online), DOI: 10.1109\/IPDPS.2013.80 (2013).","DOI":"10.1109\/IPDPS.2013.80"}],"container-title":["Journal of Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/ipsjjip\/34\/0\/34_132\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T04:00:28Z","timestamp":1771646428000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/ipsjjip\/34\/0\/34_132\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":18,"journal-issue":{"issue":"0","published-print":{"date-parts":[[2026]]}},"URL":"https:\/\/doi.org\/10.2197\/ipsjjip.34.132","relation":{},"ISSN":["1882-6652"],"issn-type":[{"value":"1882-6652","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026]]}}}