{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T06:56:41Z","timestamp":1761807401942,"version":"3.41.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2007,9,1]],"date-time":"2007-09-01T00:00:00Z","timestamp":1188604800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCR-0073491ACI-0219884","ACI-0219884EIA-0202048CCF-0541364."],"award-info":[{"award-number":["CCR-0073491ACI-0219884","ACI-0219884EIA-0202048CCF-0541364."]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007523","name":"Advanced Cyberinfrastructure","doi-asserted-by":"publisher","award":["CCR-0073491ACI-0219884","REU grant ACI-0334592","ACI-0219884EIA-0202048CCF-0541364."],"award-info":[{"award-number":["CCR-0073491ACI-0219884","REU grant ACI-0334592","ACI-0219884EIA-0202048CCF-0541364."]}],"id":[{"id":"10.13039\/100007523","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["ACI-0219884EIA-0202048CCF-0541364."],"award-info":[{"award-number":["ACI-0219884EIA-0202048CCF-0541364."]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2007,9]]},"abstract":"<jats:p>As the architectures of computers change, introducing more caches onto multicore chips, even more locality becomes necessary. With the bandwidth between caches and RAM now even more valuable, additional locality from new matrix representations will be important to keep multiple processors busy. The default storage representations of both C and Fortran, row- and column-major respectively, have fundamental deficiencies with many matrix computations. By switching the storage representation from cartesian to block indices, one is able to take better advantage of cache locality at all levels from L1 to paging. This paper only changes storage representation from row-major to Morton-hybrid, and applies it to matrix multiplication. Its purpose is to show that, even with only traditional iterative algorithms, simply changing storage representation offers significant speedups.<\/jats:p>","DOI":"10.1145\/1327312.1327315","type":"journal-article","created":{"date-parts":[[2007,12,21]],"date-time":"2007-12-21T14:52:36Z","timestamp":1198248756000},"page":"6-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Analyzing block locality in Morton-order and Morton-hybrid matrices"],"prefix":"10.1145","volume":"35","author":[{"given":"K. Patrick","family":"Lorton","sequence":"first","affiliation":[{"name":"Schrodinger, New York, NY"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David S.","family":"Wise","sequence":"additional","affiliation":[{"name":"Indiana University, Bloomington, IN"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1149982.1149987"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1178597.1178604"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/11752578_126"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2002.1058095"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/77626.79170"},{"key":"e_1_2_1_6_1","series-title":"IMA Vol","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1007\/978-1-4684-6357-6_4","volume-title":"Numerical Algorithms for Modern Parallel Architectures","author":"Fox G. C.","year":"1988"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066650.1066657"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/795665.796479"},{"volume-title":"The Opie Compiler Distribution","year":"2005","author":"Gabriel S. T.","key":"e_1_2_1_9_1"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1054943.1054962"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/358728.358741"},{"volume-title":"Matrix Computations","year":"1996","author":"Golub G. H.","key":"e_1_2_1_12_1"},{"volume-title":"On reducing TLB misses in matrix multiplication. FLAME Working Note 9","year":"2002","author":"Goto K.","key":"e_1_2_1_13_1"},{"volume-title":"Anatomy of high-performance matrix multiplication. Tech. rep","year":"2006","author":"Goto K.","key":"e_1_2_1_14_1"},{"volume-title":"TN","year":"2005","author":"Innovative Computing Laboratory","key":"e_1_2_1_15_1"},{"key":"e_1_2_1_16_1","series-title":"DIMACS Ser","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1090\/dimacs\/059\/11","volume-title":"Data Structures, Near Neighbor Searches, and Methodology: 5th & 6th DIMACS Implementation Challenges","author":"Johnson D. S.","year":"2002"},{"key":"e_1_2_1_17_1","first-page":"307","volume-title":"14th Int. Parallel and Distributed Processing Symp. (IPDPS'00)","author":"Li K.","year":"2000"},{"key":"e_1_2_1_18_1","first-page":"412","article-title":"Writing the fastest code, by hand, for fun: A human computer keeps speeding up chips","volume":"53","author":"Markoff J","year":"2005","journal-title":"The New York Times CLV"},{"volume-title":"A computer oriented geodetic data base and a new technique in file sequencing. Tech. rep","year":"1966","author":"Morton G. M.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2003.1214317"},{"volume-title":"The Design and Analysis of Spatial Data Structures","year":"1990","author":"Samet H.","key":"e_1_2_1_21_1"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2004.44"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/1049-9660(92)90022-U"},{"key":"e_1_2_1_24_1","first-page":"10","article-title":"A framework for high-performance matrix multiplication based on hierarchical abstractions, algorithms and optimized low-level kernels","volume":"14","author":"Valsalam V.","year":"2002","journal-title":"Concur. Comp. Prac. Exper."},{"key":"e_1_2_1_25_1","series-title":"Lecture Notes in Comput","doi-asserted-by":"crossref","first-page":"774","DOI":"10.1007\/3-540-44520-X_108","volume-title":"Euro-Par 2000---Parallel Processing","author":"Wise D. S.","year":"2000"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/11549468_76"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/76263.76337"}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1327312.1327315","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1327312.1327315","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:56:25Z","timestamp":1750254985000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1327312.1327315"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,9]]},"references-count":27,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,9]]}},"alternative-id":["10.1145\/1327312.1327315"],"URL":"https:\/\/doi.org\/10.1145\/1327312.1327315","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2007,9]]},"assertion":[{"value":"2007-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}