{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:19:57Z","timestamp":1750306797519,"version":"3.41.0"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"3s","license":[{"start":{"date-parts":[[2014,3,1]],"date-time":"2014-03-01T00:00:00Z","timestamp":1393632000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004963","name":"Seventh Framework Programme","doi-asserted-by":"publisher","award":["FP7-215216"],"award-info":[{"award-number":["FP7-215216"]}],"id":[{"id":"10.13039\/501100004963","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2014,3]]},"abstract":"<jats:p>When hardware cache coherence scales to many cores on chip, over saturated traffic of the shared memory system may offset the benefit from massive hardware concurrency. In this article, we investigate the cost of a write-update protocol in terms of on-chip memory network traffic and its adverse effects on the system performance based on a multithreaded many-core architecture with distributed caches. We discuss possible software and hardware solutions to alleviate the network pressure. We find that in the context of massive concurrency, by introducing a write-merging buffer with 0.46% area overhead to each core, applications with good locality and concurrency are boosted up by 18.74% in performance on average. Other applications also benefit from this addition and even achieve a throughput increase of 5.93%. In addition, this improvement indicates that higher levels of concurrency per core can be exploited without impacting performance, thus tolerating latency better and giving higher processor efficiencies compared to other solutions.<\/jats:p>","DOI":"10.1145\/2567931","type":"journal-article","created":{"date-parts":[[2014,3,25]],"date-time":"2014-03-25T13:34:12Z","timestamp":1395754452000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["On-chip traffic regulation to reduce coherence protocol cost on a microthreaded many-core architecture with distributed caches"],"prefix":"10.1145","volume":"13","author":[{"given":"Qiang","family":"Yang","sequence":"first","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Fu","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raphael","family":"Poss","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chris","family":"Jesshope","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,3,28]]},"reference":[{"volume-title":"Proceedings of the 4th Workshop on Complexity-Effective Design, held in conjunction with the 30th International Symposium on Computer Architecture.","author":"Agarwal D.","key":"e_1_2_1_1_1","unstructured":"D. Agarwal and D. Yeung . 2003. Exploiting application-level information to reduce memory bandwidth consumption . In Proceedings of the 4th Workshop on Complexity-Effective Design, held in conjunction with the 30th International Symposium on Computer Architecture. D. Agarwal and D. Yeung. 2003. Exploiting application-level information to reduce memory bandwidth consumption. In Proceedings of the 4th Workshop on Complexity-Effective Design, held in conjunction with the 30th International Symposium on Computer Architecture."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.50"},{"volume-title":"Proceedings of the International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS'08)","author":"Bernard T.","key":"e_1_2_1_3_1","unstructured":"T. Bernard , K. Bousias , L. Guang , C. Jesshope , M. Lankamp , M. Van Tol , and L. Zhang . 2008. A general model of concurrency and its implementation as many-core dynamic risc processors . In Proceedings of the International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS'08) . 1--9. T. Bernard, K. Bousias, L. Guang, C. Jesshope, M. Lankamp, M. Van Tol, and L. Zhang. 2008. A general model of concurrency and its implementation as many-core dynamic risc processors. In Proceedings of the International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS'08). 1--9."},{"key":"e_1_2_1_4_1","unstructured":"R. Bianchini T. J. Leblanc and J. Veenstra. 1994. Eliminating useless messages in write-update protocols on scalable multiprocessors. Tech. rep. University of Rochester.   R. Bianchini T. J. Leblanc and J. Veenstra. 1994. Eliminating useless messages in write-update protocols on scalable multiprocessors. Tech. rep. University of Rochester."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/bxh157"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/232973.232983"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"M. Danek L. Kafka L. Kohout J. Sykora and R. Bartosinsk. 2011. UTLEON3: Exploring Fine-Grain Multi-Threading in FPGAs. Springer.   M. Danek L. Kafka L. Kohout J. Sykora and R. Bartosinsk. 2011. UTLEON3: Exploring Fine-Grain Multi-Threading in FPGAs. Springer.","DOI":"10.1007\/978-1-4614-2410-9_1"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669150"},{"volume-title":"Proceedings of the 14th International Parallel and Distributed Processing Symposium. 181--189","author":"Ding C.","key":"e_1_2_1_9_1","unstructured":"C. Ding and K. Kennedy . 2000. The memory of bandwidth bottleneck and its amelioration by a compiler . In Proceedings of the 14th International Parallel and Distributed Processing Symposium. 181--189 . C. Ding and K. Kennedy. 2000. The memory of bandwidth bottleneck and its amelioration by a compiler. In Proceedings of the 14th International Parallel and Distributed Processing Symposium. 181--189."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSD.2008.111"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 27th Hawaii International Conference on System Sciences","volume":"1","author":"Glasco D.","unstructured":"D. Glasco , B. Delagi , and M. Flynn . 1994. Update-based cache coherence protocols for scalable shared-memory multiprocessors . In Proceedings of the 27th Hawaii International Conference on System Sciences , Vol. 1 . 534--545. D. Glasco, B. Delagi, and M. Flynn. 1994. Update-based cache coherence protocols for scalable shared-memory multiprocessors. In Proceedings of the 27th Hawaii International Conference on System Sciences, Vol. 1. 534--545."},{"volume-title":"Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA'08)","author":"Gratz P.","key":"e_1_2_1_12_1","unstructured":"P. Gratz , B. Grot , and S. W. Keckler . 2008. Regional congestion awareness for load balance in networks-on-chip . In Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA'08) . 203--214. P. Gratz, B. Grot, and S. W. Keckler. 2008. Regional congestion awareness for load balance in networks-on-chip. In Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA'08). 203--214."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024723.2000112"},{"key":"e_1_2_1_14_1","unstructured":"S. Gupta S. W. Keckler and D. Burger. 2000. Technology independent area and delay estimates for microprocessor building blocks. Tech. rep. Department of Computer Sciences The University of Texas at Austin.   S. Gupta S. W. Keckler and D. Burger. 2000. Technology independent area and delay estimates for microprocessor building blocks. Tech. rep. Department of Computer Sciences The University of Texas at Austin."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2010.5434077"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854291"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2007.10"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1785481.1785542"},{"volume-title":"Proceedings of the International Conference on Computer Design. 105--111","author":"Kondo M.","key":"e_1_2_1_19_1","unstructured":"M. Kondo , H. Okawara , H. Nakamura , and T. Boku . 2000. Scima: Software controlled integrated memory architecture for high performance computing . In Proceedings of the International Conference on Computer Design. 105--111 . M. Kondo, H. Okawara, H. Nakamura, and T. Boku. 2000. Scima: Software controlled integrated memory architecture for high performance computing. In Proceedings of the International Conference on Computer Design. 105--111."},{"key":"e_1_2_1_20_1","unstructured":"M. Lankamp R. Poss Q. Yang J. Fu I. Uddin and C. R. Jesshope. 2013. MGSim: Simulation tools for multi-core processor architectures. Tech. Rep. arXiv:1302.1390v1 {cs.AR} University of Amsterdam.  M. Lankamp R. Poss Q. Yang J. Fu I. Uddin and C. R. Jesshope. 2013. MGSim: Simulation tools for multi-core processor architectures. Tech. Rep. arXiv:1302.1390v1 {cs.AR} University of Amsterdam."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273440.1250707"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1987816.1987832"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2209249.2209269"},{"volume-title":"Proceedings of the International Conference on High Performance Computing and Networking. 1246--1249","author":"Molina C.","key":"e_1_2_1_24_1","unstructured":"C. Molina , A. Gonlaze , and J. Tubella . 1999. Reducing memory traffic via redundant store instructions . In Proceedings of the International Conference on High Performance Computing and Networking. 1246--1249 . C. Molina, A. Gonlaze, and J. Tubella. 1999. Reducing memory traffic via redundant store instructions. In Proceedings of the International Conference on High Performance Computing and Networking. 1246--1249."},{"volume-title":"Proceedings of the IEEE International Conference on Computer Design: VLSI in Computers and Processors (ICCD'95)","author":"Mounes-Toussi F.","key":"e_1_2_1_25_1","unstructured":"F. Mounes-Toussi and D. Lilja . 1995. Write buffer design for cache-coherent shared-memory multiprocessors . In Proceedings of the IEEE International Conference on Computer Design: VLSI in Computers and Processors (ICCD'95) . 506--511. F. Mounes-Toussi and D. Lilja. 1995. Write buffer design for cache-coherent shared-memory multiprocessors. In Proceedings of the IEEE International Conference on Computer Design: VLSI in Computers and Processors (ICCD'95). 506--511."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2076022.1993489"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2342356.2342436"},{"key":"e_1_2_1_28_1","volume-title":"SL: A \u201cquick and dirty","author":"Poss R.","year":"2012","unstructured":"R. Poss . 2012 . SL: A \u201cquick and dirty \u201d but working intermediate language for SVP systems. Tech. Rep. arXiv:1208.4572v1 {cs.PL}, University of Amsterdam . R. Poss. 2012. SL: A \u201cquick and dirty\u201d but working intermediate language for SVP systems. Tech. Rep. arXiv:1208.4572v1 {cs.PL}, University of Amsterdam."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSD.2012.25"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.micpro.2013.05.004"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6168950"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid.2012.53"},{"volume-title":"Proceedings of the 3rd International Symposium on High-Performance Computer Architecture. 144--155","author":"Skadron K.","key":"e_1_2_1_33_1","unstructured":"K. Skadron and D. W. Clark . 1997. Design issues and tradeoffs for write buffers . In Proceedings of the 3rd International Symposium on High-Performance Computer Architecture. 144--155 . K. Skadron and D. W. Clark. 1997. Design issues and tradeoffs for write buffers. In Proceedings of the 3rd International Symposium on High-Performance Computer Architecture. 144--155."},{"volume-title":"Proceedings of the MARC Symposium. 13--18","author":"Van Tol M. W.","key":"e_1_2_1_34_1","unstructured":"M. W. Van Tol , R. Bakker , M. Verstraaten , C. Grelck , and C. Jesshope . 2011. Efficient memory copy operations on the 48-core intel scc processor . In Proceedings of the MARC Symposium. 13--18 . M. W. Van Tol, R. Bakker, M. Verstraaten, C. Grelck, and C. Jesshope. 2011. Efficient memory copy operations on the 48-core intel scc processor. In Proceedings of the MARC Symposium. 13--18."},{"volume-title":"Proceedings of the 15th International Conference on Field Programmable Logic and Applications. IEEE, 197--202","author":"Wolkotte P. T.","key":"e_1_2_1_35_1","unstructured":"P. T. Wolkotte , G. J. Smit , and J. E. Becker . 2005. Energy-efficient noc for best-effort communication . In Proceedings of the 15th International Conference on Field Programmable Logic and Applications. IEEE, 197--202 . P. T. Wolkotte, G. J. Smit, and J. E. Becker. 2005. Energy-efficient noc for best-effort communication. In Proceedings of the 15th International Conference on Field Programmable Logic and Applications. IEEE, 197--202."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.323"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.10"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2567931","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2567931","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:34:39Z","timestamp":1750232079000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2567931"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,3]]},"references-count":37,"journal-issue":{"issue":"3s","published-print":{"date-parts":[[2014,3]]}},"alternative-id":["10.1145\/2567931"],"URL":"https:\/\/doi.org\/10.1145\/2567931","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2014,3]]},"assertion":[{"value":"2012-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-08-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-03-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}