{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,2]],"date-time":"2022-04-02T19:11:08Z","timestamp":1648926668798},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2013,8,6]],"date-time":"2013-08-06T00:00:00Z","timestamp":1375747200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2014,12]]},"DOI":"10.1007\/s11042-013-1639-x","type":"journal-article","created":{"date-parts":[[2013,8,5]],"date-time":"2013-08-05T05:02:25Z","timestamp":1375678945000},"page":"1391-1416","source":"Crossref","is-referenced-by-count":0,"title":["Demand look-ahead memory access scheduling for 3D graphics processing units"],"prefix":"10.1007","volume":"73","author":[{"given":"Chih-Chieh","family":"Hsiao","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min-Jen","family":"Lo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Slo-Li","family":"Chu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2013,8,6]]},"reference":[{"key":"1639_CR1","doi-asserted-by":"crossref","unstructured":"Ausavarungnirun R, Chang K-W, Subramanian L, Loh GH, Mutlu O (2012) Staged memory scheduling: achieving high performance and scalability in heterogeneous systems. In: Proceedings of the 39th International Symposium on Computer Architecture, pp 416\u2013427","DOI":"10.1145\/2366231.2337207"},{"key":"1639_CR2","doi-asserted-by":"crossref","unstructured":"Ebrahimi E, Miftakhutdinov R, Falling C, Lee CJ, Joao JA, Mutlu O, Patt YN (2011) Parallel application memory scheduling. In: Proceedings of the 44th Annual IEEE\/ACM International Symposium on Microarchitecture, pp 362\u2013373","DOI":"10.1145\/2155620.2155663"},{"key":"1639_CR3","doi-asserted-by":"crossref","unstructured":"Hong S, Mckee S, Salinas M, Klenke R, Aylor J, Wulf W (1999) Access order and effective bandwidth for streams on a direct rambus memory. In: Proceeding of High-Performance Computer Architecture, pp 80\u201389","DOI":"10.1109\/HPCA.1999.744337"},{"key":"1639_CR4","unstructured":"Hynix (2006) 512M (16Mx32) GDDR3 SDRAM HY5RS123235FP Specification"},{"key":"1639_CR5","doi-asserted-by":"crossref","unstructured":"Jeong MK, Erez M, Sudanthi C, Paver N (2012) A QoS-aware memory controller for dynamically balancing GPU and CPU bandwidth use in an MPSoC. In: Proceeding of Design Automation Conference, pp 850\u2013855","DOI":"10.1145\/2228360.2228513"},{"key":"1639_CR6","doi-asserted-by":"crossref","unstructured":"Joao JA, Suleman AM, Mutlu O, Patt YN (2012) Bottleneck identification and scheduling in multithreaded applications. In: Proceedings of the seventeenth international conference on Architectural Support for Programming Languages and Operating Systems, pp 223\u2013234","DOI":"10.1145\/2150976.2151001"},{"key":"1639_CR7","unstructured":"Juffa N, Coon B (2011) Maximized memory throughput using cooperative thread arrays. US Patent 7,925,860 B1 Apr 1998"},{"key":"1639_CR8","unstructured":"Kim Y, Han D, Mutlu O, Harcol-Balter M (2010) ATLAS: a scalable and high-performance scheduling algorithm for multiple memory controllers. In: Proceedings of the 16th International Symposium on High-Performance Computer Architecture, pp 1\u201312"},{"key":"1639_CR9","doi-asserted-by":"crossref","unstructured":"Kim Y, Papamichael M, Mutlu O, Harcol-Balter M (2010) Thread cluster memory scheduling: exploiting differences in memory access behavior. In: Proceedings of the 2010 43rd Annual IEEE\/ACM International Symposium on Microarchitecture, pp 65\u201376","DOI":"10.1109\/MICRO.2010.51"},{"key":"1639_CR10","doi-asserted-by":"crossref","unstructured":"Kruger F (2008) High bandwidth memory technology: system architecture implications and perspective. In: Hot chips 20","DOI":"10.1109\/HOTCHIPS.2008.7476515"},{"key":"1639_CR11","doi-asserted-by":"crossref","unstructured":"Lee J, Lakshminarayana N, Kim H, Vuduc R (2010) Many-thread aware prefetching mechanisms for GPGPU applications. In: Proceeding of International Symposium on Microarchitecture, pp 213\u2013224","DOI":"10.1109\/MICRO.2010.44"},{"key":"1639_CR12","unstructured":"Mantor M (2007) AMD\u2019s Radeon HD 2900 2nd Generation Unified Shader Architecture. In: Hot Chips 19"},{"key":"1639_CR13","unstructured":"Mizuyabu C, Chow P, Swan P, Wang C (2003) Method and apparatus for memory access scheduling in a video graphics system. US Patent 6,297,832 B1 May 2003"},{"key":"1639_CR14","unstructured":"Moya V, Gonzalez C, Roca J, Fernandez A, Espana R (2006) ATTILA: a cycle-level execution-driven simulator for modern GPU architectures. In: Proceeding of IEEE International Symposium on Performance Analysis of Systems and Software, pp 231\u2013241"},{"key":"1639_CR15","unstructured":"Moya V, Gonzalez C, Solis C, Fernandez A, Espana R (2006) Workload characterization of 3D games. In: Proceeding of IEEE International Symposium on Workload Characterization, pp 17\u201326"},{"key":"1639_CR16","doi-asserted-by":"crossref","unstructured":"Mutlu O, Moscibroda T (2007) Stall-time fair memory access scheduling for chip multiprocessors. In: Proceeding of International Symposium on Microarchitecture, pp 146\u2013160","DOI":"10.1109\/MICRO.2007.21"},{"key":"1639_CR17","doi-asserted-by":"crossref","unstructured":"Mutlu O, Moscibroda T (2008) Parallelism-aware batch scheduling: enhancing both performance and fairness of shared DRAM systems. In: Proceeding of International Symposium on Computer Architecture, pp 63\u201374","DOI":"10.1145\/1394608.1382128"},{"key":"1639_CR18","doi-asserted-by":"crossref","unstructured":"Nesbit KJ, Aggarwal N, Laudon J, Smith JE (2006) Fair queuing memory systems. In: Proceeding of International Symposium on Microarchitecture, pp 208\u2013222","DOI":"10.1109\/MICRO.2006.24"},{"issue":"2","key":"1639_CR19","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1109\/MM.2010.41","volume":"30","author":"J Nickolls","year":"2010","unstructured":"Nickolls J, Dally WJ (2010) The GPU computing era. IEEE Micro 30(2):56\u201369","journal-title":"IEEE Micro"},{"key":"1639_CR20","doi-asserted-by":"crossref","unstructured":"Rafique N, Lim W-T, Thottethodi M (2007) Effective management of DRAM bandwidth in multicore processors. In: Proceedings of the 16th International Conference on Parallel Architecture and Compilation Techniques, pp 245\u2013258","DOI":"10.1109\/PACT.2007.4336216"},{"key":"1639_CR21","doi-asserted-by":"crossref","unstructured":"Rixner S, Dally W, Kapsi U, Matton P, Owens J (2000) Memory access scheduling. In: Proceeding of International Symposium on Computer Architecture, pp 128\u2013138","DOI":"10.1145\/342001.339668"},{"key":"1639_CR22","doi-asserted-by":"crossref","unstructured":"Shao J, Davis B (2007) A burst scheduling access reordering mechanism. In: Proceeding of High-Performance Computer Architecture, pp 285\u2013294.","DOI":"10.1109\/HPCA.2007.346206"},{"issue":"4","key":"1639_CR23","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1145\/2086696.2086730","volume":"8","author":"K Therdsteerasukdi","year":"2012","unstructured":"Therdsteerasukdi K, Byun G, Cong J, Chang M-F, Reinman G (2012) Effective management of DRAM bandwidth in multicore processors utilizing RF-I and intelligent scheduling for better throughput\/watt in a mobile GPU memory system. ACM Trans Archit Code Optim 8(4):51\u201369","journal-title":"ACM Trans Archit Code Optim"},{"key":"1639_CR24","unstructured":"Van Hook T, Tang M-K (2001) Memory processing system and method for accessing memory including reordering memory requests to reduce mode switching. US Patent 6,564,304 B1 Oct 2001"},{"key":"1639_CR25","unstructured":"Wu C-C, Pean D-L, Chen C (1998) Look-ahead memory consistency model. In: Proceeding of the International Conference on Parallel and Distributed Systems, pp 504\u2013510"},{"key":"1639_CR26","doi-asserted-by":"crossref","unstructured":"Yuan G, Bakhoda A, Aamodt T (2009) Complexity effective memory access scheduling for many-core accelerator architectures. In: Proceeding of International Symposium on Microarchitecture, pp 34\u201344","DOI":"10.1145\/1669112.1669119"},{"key":"1639_CR27","doi-asserted-by":"crossref","unstructured":"Zheng H, Lin J, Zhang Z, Zhu Z (2008) Memory access scheduling schemes for systems with multi-core processors. In: Proceeding of International Conference on Parallel Processing, pp 406\u2013413","DOI":"10.1109\/ICPP.2008.53"},{"key":"1639_CR28","unstructured":"Zuravleff W, Robinson T (1997) Controller for a synchronous DRAM that maximizes throughput by allowing memory requests and commands to be issued out of order. US Patent 5,630,096 May 1997"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-013-1639-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11042-013-1639-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-013-1639-x","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,7,20]],"date-time":"2019-07-20T10:27:46Z","timestamp":1563618466000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11042-013-1639-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,8,6]]},"references-count":28,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2014,12]]}},"alternative-id":["1639"],"URL":"https:\/\/doi.org\/10.1007\/s11042-013-1639-x","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,8,6]]}}}