{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:44:35Z","timestamp":1740123875528,"version":"3.37.3"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2020,11,20]],"date-time":"2020-11-20T00:00:00Z","timestamp":1605830400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"},{"start":{"date-parts":[[2020,11,20]],"date-time":"2020-11-20T00:00:00Z","timestamp":1605830400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["146371743"],"award-info":[{"award-number":["146371743"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2021,8]]},"DOI":"10.1007\/s10766-020-00687-7","type":"journal-article","created":{"date-parts":[[2020,11,20]],"date-time":"2020-11-20T16:11:11Z","timestamp":1605888671000},"page":"506-540","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["DySHARQ: Dynamic Software-Defined Hardware-Managed Queues for Tile-Based Architectures"],"prefix":"10.1007","volume":"49","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9024-5639","authenticated-orcid":false,"given":"Sven","family":"Rheindt","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Maier","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nora","family":"Pohle","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lars","family":"Nolte","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Oliver","family":"Lenke","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Florian","family":"Schmaus","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas","family":"Wild","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wolfgang","family":"Schr\u00f6der-Preikschat","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andreas","family":"Herkersdorf","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,11,20]]},"reference":[{"key":"687_CR1","doi-asserted-by":"publisher","unstructured":"Parkhurst, J., Darringer, J., Grundmann, B.: From single core to multi-core: preparing for a new exponential. In: 2006 IEEE\/ACM International Conference on Computer Aided Design, pp. 67\u201372 (2006). https:\/\/doi.org\/10.1109\/ICCAD.2006.320067","DOI":"10.1109\/ICCAD.2006.320067"},{"issue":"1","key":"687_CR2","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1145\/216585.216588","volume":"23","author":"WA Wulf","year":"1995","unstructured":"Wulf, W.A., McKee, S.A.: Hitting the memory wall: implications of the obvious. SIGARCH Comput. Archit. News 23(1), 20\u201324 (1995). https:\/\/doi.org\/10.1145\/216585.216588","journal-title":"SIGARCH Comput. Archit. News"},{"issue":"2","key":"687_CR3","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1109\/40.592312","volume":"17","author":"DA Patterson","year":"1997","unstructured":"Patterson, D.A., Anderson, T.E., Cardwell, N., Fromm, R., Keeton, K., Kozyrakis, C.E., Thomas, R., Yelick, K.A.: A case for intelligent RAM. IEEE Micro 17(2), 34\u201344 (1997). https:\/\/doi.org\/10.1109\/40.592312","journal-title":"IEEE Micro"},{"key":"687_CR4","doi-asserted-by":"publisher","unstructured":"Teich, J., Henkel, J., Herkersdorf, A., Schmitt-Landsiedel, D., Schr\u00f6der-Preikschat, W., Snelting, G.: Invasive computing: an overview. In: Multiprocessor System-on-Chip, pp. 241\u2013268 (2011). https:\/\/doi.org\/10.1007\/978-1-4419-6460-1_11","DOI":"10.1007\/978-1-4419-6460-1_11"},{"key":"687_CR5","doi-asserted-by":"publisher","unstructured":"Wentzlaff, D., Griffin, P., Hoffmann, H., Bao, L., Edwards, B., Ramey, C., Mattina, M., Miao III, C., Brown, J.F., Agarwal, A.: On-chip interconnection architecture of the tile processor. IEEE Micro 27(5), 15\u201331 (2007). https:\/\/doi.org\/10.1109\/MM.2007.89","DOI":"10.1109\/MM.2007.89"},{"key":"687_CR6","doi-asserted-by":"publisher","unstructured":"Bell, S., Edwards, B., Amann, J., Conlin, R., Joyce, K., Leung, V., MacKay, J., Reif, M., Bao, L., III JFB, Mattina, M., Miao, C., Ramey, C., Wentzlaff, D., Anderson, W., Berger, E., Fairbanks, N., Khan, D., Montenegro, F., Stickney, J., Zook, J.: TILE64 - processor: a 64-Core SoC with mesh interconnect. In: 2008 IEEE International Solid-State Circuits Conference, ISSCC 2008, Digest of Technical Papers, San Francisco, CA, USA, February 3\u20137, 2008, IEEE, San Francisco, CA, pp 88\u201389 (2008). https:\/\/doi.org\/10.1109\/ISSCC.2008.4523070","DOI":"10.1109\/ISSCC.2008.4523070"},{"key":"687_CR7","doi-asserted-by":"crossref","unstructured":"Lotfi-Kamran, P., Grot, B., Ferdman, M., Volos, S., Kocberber, O., Picorel, J., Adileh, A., Jevdjic, D., Idgunji, S., Ozer, E., Falsafi, B.: Scale-out Processors. In: Proceedings of the 39th Annual International Symposium on Computer Architecture, IEEE Computer Society, USA, ISCA \u201912, pp. 500\u2013511 (2012)","DOI":"10.1145\/2366231.2337217"},{"key":"687_CR8","doi-asserted-by":"publisher","unstructured":"Howard, J., Dighe, S., Hoskote, Y., Vangal, S., Finan, D., Ruhl, G., Jenkins, D., Wilson, H., Borkar, N., Schrom, G., Pailet, F., Jain, S., Jacob, T., Yada, S., Marella, S., Salihundam, P., Erraguntla, V., Konow, M., Riepen, M., Droege, G., Lindemann, J., Gries, M., Apel, T., Henriss, K., Lund-Larsen, T., Steibl, S., Borkar, S., De, V., Wijngaart, R.V.D., Mattson, T.: A 48-core IA-32 message-passing processor with DVFS in 45 nm CMOS. In: 2010 IEEE International Solid-State Circuits Conference\u2014(ISSCC), pp. 108\u2013109 (2010). https:\/\/doi.org\/10.1109\/ISSCC.2010.5434077","DOI":"10.1109\/ISSCC.2010.5434077"},{"key":"687_CR9","doi-asserted-by":"crossref","unstructured":"Mittal, S.: A survey on evaluating and optimizing performance of Intel Xeon Phi. Practice and Experience, Concurrency and Computation (2020)","DOI":"10.1002\/cpe.5742"},{"key":"687_CR10","doi-asserted-by":"publisher","unstructured":"Siegl, P., Buchty, R., Berekovic, M.: Data-centric computing frontiers: a survey on processing-in-memory. In: Jacob, B. (ed) Proceedings of the Second International Symposium on Memory Systems, MEMSYS 2016, Alexandria, VA, USA, 2016, ACM, pp. 295\u2013308 (2016). https:\/\/doi.org\/10.1145\/2989081.2989087","DOI":"10.1145\/2989081.2989087"},{"key":"687_CR11","unstructured":"Kogge, P.: Memory Intensive Computing, the 3rd Wall, and the Need for Innovation in Architecture. (2017) https:\/\/memsys.io\/wp-content\/uploads\/2017\/12\/The_Wall.pdf"},{"key":"687_CR12","unstructured":"Oechslein, B., Schedel, J., Klein\u00f6der, J., Bauer, L., Henkel, J., Lohmann, D., Schr\u00f6der-Preikschat, W.: OctoPOS: a parallel operating system for invasive computing. In: Proceedings of the International Workshop on Systems for Future Multi-Core Architectures. EuroSys, pp. 9\u201314 (2011)"},{"key":"687_CR13","doi-asserted-by":"publisher","unstructured":"Kranz, D.A., Johnson, K.L., Agarwal, A., Kubiatowicz, J., Lim, B.: Integrating message-passing and shared-memory: early experience. In: Chen, M.C., Halstead, R. (eds) Proceedings of the Fourth ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPOPP), San Diego, California, USA, 1993, ACM, pp. 54\u201363, (1993). https:\/\/doi.org\/10.1145\/155332.155338","DOI":"10.1145\/155332.155338"},{"key":"687_CR14","doi-asserted-by":"crossref","unstructured":"Moir, M., Shavit, N.: Concurrent data structures. In: Handbook of Data Structures and Applications (2004)","DOI":"10.1201\/9781420035179.ch47"},{"key":"687_CR15","unstructured":"MPI Forum: MPI: A Message Passing Interface Standard Version 3.1 (2015). https:\/\/www.mpi-forum.org\/docs\/mpi-3.1\/mpi31-report.pdf"},{"key":"687_CR16","unstructured":"Corbet, J.: Ringing in a new asynchronous I\/O API. (2019) https:\/\/lwn.net\/Articles\/776703\/"},{"key":"687_CR17","doi-asserted-by":"publisher","unstructured":"Michael, M.M., Scott, M.L.: Simple, fast, and practical non-blocking and blocking concurrent queue algorithms. In: ACM Symposium on Principles of Distributed Computing, pp. 267\u2013275 (1996). https:\/\/doi.org\/10.1145\/248052.248106","DOI":"10.1145\/248052.248106"},{"key":"687_CR18","doi-asserted-by":"publisher","unstructured":"Wang, Y., Wang, R., Herdrich, A., Tsai, J., Solihin, Y.: CAF: core to core communication acceleration framework. In: Conference on Parallel Architectures and Compilation (PACT), pp. 351\u2013362 (2016). https:\/\/doi.org\/10.1145\/2967938.2967954","DOI":"10.1145\/2967938.2967954"},{"key":"687_CR19","doi-asserted-by":"publisher","unstructured":"Lee, S., Tiwari, D., Solihin, Y., Tuck, J.: HAQu: hardware-accelerated queueing for fine-grained threading on a chip multiprocessor. In: Conference on High-Performance Computer Architecture (HPCA), pp. 99\u2013110 (2011). https:\/\/doi.org\/10.1109\/HPCA.2011.5749720","DOI":"10.1109\/HPCA.2011.5749720"},{"issue":"4","key":"687_CR20","doi-asserted-by":"publisher","first-page":"24:1","DOI":"10.1145\/2858652","volume":"2","author":"D Petrovic","year":"2016","unstructured":"Petrovic, D., Ropars, T., Schiper, A.: Leveraging hardware message passing for efficient thread synchronization. TOPC 2(4), 24:1\u201324:26 (2016). https:\/\/doi.org\/10.1145\/2858652","journal-title":"TOPC"},{"key":"687_CR21","doi-asserted-by":"publisher","unstructured":"S\u00e1nchez, D., Yoo, R.M., Kozyrakis, C.: Flexible architectural support for fine-grain scheduling. In: ASPLOS Conference Proceedings, pp. 311\u2013322 (2010). https:\/\/doi.org\/10.1145\/1736020.1736055","DOI":"10.1145\/1736020.1736055"},{"issue":"6","key":"687_CR22","doi-asserted-by":"publisher","first-page":"1080","DOI":"10.1109\/TVLSI.2012.2202699","volume":"21","author":"J Lee","year":"2013","unstructured":"Lee, J., Nicopoulos, C., Lee, H.G., Panth, S., Lim, S.K., Kim, J.: IsoNet: hardware-based job queue management for many-core architectures. IEEE Trans. VLSI Syst. 21(6), 1080\u20131093 (2013). https:\/\/doi.org\/10.1109\/TVLSI.2012.2202699","journal-title":"IEEE Trans. VLSI Syst."},{"key":"687_CR23","doi-asserted-by":"publisher","unstructured":"Pujari, R.K., Wild, T., Herkersdorf, A.: TCU: a multi-objective hardware thread mapping unit for HPC clusters. In: High Performance Computing, ISC, pp. 39\u201358 (2016). https:\/\/doi.org\/10.1007\/978-3-319-41321-1_3","DOI":"10.1007\/978-3-319-41321-1_3"},{"key":"687_CR24","doi-asserted-by":"publisher","unstructured":"Kumar, S., Hughes, C.J., Nguyen, A.D.: Carbon: architectural support for fine-grained parallelism on chip multiprocessors. In: Symposium on Computer Architecture (ISCA), pp. 162\u2013173 (2007). https:\/\/doi.org\/10.1145\/1250662.1250683","DOI":"10.1145\/1250662.1250683"},{"key":"687_CR25","doi-asserted-by":"publisher","unstructured":"Sharma, R.R., Rajasekhar, Y., Sass, R.: Exploring hardware work queue support for lightweight threads in MPSoCs. In: Conference on Reconfigurable Computing and FPGAs (ReConFig), pp. 1\u20136 (2012). https:\/\/doi.org\/10.1109\/ReConFig.2012.6416747","DOI":"10.1109\/ReConFig.2012.6416747"},{"key":"687_CR26","doi-asserted-by":"publisher","unstructured":"Brewer, E.A., Chong, F.T., Liu, L.T., Sharma, S.D., Kubiatowicz, J.: Remote queues: exposing message queues for optimization and atomicity. In: ACM Symposium on Parallel Algorithms and Architectures (SPAA), pp. 42\u201353 (1995). https:\/\/doi.org\/10.1145\/215399.215416","DOI":"10.1145\/215399.215416"},{"key":"687_CR27","doi-asserted-by":"publisher","unstructured":"Rheindt, S., Schenk, A., Srivatsa, A., Wild, T., Herkersdorf, A.: CaCAO: complex and compositional atomic operations for NoC-based manycore platforms. In: Conference on Architecture of Computing Systems (ARCS), pp 139\u2013152 (2018). https:\/\/doi.org\/10.1007\/978-3-319-77610-1_11","DOI":"10.1007\/978-3-319-77610-1_11"},{"key":"687_CR28","doi-asserted-by":"publisher","unstructured":"Rheindt, S., Maier, S., Schmaus, F., Wild, T., Schr\u00f6der-Preikschat, W., Herkersdorf, A.: SHARQ: software-defined hardware-managed queues for tile-based manycore architectures. In: International Conference on Embedded Computer Systems: Architectures, Modeling, and Simulation (SAMOS XIX), Springer, Samos, Greece, pp. 212\u2013225 (2019). https:\/\/doi.org\/10.1007\/978-3-030-27562-4_15","DOI":"10.1007\/978-3-030-27562-4_15"},{"key":"687_CR29","doi-asserted-by":"publisher","unstructured":"Schmaus, F., Maier, S., Langer, T., Rabenstein, J., H\u00f6nig, T., Bauer, L., Henkel, J., Schr\u00f6der-Preikschat, W.: System software for resource arbitration on future many-* architectures. In: 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), IEEE, pp. 967\u2013975 (2020). https:\/\/doi.org\/10.1109\/IPDPSW50202.2020.00160","DOI":"10.1109\/IPDPSW50202.2020.00160"},{"key":"687_CR30","doi-asserted-by":"crossref","unstructured":"Moerman, F.: Open event machine: a multi-core run-time designed for performance. In: 2014 6th European Embedded Design in Education and Research Conference (EDERC), pp. 41\u201345 (2014)","DOI":"10.1109\/EDERC.2014.6924355"},{"key":"687_CR31","doi-asserted-by":"crossref","unstructured":"Cataldo, R., Fernandes, R., Martin, K.J.M., Sepulveda, J., Susin, A., Marcon, C., Diguet, J.: Subutai: distributed synchronization primitives in NoC interfaces for legacy parallel-applications. In: 2018 55th ACM\/ESDA\/IEEE Design Automation Conference (DAC) (2018)","DOI":"10.1109\/DAC.2018.8465806"},{"key":"687_CR32","doi-asserted-by":"publisher","first-page":"72","DOI":"10.1016\/j.sysarc.2017.03.004","volume":"77","author":"A Zaib","year":"2017","unstructured":"Zaib, A., Wild, T., Herkersdorf, A., Heisswolf, J., Becker, J., Weichslgartner, A., Teich, J.: Efficient task spawning for shared memory and message passing in many-core architectures. J. Syst. Archit. - Embed. Syst. Des. 77, 72\u201382 (2017). https:\/\/doi.org\/10.1016\/j.sysarc.2017.03.004","journal-title":"J. Syst. Archit. - Embed. Syst. Des."},{"key":"687_CR33","unstructured":"Heisswolf, J., Zaib. A., Weichslgartner, A., Karle, M., Singh, M., Wild, T., Teich, J., Herkersdorf, A., Becker, J.: The invasive network on chip: a multi-objective many-core communication infrastructure. In: Conference on Architecture of Computing Systems (ARCS), Workshop Proceedings, pp. 1\u20138 (2014)"},{"key":"687_CR34","first-page":"253","volume":"1","author":"Chu HkJ","year":"1996","unstructured":"HkJ, Chu, et al.: Zero-copy TCP in solaris. USENIX Annu. Tech. Conf. 1, 253\u2013264 (1996)","journal-title":"USENIX Annu. Tech. Conf."},{"key":"687_CR35","unstructured":"Intel Corporation: Intel 82574 GbE Controller Family\u2014Datasheet. www.intel.com\/content\/dam\/doc\/datasheet\/82574l-gbe-controller-datasheet.pdf, rev. 3.4 (2014)"},{"issue":"3","key":"687_CR36","first-page":"63","volume":"5","author":"DH Bailey","year":"1991","unstructured":"Bailey, D.H., Barszcz, E., Barton, J.T., Browning, D.S., Carter, R.L., Dagum, L., Fatoohi, R.A., Frederickson, P.O., Lasinski, T.A., Schreiber, R.S., et al.: The NAS parallel benchmarks. Int. J. Supercomput. Appl. 5(3), 63\u201373 (1991)","journal-title":"Int. J. Supercomput. Appl."},{"key":"687_CR37","doi-asserted-by":"publisher","unstructured":"Subhlok, J., Venkataramaiah, S., Singh, A.: Characterizing NAS benchmark performance on shared heterogeneous networks. In: Parallel and Distributed Processing Symposium (IPDPS) (2002). https:\/\/doi.org\/10.1109\/IPDPS.2002.1015659","DOI":"10.1109\/IPDPS.2002.1015659"},{"key":"687_CR38","doi-asserted-by":"publisher","unstructured":"Maier, S., H\u00f6nig, T., W\u00e4gemann, P., Schr\u00f6der-Preikschat, W.: Asynchronous abstract machines: anti-noise system software for many-core processors. In: Proceedings of the 9th International Workshop on Runtime and Operating Systems for Supercomputers (ROSS), ACM, pp. 19\u201326 (2019). https:\/\/doi.org\/10.1145\/3322789.3328744","DOI":"10.1145\/3322789.3328744"}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-020-00687-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-020-00687-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-020-00687-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,6,29]],"date-time":"2021-06-29T22:03:47Z","timestamp":1625004227000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-020-00687-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,20]]},"references-count":38,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,8]]}},"alternative-id":["687"],"URL":"https:\/\/doi.org\/10.1007\/s10766-020-00687-7","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"type":"print","value":"0885-7458"},{"type":"electronic","value":"1573-7640"}],"subject":[],"published":{"date-parts":[[2020,11,20]]},"assertion":[{"value":"3 April 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 November 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 November 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}