{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:38:59Z","timestamp":1740123539414,"version":"3.37.3"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"16","license":[{"start":{"date-parts":[[2022,5,20]],"date-time":"2022-05-20T00:00:00Z","timestamp":1653004800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,5,20]],"date-time":"2022-05-20T00:00:00Z","timestamp":1653004800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100014440","name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades","doi-asserted-by":"publisher","award":["PID2020-113656RB-C21\/C22"],"award-info":[{"award-number":["PID2020-113656RB-C21\/C22"]}],"id":[{"id":"10.13039\/100014440","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100011596","name":"Conselleria d\u2019Educaci\u00f3, Investigaci\u00f3, Cultura i Esport","doi-asserted-by":"publisher","award":["CDEIGENT\/2018\/014"],"award-info":[{"award-number":["CDEIGENT\/2018\/014"]}],"id":[{"id":"10.13039\/501100011596","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100014440","name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades","doi-asserted-by":"publisher","award":["FJC2019-039222-I"],"award-info":[{"award-number":["FJC2019-039222-I"]}],"id":[{"id":"10.13039\/100014440","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2022,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs.  automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of  is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the  auto-tuner delivers the same or higher performance than that achieved using a static selection approach.<\/jats:p>","DOI":"10.1007\/s11227-022-04577-2","type":"journal-article","created":{"date-parts":[[2022,5,20]],"date-time":"2022-05-20T09:03:22Z","timestamp":1653037402000},"page":"17543-17558","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["BestOf: an online implementation selector for the training and inference of deep neural networks"],"prefix":"10.1007","volume":"78","author":[{"given":"Sergio","family":"Barrachina","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adri\u00e1n","family":"Castell\u00f3","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9466-3398","authenticated-orcid":false,"given":"Manuel F.","family":"Dolz","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andr\u00e9s E.","family":"Tom\u00e1s","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,5,20]]},"reference":[{"issue":"12","key":"4577_CR1","doi-asserted-by":"publisher","first-page":"2295","DOI":"10.1109\/JPROC.2017.2761740","volume":"105","author":"V Sze","year":"2017","unstructured":"Sze V, Chen Y-H, Yang T-J, Emer JS (2017) Efficient processing of deep neural networks: a tutorial and survey. Proc IEEE 105(12):2295\u20132329","journal-title":"Proc IEEE"},{"issue":"5","key":"4577_CR2","first-page":"92:1","volume":"51","author":"S Pouyanfar","year":"2018","unstructured":"Pouyanfar S et al (2018) A survey on deep learning: algorithms, techniques, and applications. ACM Comput Surv 51(5):92:1-92:36","journal-title":"ACM Comput Surv"},{"issue":"3","key":"4577_CR3","doi-asserted-by":"publisher","first-page":"2443","DOI":"10.1007\/s00521-021-06540-3","volume":"34","author":"E Hssayni","year":"2022","unstructured":"Hssayni E, Joudar N-E, Ettaouil M (2022) KRR-CNN: kernels redundancy reduction in convolutional neural networks. Neural Comput Appl 34(3):2443\u20132454","journal-title":"Neural Comput Appl"},{"key":"4577_CR4","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1016\/j.swevo.2019.05.010","volume":"49","author":"FE Fernandes Junior","year":"2019","unstructured":"Fernandes Junior FE, Yen GG (2019) Particle swarm optimization of deep neural networks architectures for image classification. Swarm Evol Comput 49:62\u201374","journal-title":"Swarm Evol Comput"},{"key":"4577_CR5","doi-asserted-by":"crossref","unstructured":"Eddine MD, Shen Y (2022) A deep learning based approach for predicting the demand of electric vehicle charge, J Supercomput","DOI":"10.1007\/s11227-022-04428-0"},{"key":"4577_CR6","doi-asserted-by":"publisher","first-page":"70461","DOI":"10.1109\/ACCESS.2019.2918851","volume":"7","author":"M Jord\u00e0","year":"2019","unstructured":"Jord\u00e0 M, Valero-Lara P, Pe\u00f1a AJ (2019) Performance evaluation of cuDNN convolution algorithms on NVIDIA Volta GPUs. IEEE Access 7:70461\u201370473","journal-title":"IEEE Access"},{"key":"4577_CR7","first-page":"3393","volume-title":"Proceedings of the 32nd International Conference on Neural Information Processing Systems, Ser. NIPS\u201918","author":"T Chen","year":"2018","unstructured":"Chen T, Zheng L, Yan E, Jiang Z, Moreau T, Ceze L, Guestrin C, Krishnamurthy A (2018) Learning to optimize tensor programs. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, Ser. NIPS\u201918. Curran Associates Inc., Red Hook, NY, USA, pp 3393\u20133404"},{"key":"4577_CR8","unstructured":"Zheng L, Jia C, Sun M, Wu Z, Yu C.\u00a0H, Haj-Ali A, Wang Y, Yang J, Zhuo D, Sen K et\u00a0al (2020) Ansor: generating high-performance tensor programs for deep learning, In: 14th USENIX symposium on operating systems design and implementation (OSDI 20), pp 863\u2013879"},{"issue":"9","key":"4577_CR9","doi-asserted-by":"publisher","first-page":"9971","DOI":"10.1007\/s11227-021-03673-z","volume":"77","author":"S Barrachina","year":"2021","unstructured":"Barrachina S, Castell\u00f3 A, Catal\u00e1n M, Dolz MF, Mestre JI (2021) Pydtnn: a user-friendly and extensible framework for distributed deep learning. J Supercomput 77(9):9971\u20139987","journal-title":"J Supercomput"},{"key":"4577_CR10","unstructured":"Chellapilla K, Puri S, Simard P (2006) High performance convolutional neural networks for document processing,\u201d In: Tenth international workshop on frontiers in handwriting recognition"},{"key":"4577_CR11","unstructured":"Juan PS, Castell\u00f3 A, Dolz MF, Alonso-Jord\u00e1 P, Quintana-Ort\u00ed ES (2020) High performance and portable convolution operators for multicore processors, In: 32nd IEEE international symposium on computer architecture and high performance computing, SBAC-PAD (2020) Porto, Portugal, September 9\u201311. IEEE 2020:91\u201398"},{"key":"4577_CR12","doi-asserted-by":"crossref","unstructured":"Winograd S (1980) Arithmetic complexity of computations. Society for Industrial and Applied Mathematics","DOI":"10.1137\/1.9781611970364"},{"issue":"2","key":"4577_CR13","first-page":"1","volume":"43","author":"TM Low","year":"2016","unstructured":"Low TM, Igual FD, Smith TM, Quintana-Ort\u00ed ES (2016) Analytical modeling is enough for high-performance BLIS. ACM Trans Math Soft (TOMS) 43(2):1\u201318","journal-title":"ACM Trans Math Soft (TOMS)"},{"key":"4577_CR14","first-page":"1","volume-title":"Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing, Ser. SC \u201998","author":"RC Whaley","year":"1998","unstructured":"Whaley RC, Dongarra JJ (1998) Automatically tuned linear algebra software, In: Proceedings of the 1998 ACM\/IEEE Conference on Supercomputing, Ser. SC \u201998. IEEE Computer Society, USA, pp 1\u201327"},{"key":"4577_CR15","doi-asserted-by":"publisher","first-page":"170","DOI":"10.1007\/978-3-642-45293-2_13","volume-title":"Adv Parallel Process Technol","author":"U Dastgeer","year":"2013","unstructured":"Dastgeer U, Li L, Kessler C (2013) Adaptive implementation selection in the SkePU skeleton programming library. In: Wu C, Cohen A (eds) Adv Parallel Process Technol. Springer, Heidelberg, pp 170\u2013183"},{"issue":"6","key":"4577_CR16","doi-asserted-by":"publisher","first-page":"854","DOI":"10.1177\/1094342017698746","volume":"32","author":"Astorga D del Rio","year":"2018","unstructured":"del Rio Astorga D, Dolz MF, S\u00e1nchez LM, Fern\u00e1ndez J, Garc\u00eda JD (2018) An adaptive offline implementation selector for heterogeneous parallel platforms. Int J High Perform Comput Appl 32(6):854\u2013863","journal-title":"Int J High Perform Comput Appl"},{"key":"4577_CR17","doi-asserted-by":"crossref","unstructured":"Anderson A, Gregg D (2018) Optimal dnn primitive selection with partitioned boolean quadratic programming, In: Proceedings of the 2018 International symposium on code generation and optimization, ser. CGO. New York, NY, USA: Association for Computing Machinery, 2018, pp 340\u2013351. [Online]. Available: https:\/\/doi.org\/10.1145\/3168805","DOI":"10.1145\/3179541.3168805"},{"key":"4577_CR18","doi-asserted-by":"crossref","unstructured":"Fern\u00e1ndez J, Cuadrado AS, del Rio\u00a0Astorga D, Dolz MF, Daniel\u00a0Garc\u00eda J (2017) Probabilistic-based selection of alternate implementations for heterogeneous platforms, In: Algorithms and Architectures for Parallel Processing. Springer International Publishing, pp 749\u2013758","DOI":"10.1007\/978-3-319-65482-9_60"},{"key":"4577_CR19","doi-asserted-by":"crossref","unstructured":"Planas J, Badia RM, Ayguad\u00e9 E, Labarta J (2013) Self-adaptive OmpSs tasks in heterogeneous environments,\u201d In: 2013 IEEE 27th international symposium on parallel and distributed processing, pp 138\u2013149","DOI":"10.1109\/IPDPS.2013.53"},{"issue":"11","key":"4577_CR20","doi-asserted-by":"publisher","first-page":"2068","DOI":"10.1109\/JPROC.2018.2841200","volume":"106","author":"P Balaprakash","year":"2018","unstructured":"Balaprakash P, Dongarra J, Gamblin T, Hall M, Hollingsworth JK, Norris B, Vuduc R (2018) Autotuning in high-performance computing applications. Proc IEEE 106(11):2068\u20132083","journal-title":"Proc IEEE"},{"issue":"4","key":"4577_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3320060","volume":"52","author":"T Ben-Nun","year":"2019","unstructured":"Ben-Nun T, Hoefler T (2019) Demystifying parallel and distributed deep learning: an in-depth concurrency analysis. ACM Comput Surv (CSUR) 52(4):1\u201343","journal-title":"ACM Comput Surv (CSUR)"},{"key":"4577_CR22","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition, arXiv:1409.1556"},{"key":"4577_CR23","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition, In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04577-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-022-04577-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04577-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,14]],"date-time":"2022-12-14T19:05:37Z","timestamp":1671044737000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-022-04577-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,20]]},"references-count":23,"journal-issue":{"issue":"16","published-print":{"date-parts":[[2022,11]]}},"alternative-id":["4577"],"URL":"https:\/\/doi.org\/10.1007\/s11227-022-04577-2","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2022,5,20]]},"assertion":[{"value":"30 April 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 May 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}