{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,24]],"date-time":"2025-10-24T16:41:44Z","timestamp":1761324104962,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,6,14]],"date-time":"2017-06-14T00:00:00Z","timestamp":1497398400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,6,14]]},"DOI":"10.1145\/3079079.3079081","type":"proceedings-article","created":{"date-parts":[[2017,5,31]],"date-time":"2017-05-31T19:31:40Z","timestamp":1496259100000},"page":"1-10","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Compile-time optimized and statically scheduled N-D convnet primitives for multi-core and many-core (Xeon Phi) CPUs"],"prefix":"10.1145","author":[{"given":"Aleksandar","family":"Zlateski","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology and Princeton University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"H Sebastian","family":"Seung","sequence":"additional","affiliation":[{"name":"Princeton University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,6,14]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2015. Myth Busted: General Purpose CPUs Can't Tackle Deep Neural Network Training. (2015). http:\/\/itpeernetwork.intel.com\/myth-busted-general-purpose-cpus-cant-tackle-deep-neural-network-training\/  2015. Myth Busted: General Purpose CPUs Can't Tackle Deep Neural Network Training. (2015). http:\/\/itpeernetwork.intel.com\/myth-busted-general-purpose-cpus-cant-tackle-deep-neural-network-training\/"},{"key":"e_1_3_2_1_2_1","unstructured":"2015. NervanaGPU library. https:\/\/github.com\/NervanaSystems\/nervanagpu. (2015).  2015. NervanaGPU library. https:\/\/github.com\/NervanaSystems\/nervanagpu. (2015)."},{"key":"e_1_3_2_1_3_1","unstructured":"2016. Imagenet Winners Benchmarking. (2016). https:\/\/github.com\/soumith\/convnet-benchmarks  2016. Imagenet Winners Benchmarking. (2016). https:\/\/github.com\/soumith\/convnet-benchmarks"},{"key":"e_1_3_2_1_4_1","unstructured":"2016. The Intel(R) Deep Learning Framework. (2016). https:\/\/github.com\/01org\/idlf  2016. The Intel(R) Deep Learning Framework. (2016). https:\/\/github.com\/01org\/idlf"},{"key":"e_1_3_2_1_5_1","volume-title":"https:\/\/github.com\/01org\/mkl-dnn","author":"Deep Neural Networks Math Kernel","year":"2016","unstructured":"2016. Intel(R) Math Kernel Library for Deep Neural Networks . ( 2016 ). https:\/\/github.com\/01org\/mkl-dnn 2016. Intel(R) Math Kernel Library for Deep Neural Networks. (2016). https:\/\/github.com\/01org\/mkl-dnn"},{"key":"e_1_3_2_1_6_1","volume-title":"Shraman Ray Chaudhuri, and Nir Shavit","author":"Budden David","year":"2016","unstructured":"David Budden , Alexander Matveev , Shibani Santurkar , Shraman Ray Chaudhuri, and Nir Shavit . 2016 . Deep Tensor Convolution on Multicores . arXiv preprint arXiv:1611.06565 (2016). David Budden, Alexander Matveev, Shibani Santurkar, Shraman Ray Chaudhuri, and Nir Shavit. 2016. Deep Tensor Convolution on Multicores. arXiv preprint arXiv:1611.06565 (2016)."},{"key":"e_1_3_2_1_7_1","volume-title":"Tenth International Workshop on Frontiers in Handwriting Recognition. Suvisoft.","author":"Chellapilla Kumar","year":"2006","unstructured":"Kumar Chellapilla , Sidd Puri , and Patrice Simard . 2006 . High performance convolutional neural networks for document processing . In Tenth International Workshop on Frontiers in Handwriting Recognition. Suvisoft. Kumar Chellapilla, Sidd Puri, and Patrice Simard. 2006. High performance convolutional neural networks for document processing. In Tenth International Workshop on Frontiers in Handwriting Recognition. Suvisoft."},{"key":"e_1_3_2_1_8_1","volume-title":"cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur , Cliff Woolley , Philippe Vandermersch , Jonathan Cohen , John Tran , Bryan Catanzaro , and Evan Shelhamer . 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 ( 2014 ). Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014)."},{"key":"e_1_3_2_1_9_1","volume-title":"Distributed deep learning using synchronous stochastic gradient descent. arXiv preprint arXiv:1602.06709","author":"Das Dipankar","year":"2016","unstructured":"Dipankar Das , Sasikanth Avancha , Dheevatsa Mudigere , Karthikeyan Vaidynathan , Srinivas Sridharan , Dhiraj Kalamkar , Bharat Kaul , and Pradeep Dubey . 2016. Distributed deep learning using synchronous stochastic gradient descent. arXiv preprint arXiv:1602.06709 ( 2016 ). Dipankar Das, Sasikanth Avancha, Dheevatsa Mudigere, Karthikeyan Vaidynathan, Srinivas Sridharan, Dhiraj Kalamkar, Bharat Kaul, and Pradeep Dubey. 2016. Distributed deep learning using synchronous stochastic gradient descent. arXiv preprint arXiv:1602.06709 (2016)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","unstructured":"J. Deng W. Dong R. Socher L.-J. Li K. Li and L. Fei-Fei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09.  J. Deng W. Dong R. Socher L.-J. Li K. Li and L. Fei-Fei. 2009. ImageNet: A Large-Scale Hierarchical Image Database. In CVPR09 .","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_11_1","volume-title":"Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on","volume":"3","author":"Frigo Matteo","year":"1998","unstructured":"Matteo Frigo and Steven G Johnson . 1998 . FFTW: An adaptive software architecture for the FFT. In Acoustics , Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on , Vol. 3 . IEEE, 1381--1384. Matteo Frigo and Steven G Johnson. 1998. FFTW: An adaptive software architecture for the FFT. In Acoustics, Speech and Signal Processing, 1998. Proceedings of the 1998 IEEE International Conference on, Vol. 3. IEEE, 1381--1384."},{"key":"e_1_3_2_1_12_1","volume-title":"FFTW user's manual","author":"Frigo Matteo","year":"1999","unstructured":"Matteo Frigo and Steven G Johnson . 1999. FFTW user's manual . Massachusetts Institute of Technology ( 1999 ). Matteo Frigo and Steven G Johnson. 1999. FFTW user's manual. Massachusetts Institute of Technology (1999)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2799562.2799641"},{"key":"e_1_3_2_1_14_1","unstructured":"Jim Jeffers and James Reinders. 2015. High Performance Parallelism Pearls Volume Two: Multicore and Many-core Programming Approaches. Morgan Kaufmann.   Jim Jeffers and James Reinders. 2015. High Performance Parallelism Pearls Volume Two: Multicore and Many-core Programming Approaches . Morgan Kaufmann."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Jim Jeffers James Reinders and Avinash Sodani. 2016. Intel Xeon Phi Processor High Performance Programming: Knights Landing Edition Edition 2. Morgan Kaufmann.   Jim Jeffers James Reinders and Avinash Sodani. 2016. Intel Xeon Phi Processor High Performance Programming: Knights Landing Edition Edition 2 . Morgan Kaufmann.","DOI":"10.1016\/B978-0-12-809194-4.00004-1"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2011.09.001"},{"key":"e_1_3_2_1_18_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105.   Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems . 1097--1105."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.435"},{"key":"e_1_3_2_1_20_1","volume-title":"Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Detection. arXiv preprint arXiv:1508.04843","author":"Lee Kisuk","year":"2015","unstructured":"Kisuk Lee , Aleksandar Zlateski , Ashwin Vishwanathan , and H Sebastian Seung . 2015. Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Detection. arXiv preprint arXiv:1508.04843 ( 2015 ). Kisuk Lee, Aleksandar Zlateski, Ashwin Vishwanathan, and H Sebastian Seung. 2015. Recursive Training of 2D-3D Convolutional Networks for Neuronal Boundary Detection. arXiv preprint arXiv:1508.04843 (2015)."},{"key":"e_1_3_2_1_21_1","volume-title":"Fully Convolutional Networks for Semantic Segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Long Jonathan","year":"2015","unstructured":"Jonathan Long , Evan Shelhamer , and Trevor Darrell . 2015 . Fully Convolutional Networks for Semantic Segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully Convolutional Networks for Semantic Segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_1_22_1","volume-title":"International Conference on Learning Representations (ICLR2014)","author":"Mathieu Michael","year":"2014","unstructured":"Michael Mathieu , Mikael Henaff , and Yann LeCun . 2014 . Fast Training of Convolutional Networks through FFTs . In International Conference on Learning Representations (ICLR2014) . CBLS. Michael Mathieu, Mikael Henaff, and Yann LeCun. 2014. Fast Training of Convolutional Networks through FFTs. In International Conference on Learning Representations (ICLR2014). CBLS."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"D. Maturana and S. Scherer. 2015. 3D Convolutional Neural Networks for Landing Zone Detection from LiDAR. In ICRA.  D. Maturana and S. Scherer. 2015. 3D Convolutional Neural Networks for Landing Zone Detection from LiDAR. In ICRA .","DOI":"10.1109\/ICRA.2015.7139679"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"D. Maturana and S. Scherer. 2015. VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition. In IROS.  D. Maturana and S. Scherer. 2015. VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition. In IROS .","DOI":"10.1109\/IROS.2015.7353481"},{"volume-title":"https:\/\/github.com\/NervanaSystems\/neon","year":"2015","key":"e_1_3_2_1_25_1","unstructured":"Nervana. 2015. Neon. ( 2015 ). https:\/\/github.com\/NervanaSystems\/neon Nervana. 2015. Neon. (2015). https:\/\/github.com\/NervanaSystems\/neon"},{"volume-title":"Intel threading building blocks: outfitting C++ for multi-core processor parallelism. \" O'Reilly Media","author":"Reinders James","key":"e_1_3_2_1_26_1","unstructured":"James Reinders . 2007. Intel threading building blocks: outfitting C++ for multi-core processor parallelism. \" O'Reilly Media , Inc .\". James Reinders. 2007. Intel threading building blocks: outfitting C++ for multi-core processor parallelism. \" O'Reilly Media, Inc.\"."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_29_1","volume-title":"Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229","author":"Sermanet Pierre","year":"2013","unstructured":"Pierre Sermanet , David Eigen , Xiang Zhang , Micha\u00ebl Mathieu , Rob Fergus , and Yann LeCun . 2013 . Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229 (2013). Pierre Sermanet, David Eigen, Xiang Zhang, Micha\u00ebl Mathieu, Rob Fergus, and Yann LeCun. 2013. Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv preprint arXiv:1312.6229 (2013)."},{"key":"e_1_3_2_1_30_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_1_33_1","volume-title":"Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580","author":"Vasilache Nicolas","year":"2014","unstructured":"Nicolas Vasilache , Jeff Johnson , Michael Mathieu , Soumith Chintala , Serkan Piantino , and Yann LeCun . 2014. Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580 ( 2014 ). Nicolas Vasilache, Jeff Johnson, Michael Mathieu, Soumith Chintala, Serkan Piantino, and Yann LeCun. 2014. Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580 (2014)."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1370082.1370085"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2016.119"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3015002"}],"event":{"name":"ICS '17: 2017 International Conference on Supercomputing","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"],"location":"Chicago Illinois","acronym":"ICS '17"},"container-title":["Proceedings of the International Conference on Supercomputing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3079079.3079081","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3079079.3079081","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:03:25Z","timestamp":1750215805000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3079079.3079081"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,6,14]]},"references-count":36,"alternative-id":["10.1145\/3079079.3079081","10.1145\/3079079"],"URL":"https:\/\/doi.org\/10.1145\/3079079.3079081","relation":{},"subject":[],"published":{"date-parts":[[2017,6,14]]},"assertion":[{"value":"2017-06-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}