{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,21]],"date-time":"2025-06-21T11:27:17Z","timestamp":1750505237259,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":15,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,11,12]],"date-time":"2017-11-12T00:00:00Z","timestamp":1510444800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,11,12]]},"DOI":"10.1145\/3146347.3146358","type":"proceedings-article","created":{"date-parts":[[2017,10,31]],"date-time":"2017-10-31T12:31:37Z","timestamp":1509453097000},"page":"1-8","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Training distributed deep recurrent neural networks with mixed precision on GPU clusters"],"prefix":"10.1145","author":[{"given":"Alexey","family":"Svyatkovskiy","sequence":"first","affiliation":[{"name":"Princeton University, Princeton, New Jersey"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Julian","family":"Kates-Harbeck","sequence":"additional","affiliation":[{"name":"Harvard University, Cambridge, Massachusetts"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"William","family":"Tang","sequence":"additional","affiliation":[{"name":"Princeton University, Princeton, New Jersey"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,11,12]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. (2015). http:\/\/tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. (2015). http:\/\/tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_3_2_1_2_1","unstructured":"Matthieu Courbariaux Yoshua Bengio and Jean-Pierre David. 2014. Training deep neural networks with low precision multiplications. arXiv e-prints abs\/1412.7024 (Dec. 2014). http:\/\/arxiv.org\/abs\/1412.7024  Matthieu Courbariaux Yoshua Bengio and Jean-Pierre David. 2014. Training deep neural networks with low precision multiplications. arXiv e-prints abs\/1412.7024 (Dec. 2014). http:\/\/arxiv.org\/abs\/1412.7024"},{"key":"e_1_3_2_1_3_1","unstructured":"Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang Quoc V. Le and Andrew Y. Ng. 2012. Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25 F. Pereira C. J. C. Burges L. Bottou and K. Q. Weinberger (Eds.). Curran Associates Inc. 1223--1231. http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf  Jeffrey Dean Greg Corrado Rajat Monga Kai Chen Matthieu Devin Mark Mao Marc'aurelio Ranzato Andrew Senior Paul Tucker Ke Yang Quoc V. Le and Andrew Y. Ng. 2012. Large Scale Distributed Deep Networks. In Advances in Neural Information Processing Systems 25 F. Pereira C. J. C. Burges L. Bottou and K. Q. Weinberger (Eds.). Curran Associates Inc. 1223--1231. http:\/\/papers.nips.cc\/paper\/4687-large-scale-distributed-deep-networks.pdf"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Felix A Gers J\u00fcrgen Schmidhuber and Fred Cummins. 1999. Learning to forget: Continual prediction with LSTM. (1999).  Felix A Gers J\u00fcrgen Schmidhuber and Fred Cummins. 1999. Learning to forget: Continual prediction with LSTM. (1999).","DOI":"10.1049\/cp:19991218"},{"key":"e_1_3_2_1_5_1","unstructured":"Itay Hubara Matthieu Courbariaux Daniel Soudry Ran El-Yaniv and Yoshua Bengio. 2016. Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations. arXiv e-prints abs\/1609.07061 (Sept. 2016). https:\/\/arxiv.org\/abs\/1609.07061  Itay Hubara Matthieu Courbariaux Daniel Soudry Ran El-Yaniv and Yoshua Bengio. 2016. Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations. arXiv e-prints abs\/1609.07061 (Sept. 2016). https:\/\/arxiv.org\/abs\/1609.07061"},{"key":"e_1_3_2_1_6_1","unstructured":"Rafal J\u00f3zefowicz Oriol Vinyals Mike Schuster Noam Shazeer and Yonghui Wu. 2016. Exploring the Limits of Language Modeling. CoRR abs\/1602.02410 (2016). http:\/\/arxiv.org\/abs\/1602.02410  Rafal J\u00f3zefowicz Oriol Vinyals Mike Schuster Noam Shazeer and Yonghui Wu. 2016. Exploring the Limits of Language Modeling. CoRR abs\/1602.02410 (2016). http:\/\/arxiv.org\/abs\/1602.02410"},{"key":"e_1_3_2_1_7_1","unstructured":"Julian Kates-Harbeck Alexey Svyatkovskiy Kyle Felker Eliot Feibush and William Tang. 2017. Disruption Forecasting in Tokamak Fusion Plasmas using Deep Recurrent Neural Networks. Manuscript in preparation (2017).  Julian Kates-Harbeck Alexey Svyatkovskiy Kyle Felker Eliot Feibush and William Tang. 2017. Disruption Forecasting in Tokamak Fusion Plasmas using Deep Recurrent Neural Networks. Manuscript in preparation (2017)."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2009.5272262"},{"volume-title":"Advances in Neural Information Processing Systems 25","author":"Krizhevsky Alex","key":"e_1_3_2_1_9_1"},{"volume":"48","volume-title":"Proceedings of the 33rd International Conference on International Conference on Machine Learning -","author":"Lin Darryl D.","key":"e_1_3_2_1_10_1"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002491"},{"key":"e_1_3_2_1_12_1","unstructured":"Junhua Mao Wei Xu Yi Yang Jiang Wang Zhiheng Huang and Alan Yuille. 2015. Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN). ICLR (2015).  Junhua Mao Wei Xu Yi Yang Jiang Wang Zhiheng Huang and Alan Yuille. 2015. Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN). ICLR (2015)."},{"volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics","year":"2013","author":"Socher Richard","key":"e_1_3_2_1_13_1"},{"volume-title":"Advances in Neural Information Processing Systems 27","author":"Sutskever Ilya","key":"e_1_3_2_1_14_1"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.fusengdes.2013.03.003"}],"event":{"name":"SC '17: The International Conference for High Performance Computing, Networking, Storage and Analysis","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","IEEE CS"],"location":"Denver CO USA","acronym":"SC '17"},"container-title":["Proceedings of the Machine Learning on HPC Environments"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3146347.3146358","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3146347.3146358","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:13:33Z","timestamp":1750212813000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3146347.3146358"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,11,12]]},"references-count":15,"alternative-id":["10.1145\/3146347.3146358","10.1145\/3146347"],"URL":"https:\/\/doi.org\/10.1145\/3146347.3146358","relation":{},"subject":[],"published":{"date-parts":[[2017,11,12]]},"assertion":[{"value":"2017-11-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}