{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T22:51:56Z","timestamp":1778971916728,"version":"3.51.4"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2017,8]]},"abstract":"<jats:p>Training neural language models (NLMs) is very time consuming and we need parallelization for system speedup. However, standard training methods have poor scalability across multiple devices (e.g., GPUs) due to the huge time cost required to transmit data for gradient sharing in the back-propagation process. In this paper we present a sampling-based approach to reducing data transmission for better scaling of NLMs. As a ''bonus'', the resulting model also improves the training speed on a single device. Our approach yields significant speed improvements on a recurrent neural network-based language model. On four NVIDIA GTX1080 GPUs, it achieves a speedup of 2.1+ times over the standard asynchronous stochastic gradient descent baseline, yet with no increase in perplexity. This is even 4.2 times faster than the naive single GPU counterpart.<\/jats:p>","DOI":"10.24963\/ijcai.2017\/586","type":"proceedings-article","created":{"date-parts":[[2017,7,28]],"date-time":"2017-07-28T09:14:07Z","timestamp":1501233247000},"page":"4193-4199","source":"Crossref","is-referenced-by-count":3,"title":["Fast Parallel Training of Neural Language Models"],"prefix":"10.24963","author":[{"given":"Tong","family":"Xiao","sequence":"first","affiliation":[{"name":"Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingbo","family":"Zhu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tongran","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of Psychology (CAS), Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chunliang","family":"Zhang","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"name":"Twenty-Sixth International Joint Conference on Artificial Intelligence","theme":"Artificial Intelligence","location":"Melbourne, Australia","acronym":"IJCAI-2017","number":"26","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)","University of Technology Sydney (UTS)","Australian Computer Society (ACS)"],"start":{"date-parts":[[2017,8,19]]},"end":{"date-parts":[[2017,8,26]]}},"container-title":["Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2017,7,28]],"date-time":"2017-07-28T11:54:39Z","timestamp":1501242879000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2017\/586"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2017,8]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2017\/586","relation":{},"subject":[],"published":{"date-parts":[[2017,8]]}}}