{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,4,5]],"date-time":"2022-04-05T01:41:57Z","timestamp":1649122917006},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2013,10,12]],"date-time":"2013-10-12T00:00:00Z","timestamp":1381536000000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2014,4]]},"DOI":"10.1007\/s11227-013-1023-7","type":"journal-article","created":{"date-parts":[[2013,10,11]],"date-time":"2013-10-11T18:09:12Z","timestamp":1381514952000},"page":"65-86","source":"Crossref","is-referenced-by-count":1,"title":["A CUDA implementation of the Continuous Space Language Model"],"prefix":"10.1007","volume":"68","author":[{"given":"Elizabeth A.","family":"Thompson","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Timothy R.","family":"Anderson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2013,10,12]]},"reference":[{"key":"1023_CR1","volume-title":"Proceedings of the IEEE international conference on cluster computing and workshops (CLUSTER)","author":"V Allada","year":"2009","unstructured":"Allada V, Benjegerdes T, Bode B (2009) Performance analysis of memory transfers and GEMM subroutines on NVIDIA Tesla GPU cluster. In: Proceedings of the IEEE international conference on cluster computing and workshops (CLUSTER), New Orleans, LA, Aug 31\u2013Sept 4, 2009"},{"key":"1023_CR2","volume-title":"Proceedings of the 17th IEEE euromicro international conference on parallel, distributed, and network-based processing (PDP)","author":"J Franco","year":"2009","unstructured":"Franco J, Bernabe G, Fernandez J, Acacio ME (2009) A parallel implementation of the 2D wavelet transform using CUDA. In: Proceedings of the 17th IEEE euromicro international conference on parallel, distributed, and network-based processing (PDP), Weimar, Germany, Feb 18\u201320, 2009"},{"key":"1023_CR3","volume-title":"Proceedings of the 24th IEEE international symposium on parallel and distributed processing (IPDPS)","author":"EH Phillips","year":"2010","unstructured":"Phillips EH, Fatica M (2010) Implementing the Himeno benchmark with CUDA on GPU clusters. In: Proceedings of the 24th IEEE international symposium on parallel and distributed processing (IPDPS), Atlanta, GA, Apr 19\u201323, 2010"},{"key":"1023_CR4","volume-title":"Proceedings of the IEEE international symposium on parallel and distributed processing, workshops, and PhD forum (IPDPSW)","author":"Z Du","year":"2010","unstructured":"Du Z, Yin Z, Bader DA (2010) A tile-based parallel Viterbi algorithm for biological sequence alignment on GPU with CUDA. In: Proceedings of the IEEE international symposium on parallel and distributed processing, workshops, and PhD forum (IPDPSW), Atlanta, GA, Apr 19\u201323, 2010"},{"issue":"1","key":"1023_CR5","doi-asserted-by":"crossref","first-page":"132","DOI":"10.1109\/TPDS.2010.143","volume":"22","author":"WJ Der Laan Van","year":"2011","unstructured":"Van Der Laan WJ, Jalba AC, Roerdink J (2011) Accelerating wavelet lifting on graphics hardware using CUDA. IEEE Trans Parallel Distrib Syst 22(1):132\u2013146","journal-title":"IEEE Trans Parallel Distrib Syst"},{"issue":"10","key":"1023_CR6","doi-asserted-by":"crossref","first-page":"B83","DOI":"10.1364\/AO.49.000B83","volume":"49","author":"B Han","year":"2010","unstructured":"Han B, Taha TM (2010) Acceleration of spiking neural network based pattern recognition on NVIDIA graphics processors. Appl Opt 49(10):B83\u2013B91","journal-title":"Appl Opt"},{"issue":"10","key":"1023_CR7","doi-asserted-by":"crossref","first-page":"1370","DOI":"10.1016\/j.jpdc.2008.05.014","volume":"68","author":"S Che","year":"2008","unstructured":"Che S, Boyer M, Meng J, Tarjan D, Sheaffer JW, Skadron K (2008) A performance study of general-purpose applications on graphics processors using CUDA. J Parallel Distrib Comput 68(10):1370\u20131380","journal-title":"J Parallel Distrib Comput"},{"issue":"5","key":"1023_CR8","doi-asserted-by":"crossref","first-page":"451","DOI":"10.1016\/j.jpdc.2009.01.006","volume":"69","author":"D Komatitsch","year":"2009","unstructured":"Komatitsch D, Michea D, Erlebacher G (2009) Porting a high-order finite-element earthquake modeling application to NVIDIA graphics cards using CUDA. J Parallel Distrib Comput 69(5):451\u2013460","journal-title":"J Parallel Distrib Comput"},{"key":"1023_CR9","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","volume":"27","author":"CE Shannon","year":"1948","unstructured":"Shannon CE (1948) A mathematical theory of communication. Bell Syst Tech J 27:379\u2013423, 623\u2013656","journal-title":"Bell Syst Tech J"},{"issue":"3","key":"1023_CR10","doi-asserted-by":"crossref","first-page":"400","DOI":"10.1109\/TASSP.1987.1165125","volume":"35","author":"SM Katz","year":"1987","unstructured":"Katz SM (1987) Estimation of probabilities from sparse data for the language model component of a speech recognizer. IEEE Trans Acoust Speech Signal Process 35(3):400\u2013401","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"key":"1023_CR11","doi-asserted-by":"crossref","first-page":"137","DOI":"10.2478\/v10108-010-0014-6","volume":"93","author":"H Schwenk","year":"2010","unstructured":"Schwenk H (2010) Continuous-space language models for statistical machine translation. Prague Bull Math Linguist 93:137\u2013146","journal-title":"Prague Bull Math Linguist"},{"key":"1023_CR12","unstructured":"Schwenk H (2013) CSLM: Continuous Space Language Model toolkit. LIUM, University of Le Mans, France, 11 Sept (2012). www-lium.univ-lemans.fr\/cslm\/ . Accessed 3 Sept 2013"},{"key":"1023_CR13","doi-asserted-by":"crossref","first-page":"492","DOI":"10.1016\/j.csl.2006.09.003","volume":"21","author":"H Schwenk","year":"2007","unstructured":"Schwenk H (2007) Continuous space language models. Comput Speech Lang 21:492\u2013518","journal-title":"Comput Speech Lang"},{"key":"1023_CR14","volume-title":"Proceedings of the joint conference ACL\/Coling","author":"H Schwenk","year":"2006","unstructured":"Schwenk H, Dechelotte D, Gauvain J-L (2006) Continuous space language models for statistical machine translation. In: Proceedings of the joint conference ACL\/Coling, July 2006"},{"key":"1023_CR15","unstructured":"Whaley RC, Petitet A (2013) Automatically Tuned Linear Algebra Software (ATLAS). SourceForge, 10 July (2012). http:\/\/math-atlas.sourceforge.net\/ . Accessed 3 Sept 2013"},{"key":"1023_CR16","volume-title":"Proceedings of the IEEE high performance extreme computing conference (HPEC)","author":"EA Thompson","year":"2012","unstructured":"Thompson EA, Anderson T (2012) Use of CUDA for the continuous space language model. In: Proceedings of the IEEE high performance extreme computing conference (HPEC), Waltham, MA, Sept 10\u201312, 2012"},{"key":"1023_CR17","volume-title":"Proceedings of the 11th annual conference of the international speech communication association (INTERSPEECH)","author":"K Vesely","year":"2010","unstructured":"Vesely K, Burget L, Grezl F (2010) Parallel training of neural networks for speech recognition. In: Proceedings of the 11th annual conference of the international speech communication association (INTERSPEECH), Mukuhari, Chiba, Japan, Sept 26\u201330, 2010"},{"key":"1023_CR18","volume-title":"Proceedings of the 26th international conference on machine learning (ICML)","author":"R Raina","year":"2009","unstructured":"Raina R, Madhavan A, Ng AY (2009) Large-scale unsupervised learning using graphics processors. In: Proceedings of the 26th international conference on machine learning (ICML), Montreal, QC, Canada, June 14\u201318, 2009"},{"key":"1023_CR19","volume-title":"Proceedings of the 2012 annual international joint conference on neural networks (IJCNN), part of the 2012 IEEE world Congress on computational intelligence (WCCI)","author":"N Lopes","year":"2012","unstructured":"Lopes N, Ribeiro B, Goncalves J (2012) Restricted Boltzmann machines and deep belief networks on multi-core processors. In: Proceedings of the 2012 annual international joint conference on neural networks (IJCNN), part of the 2012 IEEE world Congress on computational intelligence (WCCI), Brisbane, QLD, Australia, June 10\u201315, 2012"},{"key":"1023_CR20","unstructured":"NVIDIA Performance Primitives (NPP) version 5.0. 7 Sept 2012. https:\/\/developer.nvidia.com\/sites\/default\/files\/akamai\/cuda\/files\/CUDADownloads\/NPP_Library.pdf . Accessed 3 Sept 2013"},{"key":"1023_CR21","unstructured":"OpenCL programming guide for the CUDA architecture, version 2.3. NVIDIA, 27 Aug 2009. http:\/\/www.nvidia.com\/content\/cudazone\/download\/OpenCL\/NVIDIA_OpenCL_ProgrammingGuide.pdf . Accessed 3 Sept 2013"},{"key":"1023_CR22","unstructured":"Intel Math Kernel Library 11.0 (2013). http:\/\/software.intel.com\/en-us\/intel-mkl . Accessed 3 Sept 2013"},{"key":"1023_CR23","unstructured":"BLAS (basic linear algebra subprograms). Based upon work supported by the National Science Foundation under Grant No. ASC-9313958 and DOE Grant No. DE-FG0-3-94ER25219, 29 June 2013. http:\/\/www.netlib.org\/blas\/ . Accessed 4 Sept 2013"},{"issue":"1","key":"1023_CR24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/77626.79170","volume":"16","author":"JJ Dongarra","year":"1990","unstructured":"Dongarra JJ, Du Croz J, Hammarling S, Duff I (1990) A set of level 3 basic linear algebra subprograms. ACM Trans Math Softw 16(1):1\u201317","journal-title":"ACM Trans Math Softw"},{"key":"1023_CR25","unstructured":"Multicore CPU: how to disable a core. Kioskea, Aug 2013. http:\/\/en.kioskea.net\/faq\/616-multicore-cpu-how-to-disable-a-core . Accessed 3 Sept 2013"},{"key":"1023_CR26","volume-title":"Proceedings of the 13th annual conference of the international speech communication association (INTERSPEECH)","author":"X Chen","year":"2012","unstructured":"Chen X, Eversole A, Li G, Yu D, Seide F (2012) Pipelined back-propagation for context-dependent deep neural networks. In: Proceedings of the 13th annual conference of the international speech communication association (INTERSPEECH), Portland, OR, Sept 9\u201313, 2012"},{"key":"1023_CR27","volume-title":"Proceedings of the 22nd IEEE international parallel and distributed processing symposium (IPDPS)","author":"S Barrachina","year":"2008","unstructured":"Barrachina S, Castillo M, Igual FD, Mayo R, Quintana-Orti ES (2008) Evaluation and tuning of the level 3 CUBLAS for graphics processors. In: Proceedings of the 22nd IEEE international parallel and distributed processing symposium (IPDPS), Miami, FL, Apr 14\u201318, 2008"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-1023-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-013-1023-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-013-1023-7","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,6,1]],"date-time":"2019-06-01T10:40:30Z","timestamp":1559385630000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-013-1023-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,10,12]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,4]]}},"alternative-id":["1023"],"URL":"https:\/\/doi.org\/10.1007\/s11227-013-1023-7","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,10,12]]}}}