{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,4,6]],"date-time":"2025-04-06T04:03:36Z","timestamp":1743912216253,"version":"3.40.3"},"reference-count":37,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2025,4,1]]},"DOI":"10.1587\/transinf.2024edp7074","type":"journal-article","created":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T22:12:06Z","timestamp":1731017526000},"page":"392-402","source":"Crossref","is-referenced-by-count":0,"title":["Lightweight Neural Data Sequence Modeling by Scale Causal Blocks"],"prefix":"10.1587","volume":"E108.D","author":[{"given":"Hiroaki","family":"AKUTSU","sequence":"first","affiliation":[{"name":"R&amp;D Group, Hitachi, Ltd."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ko","family":"ARAI","sequence":"additional","affiliation":[{"name":"R&amp;D Group, Hitachi, Ltd."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"unstructured":"[1] H. Akutsu and K. Arai, \u201cFast autoregressive bit sequence modeling for lossless compression,\u201d ICML 2023 Workshop Neural Compression: From Information Theory to Applications, 2023.","key":"1"},{"unstructured":"[2] Y. Wu, M. Schuster, Z. Chen, Q.V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., \u201cGoogle\u2019s neural machine translation system: Bridging the gap between human and machine translation,\u201d arXiv preprint arXiv:1609.08144, 2016. 10.48550\/arXiv.1609.08144","key":"2"},{"unstructured":"[3] A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, \u201cWaveNet: A generative model for raw audio,\u201d arXiv preprint arXiv:1609.03499, 2016. 10.48550\/arXiv.1609.03499","key":"3"},{"unstructured":"[4] A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al., \u201cConditional image generation with pixelCNN decoders,\u201d Advances in Neural Information Processing Systems, vol.29, 2016.","key":"4"},{"doi-asserted-by":"crossref","unstructured":"[5] F. Mentzer, E. Agustsson, M. Tschannen, R. Timofte, and L. Van Gool, \u201cConditional probability models for deep image compression,\u201d Proc. IEEE Conference on Computer Vision and Pattern Recognition, pp.4394-4402, 2018. 10.1109\/cvpr.2018.00462","key":"5","DOI":"10.1109\/CVPR.2018.00462"},{"unstructured":"[6] D. Minnen, J. Ball\u00e9, and G.D. Toderici, \u201cJoint autoregressive and hierarchical priors for learned image compression,\u201d Advances in Neural Information Processing Systems, vol.31, 2018.","key":"6"},{"doi-asserted-by":"crossref","unstructured":"[7] G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao, \u201cDVC: An end-to-end deep video compression framework,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp.11006-11015, 2019. 10.1109\/cvpr.2019.01126","key":"7","DOI":"10.1109\/CVPR.2019.01126"},{"unstructured":"[8] F. Mentzer, G. Toderici, D. Minnen, S. Caelles, S.J. Hwang, M. Lucic, and E. Agustsson, \u201cVCT: A video compression transformer,\u201d Advances in Neural Information Processing Systems, vol.35, 2022.","key":"8"},{"unstructured":"[9] F. Bellard, \u201cNNCP v2: Lossless data compression with transformer.\u201d https:\/\/bellard.org\/nncp\/nncp_v2.1.pdf, Feb. 2021.","key":"9"},{"unstructured":"[10] G.N.N. Martin, \u201cRange encoding: An algorithm for removing redundancy from a digitised message,\u201d Proc. Institution of Electronic and Radio Engineers International Conference on Video and Data Recording, 1979.","key":"10"},{"doi-asserted-by":"publisher","unstructured":"[11] D. Marpe, H. Schwarz, and T. Wiegand, \u201cContext-based adaptive binary arithmetic coding in the H.264\/AVC video compression standard,\u201d IEEE Trans. Circuits Syst. Video Technol., vol.13, no.7, pp.620-636, 2003. 10.1109\/tcsvt.2003.815173","key":"11","DOI":"10.1109\/TCSVT.2003.815173"},{"unstructured":"[12] J. Duda, \u201cAsymmetric numeral systems,\u201d arXiv preprint arXiv:0902.0271, 2009. 10.48550\/arXiv.0902.0271","key":"12"},{"unstructured":"[13] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, \u0141. Kaiser, and I. Polosukhin, \u201cAttention is all you need,\u201d Advances in Neural Information Processing Systems, vol.30, 2017.","key":"13"},{"unstructured":"[14] R. Child, S. Gray, A. Radford, and I. Sutskever, \u201cGenerating long sequences with sparse transformers,\u201d arXiv preprint arXiv:1904.10509, 2019. 10.48550\/arXiv.1904.10509","key":"14"},{"unstructured":"[15] N. Kitaev, \u0141. Kaiser, and A. Levskaya, \u201cReformer: The efficient transformer,\u201d arXiv preprint arXiv:2001.04451, 2020. 10.48550\/arXiv.2001.04451","key":"15"},{"unstructured":"[16] K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarl\u00f3s, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, D. Belanger, L.J. Colwell, and A. Weller, \u201cRethinking attention with performers,\u201d CoRR, vol.abs\/2009.14794, 2020.","key":"16"},{"unstructured":"[17] A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, \u201cTransformers are RNNs: Fast autoregressive transformers with linear attention,\u201d International Conference on Machine Learning, pp.5156-5165, PMLR, 2020.","key":"17"},{"unstructured":"[18] X. Ma, X. Kong, S. Wang, C. Zhou, J. May, H. Ma, and L. Zettlemoyer, \u201cLuna: Linear unified nested attention,\u201d Advances in Neural Information Processing Systems, vol.34, pp.2441-2453, 2021.","key":"18"},{"unstructured":"[19] K. Irie, I. Schlag, R. Csord\u00e1s, and J. Schmidhuber, \u201cGoing beyond linear transformers with recurrent fast weight programmers,\u201d Advances in Neural Information Processing Systems, vol.34, pp.7703-7717, 2021.","key":"19"},{"doi-asserted-by":"publisher","unstructured":"[20] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, \u201cEfficient transformers: A survey,\u201d ACM Comput. Surv., vol.55, no.6, Article No. 109, pp.1-28, 2022. 10.1145\/3530811","key":"20","DOI":"10.1145\/3530811"},{"unstructured":"[21] S. Wang, B.Z. Li, M. Khabsa, H. Fang, and H. Ma, \u201cLinformer: Self-attention with linear complexity,\u201d CoRR, vol.abs\/2006.04768, 2020.","key":"21"},{"unstructured":"[22] T. Dao, D. Fu, S. Ermon, A. Rudra, and C. R\u00e9, \u201cFlashAttention: Fast and memory-efficient exact attention with IO-awareness,\u201d Advances in Neural Information Processing Systems, vol.35, pp.16344-16359, 2022.","key":"22"},{"unstructured":"[23] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, \u201cLanguage models are unsupervised multitask learners,\u201d OpenAI, 2019.","key":"23"},{"unstructured":"[24] H. Wang, S. Ma, L. Dong, S. Huang, D. Zhang, and F. Wei, \u201cDeepNet: Scaling transformers to 1,000 layers,\u201d arXiv preprint arXiv:2203.00555, 2022. 10.48550\/arXiv.2203.00555","key":"24"},{"unstructured":"[25] T. Salimans, A. Karpathy, X. Chen, and D.P. Kingma, \u201cPixelCNN++: Improving the pixelCNN with discretized logistic mixture likelihood and other modifications,\u201d arXiv preprint arXiv:1701.05517, 2017. 10.48550\/arXiv.1701.05517","key":"25"},{"unstructured":"[26] P. Ramachandran, T.L. Paine, P. Khorrami, M. Babaeizadeh, S. Chang, Y. Zhang, M.A. Hasegawa-Johnson, R.H. Campbell, and T.S. Huang, \u201cFast generation for convolutional autoregressive models,\u201d arXiv preprint arXiv:1704.06001, 2017. 10.48550\/arXiv.1704.06001","key":"26"},{"doi-asserted-by":"crossref","unstructured":"[27] O. Ronneberger, P. Fischer, and T. Brox, \u201cU-Net: Convolutional networks for biomedical image segmentation,\u201d Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015: 18th International Conference, Munich, Germany, Oct. 5-9, 2015, Proceedings, Part III 18, pp.234-241, Springer, 2015. 10.1007\/978-3-319-24574-4_28","key":"27","DOI":"10.1007\/978-3-319-24574-4_28"},{"unstructured":"[28] D.A. Clevert, T. Unterthiner, and S. Hochreiter, \u201cFast and accurate deep network learning by exponential linear units (ELUs),\u201d arXiv preprint arXiv:1511.07289, 2015. 10.48550\/arXiv.1511.07289","key":"28"},{"doi-asserted-by":"crossref","unstructured":"[29] W. Shi, J. Caballero, F. Husz\u00e1r, J. Totz, A.P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, \u201cReal-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,\u201d Proc. IEEE Conference on Computer Vision and Pattern Recognition, pp.1874-1883, 2016. 10.1109\/cvpr.2016.207","key":"29","DOI":"10.1109\/CVPR.2016.207"},{"unstructured":"[30] J. Ho, E. Lohn, and P. Abbeel, \u201cCompression with flows via local bits-back coding,\u201d Advances in Neural Information Processing Systems, vol.32, 2019.","key":"30"},{"unstructured":"[31] EMBL, \u201cEuropean nucleotide archive: Illumina hiseq 2000 paired end sequencing; gsm1080195: mouse oocyte 1; mus musculus; rna-seq.\u201d https:\/\/www.ebi.ac.uk\/ena, 6 2016.","key":"31"},{"doi-asserted-by":"publisher","unstructured":"[32] O.I. Alomair, I.M. Brereton, M.T. Smith, G.J. Galloway, and N.D. Kurniawan, \u201cIn vivo high angular resolution diffusion-weighted imaging of mouse brain at 16.4 tesla,\u201d PloS one, vol.10, no.6, e0130133, 2015. 10.1371\/journal.pone.0130133","key":"32","DOI":"10.1371\/journal.pone.0130133"},{"doi-asserted-by":"crossref","unstructured":"[33] P. Baldi, K. Cranmer, T. Faucett, P. Sadowski, and D. Whiteson, \u201cParameterized machine learning for high-energy physics,\u201d arXiv preprint arXiv:1601.07913, 2016. 10.48550\/arXiv.1601.07913","key":"33","DOI":"10.1140\/epjc\/s10052-016-4099-4"},{"unstructured":"[34] D.P. Kingma and J. Ba, \u201cADAM: A method for stochastic optimization,\u201d arXiv preprint arXiv:1412.6980, 2014. 10.48550\/arXiv.1412.6980","key":"34"},{"doi-asserted-by":"crossref","unstructured":"[35] P. Deutsch, \u201cGzip file format specification version 4.3,\u201d Tech. Rep., 1996.","key":"35","DOI":"10.17487\/rfc1952"},{"unstructured":"[36] L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han, \u201cOn the variance of the adaptive learning rate and beyond,\u201d arXiv preprint arXiv:1908.03265, 2019. 10.48550\/arXiv.1908.03265","key":"36"},{"unstructured":"[37] A. Krizhevsky, G. Hinton, et al., \u201cLearning multiple layers of features from tiny images,\u201d 2009.","key":"37"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/4\/E108.D_2024EDP7074\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,5]],"date-time":"2025-04-05T03:24:18Z","timestamp":1743823458000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E108.D\/4\/E108.D_2024EDP7074\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,1]]},"references-count":37,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2024edp7074","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"type":"print","value":"0916-8532"},{"type":"electronic","value":"1745-1361"}],"subject":[],"published":{"date-parts":[[2025,4,1]]},"article-number":"2024EDP7074"}}