{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T10:10:37Z","timestamp":1760609437736},"reference-count":48,"publisher":"MIT Press - Journals","issue":"4","license":[{"start":{"date-parts":[[2022,3,2]],"date-time":"2022-03-02T00:00:00Z","timestamp":1646179200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,3,23]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Representations of the world environment play a crucial role in artificial intelligence. It is often inefficient to conduct reasoning and inference directly in the space of raw sensory representations, such as pixel values of images. Representation learning allows us to automatically discover suitable representations from raw sensory data. For example, given raw sensory data, a deep neural network learns nonlinear representations at its hidden layers, which are subsequently used for classification (or regression) at its output layer. This happens implicitly during training through minimizing a supervised or unsupervised loss. In this letter, we study the dynamics of such implicit nonlinear representation learning. We identify a pair of a new assumption and a novel condition, called the on-model structure assumption and the data architecture alignment condition. Under the on-model structure assumption, the data architecture alignment condition is shown to be sufficient for the global convergence and necessary for global optimality. Moreover, our theory explains how and when increasing network size does and does not improve the training behaviors in the practical regime. Our results provide practical guidance for designing a model structure; for example, the on-model structure assumption can be used as a justification for using a particular model structure instead of others. As an application, we then derive a new training framework, which satisfies the data architecture alignment condition without assuming it by automatically modifying any given training algorithm dependent on data and architecture. Given a standard training algorithm, the framework running its modified version is empirically shown to maintain competitive (practical) test performances while providing global convergence guarantees for deep residual neural networks with convolutions, skip connections, and batch normalization with standard benchmark data sets, including MNIST, CIFAR-10, CIFAR-100, Semeion, KMNIST, and SVHN.<\/jats:p>","DOI":"10.1162\/neco_a_01483","type":"journal-article","created":{"date-parts":[[2022,3,2]],"date-time":"2022-03-02T00:50:57Z","timestamp":1646182257000},"page":"991-1018","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":1,"title":["Understanding Dynamics of Nonlinear Representation Learning and Its Application"],"prefix":"10.1162","volume":"34","author":[{"given":"Kenji","family":"Kawaguchi","sequence":"first","affiliation":[{"name":"Harvard University, Cambridge, MA 02138, U.S.A. kkawaguchi@fas.harvard.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Linjun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Rutgers University, New Brunswick, NJ 08901 linjun.zhang@rutgers.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhun","family":"Deng","sequence":"additional","affiliation":[{"name":"Harvard University Cambridge, MA 02138, U.S.A. zhundeng@g.harvard.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2022,3,23]]},"reference":[{"issue":"3","key":"2022032817111023600_B1","doi-asserted-by":"publisher","first-page":"477","DOI":"10.1162\/neco_a_01164","article-title":"Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks","volume":"31","author":"Bartlett","year":"2019","journal-title":"Neural Computation"},{"issue":"8","key":"2022032817111023600_B2","doi-asserted-by":"publisher","first-page":"1798","DOI":"10.1109\/TPAMI.2013.50","article-title":"Representation learning: A review and new perspectives","volume":"35","author":"Bengio","year":"2013","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2022032817111023600_B3","volume-title":"Advances in neural information processing systems","author":"Bengio","year":"2007"},{"key":"2022032817111023600_B4","unstructured":"Bordes, A., Glorot, X., Weston, J., & Bengio, Y. (2012). Joint learning of words and meaning representations for open-text semantic parsing. In Proceedings of the 15th International Conference on Artificial Intelligence and Statistics, 22 (pp. 127\u2013135)."},{"key":"2022032817111023600_B5","doi-asserted-by":"crossref","unstructured":"Ciregan, D., Meier, U., & Schmidhuber, J. (2012). Multi-column deep neural networks for image classification. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (pp. 3642\u20133649). Piscataway, NJ: IEEE.","DOI":"10.1109\/CVPR.2012.6248110"},{"key":"2022032817111023600_B6","unstructured":"Clanuwat, T., Bober-Irizar, M., Kitamoto, A., Lamb, A., Yamamoto, K., & Ha, D. (2019). Deep learning for classical Japanese literature. In NeurIPS Creativity Workshop 2019."},{"key":"2022032817111023600_B7","first-page":"469","volume-title":"Advances in neural information processing systems","author":"Dahl","year":"2010"},{"issue":"1","key":"2022032817111023600_B8","doi-asserted-by":"publisher","first-page":"30","DOI":"10.1109\/TASL.2011.2134090","article-title":"Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition","volume":"20","author":"Dahl","year":"2011","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"2022032817111023600_B9","doi-asserted-by":"crossref","unstructured":"Deng, L., Seltzer, M. L., Yu, D., Acero, A., Mohamed, A.-r., & Hinton, G. (2010). Binary coding of speech spectrograms using a deep auto-encoder. In Proceedings of the Eleventh Annual Conference of the International Speech Communication Association. Red Hook, NY: Curran.","DOI":"10.21437\/Interspeech.2010-487"},{"key":"2022032817111023600_B10","doi-asserted-by":"crossref","unstructured":"Dong, C., Loy, C. C., He, K., & Tang, X. (2014). Learning a deep convolutional network for image super-resolution. In Proceedings of the European Conference on Computer Vision (pp. 184\u2013199). Berlin: Springer.","DOI":"10.1007\/978-3-319-10593-2_13"},{"key":"2022032817111023600_B11","doi-asserted-by":"crossref","unstructured":"Gatys, L. A., Ecker, A. S., & Bethge, M. (2016). Image style transfer using convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2414\u20132423). Piscataway, NJ: IEEE.","DOI":"10.1109\/CVPR.2016.265"},{"key":"2022032817111023600_B12","unstructured":"Glorot, X., Bordes, A., & Bengio, Y. (2011). Domain adaptation for large-scale sentiment classification: A deep learning approach. In Proceedings of the International Conference on Machine Learning.New York: ACM."},{"key":"2022032817111023600_B13","unstructured":"Golub, G. H., & Van Loan, C. F. (1996). Matrix computations. Baltimore, MD: Johns Hopkins University Press."},{"key":"2022032817111023600_B14","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision (pp. 1026\u20131034). Piscataway, NJ: IEEE.","DOI":"10.1109\/ICCV.2015.123"},{"key":"2022032817111023600_B15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Identity mappings in deep residual networks. In Proceedings of the European Conference on Computer Vision (pp. 630\u2013645). Berlin: Springer.","DOI":"10.1007\/978-3-319-46493-0_38"},{"issue":"6","key":"2022032817111023600_B16","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Processing Magazine"},{"issue":"7","key":"2022032817111023600_B17","doi-asserted-by":"publisher","first-page":"1527","DOI":"10.1162\/neco.2006.18.7.1527","article-title":"A fast learning algorithm for deep belief nets","volume":"18","author":"Hinton","year":"2006","journal-title":"Neural Computation"},{"key":"2022032817111023600_B18","first-page":"586","volume-title":"Advances in neural information processing systems","author":"Kawaguchi","year":"2016"},{"key":"2022032817111023600_B19","unstructured":"Kawaguchi, K.\n           (2021). On the theory of implicit deep learning: Global convergence with implicit layers. In Proceedings of the International Conference on Learning Representations."},{"key":"2022032817111023600_B20","first-page":"2809","volume-title":"Advances in neural information processing systems","author":"Kawaguchi","year":"2015"},{"key":"2022032817111023600_B21","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1613\/jair.4742","article-title":"Global continuous optimization with error bound and fast convergence","volume":"56","author":"Kawaguchi","year":"2016","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2022032817111023600_B22","unstructured":"Kawaguchi, K., & Sun, Q. (2021). A recipe for global convergence guarantee in deep neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, 35 (pp. 8074\u20138082). Palo Alto, CA: AAAI."},{"key":"2022032817111023600_B23","unstructured":"Krizhevsky, A., & Hinton, G. (2009). Learning multiple layers of features from tiny images (Technical report). Citeseer."},{"key":"2022032817111023600_B24","first-page":"1097","volume-title":"Advances in neural information processing systems","author":"Krizhevsky","year":"2012"},{"key":"2022032817111023600_B25","unstructured":"Laurent, T., & Brecht, J. (2018). Deep linear networks with arbitrary loss: All local minima are global. In Proceedings of the International Conference on Machine Learning (pp. 2902\u20132907)."},{"issue":"1","key":"2022032817111023600_B26","first-page":"197","article-title":"Structured output layer neural network language models for speech recognition","volume":"21","author":"Le","year":"2012","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"issue":"7553","key":"2022032817111023600_B27","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"issue":"11","key":"2022032817111023600_B28","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proceedings of the IEEE"},{"key":"2022032817111023600_B29","doi-asserted-by":"crossref","unstructured":"Luan, F., Paris, S., Shechtman, E., & Bala, K. (2017). Deep photo style transfer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4990\u20134998). Piscataway, NJ: IEEE.","DOI":"10.1109\/CVPR.2017.740"},{"key":"2022032817111023600_B30","unstructured":"Mityagin, B.\n           (2015). The zero set of a real analytic function. arXiv:1512.07276."},{"issue":"1","key":"2022032817111023600_B31","doi-asserted-by":"publisher","first-page":"14","DOI":"10.1109\/TASL.2011.2109382","article-title":"Acoustic modeling using deep belief networks","volume":"20","author":"Mohamed","year":"2011","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"2022032817111023600_B32","unstructured":"Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., & Ng, A. Y. (2011). Reading digits in natural images with unsupervised feature learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning."},{"key":"2022032817111023600_B33","first-page":"8026","volume-title":"Advances in neural information processing systems","author":"Paszke","year":"2019"},{"key":"2022032817111023600_B34","first-page":"8024","volume-title":"Advances in neural information processing systems","author":"Paszke","year":"2019"},{"key":"2022032817111023600_B35","first-page":"2825","article-title":"Scikit-learn: Machine learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"2022032817111023600_B36","unstructured":"Poggio, T., Kawaguchi, K., Liao, Q., Miranda, B., Rosasco, L., Boix, X., Hidary, J., & Mhaskar, H. (2017). Theory of deep learning III: Explaining the non-overfitting puzzle. arXiv:1801.00173."},{"key":"2022032817111023600_B37","unstructured":"Press, W. H., Teukolsky, S. A., Vetterling, W. T., & Flannery, B. P. (2007). Numerical recipes: The art of scientific computing. (3rd ed.). Cambridge: Cambridge University Press."},{"key":"2022032817111023600_B38","first-page":"2294","volume-title":"Advances in neural information processing systems","author":"Rifai","year":"2011"},{"key":"2022032817111023600_B39","unstructured":"Saxe, A. M., McClelland, J. L., & Ganguli, S. (2014). Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. In Proceedings of the International Conference on Learning Representations."},{"key":"2022032817111023600_B40","unstructured":"Schwenk, H., Rousseau, A., & Attik, M. (2012). Large, pruned or continuous space language models on a GPU for statistical machine translation. In Proceedings of the NAACL-HLT 2012 Workshop: Will We Ever Really Replace the N-gram Model? On the Future of Language Modeling for HLT (pp. 11\u201319). Stroudsburg, PA: Association for Computational Linguistics."},{"key":"2022032817111023600_B41","doi-asserted-by":"crossref","unstructured":"Seide, F., Li, G., & Yu, D. (2011). Conversational speech transcription using context-dependent deep neural networks. In Proceedings of the Twelfth Annual Conference of the International Speech Communication Association. New York: ACM.","DOI":"10.21437\/Interspeech.2011-169"},{"key":"2022032817111023600_B42","first-page":"801","volume-title":"Advances in neural information processing systems","author":"Socher","year":"2011"},{"key":"2022032817111023600_B43","unstructured":"Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., & Manning, C. D. (2011). Semi-supervised recursive autoencoders for predicting sentiment distributions. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing (pp. 151\u2013161). Stroudsburg, PA: Association for Computational Linguistics."},{"key":"2022032817111023600_B44","unstructured":"Srl, B. T., & Brescia, I. (1994). Semeion handwritten digit data set. Rome, Italy: Semeion Research Center of Sciences of Communication."},{"issue":"1","key":"2022032817111023600_B45","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/S0164-1212(99)00062-X","article-title":"A conceptual basis for feature engineering","volume":"49","author":"Turner","year":"1999","journal-title":"Journal of Systems and Software"},{"key":"2022032817111023600_B46","unstructured":"Xu, K., Zhang, M., Jegelka, S., & Kawaguchi, K. (2021). Optimization of graph neural networks: Implicit acceleration by skip connections and more depth. In Proceedings of the International Conference on Machine Learning."},{"key":"2022032817111023600_B47","unstructured":"Zheng, A., & Casari, A. (2018). Feature engineering for machine learning: Principles and techniques for data scientists. Sebastopol, CA: O'Reilly Media."},{"key":"2022032817111023600_B48","unstructured":"Zou, D., Long, P. M., & Gu, Q. (2020). On the global convergence of training deep linear ResNets. In Proceedings of the International Conference on Learning Representations."}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/34\/4\/991\/2003085\/neco_a_01483.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/34\/4\/991\/2003085\/neco_a_01483.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,3,28]],"date-time":"2022-03-28T17:12:12Z","timestamp":1648487532000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/34\/4\/991\/109667\/Understanding-Dynamics-of-Nonlinear-Representation"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,23]]},"references-count":48,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,3,23]]},"published-print":{"date-parts":[[2022,3,23]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01483","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,4]]},"published":{"date-parts":[[2022,3,23]]}}}