{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,24]],"date-time":"2024-08-24T13:31:44Z","timestamp":1724506304429},"reference-count":33,"publisher":"MIT Press","issue":"2","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,20]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Training deep learning models from a stream of nonstationary data is a critical problem to be solved to achieve general artificial intelligence. As a promising solution, the continual learning (CL) technique aims to build intelligent systems that have the plasticity to learn from new information without forgetting the previously obtained knowledge. Unfortunately, existing CL methods face two nontrivial limitations. First, when updating a model with new data, existing CL methods usually constrain the model parameters within the vicinity of the parameters optimized for old data, limiting the exploration ability of the model; second, the important strength of each parameter (used to consolidate the previously learned knowledge) is fixed and thus is suboptimal for the dynamic parameter updates. To address these limitations, we first relax the vicinity constraints with a global definition of the important strength, which allows us to explore the full parameter space. Specifically, we define the important strength as the sensitivity of the global loss function to the model parameters. Moreover, we propose adjusting the important strength adaptively to align it with the dynamic parameter updates. Through extensive experiments on popular data sets, we demonstrate that our proposed method outperforms the strong baselines by up to 24% in terms of average accuracy.<\/jats:p>","DOI":"10.1162\/neco_a_01560","type":"journal-article","created":{"date-parts":[[2022,12,22]],"date-time":"2022-12-22T00:56:30Z","timestamp":1671670590000},"page":"228-248","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":3,"title":["Dynamic Consolidation for Continual Learning"],"prefix":"10.1162","volume":"35","author":[{"given":"Hang","family":"Li","sequence":"first","affiliation":[{"name":"McGill University, Montreal, Quebec H3A 0G4, Canada hang.li3@mail.mcgill.ca"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chen","family":"Ma","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong SAR, China chenma@cityu.edu.hk"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Chen","sequence":"additional","affiliation":[{"name":"McGill University, Montreal, Quebec H3A 0G4, Canada xi.chen11@mcgill.ca"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xue","family":"Liu","sequence":"additional","affiliation":[{"name":"McGill University, Montreal, Quebec H3A 0G4, Canada xueliu@cs.mcgill.ca"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,1,20]]},"reference":[{"key":"2023012618411746200_B1","author":"Agarap","year":"2018","journal-title":"Deep learning using rectified linear units"},{"key":"2023012618411746200_B2","article-title":"Uncertainty-based continual learning with adaptive regularization","volume-title":"Advances in neural information processing systems","author":"Ahn","year":"2019"},{"key":"2023012618411746200_B3","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01219-9_9","article-title":"Memory aware synapses: Learning what (not) to forget","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Aljundi","year":"2018"},{"key":"2023012618411746200_B4","article-title":"Online continual learning with maximal interfered retrieval","volume-title":"Advances in neural information processing systems","author":"Aljundi","year":"2019"},{"key":"2023012618411746200_B5","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.753","article-title":"Expert gate: Lifelong learning with a network of experts","volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition","author":"Aljundi","year":"2017"},{"key":"2023012618411746200_B6","article-title":"Gradient based sample selection for online continual learning","volume-title":"Advances in neural information processing systems","author":"Aljundi","year":"2019"},{"key":"2023012618411746200_B7","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2019.00067","article-title":"IL2m: Class incremental learning with dual memory","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Belouadah","year":"2019"},{"key":"2023012618411746200_B8","article-title":"Coresets via bilevel optimization for continual learning and streaming","volume-title":"Advances in neural information processing systems","author":"Borsos","year":"2020"},{"key":"2023012618411746200_B9","article-title":"Online continual learning from imbalanced data","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chrysakis","year":"2020"},{"key":"2023012618411746200_B10","author":"De Lange","year":"2019","journal-title":"A continual learning survey: Defying forgetting in classification tasks"},{"key":"2023012618411746200_B11","first-page":"9285","article-title":"Dytox: Transformers for continual learning with dynamic token expansion","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Douillard","year":"2022"},{"key":"2023012618411746200_B12","author":"Farquhar","year":"2018","journal-title":"Towards robust evaluations of continual learning"},{"key":"2023012618411746200_B13","author":"Goodfellow","year":"2014","journal-title":"An empirical investigation of catastrophic forgetting in gradient-based neural networks"},{"key":"2023012618411746200_B14","first-page":"770","article-title":"Deep residual learning for image recognition","author":"He","year":"2016","journal-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition"},{"key":"2023012618411746200_B15","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.11595","volume-title":"Selective experience replay for lifelong learning","author":"Isele","year":"2018"},{"key":"2023012618411746200_B16","volume-title":"Adam: A method for stochastic optimization","author":"Kingma","year":"2015"},{"issue":"13","key":"2023012618411746200_B17","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2023012618411746200_B18","article-title":"Sliced Cram\u00e9r synaptic consolidation for preserving deeply learned representations","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kolouri","year":"2020"},{"key":"2023012618411746200_B19","first-page":"1097","article-title":"ImageNet classification with deep convolutional neural networks","volume-title":"Advances in neural information processing systems","author":"Krizhevsky","year":"2012"},{"issue":"12","key":"2023012618411746200_B20","doi-asserted-by":"publisher","first-page":"2935","DOI":"10.1109\/TPAMI.2017.2773081","article-title":"Learning without forgetting","volume":"40","author":"Li","year":"2017","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2023012618411746200_B21","first-page":"6467","article-title":"Gradient episodic memory for continual learning","volume-title":"Advances in neural information processing systems","author":"Lopez-Paz","year":"2017"},{"key":"2023012618411746200_B22","first-page":"7765","article-title":"PackNet: Adding multiple tasks to a single network by iterative pruning","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mallya","year":"2018"},{"key":"2023012618411746200_B23","article-title":"Reading digits in natural images with unsupervised feature learning","author":"Netzer","year":"2011","journal-title":"NIPS Workshop on Deep Learning and Unsupervised Feature Learning"},{"key":"2023012618411746200_B24","first-page":"2001","article-title":"iCaRL: Incremental classifier and representation learning","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"Rebuffi","year":"2017"},{"issue":"2","key":"2023012618411746200_B25","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1080\/09540099550039318","article-title":"Catastrophic forgetting, rehearsal and pseudorehearsal","volume":"7","author":"Robins","year":"1995","journal-title":"Connection Science"},{"key":"2023012618411746200_B26","author":"Rusu","year":"2016","journal-title":"Progressive neural networks"},{"key":"2023012618411746200_B27","first-page":"4555","article-title":"Overcoming catastrophic forgetting with hard attention to the task","volume-title":"Proceedings of Machine Learning Research","author":"Serr\u00e0","year":"2018"},{"key":"2023012618411746200_B28","first-page":"2990","article-title":"Continual learning with deep generative replay","volume-title":"Advances in neural information processing systems","author":"Shin","year":"2017"},{"key":"2023012618411746200_B29","article-title":"Rectification-based knowledge retention for continual learning","author":"Singh","year":"2021","journal-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition"},{"key":"2023012618411746200_B30","article-title":"Tiny ImageNet Challenge course (CS231n)","author":"Stanford University"},{"key":"2023012618411746200_B31","author":"van de Ven","year":"2019","journal-title":"Three scenarios for continual learning"},{"key":"2023012618411746200_B32","first-page":"374","article-title":"Large scale incremental learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wu","year":"2019"},{"key":"2023012618411746200_B33","first-page":"3987","article-title":"Continual learning through synaptic intelligence","volume-title":"Proceedings of Machine Learning Research","author":"Zenke","year":"2017"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/2\/228\/2067674\/neco_a_01560.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/2\/228\/2067674\/neco_a_01560.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,26]],"date-time":"2023-01-26T18:41:50Z","timestamp":1674758510000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/35\/2\/228\/114139\/Dynamic-Consolidation-for-Continual-Learning"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,20]]},"references-count":33,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2023,1,20]]},"published-print":{"date-parts":[[2023,1,20]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01560","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,2]]},"published":{"date-parts":[[2023,1,20]]}}}