{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T06:45:58Z","timestamp":1757313958872,"version":"3.37.3"},"reference-count":84,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T00:00:00Z","timestamp":1601424000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T00:00:00Z","timestamp":1601424000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int. J. Mach. Learn. &amp; Cyber."],"published-print":{"date-parts":[[2021,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We propose a modular architecture of Deep Neural Network (DNN) for multi-class classification task. The architecture consists of two parts, a router network and a set of expert networks. In this architecture, for a <jats:italic>C<\/jats:italic>-class classification problem, we have exactly <jats:italic>C<\/jats:italic> experts. The backbone network for these experts and the router are built with simple and identical DNN architecture. For each class, the modular network has a certain number <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\rho$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mi>\u03c1<\/mml:mi>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> of expert networks specializing in that particular class, where <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\rho$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mi>\u03c1<\/mml:mi>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> is called the redundancy rate in this study. We demonstrate that <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\rho$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mi>\u03c1<\/mml:mi>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> plays a vital role in the performance of the network. Although these experts are light weight and weak learners alone, together they match the performance of more complex DNNs. We train the network in two phase wherein, first the router is trained on the whole set of training data followed by training each expert network enforced by a new stochastic objective function that facilitates alternative training on a small subset of expert data and the whole set of data. This alternative training provides an additional form of regularization and avoids over-fitting the expert network on subset data. During the testing phase, the router dynamically selects a fixed number of experts for further evaluation of the input datum. The modular nature and low parameter requirement of the network makes it very suitable in distributed and low computational environments. Extensive empirical study and theoretical analysis on CIFAR-10, CIFAR-100 and F-MNIST substantiate the effectiveness and efficiency of our proposed modular network.<\/jats:p>","DOI":"10.1007\/s13042-020-01201-8","type":"journal-article","created":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T08:05:14Z","timestamp":1601453114000},"page":"763-781","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["MS-NET: modular selective network"],"prefix":"10.1007","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2483-577X","authenticated-orcid":false,"given":"Intisar Md","family":"Chowdhury","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2044-0671","authenticated-orcid":false,"given":"Kai","family":"Su","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3101-749X","authenticated-orcid":false,"given":"Qiangfu","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,9,30]]},"reference":[{"key":"1201_CR1","unstructured":"URL https:\/\/www.cs.toronto.edu\/~kriz\/cifar.html"},{"key":"1201_CR2","first-page":"469","volume-title":"European conference on computer vision","author":"T Ahonen","year":"2004","unstructured":"Ahonen T, Hadid A, Pietik\u00e4inen M (2004) Face recognition with local binary patterns. European conference on computer vision. Springer, Berlin, pp 469\u2013481"},{"key":"1201_CR3","unstructured":"Amodei D, Ananthanarayanan S, Anubhai R, Bai J, Battenberg E, Case C, Casper J, Catanzaro B, Cheng Q, Chen G, et al. (2016) Deep speech 2: End-to-end speech recognition in english and mandarin. In: International conference on machine learning, pp 173\u2013182"},{"issue":"1","key":"1201_CR4","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1109\/72.363444","volume":"6","author":"R Anand","year":"1995","unstructured":"Anand R, Mehrotra K, Mohan CK, Ranka S (1995) Efficient classification for multiclass problems using modular neural networks. IEEE Trans Neural Netw 6(1):117\u2013124","journal-title":"IEEE Trans Neural Netw"},{"issue":"12","key":"1201_CR5","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V Badrinarayanan","year":"2017","unstructured":"Badrinarayanan V, Kendall A, Cipolla R (2017) Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans Pattern Anal Mach Intell 39(12):2481\u20132495","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"2","key":"1201_CR6","first-page":"123","volume":"24","author":"L Breiman","year":"1996","unstructured":"Breiman L (1996) Bagging predictors. Mach Learn 24(2):123\u2013140","journal-title":"Mach Learn"},{"issue":"1","key":"1201_CR7","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L (2001) Random forests. Mach Learn 45(1):5\u201332","journal-title":"Mach Learn"},{"key":"1201_CR8","doi-asserted-by":"crossref","unstructured":"Bucilu\u01ce C, Caruana R, Niculescu-Mizil A (2006) Model compression. In: Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, pp 535\u2013541. ACM","DOI":"10.1145\/1150402.1150464"},{"issue":"4","key":"1201_CR9","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"LC Chen","year":"2017","unstructured":"Chen LC, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2017) Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans Pattern Anal Mach Intell 40(4):834\u2013848","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1201_CR10","first-page":"151","volume-title":"European Working Session on Learning","author":"P Clark","year":"1991","unstructured":"Clark P, Boswell R (1991) Rule induction with cn2: Some recent improvements. European Working Session on Learning. Springer, Berlin, pp 151\u2013163"},{"key":"1201_CR11","doi-asserted-by":"crossref","unstructured":"Collobert R, Weston J (2008) A unified architecture for natural language processing: Deep neural networks with multitask learning. In: Proceedings of the 25th international conference on Machine learning, pp 160\u2013167. ACM","DOI":"10.1145\/1390156.1390177"},{"issue":"3","key":"1201_CR12","first-page":"273","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes C, Vapnik V (1995) Support-vector networks. Mach Learn 20(3):273\u2013297","journal-title":"Mach Learn"},{"key":"1201_CR13","doi-asserted-by":"crossref","unstructured":"Cubuk ED, Zoph B, Mane D, Vasudevan V, Le QV (2019) Autoaugment: Learning augmentation strategies from data. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 113\u2013123","DOI":"10.1109\/CVPR.2019.00020"},{"key":"1201_CR14","doi-asserted-by":"crossref","unstructured":"Cubuk ED, Zoph B, Shlens J, Le QV (2019) Randaugment: Practical data augmentation with no separate search. arXiv preprint arXiv:1909.13719","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"1201_CR15","unstructured":"Dumoulin V, Visin F (2016) A guide to convolution arithmetic for deep learning. arXiv preprint arXiv:1603.07285"},{"key":"1201_CR16","doi-asserted-by":"crossref","unstructured":"Dutt A, Pellerin D, Qu\u00e9not G (2019) Coupled ensembles of neural networks. Neurocomputing","DOI":"10.1109\/CBMI.2018.8516453"},{"key":"1201_CR17","doi-asserted-by":"crossref","unstructured":"Elsken T, Metzen JH, Hutter F (2018) Neural architecture search: A survey. arXiv preprint arXiv:1808.05377","DOI":"10.1007\/978-3-030-05318-5_3"},{"key":"1201_CR18","doi-asserted-by":"publisher","first-page":"23","DOI":"10.1007\/3-540-59119-2_166","volume-title":"European conference on computational learning theory","author":"Y Freund","year":"1995","unstructured":"Freund Y, Schapire RE (1995) A desicion-theoretic generalization of on-line learning and an application to boosting. European conference on computational learning theory. Springer, Berlin, pp 23\u201337"},{"issue":"4","key":"1201_CR19","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1016\/S0167-9473(01)00065-2","volume":"38","author":"JH Friedman","year":"2002","unstructured":"Friedman JH (2002) Stochastic gradient boosting. Comput Statist Data Analy 38(4):367\u2013378","journal-title":"Comput Statist Data Analy"},{"key":"1201_CR20","unstructured":"Furlanello T, Lipton ZC, Tschannen M, Itti L, Anandkumar A (2018) Born again neural networks. arXiv preprint arXiv:1805.04770"},{"issue":"1","key":"1201_CR21","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1023\/A:1006524209794","volume":"13","author":"J F\u00fcrnkranz","year":"1999","unstructured":"F\u00fcrnkranz J (1999) Separate-and-conquer rule learning. Artif Intell Rev 13(1):3\u201354","journal-title":"Artif Intell Rev"},{"key":"1201_CR22","first-page":"97","volume-title":"European Conference on Machine Learning","author":"J F\u00fcrnkranz","year":"2002","unstructured":"F\u00fcrnkranz J (2002) Pairwise classification as an ensemble technique. European Conference on Machine Learning. Springer, Berlin, pp 97\u2013110"},{"key":"1201_CR23","first-page":"721","volume":"2","author":"J F\u00fcrnkranz","year":"2002","unstructured":"F\u00fcrnkranz J (2002) Round robin classification. J Mach Learn Res 2:721\u2013747 (Mar)","journal-title":"J Mach Learn Res"},{"issue":"5","key":"1201_CR24","doi-asserted-by":"publisher","first-page":"385","DOI":"10.3233\/IDA-2003-7502","volume":"7","author":"J F\u00fcrnkranz","year":"2003","unstructured":"F\u00fcrnkranz J (2003) Round robin ensembles. Intell Data Anal 7(5):385\u2013403","journal-title":"Intell Data Anal"},{"key":"1201_CR25","unstructured":"Gao S, Cheng MM, Zhao K, Zhang XY, Yang MH, Torr PH (2019) Res2net: A new multi-scale backbone architecture. IEEE transactions on pattern analysis and machine intelligence"},{"key":"1201_CR26","doi-asserted-by":"crossref","unstructured":"Girshick R (2015) Fast r-cnn. In: Proceedings of the IEEE international conference on computer vision, pp 1440\u20131448","DOI":"10.1109\/ICCV.2015.169"},{"key":"1201_CR27","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. In: Advances in neural information processing systems, pp 2672\u20132680"},{"key":"1201_CR28","unstructured":"Goodfellow IJ, Mirza M, Xiao D, Courville A, Bengio Y (2013) An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211"},{"key":"1201_CR29","doi-asserted-by":"crossref","unstructured":"Graves A, Mohamed Ar, Hinton G (2013) Speech recognition with deep recurrent neural networks. In: 2013 IEEE international conference on acoustics, speech and signal processing, pp 6645\u20136649. IEEE","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"1201_CR30","doi-asserted-by":"publisher","first-page":"751","DOI":"10.1109\/34.142911","volume":"7","author":"JB Hampshire II","year":"1992","unstructured":"Hampshire JB II, Waibel A (1992) The meta-pi network: Building distributed knowledge representations for robust multisource pattern recognition. IEEE Trans Pattern Anal Mach Intell 7:751\u2013769","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1201_CR31","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"1201_CR32","first-page":"630","volume-title":"European conference on computer vision","author":"K He","year":"2016","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. European conference on computer vision. Springer, Berlin, pp 630\u2013645"},{"key":"1201_CR33","unstructured":"Hinton G, Vinyals O, Dean J (2015) Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531"},{"issue":"3","key":"1201_CR34","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1023\/A:1021251113462","volume":"115","author":"YC Ho","year":"2002","unstructured":"Ho YC, Pepyne DL (2002) Simple explanation of the no-free-lunch theorem and its implications. J Optimiz Theory Appl 115(3):549\u2013570","journal-title":"J Optimiz Theory Appl"},{"key":"1201_CR35","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861"},{"key":"1201_CR36","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Van Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4700\u20134708","DOI":"10.1109\/CVPR.2017.243"},{"key":"1201_CR37","unstructured":"Huang Y, Cheng Y, Bapna A, Firat O, Chen D, Chen M, Lee H, Ngiam J, Le QV, Wu Y, et al. (2019) Gpipe: Efficient training of giant neural networks using pipeline parallelism. In: Advances in Neural Information Processing Systems, pp 103\u2013112"},{"key":"1201_CR38","doi-asserted-by":"crossref","unstructured":"Intisar CM, Watanobe Y (2018) Classification of online judge programmers based on rule extraction from self organizing feature map. In: 2018 9th International Conference on Awareness Science and Technology (iCAST), pp 313\u2013318. IEEE","DOI":"10.1109\/ICAwST.2018.8517222"},{"key":"1201_CR39","doi-asserted-by":"crossref","unstructured":"Intisar CM, Zhao Q (2019) A selective modular neural network framework. In: 2019 IEEE 10th International Conference on Awareness Science and Technology (iCAST), pp 1\u20136. IEEE","DOI":"10.1109\/ICAwST.2019.8923334"},{"issue":"1","key":"1201_CR40","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1162\/neco.1991.3.1.79","volume":"3","author":"RA Jacobs","year":"1991","unstructured":"Jacobs RA, Jordan MI, Nowlan SJ, Hinton GE et al (1991) Adaptive mixtures of local experts. Neural Comput 3(1):79\u201387","journal-title":"Neural Comput"},{"issue":"12","key":"1201_CR41","doi-asserted-by":"publisher","first-page":"1167","DOI":"10.1016\/0031-3203(91)90143-S","volume":"24","author":"AK Jain","year":"1991","unstructured":"Jain AK, Farrokhnia F (1991) Unsupervised texture segmentation using gabor filters. Pattern Recogn 24(12):1167\u20131186","journal-title":"Pattern Recogn"},{"issue":"2","key":"1201_CR42","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1162\/neco.1994.6.2.181","volume":"6","author":"MI Jordan","year":"1994","unstructured":"Jordan MI, Jacobs RA (1994) Hierarchical mixtures of experts and the em algorithm. Neural Comput 6(2):181\u2013214","journal-title":"Neural Comput"},{"key":"1201_CR43","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1613\/jair.301","volume":"4","author":"LP Kaelbling","year":"1996","unstructured":"Kaelbling LP, Littman ML, Moore AW (1996) Reinforcement learning: A survey. J Artif Intell Res 4:237\u2013285","journal-title":"J Artif Intell Res"},{"key":"1201_CR44","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1007\/978-1-4842-2766-4_12","volume-title":"Deep learning with python","author":"N Ketkar","year":"2017","unstructured":"Ketkar N (2017) Introduction to pytorch. Deep learning with python. Springer, Berlin, pp 195\u2013208"},{"key":"1201_CR45","unstructured":"Kim T, Cha M, Kim H, Lee JK, Kim J (2017) Learning to discover cross-domain relations with generative adversarial networks. In: Proceedings of the 34th International Conference on Machine Learning-Volume 70, pp 1857\u20131865. JMLR. org"},{"key":"1201_CR46","doi-asserted-by":"crossref","unstructured":"Kolesnikov A, Beyer L, Zhai X, Puigcerver J, Yung J, Gelly S, Houlsby N (2019) Big transfer (bit): General visual representation learning. arXiv preprint arXiv:1912.11370","DOI":"10.1007\/978-3-030-58558-7_29"},{"key":"1201_CR47","unstructured":"Larsson G, Maire M, Shakhnarovich G (2016) Fractalnet: Ultra-deep neural networks without residuals. arXiv preprint arXiv:1605.07648"},{"key":"1201_CR48","unstructured":"LeCun Y, Denker JS, Solla SA (1990) Optimal brain damage. In: Advances in neural information processing systems, pp 598\u2013605"},{"key":"1201_CR49","doi-asserted-by":"crossref","unstructured":"Lee H, Lee JS (2018) Local critic training for model-parallel learning of deep neural networks. arXiv preprint arXiv:1805.01128","DOI":"10.1109\/IJCNN.2019.8851709"},{"key":"1201_CR50","unstructured":"Leung HC, Zue VW (1989) Applications of error back-propagation to phonetic classification. In: Advances in neural information processing systems, pp 206\u2013214"},{"key":"1201_CR51","unstructured":"Li H, Kadav A, Durdanovic I, Samet H, Graf HP (2016) Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710"},{"key":"1201_CR52","unstructured":"Lim S, Kim I, Kim T, Kim C, Kim S (2019) Fast autoaugment. arXiv preprint arXiv:1905.00397"},{"key":"1201_CR53","unstructured":"Loshchilov I, Hutter F (2016) Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983"},{"key":"1201_CR54","doi-asserted-by":"crossref","unstructured":"Lowe DG, et al. (1999) Object recognition from local scale-invariant features. In: iccv, vol. 99, pp 1150\u20131157","DOI":"10.1109\/ICCV.1999.790410"},{"key":"1201_CR55","volume-title":"Society of mind","author":"M Minsky","year":"1988","unstructured":"Minsky M (1988) Society of mind. Simon and Schuster, New York"},{"key":"1201_CR56","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602"},{"issue":"7540","key":"1201_CR57","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Rusu AA, Veness J, Bellemare MG, Graves A, Riedmiller M, Fidjeland AK, Ostrovski G et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529","journal-title":"Nature"},{"key":"1201_CR58","unstructured":"Molchanov P, Tyree S, Karras T, Aila T, Kautz J (2016) Pruning convolutional neural networks for resource efficient inference. arXiv preprint arXiv:1611.06440"},{"key":"1201_CR59","unstructured":"Nayman N, Noy A, Ridnik T, Friedman I, Jin R, Zelnik L (2019) Xnas: Neural architecture search with expert advice. In: Advances in Neural Information Processing Systems, pp 1977\u20131987"},{"issue":"1","key":"1201_CR60","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1007\/s10618-011-0219-9","volume":"24","author":"SH Park","year":"2012","unstructured":"Park SH, F\u00fcrnkranz J (2012) Efficient prediction algorithms for binary decomposition techniques. Data Min Knowl Discovery 24(1):40\u201377","journal-title":"Data Min Knowl Discovery"},{"key":"1201_CR61","unstructured":"Pham H, Guan MY, Zoph B, Le QV, Dean J (2018) Efficient neural architecture search via parameter sharing. arXiv preprint arXiv:1802.03268"},{"key":"1201_CR62","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in neural information processing systems, pp 91\u201399"},{"key":"1201_CR63","first-page":"234","volume-title":"International Conference on Medical image computing and computer-assisted intervention","author":"O Ronneberger","year":"2015","unstructured":"Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. International Conference on Medical image computing and computer-assisted intervention. Springer, Berlin, pp 234\u2013241"},{"issue":"3","key":"1201_CR64","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M et al (2015) Imagenet large scale visual recognition challenge. Int J Computer Vision 115(3):211\u2013252","journal-title":"Int J Computer Vision"},{"key":"1201_CR65","unstructured":"Savarese P, Maire M (2019) Learning implicitly recurrent cnns through parameter sharing. arXiv preprint arXiv:1902.09701"},{"key":"1201_CR66","unstructured":"Savarese PH, Mazza LO, Figueiredo DR (2016) Learning identity mappings with residual gates. arXiv preprint arXiv:1611.01260"},{"key":"1201_CR67","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"},{"key":"1201_CR68","unstructured":"Socher R, Bengio Y, Manning C (2012) Deep learning for nlp. Tutorial at Association of Computational Logistics (ACL)"},{"key":"1201_CR69","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"1201_CR70","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"1201_CR71","unstructured":"Tan M, Le QV (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946"},{"issue":"12","key":"1201_CR72","doi-asserted-by":"publisher","first-page":"1888","DOI":"10.1109\/29.45535","volume":"37","author":"A Waibel","year":"1989","unstructured":"Waibel A, Sawai H, Shikano K (1989) Modularity and scaling in large phonemic neural networks. IEEE Trans Acoust Speech Signal Process 37(12):1888\u20131898","journal-title":"IEEE Trans Acoust Speech Signal Process"},{"issue":"1","key":"1201_CR73","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1108\/14676370310455332","volume":"4","author":"K Warburton","year":"2003","unstructured":"Warburton K (2003) Deep learning and education for sustainability. Int J Sustainab Higher Educ 4(1):44\u201356","journal-title":"Int J Sustainab Higher Educ"},{"key":"1201_CR74","doi-asserted-by":"publisher","first-page":"62","DOI":"10.1016\/j.neunet.2017.09.017","volume":"97","author":"C Watanabe","year":"2018","unstructured":"Watanabe C, Hiramatsu K, Kashino K (2018) Modular representation of layered neural networks. Neural Netw 97:62\u201373","journal-title":"Neural Netw"},{"key":"1201_CR75","unstructured":"Wong C, Houlsby N, Lu Y, Gesmundo A (2018) Transfer learning with neural automl. In: Advances in Neural Information Processing Systems, pp 8356\u20138365"},{"key":"1201_CR76","unstructured":"Wu H, Zhang J, Huang K, Liang K, Yu Y (2019) Fastfcn: Rethinking dilated convolution in the backbone for semantic segmentation. arXiv preprint arXiv:1903.11816"},{"key":"1201_CR77","doi-asserted-by":"crossref","unstructured":"Xie S, Girshick R, Doll\u00e1r P, Tu Z, He K (2017) Aggregated residual transformations for deep neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1492\u20131500","DOI":"10.1109\/CVPR.2017.634"},{"key":"1201_CR78","doi-asserted-by":"crossref","unstructured":"Yang C, Xie L, Qiao S, Yuille A (2018) Knowledge distillation in generations: More tolerant teachers educate better students. arXiv preprint arXiv:1805.05551","DOI":"10.1609\/aaai.v33i01.33015628"},{"key":"1201_CR79","unstructured":"Zalandoresearch: zalandoresearch\/fashion-mnist (2019). URL https:\/\/github.com\/zalandoresearch\/fashion-mnist"},{"key":"1201_CR80","doi-asserted-by":"crossref","unstructured":"Zhang Y, Xiang T, Hospedales TM, Lu H (2018) Deep mutual learning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp 4320\u20134328","DOI":"10.1109\/CVPR.2018.00454"},{"issue":"6","key":"1201_CR81","doi-asserted-by":"publisher","first-page":"1371","DOI":"10.1109\/72.641460","volume":"8","author":"Q Zhao","year":"1997","unstructured":"Zhao Q (1997) Stable online evolutionary learning of nn-mlp. IEEE Trans Neural Netw 8(6):1371\u20131378","journal-title":"IEEE Trans Neural Netw"},{"issue":"3","key":"1201_CR82","doi-asserted-by":"publisher","first-page":"762","DOI":"10.1109\/72.501733","volume":"7","author":"Q Zhao","year":"1996","unstructured":"Zhao Q, Higuchi T (1996) Evolutionary learning of nearest-neighbor mlp. IEEE Trans Neural Netw 7(3):762\u2013767","journal-title":"IEEE Trans Neural Netw"},{"key":"1201_CR83","unstructured":"Zoph B, Le QV (2016) Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578"},{"key":"1201_CR84","doi-asserted-by":"crossref","unstructured":"Zoph B, Vasudevan V, Shlens J, Le QV (2018) Learning transferable architectures for scalable image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 8697\u20138710","DOI":"10.1109\/CVPR.2018.00907"}],"container-title":["International Journal of Machine Learning and Cybernetics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13042-020-01201-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13042-020-01201-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13042-020-01201-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,30]],"date-time":"2021-09-30T01:26:44Z","timestamp":1632965204000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13042-020-01201-8"}},"subtitle":["Round robin based modular neural network architecture with limited redundancy"],"short-title":[],"issued":{"date-parts":[[2020,9,30]]},"references-count":84,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,3]]}},"alternative-id":["1201"],"URL":"https:\/\/doi.org\/10.1007\/s13042-020-01201-8","relation":{},"ISSN":["1868-8071","1868-808X"],"issn-type":[{"type":"print","value":"1868-8071"},{"type":"electronic","value":"1868-808X"}],"subject":[],"published":{"date-parts":[[2020,9,30]]},"assertion":[{"value":"29 March 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 September 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with ethical standards"}},{"value":"There is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}