{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T20:07:27Z","timestamp":1780085247845,"version":"3.54.0"},"reference-count":54,"publisher":"World Scientific Pub Co Pte Ltd","issue":"14","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2022,11]]},"abstract":"<jats:p>One of the most common approaches through which people communicate is facial expressions. A large number of features documented in the literature were created by hand, with the goal of overcoming specific challenges such as occlusions, scale, and illumination variations. These classic methods are then applied to a dataset of facial images or frames in order to train a classifier. The majority of these studies perform admirably on datasets of images shot in a controlled environment, but they struggle with more difficult datasets (FER-2013) that have higher image variation and partial faces. The nonuniform features of the human face as well as changes in lighting, shadows, facial posture, and direction are the key obstacles. Techniques of deep learning have been studied as a set of methodologies for gaining scalability and robustness on new forms of data. In this paper, we look at how well-known deep learning techniques (e.g. GoogLeNet, AlexNet) perform when it comes to facial expression identification, and propose an enhanced hybrid deep learning model based on STN for facial emotion recognition, which gives the best feature extraction and classification in one go and maximizes the accuracy for a large number of samples on FERG, JAFFE, FER-2013, and CK+ datasets. It is capable of focusing on the main parts of the face and attaining extensive development over preceding fashions on the FERG, JAFFE, CK+ datasets, and the more challenging one namely FER-2013.<\/jats:p>","DOI":"10.1142\/s0218001422520280","type":"journal-article","created":{"date-parts":[[2022,10,30]],"date-time":"2022-10-30T07:14:00Z","timestamp":1667114040000},"source":"Crossref","is-referenced-by-count":16,"title":["Enhanced Deep Learning Hybrid Model of CNN Based on Spatial Transformer Network for Facial Expression Recognition"],"prefix":"10.1142","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9395-0050","authenticated-orcid":false,"given":"Nizamuddin","family":"Khan","sequence":"first","affiliation":[{"name":"Amity Institute of Information Technology, Amity University, Noida 201303, Uttar Pradesh, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ajay Vikram","family":"Singh","sequence":"additional","affiliation":[{"name":"Amity Institute of Information Technology, Amity University, Noida 201303, Uttar Pradesh, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rajeev","family":"Agrawal","sequence":"additional","affiliation":[{"name":"Lloyd Institute of Engineering & Technology, Greater Noida 201306, Uttar Pradesh, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"219","published-online":{"date-parts":[[2022,10,29]]},"reference":[{"key":"S0218001422520280BIB001","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1007\/s10803-009-0884-3","volume":"40","author":"Bal E.","year":"2010","journal-title":"J. Autism Dev. Disord."},{"key":"S0218001422520280BIB002","first-page":"568","volume-title":"Proc. IEEE Computer Society Conf. Computer Vision and Pattern Recognition (CVPR 2005)","volume":"2","author":"Bartlett M. S.","year":"2005"},{"key":"S0218001422520280BIB003","doi-asserted-by":"crossref","first-page":"8828245","DOI":"10.1155\/2021\/8828245","volume":"2021","author":"Ben N.","year":"2021","journal-title":"Comput. Intell. Neurosci."},{"key":"S0218001422520280BIB005","first-page":"884","volume-title":"Proc. Int. Workshops on Electrical and Computer Engineering Subfields","author":"Chen J.","year":"2014"},{"key":"S0218001422520280BIB006","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2008.03.012"},{"key":"S0218001422520280BIB008","first-page":"6","volume":"2","author":"Cohn J. F.","year":"1995","journal-title":"Am. Psychol. Soc."},{"key":"S0218001422520280BIB009","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1109\/79.911197","volume":"18","author":"Cowie R.","year":"2011","journal-title":"IEEE Signal Process. Mag."},{"key":"S0218001422520280BIB010","first-page":"1223","volume-title":"Proc. 25th Int. Conf. Neural Information Processing Systems","volume":"1","author":"Dean J.","year":"2012"},{"key":"S0218001422520280BIB011","series-title":"Lecture Notes in Computer Science","first-page":"136","volume-title":"ACCV 2016: Computer Vision","volume":"10112","author":"Deepali A.","year":"2016"},{"key":"S0218001422520280BIB012","doi-asserted-by":"publisher","DOI":"10.1037\/h0030377"},{"key":"S0218001422520280BIB013","doi-asserted-by":"crossref","first-page":"6295","DOI":"10.1007\/s00521-019-04138-4","volume":"32","author":"Eleyan A. M. A.","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"S0218001422520280BIB014","volume-title":"Facial Action Coding System: A Technique for the Measurement of Facial Movement","author":"Friesen E.","year":"1978"},{"key":"S0218001422520280BIB016","series-title":"Smart Innovation, Systems and Technologies","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-319-66790-4_1","volume-title":"Advances in Hybridization of Intelligent Methods","volume":"85","author":"Giannopoulos P.","year":"2018"},{"key":"S0218001422520280BIB017","first-page":"109","volume-title":"Proc. 30th Int. Conf. Neural Information Processing Systems","author":"Han S.","year":"2016"},{"key":"S0218001422520280BIB018","volume-title":"Proc. Fifteenth Annu. Conf. Int. Speech Communication Association","author":"Han K.","year":"2014"},{"key":"S0218001422520280BIB019","volume-title":"Proc. IEEE 2018 IEEE 16th Int. Conf. Software Engineering Research, Management and Applications (SERA)","author":"Hang Z.","year":"2018"},{"key":"S0218001422520280BIB020","doi-asserted-by":"crossref","first-page":"3211","DOI":"10.1007\/s11227-018-2554-8","volume":"76","author":"Hao W.","year":"2020","journal-title":"J. Supercomput."},{"key":"S0218001422520280BIB021","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2014.2386334"},{"key":"S0218001422520280BIB022","first-page":"2983","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Heechul J.","year":"2015"},{"key":"S0218001422520280BIB023","doi-asserted-by":"crossref","first-page":"2056010","DOI":"10.1142\/S0218001420560108","volume":"34","author":"Huang L.","year":"2021","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"S0218001422520280BIB024","volume-title":"Proc. Workshop on Challenges in Representation Learning (ICML)","author":"Ionescu R. T.","year":"2013"},{"key":"S0218001422520280BIB025","first-page":"2017","volume-title":"Advances in Neural Information Processing Systems","volume":"28","author":"Jaderberg M.","year":"2015"},{"key":"S0218001422520280BIB026","volume-title":"Proc. IEEE Int. Conf. Computer Vision Workshops","author":"Khorrami P.","year":"2015"},{"key":"S0218001422520280BIB027","first-page":"662","volume-title":"Proc. IEEE Int. Conf. Systems Man and Cybernetics (SMC)","author":"Koutlas A.","year":"2008"},{"key":"S0218001422520280BIB028","first-page":"1097","volume-title":"Advances in Neural Information Processing Systems","volume":"25","author":"Krizhevsky A.","year":"2012"},{"key":"S0218001422520280BIB029","first-page":"143","volume-title":"Connectionism in Perspective","author":"LeCun Y.","year":"1989"},{"key":"S0218001422520280BIB031","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1007\/978-3-319-16817-3_10","volume-title":"ACCV 2014: Computer Vision","volume":"9006","author":"Liu M.","year":"2014"},{"key":"S0218001422520280BIB032","doi-asserted-by":"publisher","DOI":"10.1142\/S021800142152008X"},{"key":"S0218001422520280BIB033","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2016.07.026"},{"key":"S0218001422520280BIB034","volume-title":"Proc. Int. Conf. Learning Representations (ICLR)","author":"Loshchilov I.","year":"2019"},{"key":"S0218001422520280BIB035","volume-title":"Proc. IEEE Computer Society Conf. Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Lucey P.","year":"2010"},{"key":"S0218001422520280BIB036","first-page":"14","volume-title":"Proc. Third Int. Conf. Automatic Face and Gesture Recognition","author":"Lyons M. J.","year":"1998"},{"key":"S0218001422520280BIB037","doi-asserted-by":"crossref","first-page":"6698","DOI":"10.1007\/s10489-021-02219-3","volume":"51","author":"Ma X.","year":"2021","journal-title":"Appl. Intell."},{"key":"S0218001422520280BIB038","volume-title":"Proc. 2019 IEEE\/CVF Conf. Computer Vision and Pattern Recognition Workshops","author":"Marrero Fernandez P. D.","year":"2019"},{"key":"S0218001422520280BIB039","first-page":"558","volume-title":"Proc. IEEE Int. Conf. Automatic Face & Gesture Recognition","author":"Meng Z.","year":"2017"},{"key":"S0218001422520280BIB040","doi-asserted-by":"crossref","first-page":"3046","DOI":"10.3390\/s21093046","volume":"21","author":"Minaei S.","year":"2021","journal-title":"Sensors"},{"key":"S0218001422520280BIB041","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2016.7477450"},{"key":"S0218001422520280BIB042","doi-asserted-by":"publisher","DOI":"10.1109\/TITB.2009.2034649"},{"key":"S0218001422520280BIB044","doi-asserted-by":"crossref","first-page":"25241","DOI":"10.1007\/s11042-021-10918-9","volume":"80","author":"Said Y.","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"S0218001422520280BIB045","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"808","DOI":"10.1007\/978-3-642-33783-3_58","volume-title":"ECCV 2012: Computer Vision","volume":"7577","author":"Salah R.","year":"2012"},{"key":"S0218001422520280BIB046","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2008.08.005"},{"key":"S0218001422520280BIB048","doi-asserted-by":"publisher","DOI":"10.1142\/S0218001421520169"},{"key":"S0218001422520280BIB051","doi-asserted-by":"crossref","first-page":"1301","DOI":"10.1109\/JSTSP.2017.2764438","volume":"11","author":"Tzirakis P.","year":"2017","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"S0218001422520280BIB052","volume-title":"Proc. 2019 IEEE Int. Symp. Electrical and Electronics Engineering (ISEE)","author":"Vin T. P.","year":"2019"},{"key":"S0218001422520280BIB053","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-013-0861-x"},{"key":"S0218001422520280BIB054","volume-title":"Proc. 2020 IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Wang K.","year":"2020"},{"key":"S0218001422520280BIB055","first-page":"7208794","volume":"2018","author":"Wang X.","year":"2018","journal-title":"Comput. Intell. Neurosci."},{"key":"S0218001422520280BIB056","volume-title":"Proc. 7th Int. Conf. Automatic Face and Gesture Recognition (FGR 2006)","author":"Whitehill J.","year":"2006"},{"key":"S0218001422520280BIB057","doi-asserted-by":"publisher","DOI":"10.1145\/1165255.1165259"},{"key":"S0218001422520280BIB058","volume-title":"Proc. ACM 3rd Int. Conf. Robotics, Control and Automation","author":"Yoshihiro S.","year":"2018"},{"key":"S0218001422520280BIB059","first-page":"30","volume":"59","author":"Zaenal A.","year":"2012","journal-title":"Int. J. Comput. Appl."},{"key":"S0218001422520280BIB060","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"S0218001422520280BIB061","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2017.2788081"},{"key":"S0218001422520280BIB062","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"crossref","first-page":"425","DOI":"10.1007\/978-3-319-46475-6_27","volume-title":"ECCV 2016: Computer Vision","volume":"9006","author":"Zhao X.","year":"2016"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001422520280","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,6]],"date-time":"2024-10-06T21:22:32Z","timestamp":1728249752000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218001422520280"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,29]]},"references-count":54,"journal-issue":{"issue":"14","published-print":{"date-parts":[[2022,11]]}},"alternative-id":["10.1142\/S0218001422520280"],"URL":"https:\/\/doi.org\/10.1142\/s0218001422520280","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"value":"0218-0014","type":"print"},{"value":"1793-6381","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,29]]},"article-number":"2252028"}}