{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T21:31:34Z","timestamp":1775251894943,"version":"3.50.1"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"1s","license":[{"start":{"date-parts":[[2018,3,6]],"date-time":"2018-03-06T00:00:00Z","timestamp":1520294400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the General Financial Grant from the China Postdoctoral Science Foundation","award":["2015M570413"],"award-info":[{"award-number":["2015M570413"]}]},{"name":"the Graduate Student Scientific Research Innovation Projects in Jiangsu Province","award":["KYCX17_1811"],"award-info":[{"award-number":["KYCX17_1811"]}]},{"name":"the Open Project Program of the National Laboratory of Pattern Recognition","award":["201700022"],"award-info":[{"award-number":["201700022"]}]},{"name":"the Innovation Project of Undergraduate Students in Jiangsu University","award":["16A235"],"award-info":[{"award-number":["16A235"]}]},{"DOI":"10.13039\/501100001809","name":"the National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61272211,61672267"],"award-info":[{"award-number":["61272211,61672267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2018,3,31]]},"abstract":"<jats:p>Feature learning has enjoyed much attention and achieved good performance in recent studies of image processing. Unlike the required training conditions often assumed there, far less labeled data is available for training emotion classification systems. In addition, current feature learning is typically performed on an entire face image without considering the dependency between features. These approaches ignore the fact that faces are structured and the neighboring features are dependent. Thus, the learned features lack the power to describe visually coherent facial images. Our method is therefore designed with the goal of simplifying the problem domain by removing expression-irrelevant factors from the input images, with a key region-based mechanism, which is an effort to reduce the amount of data required to effectively train the feature-learning methods. Meanwhile, we can construct geometric constraints between the key regions and its detected positions. To this end, we introduce a Spatially Coherent featurelearning method for Pose-invariant Facial Expression Recognition (SC-PFER). In our model, we first perform face frontalization through a 3D pose-normalization technique, which could normalize poses while preserving the identity information through synthesizing frontal faces for facial images with arbitrary views. Subsequently, we select a sequence of key regions around 51 key points in the synthetic frontal face images for efficient unsupervised feature learning. Finally, we introduce a linkage structure over the learning-based features and the corresponding geometry information of each key region to encode the dependencies of the regions. Our method, on the whole, does not require training multiple models for each specific pose and avoids separating training and parameter tuning for each pose. The proposed framework has been evaluated on two benchmark databases, BU-3DFE and SFEW, for pose-invariant Facial Expression Recognition (FER). The experimental results demonstrate that our algorithm outperforms current state-of-the-art FER methods. Specifically, our model achieves an improvement of 1.72% and 1.11% FER accuracy, on average, on BU-3DFE and SFEW, respectively.<\/jats:p>","DOI":"10.1145\/3176646","type":"journal-article","created":{"date-parts":[[2018,3,7]],"date-time":"2018-03-07T19:00:36Z","timestamp":1520449236000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Spatially Coherent Feature Learning for Pose-Invariant Facial Expression Recognition"],"prefix":"10.1145","volume":"14","author":[{"given":"Feifei","family":"Zhang","sequence":"first","affiliation":[{"name":"Jiangsu University, Zhenjiang, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qirong","family":"Mao","sequence":"additional","affiliation":[{"name":"Jiangsu University, Zhenjiang, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiangjun","family":"Shen","sequence":"additional","affiliation":[{"name":"Jiangsu University, Zhenjiang, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongzhao","family":"Zhan","sequence":"additional","affiliation":[{"name":"Jiangsu University, Zhenjiang, Jiangsu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ming","family":"Dong","sequence":"additional","affiliation":[{"name":"Corresponding Author B, Wayne State University, Detroit, Michigan, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,3,6]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of IEEE International Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Benitez Quiroz Fabian C.","unstructured":"Fabian C. Benitez Quiroz , Ramprakash Srinivasan , and Aleix M. Martinez . 2016. EmotioNet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild . In Proceedings of IEEE International Conference on Computer Vision and Pattern Recognition (CVPR\u201916) . 5562--5570. Fabian C. Benitez Quiroz, Ramprakash Srinivasan, and Aleix M. Martinez. 2016. EmotioNet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. In Proceedings of IEEE International Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 5562--5570."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502173"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the International Conference on Artificial Intelligence and Statistics. 215--223","author":"Coates Adam","unstructured":"Adam Coates , Honglak Lee , and Andrew Y. Ng . 2011. An analysis of single-layer networks in unsupervised feature learning . In Proceedings of the International Conference on Artificial Intelligence and Statistics. 215--223 . Adam Coates, Honglak Lee, and Andrew Y. Ng. 2011. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the International Conference on Artificial Intelligence and Statistics. 215--223."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG.2011.5771368"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461466.2461529"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2011.6130508"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2394777"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2845089"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2390959"},{"key":"e_1_2_1_10_1","volume-title":"Friesen","author":"Ekman Paul","year":"1977","unstructured":"Paul Ekman and Wallace V . Friesen . 1977 . Facial action coding system. Consulting Psychologists Press , Stanford University, Palo Alto, CA. Paul Ekman and Wallace V. Friesen. 1977. Facial action coding system. Consulting Psychologists Press, Stanford University, Palo Alto, CA."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2014.2375634"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0031-3203(02)00052-3"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.169"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/SITIS.2011.64"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299058"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 21st International Conference on Pattern Recognition (ICPR\u201912)","author":"Hesse Nikolas","year":"2012","unstructured":"Nikolas Hesse , Tobias Gehrig , Hua Gao , and Haz\u0131m Kemal Ekenel . 2012 . Multi-view facial expression recognition using local appearance features . In Proceedings of the 21st International Conference on Pattern Recognition (ICPR\u201912) . IEEE, 3533--3536. Nikolas Hesse, Tobias Gehrig, Hua Gao, and Haz\u0131m Kemal Ekenel. 2012. Multi-view facial expression recognition using local appearance features. In Proceedings of the 21st International Conference on Pattern Recognition (ICPR\u201912). IEEE, 3533--3536."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 19th International Conference on Pattern Recognition (ICPR\u201908)","author":"Hu Yuxiao","unstructured":"Yuxiao Hu , Zhihong Zeng , Lijun Yin , Xiaozhou Wei , Jilin Tu , and Thomas S. Huang . 2008. A study of non-frontal-view facial expressions recognition . In Proceedings of the 19th International Conference on Pattern Recognition (ICPR\u201908) . IEEE, 1--4. Yuxiao Hu, Zhihong Zeng, Lijun Yin, Xiaozhou Wei, Jilin Tu, and Thomas S. Huang. 2008. A study of non-frontal-view facial expressions recognition. In Proceedings of the 19th International Conference on Pattern Recognition (ICPR\u201908). IEEE, 1--4."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 20th Computer Vision Winter Workshop Paul Wohlhart.","author":"Jampour Mahdi","year":"2015","unstructured":"Mahdi Jampour , Thomas Mauthner , and Horst Bischof . 2015 . Multi-view facial expressions recognition using local linear regression of sparse codes . In Proceedings of the 20th Computer Vision Winter Workshop Paul Wohlhart. Mahdi Jampour, Thomas Mauthner, and Horst Bischof. 2015. Multi-view facial expressions recognition using local linear regression of sparse codes. In Proceedings of the 20th Computer Vision Winter Workshop Paul Wohlhart."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.341"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33718-5_58"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3092840"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2808204"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NIPS\u201912)","author":"Krizhevsky Alex","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E. Hinton . 2012. ImageNet classification with deep convolutional neural networks . In Proceedings of Advances in Neural Information Processing Systems (NIPS\u201912) . 1097--1105. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Proceedings of Advances in Neural Information Processing Systems (NIPS\u201912). 1097--1105."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2009.2026826"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818346.2830587"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.233"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2010.12.001"},{"key":"e_1_2_1_29_1","volume-title":"Prediction as a candidate for learning deep hierarchical models of data. Thesis","author":"Palm Rasmus Berg","unstructured":"Rasmus Berg Palm . 2012. Prediction as a candidate for learning deep hierarchical models of data. Thesis , Technical University of Denmark , Kongens Lyngby , 5. Rasmus Berg Palm. 2012. Prediction as a candidate for learning deep hierarchical models of data. Thesis, Technical University of Denmark, Kongens Lyngby, 5."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2482228"},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Maja Pantic and Marian Stewart Bartlett. 2007. Machine analysis of facial expressions. In Face Recognition. InTech.  Maja Pantic and Marian Stewart Bartlett. 2007. Machine analysis of facial expressions. In Face Recognition. InTech.","DOI":"10.5772\/4847"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.41"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NIPS\u201915)","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015 . Faster r-CNN: Towards real-time object detection with region proposal networks . In Proceedings of Advances in Neural Information Processing Systems (NIPS\u201915) . 91--99. Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-CNN: Towards real-time object detection with region proposal networks. In Proceedings of Advances in Neural Information Processing Systems (NIPS\u201915). 91--99."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_58"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML\u201911)","author":"Rifai Salah","year":"2011","unstructured":"Salah Rifai , Pascal Vincent , Xavier Muller , Xavier Glorot , and Yoshua Bengio . 2011 . Contractive auto-encoders: Explicit invariance during feature extraction . In Proceedings of the 28th International Conference on Machine Learning (ICML\u201911) . 833--840. Salah Rifai, Pascal Vincent, Xavier Muller, Xavier Glorot, and Yoshua Bengio. 2011. Contractive auto-encoders: Explicit invariance during feature extraction. In Proceedings of the 28th International Conference on Machine Learning (ICML\u201911). 833--840."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.233"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2010.1001"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654916"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2014.2366127"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912)","author":"Sharma Abhishek","unstructured":"Abhishek Sharma , Abhishek Kumar , Hal Daume , and David W. Jacobs . 2012. Generalized multiview analysis: A discriminative latent space . In Proceedings of 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912) . 2160--2167. Abhishek Sharma, Abhishek Kumar, Hal Daume, and David W. Jacobs. 2012. Generalized multiview analysis: A discriminative latent space. In Proceedings of 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912). 2160--2167."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2014.2337845"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33868-7_25"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.244"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995710"},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo (ICME\u201910)","author":"Tang Hao","unstructured":"Hao Tang , Mark Hasegawa-Johnson , and Thomas S. Huang . 2010. Non-frontal view facial expression recognition based on ergodic hidden Markov model supervectors . In Proceedings of IEEE International Conference on Multimedia and Expo (ICME\u201910) . IEEE, 1202--1207. Hao Tang, Mark Hasegawa-Johnson, and Thomas S. Huang. 2010. Non-frontal view facial expression recognition based on ergodic hidden Markov model supervectors. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME\u201910). IEEE, 1202--1207."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33885-4_58"},{"key":"e_1_2_1_47_1","volume-title":"Proceedings of the 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG\u201913)","author":"Tariq Usman","unstructured":"Usman Tariq , Jianchao Yang , and Thomas S. Huang . 2013. Maximum margin GMM learning for facial expression recognition . In Proceedings of the 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG\u201913) . IEEE, 1--6. Usman Tariq, Jianchao Yang, and Thomas S. Huang. 2013. Maximum margin GMM learning for facial expression recognition. In Proceedings of the 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG\u201913). IEEE, 1--6."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2014.05.011"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2002.1017623"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCB.2012.2192269"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.5555\/1126250.1126340"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.52"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2008.921737"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967240"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.377"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2014.2304712"},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the IEEE 12th International Conference on Computer Vision (ICCV\u201909)","author":"Zheng Wenming","year":"1901","unstructured":"Wenming Zheng , Hao Tang , Zhouchen Lin , and Thomas S. Huang . 2009. A novel approach to expression recognition from non-frontal face images . In Proceedings of the IEEE 12th International Conference on Computer Vision (ICCV\u201909) . IEEE, 1901 --1908. Wenming Zheng, Hao Tang, Zhouchen Lin, and Thomas S. Huang. 2009. A novel approach to expression recognition from non-frontal face images. In Proceedings of the IEEE 12th International Conference on Computer Vision (ICCV\u201909). IEEE, 1901--1908."},{"key":"e_1_2_1_58_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Zhu Xiangyu","unstructured":"Xiangyu Zhu , Zhen Lei , Junjie Yan , Dong Yi , and Stan Z. Li . 2015. High-fidelity pose and expression normalization for face recognition in the wild . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915) . 787--796. Xiangyu Zhu, Zhen Lei, Junjie Yan, Dong Yi, and Stan Z. Li. 2015. High-fidelity pose and expression normalization for face recognition in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915). 787--796."},{"key":"e_1_2_1_59_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912)","author":"Zhu Xiangxin","year":"2012","unstructured":"Xiangxin Zhu and Deva Ramanan . 2012 . Face detection, pose estimation, and landmark localization in the wild . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912) . 2879--2886. Xiangxin Zhu and Deva Ramanan. 2012. Face detection, pose estimation, and landmark localization in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201912). 2879--2886."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.259"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3176646","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3176646","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:11:32Z","timestamp":1750212692000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3176646"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,3,6]]},"references-count":60,"journal-issue":{"issue":"1s","published-print":{"date-parts":[[2018,3,31]]}},"alternative-id":["10.1145\/3176646"],"URL":"https:\/\/doi.org\/10.1145\/3176646","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,3,6]]},"assertion":[{"value":"2016-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-03-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}