{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T06:30:22Z","timestamp":1782282622314,"version":"3.54.5"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2018,12,4]],"date-time":"2018-12-04T00:00:00Z","timestamp":1543881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2018,12,31]]},"abstract":"<jats:p>Despite the popularity of real-time monocular face tracking systems in many successful applications, one overlooked problem with these systems is rigid instability. It occurs when the input facial motion can be explained by either head pose change or facial expression change, creating ambiguities that often lead to jittery and unstable rigid head poses under large expressions. Existing rigid stabilization methods either employ a heavy anatomically-motivated approach that are unsuitable for real-time applications, or utilize heuristic-based rules that can be problematic under certain expressions. We propose the first rigid stabilization method for real-time monocular face tracking using a dynamic rigidity prior learned from realistic datasets. The prior is defined on a region-based face model and provides dynamic region-based adaptivity for rigid pose optimization during real-time performance. We introduce an effective offline training scheme to learn the dynamic rigidity prior by optimizing the convergence of the rigid pose optimization to the ground-truth poses in the training data. Our real-time face tracking system is an optimization framework that alternates between rigid pose optimization and expression optimization. To ensure tracking accuracy, we combine both robust, drift-free facial landmarks and dense optical flow into the optimization objectives. We evaluate our system extensively against state-of-the-art monocular face tracking systems and achieve significant improvement in tracking accuracy on the high-quality face tracking benchmark. Our system can improve facial-performance-based applications such as facial animation retargeting and virtual face makeup with accurate expression and stable pose. We further validate the dynamic rigidity prior by comparing it against other variants on the tracking accuracy.<\/jats:p>","DOI":"10.1145\/3272127.3275093","type":"journal-article","created":{"date-parts":[[2018,11,28]],"date-time":"2018-11-28T19:16:10Z","timestamp":1543432570000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["Stabilized real-time face tracking via a learned dynamic rigidity prior"],"prefix":"10.1145","volume":"37","author":[{"given":"Chen","family":"Cao","sequence":"first","affiliation":[{"name":"Snap Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Menglei","family":"Chai","sequence":"additional","affiliation":[{"name":"Snap Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Oliver","family":"Woodford","sequence":"additional","affiliation":[{"name":"Snap Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Linjie","family":"Luo","sequence":"additional","affiliation":[{"name":"Snap Inc."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,12,4]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Sameer Agarwal Keir Mierle and Others. 2016. Ceres Solver. http:\/\/ceres-solver.org. (2016).  Sameer Agarwal Keir Mierle and Others. 2016. Ceres Solver. http:\/\/ceres-solver.org. (2016)."},{"key":"e_1_2_2_2_1","unstructured":"Apple. 2017. Animoji. A new way to get into character. (2017). https:\/\/www.apple.com\/iphone-x\/#truedepth-camera  Apple. 2017. Animoji. A new way to get into character. (2017). https:\/\/www.apple.com\/iphone-x\/#truedepth-camera"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1833349.1778777"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601182"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964970"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/311535.311556"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461976"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1833349.1778778"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766943"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601204"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462012"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2013.249"},{"key":"e_1_2_2_13_1","unstructured":"Jin-Xiang Chai Jing Xiao and Jessica Hodgins. 2003. Vision-based Control of 3D Facial Animation. In SCA.   Jin-Xiang Chai Jing Xiao and Jessica Hodgins. 2003. Vision-based Control of 3D Facial Animation. In SCA."},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.449"},{"key":"e_1_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Yasutaka Furukawa and Jean Ponce. 2009. Dense 3D Motion Capture for Human Faces. In CVPR.  Yasutaka Furukawa and Jean Ponce. 2009. Dense 3D Motion Capture for Human Faces. In CVPR.","DOI":"10.1109\/CVPR.2009.5206868"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508380"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2890493"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2982419"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2070781.2024163"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964969"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/846276.846303"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.241"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DIMPVT.2012.67"},{"key":"e_1_2_2_24_1","volume-title":"Luc\", Jiri Matas, Nicu Sebe, and Max Welling.","author":"Kroeger Till","year":"2016","unstructured":"Till Kroeger , Radu Timofte , Dengxin Dai , editor =\"Leibe Bastian Van Gool , Luc\", Jiri Matas, Nicu Sebe, and Max Welling. 2016 . Fast Optical Flow Using Dense Inverse Search. Springer International Publishing , Cham, 471--488. Till Kroeger, Radu Timofte, Dengxin Dai, editor=\"Leibe Bastian Van Gool, Luc\", Jiri Matas, Nicu Sebe, and Max Welling. 2016. Fast Optical Flow Using Dense Inverse Search. Springer International Publishing, Cham, 471--488."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462019"},{"key":"e_1_2_2_26_1","volume-title":"Eurographics Symposium on Rendering. 183--194","author":"Ma Wan-Chun","year":"2007","unstructured":"Wan-Chun Ma , Tim Hawkins , Pieter Peers , Charles-Felix Chabert , Malte Weiss , and Paul Debevec . 2007 . Rapid Acquisition of Specular and Diffuse Normal Maps from Polarized Spherical Gradient Illumination . In Eurographics Symposium on Rendering. 183--194 . Wan-Chun Ma, Tim Hawkins, Pieter Peers, Charles-Felix Chabert, Malte Weiss, and Paul Debevec. 2007. Rapid Acquisition of Specular and Diffuse Normal Maps from Polarized Spherical Gradient Illumination. In Eurographics Symposium on Rendering. 183--194."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508417"},{"key":"e_1_2_2_28_1","volume-title":"Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic (NIPS'01)","author":"Ng Andrew Y.","year":"2001","unstructured":"Andrew Y. Ng , Michael I. Jordan , and Yair Weiss . 2001 . On Spectral Clustering: Analysis and an Algorithm . In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic (NIPS'01) . MIT Press, Cambridge, MA, USA, 849--856. http:\/\/dl.acm.org\/citation.cfm?id=2980539.2980649 Andrew Y. Ng, Michael I. Jordan, and Yair Weiss. 2001. On Spectral Clustering: Analysis and an Algorithm. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic (NIPS'01). MIT Press, Cambridge, MA, USA, 849--856. http:\/\/dl.acm.org\/citation.cfm?id=2980539.2980649"},{"key":"e_1_2_2_29_1","volume-title":"Segmented AAMs Improve Person-Independent Face Fitting. In In BMVC'07 - Proceedings of the 18th British Machine Vision Conference.","author":"Peyras Julien","year":"2007","unstructured":"Julien Peyras , Adrien Bartoli , Hugo Mercier , and Patrice Dalle . 2007 . Segmented AAMs Improve Person-Independent Face Fitting. In In BMVC'07 - Proceedings of the 18th British Machine Vision Conference. Julien Peyras, Adrien Bartoli, Hugo Mercier, and Patrice Dalle. 2007. Segmented AAMs Improve Person-Independent Face Fitting. In In BMVC'07 - Proceedings of the 18th British Machine Vision Conference."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2019406.2019435"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661290"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661290"},{"key":"e_1_2_2_33_1","volume-title":"Seitz","author":"Suwajanakorn Supasorn","year":"2014","unstructured":"Supasorn Suwajanakorn , Ira Kemelmacher-Shlizerman , and Steven M . Seitz . 2014 . Total Moving Face Reconstruction. In ECCV. Supasorn Suwajanakorn, Ira Kemelmacher-Shlizerman, and Steven M. Seitz. 2014. Total Moving Face Reconstruction. In ECCV."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964971"},{"key":"e_1_2_2_35_1","volume-title":"MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction. In The IEEE International Conference on Computer Vision (ICCV).","author":"Tewari Ayush","year":"2017","unstructured":"Ayush Tewari , Michael Zoll\u00f6fer , Hyeongwoo Kim , Pablo Garrido , Florian Bernard , Patrick Perez , and Theobalt Christian . 2017 . MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction. In The IEEE International Conference on Computer Vision (ICCV). Ayush Tewari, Michael Zoll\u00f6fer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Perez, and Theobalt Christian. 2017. MoFA: Model-based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction. In The IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_2_36_1","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2387--2395","author":"Thies J.","unstructured":"J. Thies , M. Zollh\u00f6fer , M. Stamminger , C. Theobalt , and M. Nie\u00dfner . 2016. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos . In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2387--2395 . J. Thies, M. Zollh\u00f6fer, M. Stamminger, C. Theobalt, and M. Nie\u00dfner. 2016. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2387--2395."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366206"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925947"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964972"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1599470.1599472"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925882"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015759"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2006.9"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3272127.3275093","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3272127.3275093","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:44:26Z","timestamp":1750207466000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3272127.3275093"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,12,4]]},"references-count":43,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2018,12,31]]}},"alternative-id":["10.1145\/3272127.3275093"],"URL":"https:\/\/doi.org\/10.1145\/3272127.3275093","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,12,4]]},"assertion":[{"value":"2018-12-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}