{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T16:19:52Z","timestamp":1781367592989,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":54,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3414034","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:53Z","timestamp":1602505613000},"page":"1864-1872","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":153,"title":["Sharp Multiple Instance Learning for DeepFake Video Detection"],"prefix":"10.1145","author":[{"given":"Xiaodan","family":"Li","sequence":"first","affiliation":[{"name":"Alibaba Group, China, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yining","family":"Lang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuefeng","family":"Chen","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaofeng","family":"Mao","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuan","family":"He","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuhui","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hui","family":"Xue","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Quan","family":"Lu","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/WIFS.2018.8630761"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2013.06.003"},{"key":"e_1_3_2_2_3_1","unstructured":"Anonymous. [n.d.]. Example of Partially attacked DeepFake video on Youtube. [EB\/OL]. https:\/\/www.youtube.com\/watch?v=BU9YAHigNx8 Accessed April 4 2020.  Anonymous. [n.d.]. Example of Partially attacked DeepFake video on Youtube. [EB\/OL]. https:\/\/www.youtube.com\/watch?v=BU9YAHigNx8 Accessed April 4 2020."},{"key":"e_1_3_2_2_4_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. In arXiv preprint:1409.0473.  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. In arXiv preprint:1409.0473."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2909827.2930786"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR. 6299--6308.  Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR. 6299--6308.","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Davide Cozzolino Giovanni Poggi and Luisa Verdoliva. 2017. Recasting Residual-based Local Descriptors as Convolutional Neural Networks: an Application to Image Forgery Detection. (2017) 159--164.  Davide Cozzolino Giovanni Poggi and Luisa Verdoliva. 2017. Recasting Residual-based Local Descriptors as Convolutional Neural Networks: an Application to Image Forgery Detection. (2017) 159--164.","DOI":"10.1145\/3082031.3083247"},{"key":"e_1_3_2_2_9_1","unstructured":"DeepFakes. 2019 a. www.github.com\/deepfakes\/ faceswap. Accessed (2019).  DeepFakes. 2019 a. www.github.com\/deepfakes\/ faceswap. Accessed (2019)."},{"key":"e_1_3_2_2_10_1","unstructured":"DeepFakes. 2019 b. www.github.com\/MarekKowalski\/. Accessed (2019).  DeepFakes. 2019 b. www.github.com\/MarekKowalski\/. Accessed (2019)."},{"key":"e_1_3_2_2_11_1","first-page":"1","article-title":"Solving the multiple-instance problem with axis-parallel rectangles","volume":"89","author":"Dietterich Thomas G.","year":"2001","journal-title":"Artificial Intelligence"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Xinyi Ding Zohreh Raziei Eric C Larson Eli V Olinick Paul S Krueger and Michael Hahsler. 2019. Swapped Face Detection using Deep Learning and Subjective Assessment. arXiv: Learning (2019).  Xinyi Ding Zohreh Raziei Eric C Larson Eli V Olinick Paul S Krueger and Michael Hahsler. 2019. Swapped Face Detection using Deep Learning and Subjective Assessment. arXiv: Learning (2019).","DOI":"10.1186\/s13635-020-00109-8"},{"key":"e_1_3_2_2_13_1","unstructured":"Brian Dolhansky Russ Howes Ben Pflaum Nicole Baram and Cristian Canton Ferrer. 2019. The Deepfake Detection Challenge (DFDC) Preview Dataset. arXiv preprint:1910.08854 (2019).  Brian Dolhansky Russ Howes Ben Pflaum Nicole Baram and Cristian Canton Ferrer. 2019. The Deepfake Detection Challenge (DFDC) Preview Dataset. arXiv preprint:1910.08854 (2019)."},{"key":"e_1_3_2_2_14_1","unstructured":"Facebook. [n.d.]. DeepFake Detection Challenge. [EB\/OL]. https:\/\/www.kaggle.com\/c\/deepfake-detection-challenge Accessed May 20 2020.  Facebook. [n.d.]. DeepFake Detection Challenge. [EB\/OL]. https:\/\/www.kaggle.com\/c\/deepfake-detection-challenge Accessed May 20 2020."},{"key":"e_1_3_2_2_15_1","unstructured":"Jiashi Feng Bingbing Ni Qi Tian and Shuicheng Yan. 2011. Geometric lp-norm feature pooling for image classification. In CVPR.  Jiashi Feng Bingbing Ni Qi Tian and Shuicheng Yan. 2011. Geometric lp-norm feature pooling for image classification. In CVPR."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Ji Feng and Zhihua Zhou. 2017. Deep MIML Network. In AAAI.  Ji Feng and Zhihua Zhou. 2017. Deep MIML Network. In AAAI.","DOI":"10.1609\/aaai.v31i1.10890"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2012.2190402"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Miroslav Goljan and Jessica Fridrich. 2015. CFA-aware features for steganalysis of color images. electronic imaging Vol. 9409.  Miroslav Goljan and Jessica Fridrich. 2015. CFA-aware features for steganalysis of color images. electronic imaging Vol. 9409.","DOI":"10.1117\/12.2078399"},{"key":"e_1_3_2_2_19_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In NIPS. 2672--2680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In NIPS. 2672--2680."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"David Guera and Edward J Delp. 2018. Deepfake Video Detection Using Recurrent Neural Networks. In AVSS. 1--6.  David Guera and Edward J Delp. 2018. Deepfake Video Detection Using Recurrent Neural Networks. In AVSS. 1--6.","DOI":"10.1109\/AVSS.2018.8639163"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation Vol. 9 8 (1997) 1735--1780.  Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation Vol. 9 8 (1997) 1735--1780.","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_2_22_1","unstructured":"Maximilian Ilse Jakub M Tomczak and Max Welling. 2018. Attention-based deep multiple instance learning. arXiv preprint:1802.04712 (2018).  Maximilian Ilse Jakub M Tomczak and Max Welling. 2018. Attention-based deep multiple instance learning. arXiv preprint:1802.04712 (2018)."},{"key":"e_1_3_2_2_23_1","unstructured":"James D. Keeler David E. Rumelhart and Wee Kheng Leow. 1990. Integrated segmentation and recognition of hand-printed numerals. In NIPS.  James D. Keeler David E. Rumelhart and Wee Kheng Leow. 1990. Integrated segmentation and recognition of hand-printed numerals. In NIPS."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Yoon Kim. 2014. Convolutional neural networks for sentence classification. arXiv preprint:1408.5882 (2014).  Yoon Kim. 2014. Convolutional neural networks for sentence classification. arXiv preprint:1408.5882 (2014).","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_2_25_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint:1412.6980","author":"Kingma Diederik P","year":"2014"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"crossref","unstructured":"Lingzhi Li Jianmin Bao Ting Zhang Hao Yang Dong Chen Fang Wen and Baining Guo. 2019 a. Face X-ray for More General Face Forgery Detection. CVPR.  Lingzhi Li Jianmin Bao Ting Zhang Hao Yang Dong Chen Fang Wen and Baining Guo. 2019 a. Face X-ray for More General Face Forgery Detection. CVPR.","DOI":"10.1109\/CVPR42600.2020.00505"},{"key":"e_1_3_2_2_27_1","unstructured":"Yuezun Li and Siwei Lyu. 2019. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In CVPRW.  Yuezun Li and Siwei Lyu. 2019. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In CVPRW."},{"key":"e_1_3_2_2_28_1","unstructured":"Yuezun Li Xin Yang Pu Sun Honggang Qi and Siwei Lyu. 2019 b. Celeb-df: A new dataset for deepfake forensics. arXiv preprint:1909.12962 (2019).  Yuezun Li Xin Yang Pu Sun Honggang Qi and Siwei Lyu. 2019 b. Celeb-df: A new dataset for deepfake forensics. arXiv preprint:1909.12962 (2019)."},{"key":"e_1_3_2_2_29_1","unstructured":"Tsung-Yi Lin Priya Goyal Ross Girshick Kaiming He and Piotr Doll\u00e1r. 2017b. Focal loss for dense object detection. In ICCV. 2980--2988.  Tsung-Yi Lin Priya Goyal Ross Girshick Kaiming He and Piotr Doll\u00e1r. 2017b. Focal loss for dense object detection. In ICCV. 2980--2988."},{"key":"e_1_3_2_2_30_1","unstructured":"Zhouhan Lin Minwei Feng Cicero Nogueira Dos Santos Mo Yu Bing Xiang Bowen Zhou and Yoshua Bengio. 2017a. A Structured Self-attentive Sentence Embedding. arXiv: Computation and Language.  Zhouhan Lin Minwei Feng Cicero Nogueira Dos Santos Mo Yu Bing Xiang Bowen Zhou and Yoshua Bengio. 2017a. A Structured Self-attentive Sentence Embedding. arXiv: Computation and Language."},{"key":"e_1_3_2_2_31_1","unstructured":"Oded Maron and Tomas Lozano-Perez. 1998. A framework for multiple-instance learning. In NIPS.  Oded Maron and Tomas Lozano-Perez. 1998. A framework for multiple-instance learning. In NIPS."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"crossref","unstructured":"Francesco Marra Diego Gragnaniello Davide Cozzolino and Luisa Verdoliva. 2018. Detection of GAN-Generated Fake Images over Social Networks. In IEEE MIPR.  Francesco Marra Diego Gragnaniello Davide Cozzolino and Luisa Verdoliva. 2018. Detection of GAN-Generated Fake Images over Social Networks. In IEEE MIPR.","DOI":"10.1109\/MIPR.2018.00084"},{"key":"e_1_3_2_2_33_1","unstructured":"Maxime Oquab Leon Bottou Ivan Laptev and Josef Sivic. 2014. Weakly supervised object recognition with convolutional neural networks. In NIPS.  Maxime Oquab Leon Bottou Ivan Laptev and Josef Sivic. 2014. Weakly supervised object recognition with convolutional neural networks. In NIPS."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Xunyu Pan Xing Zhang and Siwei Lyu. 2012. Exposing image splicing with inconsistent local noise variances. In ICCP. 1--10.  Xunyu Pan Xing Zhang and Siwei Lyu. 2012. Exposing image splicing with inconsistent local noise variances. In ICCP. 1--10.","DOI":"10.1109\/ICCPhot.2012.6215223"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Nikolaos Pappas and Andrei Popescu-Belis. 2014. Explaining the stars: Weighted multiple-instance learning for aspectbased sentiment analysis. In EMNLP.  Nikolaos Pappas and Andrei Popescu-Belis. 2014. Explaining the stars: Weighted multiple-instance learning for aspectbased sentiment analysis. In EMNLP.","DOI":"10.3115\/v1\/D14-1052"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Nikolaos Pappas and Andrei Popescu-Belis. 2017. Explicit Document Modeling through Weighted Multiple-Instance Learning. In Journal of Artificial Intelligence Research.  Nikolaos Pappas and Andrei Popescu-Belis. 2017. Explicit Document Modeling through Weighted Multiple-Instance Learning. In Journal of Artificial Intelligence Research.","DOI":"10.1613\/jair.5240"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"crossref","unstructured":"Pedro O Pinheiro and Ronan Collobert. 2015. From image-level to pixel-level labeling with convolutional networks. In CVPR.  Pedro O Pinheiro and Ronan Collobert. 2015. From image-level to pixel-level labeling with convolutional networks. In CVPR.","DOI":"10.1109\/CVPR.2015.7298780"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/RBME.2017.2651164"},{"key":"e_1_3_2_2_39_1","unstructured":"Alec Radford Luke Metz and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint:1511.06434 (2015).  Alec Radford Luke Metz and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint:1511.06434 (2015)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/WIFS.2017.8267647"},{"key":"e_1_3_2_2_41_1","unstructured":"Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Niesner. 2018. FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces. CVPR.  Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Niesner. 2018. FaceForensics: A Large-scale Video Dataset for Forgery Detection in Human Faces. CVPR."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Niesner. 2019. Faceforensics+: Learning to detect manipulated facial images. In arXiv preprint arXiv:1901.08971.  Andreas Rossler Davide Cozzolino Luisa Verdoliva Christian Riess Justus Thies and Matthias Niesner. 2019. Faceforensics+: Learning to detect manipulated facial images. In arXiv preprint arXiv:1901.08971.","DOI":"10.1109\/ICCV.2019.00009"},{"key":"e_1_3_2_2_43_1","first-page":"1","article-title":"Recurrent convolutional strategies for face manipulation detection in videos","volume":"3","author":"Sabir Ekraam","year":"2019","journal-title":"Interfaces (GUI)"},{"key":"e_1_3_2_2_44_1","volume-title":"Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV. 618--626.","author":"Selvaraju Ramprasaath R","year":"2017"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMI.2016.2525803"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267357.3267367"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323035"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292039"},{"key":"e_1_3_2_2_49_1","volume-title":"Fakespotter: A simple baseline for spotting ai-synthesized fake faces. In arXiv preprint arXiv:1909.06122.","author":"Wang Run","year":"2019"},{"key":"e_1_3_2_2_50_1","unstructured":"Xinggang Wang Yongluan Yan Peng Tang Xiang Bai and Wenyu Liu. 2016. Revisiting multiple instance neural networks. In Pattern Recognition.  Xinggang Wang Yongluan Yan Peng Tang Xiang Bai and Wenyu Liu. 2016. Revisiting multiple instance neural networks. In Pattern Recognition."},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.08.026"},{"key":"e_1_3_2_2_52_1","unstructured":"Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show attend and tell: Neural image caption generation with visual attention. In ICML.  Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show attend and tell: Neural image caption generation with visual attention. In ICML."},{"key":"e_1_3_2_2_53_1","unstructured":"Sergey Zagoruyko and Nikos Komodakis. 2016. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In arXiv preprint.  Sergey Zagoruyko and Nikos Komodakis. 2016. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In arXiv preprint."},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"crossref","unstructured":"Wentao Zhu Qi Lou Yeeleng Scott Vang and Xiaohui Xie. 2017. Deep multi-instance networks with sparse label assignment for whole mammogram classification. In MICCAI.  Wentao Zhu Qi Lou Yeeleng Scott Vang and Xiaohui Xie. 2017. Deep multi-instance networks with sparse label assignment for whole mammogram classification. In MICCAI.","DOI":"10.1101\/095794"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","location":"Seattle WA USA","acronym":"MM '20","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414034","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3414034","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:23Z","timestamp":1750197683000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414034"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":54,"alternative-id":["10.1145\/3394171.3414034","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3414034","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}