{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T14:50:27Z","timestamp":1784904627234,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413805","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:25Z","timestamp":1602505585000},"page":"4144-4152","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":58,"title":["Weakly Supervised 3D Object Detection from Point Clouds"],"prefix":"10.1145","author":[{"given":"Zengyi","family":"Qin","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinglu","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Lu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Mart'in","year":"2016","unstructured":"Mart'in Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . Tensorflow: A system for large-scale machine learning . In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . 265--283. Mart'in Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. Tensorflow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 265--283."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2774007"},{"key":"e_1_3_2_2_3_1","volume-title":"Andreas Damianou, Neil D Lawrence, and Zhenwen Dai.","author":"Ahn Sungsoo","year":"2019","unstructured":"Sungsoo Ahn , Shell Xu Hu , Andreas Damianou, Neil D Lawrence, and Zhenwen Dai. 2019 . Variational Information Distillation for Knowledge Transfer. arXiv: Computer Vision and Pattern Recognition ( 2019). Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D Lawrence, and Zhenwen Dai. 2019. Variational Information Distillation for Knowledge Transfer. arXiv: Computer Vision and Pattern Recognition (2019)."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344972"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.311"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00567"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.236"},{"key":"e_1_3_2_2_8_1","unstructured":"Xiaozhi Chen Kaustav Kundu Yukun Zhu Andrew G Berneshawi Huimin Ma Sanja Fidler and Raquel Urtasun. 2015. 3d object proposals for accurate object class detection. In Advances in Neural Information Processing Systems. 424--432.  Xiaozhi Chen Kaustav Kundu Yukun Zhu Andrew G Berneshawi Huimin Ma Sanja Fidler and Raquel Urtasun. 2015. 3d object proposals for accurate object class detection. In Advances in Neural Information Processing Systems. 424--432."},{"key":"e_1_3_2_2_9_1","first-page":"3","article-title":"Multi-view 3d object detection network for autonomous driving","volume":"1","author":"Chen Xiaozhi","year":"2017","unstructured":"Xiaozhi Chen , Huimin Ma , Ji Wan , Bo Li , and Tian Xia . 2017 . Multi-view 3d object detection network for autonomous driving . In IEEE CVPR , Vol. 1. 3 . Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. 2017. Multi-view 3d object detection network for autonomous driving. In IEEE CVPR, Vol. 1. 3.","journal-title":"IEEE CVPR"},{"key":"e_1_3_2_2_10_1","volume-title":"Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals. computer vision and pattern recognition","author":"Cho Minsu","year":"2015","unstructured":"Minsu Cho , Suha Kwak , Cordelia Schmid , and Jean Ponce . 2015. Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals. computer vision and pattern recognition ( 2015 ), 1201--1210. Minsu Cho, Suha Kwak, Cordelia Schmid, and Jean Ponce. 2015. Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals. computer vision and pattern recognition (2015), 1201--1210."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_3_2_2_13_1","unstructured":"Huan Fu Mingming Gong Chaohui Wang Kayhan Batmanghelich and Dacheng Tao. 2018. Deep Ordinal Regression Network for Monocular Depth Estimation. In Computer Vision and Pattern Recognition (CVPR).  Huan Fu Mingming Gong Chaohui Wang Kayhan Batmanghelich and Dacheng Tao. 2018. Deep Ordinal Regression Network for Monocular Depth Estimation. In Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_7"},{"key":"e_1_3_2_2_15_1","volume-title":"Computer Vision and Pattern Recognition (CVPR)","author":"Geiger Andreas","unstructured":"Andreas Geiger , Philip Lenz , and Raquel Urtasun . 2012. Are we ready for autonomous driving? the kitti vision benchmark suite . In Computer Vision and Pattern Recognition (CVPR) . IEEE , 3354--3361. Andreas Geiger, Philip Lenz, and Raquel Urtasun. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In Computer Vision and Pattern Recognition (CVPR). IEEE, 3354--3361."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.309"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2014.2374218"},{"key":"e_1_3_2_2_18_1","volume-title":"Piotr Doll\u00e1 r, and Ross Girshick","author":"He Kaiming","year":"2017","unstructured":"Kaiming He , Georgia Gkioxari , Piotr Doll\u00e1 r, and Ross Girshick . 2017 . Mask R-CNN. arXiv preprint arXiv:1703.06870 (2017). Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1 r, and Ross Girshick. 2017. Mask R-CNN. arXiv preprint arXiv:1703.06870 (2017)."},{"key":"e_1_3_2_2_19_1","volume-title":"Distilling the Knowledge in a Neural Network. arXiv: Machine Learning","author":"Hinton Geoffrey E","year":"2015","unstructured":"Geoffrey E Hinton , Oriol Vinyals , and Jeffrey Dean . 2015. Distilling the Knowledge in a Neural Network. arXiv: Machine Learning ( 2015 ). Geoffrey E Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv: Machine Learning (2015)."},{"key":"e_1_3_2_2_20_1","volume-title":"Joint Monocular 3D Vehicle Detection and Tracking. arXiv preprint arXiv:1811.10742","author":"Hu Hou-Ning","year":"2018","unstructured":"Hou-Ning Hu , Qi-Zhi Cai , Dequan Wang , Ji Lin , Min Sun , Philipp Krahenb\u00fchl , Trevor Darrell , and Fisher Yu. 2018. Joint Monocular 3D Vehicle Detection and Tracking. arXiv preprint arXiv:1811.10742 ( 2018 ). Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin, Min Sun, Philipp Krahenb\u00fchl, Trevor Darrell, and Fisher Yu. 2018. Joint Monocular 3D Vehicle Detection and Tracking. arXiv preprint arXiv:1811.10742 (2018)."},{"key":"e_1_3_2_2_21_1","volume-title":"Like What You Like: Knowledge Distill via Neuron Selectivity Transfer. arXiv: Computer Vision and Pattern Recognition","author":"Huang Zehao","year":"2017","unstructured":"Zehao Huang and Naiyan Wang . 2017. Like What You Like: Knowledge Distill via Neuron Selectivity Transfer. arXiv: Computer Vision and Pattern Recognition ( 2017 ). Zehao Huang and Naiyan Wang. 2017. Like What You Like: Knowledge Distill via Neuron Selectivity Transfer. arXiv: Computer Vision and Pattern Recognition (2017)."},{"key":"e_1_3_2_2_22_1","volume-title":"Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Inoue Naoto","year":"2018","unstructured":"Naoto Inoue , Ryosuke Furuta , Toshihiko Yamasaki , and Kiyoharu Aizawa . 2018 . Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, and Kiyoharu Aizawa. 2018. Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_23_1","volume-title":"ContextLocNet: Context-aware Deep Network Models for Weakly Supervised Localization. In Proc. European Conference on Computer Vision (ECCV)","author":"Kantorov V.","year":"2016","unstructured":"V. Kantorov , M. Oquab , Cho M., and I. Laptev . 2016 . ContextLocNet: Context-aware Deep Network Models for Weakly Supervised Localization. In Proc. European Conference on Computer Vision (ECCV) , 2016 . V. Kantorov, M. Oquab, Cho M., and I. Laptev. 2016. ContextLocNet: Context-aware Deep Network Models for Weakly Supervised Localization. In Proc. European Conference on Computer Vision (ECCV), 2016."},{"key":"e_1_3_2_2_24_1","volume-title":"Adam: A Method for Stochastic Optimization. In International Conference for Learning Representations.","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015 . Adam: A Method for Stochastic Optimization. In International Conference for Learning Representations. Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In International Conference for Learning Representations."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.14569\/IJACSA.2016.070180"},{"key":"e_1_3_2_2_26_1","unstructured":"Chenhao Lin Siwen Wang Dongqi Xu Yu Lu and Wayne Zhang. 2020. Object Instance Mining for Weakly Supervised Object Detection. AAAI.  Chenhao Lin Siwen Wang Dongqi Xu Yu Lu and Wayne Zhang. 2020. Object Instance Mining for Weakly Supervised Object Detection. AAAI."},{"key":"e_1_3_2_2_27_1","volume-title":"Visualizing and Understanding Convolutional Networks. In European Conference on Computer Vision. 818--833","author":"Zeiler Matthew D","year":"2014","unstructured":"D Zeiler Matthew and Fergus Rob . 2014 . Visualizing and Understanding Convolutional Networks. In European Conference on Computer Vision. 818--833 . D Zeiler Matthew and Fergus Rob. 2014. Visualizing and Understanding Convolutional Networks. In European Conference on Computer Vision. 818--833."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.597"},{"key":"e_1_3_2_2_29_1","volume-title":"Frustum PointNets for 3D Object Detection from RGB-D Data. arXiv preprint arXiv:1711.08488","author":"Qi Charles R","year":"2017","unstructured":"Charles R Qi , Wei Liu , Chenxia Wu , Hao Su , and Leonidas J Guibas . 2017. Frustum PointNets for 3D Object Detection from RGB-D Data. arXiv preprint arXiv:1711.08488 ( 2017 ). Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. 2017. Frustum PointNets for 3D Object Detection from RGB-D Data. arXiv preprint arXiv:1711.08488 (2017)."},{"key":"e_1_3_2_2_30_1","volume-title":"2019 a. MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization. AAAI","author":"Qin Zengyi","year":"2019","unstructured":"Zengyi Qin , Jinglu Wang , and Yan Lu . 2019 a. MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization. AAAI ( 2019 ). Zengyi Qin, Jinglu Wang, and Yan Lu. 2019 a. MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization. AAAI (2019)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00780"},{"key":"e_1_3_2_2_32_1","volume-title":"Faster R-CNN: towards real-time object detection with region proposal networks","author":"Ren Shaoqing","year":"2017","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2017. Faster R-CNN: towards real-time object detection with region proposal networks . IEEE Transactions on Pattern Analysis & Machine Intelligence ( 2017 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2017. Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis & Machine Intelligence (2017)."},{"key":"e_1_3_2_2_33_1","volume-title":"Orthographic feature transform for monocular 3D object detection. arXiv preprint arXiv:1811.08188","author":"Roddick Thomas","year":"2018","unstructured":"Thomas Roddick , Alex Kendall , and Roberto Cipolla . 2018. Orthographic feature transform for monocular 3D object detection. arXiv preprint arXiv:1811.08188 ( 2018 ). Thomas Roddick, Alex Kendall, and Roberto Cipolla. 2018. Orthographic feature transform for monocular 3D object detection. arXiv preprint arXiv:1811.08188 (2018)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_2_35_1","volume-title":"Self paced deep learning for weakly supervised object detection","author":"Sangineto Enver","year":"2019","unstructured":"Enver Sangineto , Moin Nabi , Dubravko Culibrk , and Nicu Sebe . 2019. Self paced deep learning for weakly supervised object detection . IEEE transactions on pattern analysis and machine intelligence, Vol. 41 , 3 ( 2019 ), 712--725. Enver Sangineto, Moin Nabi, Dubravko Culibrk, and Nicu Sebe. 2019. Self paced deep learning for weakly supervised object detection. IEEE transactions on pattern analysis and machine intelligence, Vol. 41, 3 (2019), 712--725."},{"key":"e_1_3_2_2_36_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSEN.2018.2888815"},{"key":"e_1_3_2_2_38_1","volume-title":"Pcl: Proposal cluster learning for weakly supervised object detection","author":"Tang Peng","year":"2018","unstructured":"Peng Tang , Xinggang Wang , Song Bai , Wei Shen , Xiang Bai , Wenyu Liu , and Alan Loddon Yuille . 2018 . Pcl: Proposal cluster learning for weakly supervised object detection . IEEE transactions on pattern analysis and machine intelligence (2018). Peng Tang, Xinggang Wang, Song Bai, Wei Shen, Xiang Bai, Wenyu Liu, and Alan Loddon Yuille. 2018. Pcl: Proposal cluster learning for weakly supervised object detection. IEEE transactions on pattern analysis and machine intelligence (2018)."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Peng Tang Xinggang Wang Xiang Bai and Wenyu Liu. 2017. Multiple Instance Detection Network with Online Instance Classifier Refinement. In CVPR.  Peng Tang Xinggang Wang Xiang Bai and Wenyu Liu. 2017. Multiple Instance Detection Network with Online Instance Classifier Refinement. In CVPR.","DOI":"10.1109\/CVPR.2017.326"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00202"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0620-5"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00141"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Yan Wang Wei-Lun Chao Divyansh Garg Bharath Hariharan Mark Campbell and Kilian Weinberger. 2019. Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving. In CVPR.  Yan Wang Wei-Lun Chao Divyansh Garg Bharath Hariharan Mark Campbell and Kilian Weinberger. 2019. Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving. In CVPR.","DOI":"10.1109\/CVPR.2019.00864"},{"key":"e_1_3_2_2_44_1","volume-title":"Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud. ArXiv","author":"Weng Xinshuo","year":"2019","unstructured":"Xinshuo Weng and Kris Makoto Kitani . 2019. Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud. ArXiv , Vol. abs\/ 1903 .09847 ( 2019 ). Xinshuo Weng and Kris Makoto Kitani. 2019. Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud. ArXiv, Vol. abs\/1903.09847 (2019)."},{"key":"e_1_3_2_2_45_1","volume-title":"Small, Low Power Fully Convolutional Neural Networks for Real-Time Object Detection for Autonomous Driving. CoRR","author":"Wu Bichen","year":"2016","unstructured":"Bichen Wu , Forrest N. Iandola , Peter H. Jin , and Kurt Keutzer . 2016. SqueezeDet : Unified , Small, Low Power Fully Convolutional Neural Networks for Real-Time Object Detection for Autonomous Driving. CoRR ( 2016 ). Bichen Wu, Forrest N. Iandola, Peter H. Jin, and Kurt Keutzer. 2016. SqueezeDet: Unified, Small, Low Power Fully Convolutional Neural Networks for Real-Time Object Detection for Autonomous Driving. CoRR (2016)."},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2014.6836101"},{"key":"e_1_3_2_2_47_1","volume-title":"STD: Sparse-to-Dense 3D Object Detector for Point Cloud. ICCV","author":"Yang Zetong","year":"2019","unstructured":"Zetong Yang , Yanan Sun , Shu Liu , Xiaoyong Shen , and Jiaya Jia . 2019 . STD: Sparse-to-Dense 3D Object Detector for Point Cloud. ICCV (2019). arxiv: 1907.10471 http:\/\/arxiv.org\/abs\/1907.10471 Zetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen, and Jiaya Jia. 2019. STD: Sparse-to-Dense 3D Object Detector for Point Cloud. ICCV (2019). arxiv: 1907.10471 http:\/\/arxiv.org\/abs\/1907.10471"},{"key":"e_1_3_2_2_48_1","volume-title":"VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. computer vision and pattern recognition","author":"Zhou Yin","year":"2018","unstructured":"Yin Zhou and Oncel Tuzel . 2018. VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. computer vision and pattern recognition ( 2018 ), 4490--4499. Yin Zhou and Oncel Tuzel. 2018. VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. computer vision and pattern recognition (2018), 4490--4499."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","location":"Seattle WA USA","acronym":"MM '20","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413805","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413805","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:17Z","timestamp":1750197677000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413805"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":48,"alternative-id":["10.1145\/3394171.3413805","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413805","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}