{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T14:32:49Z","timestamp":1786199569063,"version":"3.56.0"},"reference-count":811,"publisher":"Emerald","issue":"1-3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7,6]]},"abstract":"<jats:p>Recent years have witnessed enormous progress in AI-related fields such as computer vision, machine learning, and autonomous vehicles. As with any rapidly growing field, it becomes increasingly difficult to stay up-to-date or enter the field as a beginner. While several survey papers on particular sub-problems have appeared, no comprehensive survey on problems, datasets, and methods in computer vision for autonomous vehicles has been published. This monograph attempts to narrow this gap by providing a survey on the state-of-the-art datasets and techniques. Our survey includes both the historically most relevant literature as well as the current state of the art on several specific topics, including recognition, reconstruction, motion estimation, tracking, scene understanding, and end-to-end learning for autonomous driving. Towards this goal, we analyze the performance of the state of the art on several challenging benchmarking datasets, including KITTI, MOT, and Cityscapes. Besides, we discuss open problems and current research challenges. To ease accessibility and accommodate missing references, we also provide a website that allows navigating topics as well as methods and provides additional information.<\/jats:p>","DOI":"10.1561\/0600000079","type":"journal-article","created":{"date-parts":[[2020,7,6]],"date-time":"2020-07-06T04:18:31Z","timestamp":1594009111000},"page":"1-308","source":"Crossref","is-referenced-by-count":461,"title":["Computer Vision for Autonomous Vehicles"],"prefix":"10.1108","volume":"12","author":[{"given":"Joel","family":"Janai","sequence":"first","affiliation":[{"name":"Max-Planck-Institute for Intelligent Systems T\u00fcbingen, Germany University of T\u00fcbingen ,","place":["Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fatma","family":"G\u00fcney","sequence":"additional","affiliation":[{"name":"College of Engineering, Ko\u00e7 University ,","place":["Turkey"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aseem","family":"Behl","sequence":"additional","affiliation":[{"name":"Max-Planck-Institute for Intelligent Systems T\u00fcbingen, Germany University of T\u00fcbingen ,","place":["Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas","family":"Geiger","sequence":"additional","affiliation":[{"name":"Max-Planck-Institute for Intelligent Systems T\u00fcbingen, Germany University of T\u00fcbingen ,","place":["Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2020,7,6]]},"reference":[{"key":"2026032615364252300_ref001","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00096","article-title":"Efficient interactive annotation of segmentation datasets with polygon-RNN++","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Acuna","year":"2018"},{"key":"2026032615364252300_ref002","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2009.5459148","article-title":"Building Rome in a day","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Agarwal","year":"2009"},{"key":"2026032615364252300_ref003","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1016\/j.robot.2016.07.003","article-title":"A practical approach for detection and classification of traffic signs using Convolutional Neural Networks","volume":"84","author":"Aghdam","year":"2016","journal-title":"Robotics and Autonomous Systems (RAS)"},{"key":"2026032615364252300_ref004","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5540228","article-title":"3D scene priors for road detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Alvarez","year":"2010"},{"key":"2026032615364252300_ref005","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33786-4_28","article-title":"Road scene segmentation from a single image","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"\u00c1lvarez","year":"2012"},{"issue":"1","key":"2026032615364252300_ref006","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1109\/TITS.2010.2076349","article-title":"Road detection based on illuminant invariance","volume":"12","author":"\u00c1lvarez","year":"2011","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref007","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2008.4621152","article-title":"Real time detection of lane markers in urban streets","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Aly","year":"2008"},{"issue":"3","key":"2026032615364252300_ref008","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1007\/BF00158167","article-title":"A computational framework and an algorithm for the measurement of visual motion","volume":"2","author":"Anandan","year":"1989","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref009","doi-asserted-by":"crossref","DOI":"10.1016\/j.robot.2009.09.011","article-title":"6D scan registration using depth-interpolated local image features","volume-title":"Robotics and Autonomous Systems (RAS)","author":"Andreasson","year":"2010"},{"key":"2026032615364252300_ref010","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587583","article-title":"People-tracking-by-detection and people-detection-by-tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Andriluka","year":"2008"},{"key":"2026032615364252300_ref011","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5540156","article-title":"Monocular 3D pose estimation and tracking by detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Andriluka","year":"2010"},{"key":"2026032615364252300_ref012","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995311","article-title":"Multi-target tracking by continuous energy minimization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Andriyenko","year":"2011"},{"key":"2026032615364252300_ref013","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2012.6247893","article-title":"Discrete-continuous optimization for multi-target tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Andriyenko","year":"2012"},{"issue":"6","key":"2026032615364252300_ref014","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1109\/MC.2010.170","article-title":"Google street view: Capturing the world at street level","volume":"43","author":"Anguelov","year":"2010","journal-title":"IEEE Computer"},{"key":"2026032615364252300_ref015","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.572","article-title":"NetVLAD: CNN architecture for weakly supervised place recognition","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Arandjelovi\u0107","year":"2016"},{"key":"2026032615364252300_ref016","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.49","article-title":"Multiscale combinatorial grouping","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Arbel\u00e1ez","year":"2014"},{"key":"2026032615364252300_ref017","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.100","article-title":"Pixelwise instance segmentation with a dynamically instantiated network","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Arnab","year":"2017"},{"key":"2026032615364252300_ref018","unstructured":"Babaee, M., A.Athar, and G.Rigoll (2018). \u201cMultiple people tracking using hierarchical deep tracklet re-identification\u201d. arXiv: 1811.04091[cs.CV]."},{"key":"2026032615364252300_ref019","article-title":"Free space computation using stochastic occupancy grids and dynamic programming","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV) Workshops","author":"Badino","year":"2007"},{"key":"2026032615364252300_ref020","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-03798-6_6","article-title":"The stixel world - A compact medium level representation of the 3D-world","volume-title":"Proc. of the DAGM Symposium on Pattern Recognition (DAGM)","author":"Badino","year":"2009"},{"key":"2026032615364252300_ref021","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2012.6224716","article-title":"Real-time topometric localization","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Badino","year":"2012"},{"issue":"1","key":"2026032615364252300_ref022","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1007\/s11263-013-0673-5","article-title":"Mixture of trees probabilistic graphical model for video segmentation","volume":"110","author":"Badrinarayanan","year":"2014","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref023","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5540054","article-title":"Label propagation in video sequences","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Badrinarayanan","year":"2010"},{"issue":"12","key":"2026032615364252300_ref024","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref025","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46466-4_10","article-title":"Exploiting semantic information and deep matching for optical flow","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Bai","year":"2016"},{"key":"2026032615364252300_ref026","first-page":"2858","article-title":"Deep Watershed transform for instance segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Bai","year":"2017"},{"key":"2026032615364252300_ref027","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.457","article-title":"Flow fields: Dense correspondence fields for highly accurate large displacement optical flow estimation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Bailer","year":"2015"},{"key":"2026032615364252300_ref028","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11263-010-0390-2","article-title":"A database and evaluation methodology for optical flow","volume":"92","author":"Baker","year":"2011","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref029","unstructured":"Balntas, V.\n           (2019). \u201cSILDa: A multi-task dataset for evaluating visual localization\u201d. https:\/\/medium.com\/scape-technologies\/silda-a-multi-task-dataset-for-evaluating-visual-localization-7fc6c2c56c74. Online: accessed 17-June-2019."},{"key":"2026032615364252300_ref030","unstructured":"Balntas, V., L.Hammarstrand, H.Heijnen, F.Kahl, W.Maddern, K.Mikolajczyk, T.Pajdla, M.Pollefeys, T.Sattler, J. L.Sch\u00f6nberger, P.Speciale, J.Sivic, C.Toft, and A.Torii (2019). \u201cWorkshop on long-term visual localization under changing conditions\u201d. https:\/\/sites.google.com\/view\/ltvl2019\/. Online: accessed 17-June-2019."},{"key":"2026032615364252300_ref031","first-page":"782","article-title":"RelocNet: Continuous metric learning relocalisation using neural nets","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Balntas","year":"2018"},{"key":"2026032615364252300_ref032","doi-asserted-by":"crossref","unstructured":"Bansal, M., A.Krizhevsky, and A.Ogaler (2018). \u201cChauffeurNet: Learning to drive by imitating the best and synthesizing the worst\u201d. arXiv: 1812.03079.","DOI":"10.15607\/RSS.2019.XV.031"},{"key":"2026032615364252300_ref033","doi-asserted-by":"crossref","DOI":"10.1145\/2072298.2071954","article-title":"Geo-localization of street views with aerial image databases","volume-title":"Proc. of the International Conf. on Multimedia (ICM)","author":"Bansal","year":"2011"},{"key":"2026032615364252300_ref034","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.167","article-title":"Dense object reconstruction with semantic priors","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Bao","year":"2013"},{"issue":"2","key":"2026032615364252300_ref035","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1109\/TITS.2008.922935","article-title":"Real-time speed sign detection using the radial symmetry detector","volume":"9","author":"Barnes","year":"2008","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref036","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.281","article-title":"Bounding boxes, segmentations and object coordinates: How important is recognition for 3D scene flow estimation in autonomous driving scenarios?","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Behl","year":"2017"},{"key":"2026032615364252300_ref037","article-title":"Point-FlowNet: Learning representations for rigid motion estimation from point clouds","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Behl","year":"2019"},{"key":"2026032615364252300_ref038","doi-asserted-by":"crossref","unstructured":"Behley, J., M.Garbade, A.Milioto, J.Quenzel, S.Behnke, C.Stachniss, and J.Gall (2019). \u201cSemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences\u201d. arXiv: 1904.01416[cs.CV].","DOI":"10.1109\/ICCV.2019.00939"},{"key":"2026032615364252300_ref039","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2012.6225003","article-title":"Performance of histogram descriptors for the classification of 3D laser range data in urban environments","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Behley","year":"2012"},{"key":"2026032615364252300_ref040","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2013.6696957","article-title":"Laser-based segment classification using a mixture of bag-of-words","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Behley","year":"2013"},{"key":"2026032615364252300_ref041","first-page":"3517","article-title":"BirdNet: A 3D object detection framework from LiDAR information","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Beltr\u00e1n","year":"2018"},{"key":"2026032615364252300_ref042","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2012.6248017","article-title":"Pedestrian detection at 100 frames per second","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Benenson","year":"2012"},{"key":"2026032615364252300_ref043","article-title":"Ten years of pedestrian detection, what have we learned?","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Benenson","year":"2014"},{"key":"2026032615364252300_ref044","doi-asserted-by":"crossref","DOI":"10.1109\/PETS-WINTER.2009.5399488","article-title":"Multiple object tracking using flow linear programming","volume-title":"Performance Evaluation of Tracking and Surveillance","author":"Berclaz","year":"2009"},{"issue":"9","key":"2026032615364252300_ref045","doi-asserted-by":"crossref","first-page":"1806","DOI":"10.1109\/TPAMI.2011.21","article-title":"Multiple object tracking using K-shortest paths optimization","volume":"33","author":"Berclaz","year":"2011","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref046","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01258-8_21","article-title":"Object detection in video with spatiotemporal sampling networks","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Bertasius","year":"2018"},{"key":"2026032615364252300_ref047","first-page":"175","article-title":"VIAC: An out of ordinary experiment","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Bertozzi","year":"2011"},{"issue":"1","key":"2026032615364252300_ref048","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0921-8890(99)00125-6","article-title":"Vision-based intelligent vehicles: State of the art and perspectives","volume":"32","author":"Bertozzi","year":"2000","journal-title":"Robotics and Autonomous Systems (RAS)"},{"key":"2026032615364252300_ref049","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1109\/34.121791","article-title":"A method for registration of 3D shapes","volume":"14","author":"Besl","year":"1992","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref050","doi-asserted-by":"crossref","DOI":"10.1109\/ICIP.2016.7533003","article-title":"Simple online and realtime tracking","volume-title":"Proc. IEEE International Conf. on Image Processing (ICIP)","author":"Bewley","year":"2016"},{"key":"2026032615364252300_ref051","doi-asserted-by":"crossref","unstructured":"Bewley, A., J.Rigley, Y.Liu, J.Hawke, R.Shen, V. D.Lam, and A.Kendall (2018). \u201cLearning to drive from simulation without real world labels\u201d. arXiv: 1812.03823[cs.CV].","DOI":"10.1109\/ICRA.2019.8793668"},{"key":"2026032615364252300_ref052","doi-asserted-by":"crossref","DOI":"10.1007\/3-540-47977-5_8","article-title":"A probabilistic theory of occupancy and emptiness","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Bhotika","year":"2002"},{"key":"2026032615364252300_ref053","article-title":"DDD17: End-To-end DAVIS driving dataset","volume-title":"Proc. of the International Conf. on Machine Learning (ICML) Workshops","author":"Binas","year":"2017"},{"key":"2026032615364252300_ref054","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.1993.378214","article-title":"A framework for the robust estimation of optical flow","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Black","year":"1993"},{"key":"2026032615364252300_ref055","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.346","article-title":"Large-scale semantic 3D reconstruction: An adaptive multi-resolution model for multi-class volumetric labeling","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Blaha","year":"2016"},{"issue":"2","key":"2026032615364252300_ref056","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1177\/0278364913507326","article-title":"The M\u00e1laga urban dataset: High-rate stereo and LiDAR in a realistic urban scenario","volume":"33","author":"Blanco-Claraco","year":"2014","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref057","first-page":"1","article-title":"High-speed tracking-by-detection without using image information","volume-title":"Proc. of International Conf. on Advanced Video and Signal Based Surveillance (AVSS)","author":"Bochinski","year":"2017"},{"key":"2026032615364252300_ref058","doi-asserted-by":"crossref","DOI":"10.1109\/ICPR.2016.7900128","article-title":"Efficient volumetric fusion of airborne and street-side data for urban reconstruction","volume-title":"Proc. of the International Conf. on Pattern Recognition (ICPR)","author":"B\u00f3dis-Szomor\u00fa","year":"2016"},{"key":"2026032615364252300_ref059","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46454-1_34","article-title":"Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Bogo","year":"2016"},{"key":"2026032615364252300_ref060","unstructured":"Bojarski, M., D. D.Testa, D.Dworakowski, B.Firner, B.Flepp, P.Goyal, L. D.Jackel, M.Monfort, U.Muller, J.Zhang, X.Zhang, J.Zhao, and K.Zieba (2016). \u201cEnd to end learning for self-driving cars\u201d. arXiv: 1604.07316[cs.CV]."},{"issue":"1","key":"2026032615364252300_ref061","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1109\/TITS.2011.2173196","article-title":"A novel lane detection system with efficient ground truth generation","volume":"13","author":"Borkar","year":"2012","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref062","unstructured":"Bouguet, J.-Y.\n           (2010). \u201cCamera calibration toolbox for Matlab\u201d. url: http:\/\/www.vision.caltech.edu\/bouguetj\/calib_doc."},{"issue":"13","key":"2026032615364252300_ref063","doi-asserted-by":"crossref","first-page":"16199","DOI":"10.1007\/s11042-017-5195-7","article-title":"Less restrictive camera odometry estimation from monocular camera","volume":"77","author":"Boukhers","year":"2018","journal-title":"Multimedia Tools Appl"},{"key":"2026032615364252300_ref064","article-title":"Fast approximate energy minimization via graph cuts","volume":"23","author":"Boykov","year":"1999","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref065","article-title":"Semantic instance segmentation with a discriminative loss function","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Brabandere","year":"2017"},{"key":"2026032615364252300_ref066","first-page":"2616","article-title":"Geometry-aware learning of maps for camera localization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Brahmbhatt","year":"2018"},{"issue":"9","key":"2026032615364252300_ref067","doi-asserted-by":"crossref","first-page":"693","DOI":"10.1002\/rob.20140","article-title":"The TerraMax autonomous vehicle","volume":"23","author":"Braid","year":"2006","journal-title":"Journal of Field Robotics (JFR)"},{"key":"2026032615364252300_ref068","unstructured":"Braun, M., S.Krebs, F.Flohr, and D. M.Gavrila (2019). \u201cThe EuroCity persons dataset: A novel benchmark for object detection\u201d. arXiv: 1805.07193[cs.CV]."},{"key":"2026032615364252300_ref069","first-page":"1546","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Braun","year":"2016"},{"key":"2026032615364252300_ref070","unstructured":"Braunschweig, T. U.\n           (2010). \u201cProject Stadtpilot\u201d. https:\/\/www.tu-braunschweig.de\/stadtpilot. Online: accessed 18-October-2019."},{"issue":"3","key":"2026032615364252300_ref071","doi-asserted-by":"crossref","first-page":"492","DOI":"10.1137\/090769521","article-title":"Total generalized variation","volume":"3","author":"Bredies","year":"2010","journal-title":"Journal of Imaging Sciences (SIAM)"},{"key":"2026032615364252300_ref072","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2009.5459278","article-title":"Robust tracking-by-detection using a detector confidence particle filter","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Breitenstein","year":"2009"},{"issue":"9","key":"2026032615364252300_ref073","doi-asserted-by":"crossref","first-page":"1820","DOI":"10.1109\/TPAMI.2010.232","article-title":"Online multiperson tracking-by-detection from a single, uncalibrated camera","volume":"33","author":"Breitenstein","year":"2011","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref074","article-title":"Multi-object tracking as maximum weight independent set","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Brendel","year":"2011"},{"key":"2026032615364252300_ref075","doi-asserted-by":"crossref","DOI":"10.1142\/3986","volume-title":"Automatic Vehicle Guidance: The Experience of the Argo Vehicle","author":"Broggi","year":"1999"},{"key":"2026032615364252300_ref076","article-title":"Shapebased pedestrian detection","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Broggi","year":"2000"},{"issue":"6","key":"2026032615364252300_ref077","doi-asserted-by":"crossref","first-page":"3508","DOI":"10.1109\/TITS.2015.2477556","article-title":"PROUD \u2013 Public Road Urban Driverless-Car Test","volume":"16","author":"Broggi","year":"2015","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref078","first-page":"981","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Broggi","year":"2007"},{"key":"2026032615364252300_ref079","first-page":"635","article-title":"Development of the control system for the Vislab Intercontinental Autonomous Challenge","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Broggi","year":"2010"},{"key":"2026032615364252300_ref080","first-page":"619","article-title":"Model-based three dimensional interpretations of two dimensional images","volume-title":"Proc. of the International Joint Conf. on Artificial Intelligence (IJCAI)","author":"Brooks","year":"1981"},{"key":"2026032615364252300_ref081","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1109\/TPAMI.2010.143","article-title":"Large displacement optical flow: Descriptor matching in variational motion estimation","volume":"33","author":"Brox","year":"2011","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"4","key":"2026032615364252300_ref082","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2015.2453975","article-title":"Map-based probabilistic visual self-localization","volume":"38","author":"Brubaker","year":"2016","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref083","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1007\/1-4020-3858-8_15","volume-title":"Geometric Properties for Incomplete Data","author":"Bruhn","year":"2006"},{"key":"2026032615364252300_ref084","doi-asserted-by":"crossref","DOI":"10.1109\/ITSC.2016.7795703","article-title":"Flow-decoupled normalized reprojection error for visual odometry","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Buczko","year":"2016"},{"key":"2026032615364252300_ref085","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2016.7535429","article-title":"How to distinguish inliers from outliers in visual odometry for high-speed automotive applications","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Buczko","year":"2016"},{"key":"2026032615364252300_ref086","first-page":"739","article-title":"Monocular outlier detection for visual odometry","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Buczko","year":"2017"},{"key":"2026032615364252300_ref087","first-page":"1","article-title":"Selfvalidation for automotive visual odometry","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Buczko","year":"2018"},{"key":"2026032615364252300_ref088","doi-asserted-by":"crossref","DOI":"10.5244\/C.24.27","article-title":"Label propagation in complex video sequences using semi-supervised learning","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Budvytis","year":"2010"},{"key":"2026032615364252300_ref089","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-73429-1","volume-title":"The 2005 Darpa Grand Challenge: The Great Robot Race","author":"Buehler","year":"2007"},{"key":"2026032615364252300_ref090","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-03991-1","article-title":"The DARPA urban challenge","volume-title":"DARPA Challenge","author":"Buehler","year":"2009"},{"issue":"2","key":"2026032615364252300_ref091","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1109\/TRA.2003.808850","article-title":"Robust scene reconstruction from an omnidirectional vision system","volume":"19","author":"Bunschoten","year":"2003","journal-title":"IEEE Trans. on Robotics and Automation (TRA)"},{"key":"2026032615364252300_ref092","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33783-3_44","article-title":"A naturalistic open source movie for optical flow evaluation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Butler","year":"2012"},{"key":"2026032615364252300_ref093","doi-asserted-by":"crossref","unstructured":"Caesar, H., V.Bankiti, A. H.Lang, S.Vora, V. E.Liong, Q.Xu, A.Krishnan, Y.Pan, G.Baldan, and O.Beijbom (2019). \u201cnuScenes: A multimodal dataset for autonomous driving\u201d. arXiv: 1903.11027.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"2026032615364252300_ref094","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46493-0_22","article-title":"A unified multi-scale deep convolutional neural network for fast object detection","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Cai","year":"2016"},{"key":"2026032615364252300_ref095","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00784","article-title":"Hybrid scene compression for visual localization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Camposeco","year":"2019"},{"issue":"9","key":"2026032615364252300_ref096","doi-asserted-by":"crossref","first-page":"1023","DOI":"10.1177\/0278364915614638","article-title":"University of Michigan North Campus long-term vision and lidar dataset","volume":"35","author":"Carlevaris-Bianco","year":"2016","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref097","first-page":"430","article-title":"Semantic segmentation with second-order pooling","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Carreira","year":"2012"},{"issue":"7","key":"2026032615364252300_ref098","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1109\/TPAMI.2011.231","article-title":"CPMC: Automatic object segmentation using constrained parametric min-cuts","volume":"34","author":"Carreira","year":"2012","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref099","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995442","article-title":"Scene flow estimation by growing correspondence seeds","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Cech","year":"2011"},{"key":"2026032615364252300_ref100","unstructured":"Center, G. H.\n           (2017). \u201cSelf-driving cars, in 1956?\u201d https:\/\/www.gmheritagecenter.com\/featured\/Autonomous_Vehicles.html. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref101","unstructured":"Cernea, D.\n           (2015). \u201cOpenMVS: open multiple view stereovision\u201d. http:\/\/cdcseacave.github.io\/openMVS. Online: accessed 23-April-2019."},{"key":"2026032615364252300_ref102","first-page":"2040","article-title":"Deep MANTA: A coarse-to-fine many-task network for joint 2D and 3D vehicle analysis from monocular image","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chabot","year":"2017"},{"key":"2026032615364252300_ref103","first-page":"5410","article-title":"Pyramid stereo matching network","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chang","year":"2018"},{"key":"2026032615364252300_ref104","first-page":"8748","article-title":"Argoverse: 3D tracking and forecasting with rich maps","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chang","year":"2019"},{"key":"2026032615364252300_ref105","first-page":"2722","article-title":"Deep-Driving: Learning affordance for direct perception in autonomous driving","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Chen","year":"2015"},{"key":"2026032615364252300_ref106","article-title":"Learning by cheating","volume-title":"Proc. Conf. on Robot Learning (CoRL)","author":"Chen","year":"2019"},{"key":"2026032615364252300_ref107","first-page":"8713","article-title":"Searching for efficient multi-scale architectures for dense image prediction","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Chen","year":"2018"},{"key":"2026032615364252300_ref108","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.409","article-title":"Beat the MTurkers: Automatic image labeling from weak 3D supervision","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chen","year":"2014"},{"key":"2026032615364252300_ref109","first-page":"4013","article-title":"MaskLab: Instance segmentation by refining object detection with semantic and direction features","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chen","year":"2018"},{"key":"2026032615364252300_ref110","article-title":"Semantic image segmentation with deep convolutional nets and fully connected CRFs","volume-title":"Proc. of the International Conf. on Learning Representations (ICLR)","author":"Chen","year":"2015"},{"issue":"4","key":"2026032615364252300_ref111","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref112","unstructured":"Chen, L., G.Papandreou, F.Schroff, and H.Adam (2017a). \u201cRethinking atrous convolution for semantic image segmentation\u201d. arXiv: 1706.05587[cs.CV]."},{"key":"2026032615364252300_ref113","first-page":"833","article-title":"Enco der-decoder with atrous separable convolution for semantic image segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chen","year":"2018"},{"key":"2026032615364252300_ref114","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.509","article-title":"Full flow: Optical flow estimation by global optimization over regular grids","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chen","year":"2016"},{"key":"2026032615364252300_ref115","article-title":"3D Object proposals for accurate object class detection","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Chen","year":"2015"},{"issue":"5","key":"2026032615364252300_ref116","doi-asserted-by":"crossref","first-page":"1259","DOI":"10.1109\/TPAMI.2017.2706685","article-title":"3D object proposals using stereo imagery for accurate object class detection","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref117","article-title":"Multiview 3D object detection network for autonomous driving","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chen","year":"2017"},{"key":"2026032615364252300_ref118","doi-asserted-by":"crossref","DOI":"10.1109\/3DV.2016.68","article-title":"Multi-label semantic 3D reconstruction using voxel blocks","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Cherabier","year":"2016"},{"key":"2026032615364252300_ref119","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01258-8_20","article-title":"Learning priors for semantic 3D reconstruction","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Cherabier","year":"2018"},{"issue":"7","key":"2026032615364252300_ref120","doi-asserted-by":"crossref","first-page":"1577","DOI":"10.1109\/TPAMI.2012.248","article-title":"A general framework for tracking multiple people from a moving camera","volume":"35","author":"Choi","year":"2013","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref121","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.347","article-title":"Near-online multi-target tracking with aggregated local flow descriptor","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Choi","year":"2015"},{"key":"2026032615364252300_ref122","article-title":"A unified framework for multitarget tracking and collective activity recognition","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Choi","year":"2012"},{"key":"2026032615364252300_ref123","first-page":"1800","article-title":"Xception: Deep learning with depthwise separable convolutions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Chollet","year":"2017"},{"key":"2026032615364252300_ref124","first-page":"161","article-title":"Online multiobject tracking with instance-aware tracker and dynamic model refreshment","volume-title":"Proc. of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Chu","year":"2019"},{"key":"2026032615364252300_ref125","doi-asserted-by":"crossref","unstructured":"Chu, P. and H.Ling (2019). \u201cFAMNet: Joint learning of feature, affinity and multi-dimensional assignment for online multiple object tracking\u201d. arXiv: 1904.04989[cs.CV].","DOI":"10.1109\/ICCV.2019.00627"},{"key":"2026032615364252300_ref126","first-page":"4846","article-title":"Online multi-object tracking using cnn-based single object tracker with spatial-temporal attention mechanism","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Chu","year":"2017"},{"key":"2026032615364252300_ref127","first-page":"1918","article-title":"A committee of neural networks for traffic sign classification","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Ciresan","year":"2011"},{"key":"2026032615364252300_ref128","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1016\/j.neunet.2012.02.023","article-title":"Multi-column deep neural network for traffic sign classification","volume":"32","author":"Ciresan","year":"2012","journal-title":"Neural Networks"},{"key":"2026032615364252300_ref129","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01267-0_15","article-title":"On offline evaluation of vision-based driving models","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Codevilla","year":"2018"},{"key":"2026032615364252300_ref130","first-page":"1","article-title":"End-to-end driving via conditional imitation learning","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Codevilla","year":"2018"},{"key":"2026032615364252300_ref131","doi-asserted-by":"crossref","unstructured":"Codevilla, F., E.Santana, A. M.L\u00f3pez, and A.Gaidon (2019). \u201cExploring the limitations of behavior cloning for autonomous driving\u201d. arXiv: 1904.08980[cs.CV].","DOI":"10.1109\/ICCV.2019.00942"},{"key":"2026032615364252300_ref132","first-page":"358","article-title":"A space-sweep approach to true multiimage matching","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Collins","year":"1996"},{"issue":"1","key":"2026032615364252300_ref133","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1006\/cviu.1995.1004","article-title":"Active shape models-their training and application","volume":"61","author":"Cootes","year":"1995","journal-title":"Computer Vision and Image Understanding (CVIU)"},{"key":"2026032615364252300_ref134","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.350","article-title":"The cityscapes dataset for semantic urban scene understanding","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Cordts","year":"2016"},{"key":"2026032615364252300_ref135","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-11752-2_14","article-title":"Object-level priors for stixel generation","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Cordts","year":"2014"},{"issue":"2-3","key":"2026032615364252300_ref136","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1007\/s11263-007-0081-9","article-title":"3D urban scene modeling integrating recognition and reconstruction","volume":"78","author":"Cornelis","year":"2008","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref137","unstructured":"Corporation, M. M.\n           (1998). \u201cMitsubishi motors develops \u2018new driver support system\u2019\u201d. https:\/\/www.mitsubishi-motors.com\/en\/corporate\/pressrelease\/corporate\/detail429.html. Online: accessed 17-May-2019."},{"key":"2026032615364252300_ref138","first-page":"3469","article-title":"Fusion scheme for semantic and instance-level segmentation","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Costea","year":"2018"},{"key":"2026032615364252300_ref139","first-page":"993","article-title":"Fast boosting based detection using scale invariant multimodal multiresolution filtered features","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Costea","year":"2017"},{"key":"2026032615364252300_ref140","first-page":"3001","article-title":"Discrete-continuous optimization for large-scale structure from motion","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Crandall","year":"2011"},{"issue":"6","key":"2026032615364252300_ref141","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1177\/0278364908090961","article-title":"FAB-MAP: Probabilistic localization and mapping in the space of appearance","volume":"27","author":"Cummins","year":"2008","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref142","first-page":"303","article-title":"A volumetric method for building complex models from range images","volume-title":"ACM Trans. on Graphics","author":"Curless","year":"1996"},{"key":"2026032615364252300_ref143","article-title":"Soft-slam: Computationally efficient stereo visual slam for autonomous uavs","volume-title":"Journal of Field Robotics (JFR)","author":"Cvi\u0161ic","year":"2017"},{"issue":"4","key":"2026032615364252300_ref144","first-page":"578","article-title":"Stereo odometry based on careful feature selection and tracking","volume":"35","author":"Cvisic","year":"2015","journal-title":"Proc. European Conf. on Mobile Robotics (ECMR)"},{"key":"2026032615364252300_ref145","first-page":"458","article-title":"3DMV: Joint 3D-multi-view prediction for 3D semantic scene segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Dai","year":"2018"},{"key":"2026032615364252300_ref146","first-page":"534","article-title":"Instancesensitive fully convolutional networks","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Dai","year":"2016"},{"key":"2026032615364252300_ref147","first-page":"3992","article-title":"Convolutional feature masking for joint object and stuff segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dai","year":"2015"},{"key":"2026032615364252300_ref148","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.343","article-title":"Instance-aware semantic segmentation via multi-task network cascades","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dai","year":"2016"},{"key":"2026032615364252300_ref149","article-title":"R-FCN: Object detection via region-based fully convolutional networks","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Dai","year":"2016"},{"key":"2026032615364252300_ref150","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.89","article-title":"Deformable convolutional networks","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Dai","year":"2017"},{"key":"2026032615364252300_ref151","unstructured":"Daimler\n              AG\n            \n           (2019). \u201cBosch and Daimler. Metropolis in California to become a pilot city for automated driving\u201d. https:\/\/www.daimler.com\/innovation\/case\/autonomous\/pilot-city-for-automated-driving.html. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref152","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2005.177","article-title":"Histograms of oriented gradients for human detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dalal","year":"2005"},{"key":"2026032615364252300_ref153","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.170","article-title":"Dense reconstruction using 3D object shape priors","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dame","year":"2013"},{"key":"2026032615364252300_ref154","unstructured":"DARPA\n           (2014). \u201cThe DARPA grand challenge: Ten years later\u201d. https:\/\/www.darpa.mil\/news-events\/2014-03-13. Online: accessed 18-June-2019."},{"key":"2026032615364252300_ref155","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7299036","article-title":"GMMCP Tracker: Globally optimal generalized maximum multi clique problem for multiple object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dehghan","year":"2015"},{"key":"2026032615364252300_ref156","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298718","article-title":"Target identity-aware network flow for online multiple target tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dehghan","year":"2015"},{"key":"2026032615364252300_ref157","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-45886-1_14","article-title":"Stereo visual odometry without temporal filtering","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Deigmoeller","year":"2016"},{"key":"2026032615364252300_ref158","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.193","article-title":"Photometric bundle adjustment for dense multi-view 3D modeling","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Delaunoy","year":"2014"},{"issue":"2","key":"2026032615364252300_ref159","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1007\/s11263-010-0408-9","article-title":"Gradient flows for optimizing triangular mesh-based surfaces: Applications to 3D reconstruction problems dealing with visibility","volume":"95","author":"Delaunoy","year":"2011","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref160","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.1999.772544","article-title":"Monte carlo localization for mobile robots","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Dellaert","year":"1999"},{"issue":"12","key":"2026032615364252300_ref161","doi-asserted-by":"crossref","first-page":"1181","DOI":"10.1177\/0278364906072768","article-title":"Square root SAM: Simultaneous localization and mapping via square root information smoothing","volume":"25","author":"Dellaert","year":"2006","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref162","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206848","article-title":"Imagenet: A large-scale hierarchical image database","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Deng","year":"2009"},{"key":"2026032615364252300_ref163","unstructured":"Department of Motor Vehicles CA\n           (2019). \u201cReport of traffic collision involving an autonomous vehicle\u201d. https:\/\/www.dmv.ca.gov\/portal\/dmv\/detail\/vr\/autonomous\/autonomousveh_ol316. Online: accessed 17-May-2019."},{"issue":"1","key":"2026032615364252300_ref164","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11263-011-0439-x","article-title":"Discriminative models for multi-class object layout","volume":"95","author":"Desai","year":"2011","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref165","first-page":"2480","article-title":"IMLS-SLAM: Scan-to-model matching based on 3D data","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Deschaud","year":"2018"},{"key":"2026032615364252300_ref166","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2016.7487649","article-title":"Motion-based detection and tracking in 3D LiDAR scans","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Dewan","year":"2016"},{"key":"2026032615364252300_ref167","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2016.7759282","article-title":"Rigid scene flow for 3D LiDAR scans","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Dewan","year":"2016"},{"key":"2026032615364252300_ref168","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.1994.639472","article-title":"The seeing passenger car \u2018VaMoRs-P\u2019","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Dickmanns","year":"1994"},{"issue":"2","key":"2026032615364252300_ref169","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1109\/34.121789","article-title":"Recursive 3-D road and relative ego-state recognition","volume":"14","author":"Dickmanns","year":"1992","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref170","unstructured":"Dickmanns, E. D.\n           (1995). \u201cDynamic machine vision\u201d. http:\/\/dyna-vision.de\/. Online: accessed 18-June-2019."},{"issue":"4","key":"2026032615364252300_ref171","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1007\/BF01212361","article-title":"Dynamic monocular machine vision","volume":"1","author":"Dickmanns","year":"1988","journal-title":"Machine Vision and Applications (MVA)"},{"issue":"6","key":"2026032615364252300_ref172","doi-asserted-by":"crossref","first-page":"1273","DOI":"10.1109\/21.61200","article-title":"An integrated spatio-temporal approach to automatic visual guidance of autonomous vehicles","volume":"20","author":"Dickmanns","year":"1990","journal-title":"IEEE Trans. on Systems, Man and Cybernetics (TSMC)"},{"issue":"4","key":"2026032615364252300_ref173","doi-asserted-by":"crossref","first-page":"743","DOI":"10.1109\/TPAMI.2011.155","article-title":"Pedestrian detection: An evaluation of the state of the art","volume":"34","author":"Dollar","year":"2011","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"8","key":"2026032615364252300_ref174","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.1109\/TPAMI.2014.2300479","article-title":"Fast feature pyramids for object detection","volume":"36","author":"Doll\u00e1r","year":"2014","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref175","doi-asserted-by":"crossref","DOI":"10.5244\/C.23.91","article-title":"Integral channel features","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Doll\u00e1r","year":"2009"},{"key":"2026032615364252300_ref176","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206631","article-title":"Pedestrian detection: A benchmark","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Dollar","year":"2009"},{"issue":"4","key":"2026032615364252300_ref177","doi-asserted-by":"crossref","first-page":"743","DOI":"10.1109\/TPAMI.2011.155","article-title":"Pedestrian detection: An evaluation of the state of the art","volume":"34","author":"Doll\u00e1r","year":"2012","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref178","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.316","article-title":"FlowNet: Learning optical flow with convolutional networks","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Dosovitskiy","year":"2015"},{"key":"2026032615364252300_ref179","unstructured":"Dosovitskiy, A., G.Ros, F.Codevilla, A.Lopez, and V.Koltun (2017a). \u201cCARLA\u201d. http:\/\/carla.org\/. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref180","article-title":"CARLA: An open urban driving simulator","volume-title":"Proc. Conf. on Robot Learning (CoRL)","author":"Dosovitskiy","year":"2017"},{"key":"2026032615364252300_ref181","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-11752-2_4","article-title":"Semi-global matching: A principled derivation in terms of message passing","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Drory","year":"2014"},{"key":"2026032615364252300_ref182","first-page":"3194","article-title":"A general pipeline for 3D detection of vehicles","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Du","year":"2018"},{"key":"2026032615364252300_ref183","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46454-1_6","article-title":"Towards large-scale city reconstruction from satellites","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Duan","year":"2016"},{"key":"2026032615364252300_ref184","first-page":"5266","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Dub\u00e9","year":"2017"},{"key":"2026032615364252300_ref185","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2013.183","article-title":"Semi-dense visual odometry for a monocular camera","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Engel","year":"2013"},{"key":"2026032615364252300_ref186","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1109\/TPAMI.2017.2658577","article-title":"Direct sparse odometry","volume":"40","author":"Engel","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref187","article-title":"LSD-SLAM: Largescale direct monocular SLAM","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Engel","year":"2014"},{"key":"2026032615364252300_ref188","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2015.7353631","article-title":"Large-scale direct SLAM with stereo cameras","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Engel","year":"2015"},{"key":"2026032615364252300_ref189","first-page":"1355","article-title":"Vote3Deep: Fast object detection in 3D point clouds using efficient convolutional neural networks","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Engelcke","year":"2017"},{"key":"2026032615364252300_ref190","doi-asserted-by":"crossref","first-page":"2179","DOI":"10.1109\/TPAMI.2008.260","article-title":"Monocular pedestrian detection: Survey and experiments","volume":"31","author":"Enzweiler","year":"2009","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref191","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587592","article-title":"A mixed generative-discriminative framework for pedestrian classification","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Enzweiler","year":"2008"},{"issue":"10","key":"2026032615364252300_ref192","doi-asserted-by":"crossref","first-page":"2967","DOI":"10.1109\/TIP.2011.2142006","article-title":"A multilevel mixture-of-experts framework for pedestrian classification","volume":"20","author":"Enzweiler","year":"2011","journal-title":"IEEE Trans. on Image Processing (TIP)"},{"key":"2026032615364252300_ref193","doi-asserted-by":"crossref","DOI":"10.5244\/C.26.71","article-title":"Stixmentation \u2013 Probabilistic stixel based traffic scene labeling","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Erbs","year":"2012"},{"key":"2026032615364252300_ref194","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2013.6629530","article-title":"From stixels to objects \u2013 A conditional random field based approach","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Erbs","year":"2013"},{"key":"2026032615364252300_ref195","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587581","article-title":"A mobile vision system for robust multi-person tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ess","year":"2008"},{"key":"2026032615364252300_ref196","doi-asserted-by":"crossref","first-page":"1831","DOI":"10.1109\/TPAMI.2009.109","article-title":"Robust multi-person tracking from a mobile platform","volume":"31","author":"Ess","year":"2009","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref197","doi-asserted-by":"crossref","DOI":"10.5244\/C.23.84","article-title":"Segmentation-based urban traffic scene understanding","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Ess","year":"2009"},{"issue":"2","key":"2026032615364252300_ref198","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (VOC) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref199","first-page":"933","article-title":"Keypoint trajectory estimation using propagation based tracking","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Fanani","year":"2016"},{"key":"2026032615364252300_ref200","first-page":"1714","article-title":"Multimodal scale estimation for monocular visual odometry","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Fanani","year":"2017"},{"key":"2026032615364252300_ref201","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.imavis.2017.08.002","article-title":"Predictive monocular odometry (PMO): What is possible without RANSAC and multiframe bundle adjustment?","volume":"68","author":"Fanani","year":"2017","journal-title":"Image and Vision Computing (IVC)"},{"key":"2026032615364252300_ref202","doi-asserted-by":"crossref","DOI":"10.1007\/3-540-45103-X_50","article-title":"Two-frame motion estimation based on polynomial expansion","volume-title":"Scandinavian Conference on Image Analysis (SCIA)","author":"Farneback","year":"2003"},{"issue":"3","key":"2026032615364252300_ref203","doi-asserted-by":"crossref","first-page":"336","DOI":"10.1109\/83.661183","article-title":"Variational principles, surface evolution, PDEs, level set methods, and the stereo problem","volume":"7","author":"Faugeras","year":"1998","journal-title":"IEEE Trans. on Image Processing (TIP)"},{"key":"2026032615364252300_ref204","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.330","article-title":"Detect to track and track to detect","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Feichtenhofer","year":"2017"},{"key":"2026032615364252300_ref205","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587597","article-title":"A discriminatively trained, multiscale, deformable part model","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Felzenszwalb","year":"2008"},{"issue":"1","key":"2026032615364252300_ref206","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1007\/s11263-006-7899-4","article-title":"Efficient belief propagation for early vision","volume":"70","author":"Felzenszwalb","year":"2006","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref207","first-page":"1","article-title":"PETS2009: Dataset and challenge","volume-title":"Performance Evaluation of Tracking and Surveillance","author":"Ferryman","year":"2009"},{"key":"2026032615364252300_ref208","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2012.6248007","article-title":"Joint 2D-3D temporally consistent semantic segmentation of street scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Floros","year":"2012"},{"key":"2026032615364252300_ref209","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15561-1_27","article-title":"Building Rome on a cloudless day","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Frahm","year":"2010"},{"key":"2026032615364252300_ref210","doi-asserted-by":"crossref","first-page":"538","DOI":"10.1016\/j.isprsjprs.2010.08.009","article-title":"Fast robust large-scale mapping from video and internet photo collections","volume":"65","author":"Frahm","year":"2010","journal-title":"ISPRS Journal of Photogrammetry and Remote Sensing (JPRS)"},{"key":"2026032615364252300_ref211","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.1994.639486","article-title":"The Daimler-Benz steering assistant: A spin-off from autonomous driving","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Franke","year":"1994"},{"issue":"6","key":"2026032615364252300_ref212","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/5254.736001","article-title":"Autonomous driving goes downtown","volume":"13","author":"Franke","year":"1998","journal-title":"Intelligent Systems (IS)"},{"key":"2026032615364252300_ref213","doi-asserted-by":"crossref","DOI":"10.1007\/11550518_27","article-title":"6D-Vision: Fusion of stereo and motion for robust environment perception","volume-title":"Proc. of the DAGM Symposium on Pattern Recognition (DAGM)","author":"Franke","year":"2005"},{"key":"2026032615364252300_ref214","article-title":"Visual odometry: Part II \u2013 Matching, robustness, and applications","volume-title":"Robotics and Automation Magazine (RAM)","author":"Fraundorfer","year":"2011"},{"key":"2026032615364252300_ref215","doi-asserted-by":"crossref","DOI":"10.1109\/ITSC.2013.6728473","article-title":"A new performance measure and evaluation benchmark for road detection algorithms","volume-title":"Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC)","author":"Fritsch","year":"2013"},{"key":"2026032615364252300_ref216","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2018.8462884","article-title":"End-to-end learning of multi-sensor 3D tracking by detection","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Frossard","year":"2018"},{"key":"2026032615364252300_ref217","article-title":"Objectaware bundle adjustment for correcting monocular scale drift","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Frost","year":"2016"},{"issue":"2","key":"2026032615364252300_ref218","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1023\/B:VISI.0000043756.03810.dd","article-title":"Data processing algorithms for generating textured 3D building facade meshes from laser scans and camera images","volume":"61","author":"Fr\u00fch","year":"2005","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref219","first-page":"11","article-title":"MVE \u2013 A multi-view reconstruction environment","volume-title":"Eurographics Workshop on Graphics and Cultural Heritage (GCH)","author":"Fuhrmann","year":"2014"},{"key":"2026032615364252300_ref220","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2013.6629566","article-title":"Toward automated driving in cities using close-to-market sensors: An overview of the V-Charge Project","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Furgale","year":"2013"},{"issue":"8","key":"2026032615364252300_ref221","doi-asserted-by":"crossref","first-page":"1362","DOI":"10.1109\/TPAMI.2009.161","article-title":"Accurate, dense, and robust multi-view stereopsis","volume":"32","author":"Furukawa","year":"2010","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref222","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46448-0_36","article-title":"Superpixel convolutional networks using bilateral inceptions","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Gadde","year":"2016"},{"issue":"5","key":"2026032615364252300_ref223","doi-asserted-by":"crossref","first-page":"1273","DOI":"10.1109\/TPAMI.2017.2696526","article-title":"Efficient 2D and 3D facade segmentation using auto-context","volume":"40","author":"Gadde","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref224","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.459","article-title":"PatchBatch: A batch augmented loss for optical flow","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gadot","year":"2016"},{"key":"2026032615364252300_ref225","unstructured":"Gaidon, A., Q.Wang, Y.Cabon, and E.Vig (2016a). \u201cVirtual KITTI\u201d. https:\/\/europe.naverlabs.com\/research\/computer-vision\/proxy-virtual-worlds. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref226","article-title":"Virtual worlds as proxy for multi-object tracking analysis","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gaidon","year":"2016"},{"key":"2026032615364252300_ref227","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00407","article-title":"A unifying contrast maximization framework for event cameras, with applications to motion, depth, and optical flow estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gallego","year":"2018"},{"key":"2026032615364252300_ref228","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587799","article-title":"Object categorization using co-occurrence, location and appearance","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Galleguillos","year":"2008"},{"key":"2026032615364252300_ref229","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.106","article-title":"Massively parallel multiview stereopsis by surface normal diffusion","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Galliani","year":"2015"},{"key":"2026032615364252300_ref230","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587671","article-title":"Variable baseline\/resolution stereo","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gallup","year":"2008"},{"key":"2026032615364252300_ref231","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2007.383245","article-title":"Real-time plane-sweeping stereo with multiple sweeping directions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gallup","year":"2007"},{"key":"2026032615364252300_ref232","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5539804","article-title":"Piecewise planar and non-planar stereo for urban scene reconstruction","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gallup","year":"2010"},{"issue":"5","key":"2026032615364252300_ref233","doi-asserted-by":"crossref","first-page":"1188","DOI":"10.1109\/TRO.2012.2197158","article-title":"Bags of binary words for fast place recognition in image sequences","volume":"28","author":"G\u00e1lvez-L\u00f3pez","year":"2012","journal-title":"IEEE Trans. on Robotics"},{"key":"2026032615364252300_ref234","doi-asserted-by":"crossref","first-page":"332","DOI":"10.1016\/j.neucom.2018.08.009","article-title":"Evaluation of deep neural networks for traffic sign detection systems","volume":"316","author":"Garc\u00eda","year":"2018","journal-title":"Neurocomputing"},{"key":"2026032615364252300_ref235","doi-asserted-by":"crossref","DOI":"10.1109\/IJCNN.2016.7727386","article-title":"PointNet: A 3D convolutional neural network for real-time object class recognition","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Garcia-Garcia","year":"2016"},{"key":"2026032615364252300_ref236","first-page":"3369","article-title":"Lightweight probabilistic deep networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gast","year":"2018"},{"key":"2026032615364252300_ref237","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1007\/s11263-006-9038-7","article-title":"Multi-cue pedestrian detection and tracking from a moving vehicle","volume":"73","author":"Gavrila","year":"2007","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref238","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01258-8_46","article-title":"Asynchronous, photometric feature tracking using events and frames","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Gehrig","year":"2018"},{"key":"2026032615364252300_ref239","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-04667-4_14","article-title":"A real-time low-power stereo vision engine using semi-global matching","volume-title":"Proc. of the International Conf. on Computer Vision Systems (ICVS)","author":"Gehrig","year":"2009"},{"key":"2026032615364252300_ref240","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2009.5164267","article-title":"Monocular road mosaicing for urban environments","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Geiger","year":"2009"},{"issue":"3","key":"2026032615364252300_ref241","doi-asserted-by":"crossref","first-page":"1008","DOI":"10.1109\/TITS.2012.2189882","article-title":"Team AnnieWAY\u2019s entry to the grand cooperative driving challenge 2011","volume":"13","author":"Geiger","year":"2012","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"issue":"5","key":"2026032615364252300_ref242","doi-asserted-by":"crossref","first-page":"1012","DOI":"10.1109\/TPAMI.2013.185","article-title":"3D traffic scene understanding from movable platforms","volume":"36","author":"Geiger","year":"2014","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"11","key":"2026032615364252300_ref243","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref244","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2012.6248074","article-title":"Are we ready for autonomous driving? The KITTI vision benchmark suite","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Geiger","year":"2012"},{"key":"2026032615364252300_ref245","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2012.6224570","article-title":"Automatic calibration of range and camera sensors using a single shot","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Geiger","year":"2012"},{"key":"2026032615364252300_ref246","article-title":"Efficient largescale stereo matching","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Geiger","year":"2010"},{"key":"2026032615364252300_ref247","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2011.5940405","article-title":"StereoScan: Dense 3D reconstruction in real-time","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Geiger","year":"2011"},{"issue":"7","key":"2026032615364252300_ref248","doi-asserted-by":"crossref","first-page":"1239","DOI":"10.1109\/TPAMI.2009.122","article-title":"Survey on pedestrian detection for advanced driver assistance systems","volume":"32","author":"Geronimo","year":"2010","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref249","doi-asserted-by":"crossref","DOI":"10.1007\/3-540-45053-X_29","article-title":"A unifying theory for central panoramic systems and practical implications","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Geyer","year":"2000"},{"key":"2026032615364252300_ref250","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46487-9_32","article-title":"Laplacian pyramid reconstruction and refinement for semantic segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Ghiasi","year":"2016"},{"key":"2026032615364252300_ref251","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-24673-2_20","article-title":"A Bayesian framework for multi-cue 3D object tracking","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Giebel","year":"2004"},{"key":"2026032615364252300_ref252","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.169","article-title":"Fast R-CNN","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Girshick","year":"2015"},{"key":"2026032615364252300_ref253","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.81","article-title":"Rich feature hierarchies for accurate object detection and semantic segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Girshick","year":"2014"},{"key":"2026032615364252300_ref254","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.458","article-title":"Using k-poselets for detecting people and localizing their keypoints","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Gkioxari","year":"2014"},{"issue":"11","key":"2026032615364252300_ref255","doi-asserted-by":"crossref","first-page":"3980","DOI":"10.1109\/TCYB.2016.2593940","article-title":"On-board object detection: Multicue, multimodal, and multiview random forest of local experts","volume":"47","author":"Gonz\u00e1lez","year":"2016","journal-title":"IEEE Trans. on Cybernetics"},{"key":"2026032615364252300_ref256","first-page":"1210","article-title":"Fast dense panoramic stereovision","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Gonz\u00e1lez-Barbosa","year":"2005"},{"key":"2026032615364252300_ref257","first-page":"7872","article-title":"LIMO: Lidarmonocular visual odometry","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Gr\u00e4ter","year":"2018"},{"key":"2026032615364252300_ref258","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2015.7139484","article-title":"Integrating metric and semantic maps for vision-only automated parking","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Grimmett","year":"2015"},{"key":"2026032615364252300_ref259","article-title":"The BRAiVE platform","volume-title":"Proc. of the IFAC Symposium on Intelligent Autonomous Vehicles (IFAC)","author":"Grisleri","year":"2010"},{"issue":"4","key":"2026032615364252300_ref260","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1109\/MITS.2018.2867526","article-title":"Fast joint object detection and viewpoint estimation for traffic scene understanding","volume":"10","author":"Guindel","year":"2018","journal-title":"IEEE Intelligent Transportation Systems Magazine (ITSM)"},{"key":"2026032615364252300_ref261","unstructured":"Guivant, J. and E.Nebot (2006). \u201cVictoria park dataset\u201d. http:\/\/www-personal.acfr.usyd.edu.au\/nebot\/victoria_park.htm. Online: accessed 8-April-2019."},{"key":"2026032615364252300_ref262","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7299044","article-title":"Displets: Resolving stereo ambiguities using object knowledge","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"G\u00fcney","year":"2015"},{"key":"2026032615364252300_ref263","article-title":"Deep discrete flow","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"G\u00fcney","year":"2016"},{"key":"2026032615364252300_ref264","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33783-3_38","article-title":"Stixels motion estimation without optical flow computation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"G\u00fcnyel","year":"2012"},{"key":"2026032615364252300_ref265","first-page":"212","volume-title":"Proc. of the SPIE Conf. Integrating Photogrammetric Techniques with Scene Analysis and Machine Vision IIII","author":"Haala","year":"1997"},{"key":"2026032615364252300_ref266","doi-asserted-by":"crossref","DOI":"10.5194\/isprs-annals-IV-1-W1-91-2017","article-title":"Semantic3D.net: A new large-scale point cloud classification benchmark","volume-title":"ISPRS Annals of Photogrammetry, Remote Sensing and Spatial Information Sciences (APRS)","author":"Hackel","year":"2017"},{"key":"2026032615364252300_ref267","first-page":"177","article-title":"Fast semantic segmentation of 3D point clouds with strongly varying density","volume-title":"ISPRS Annals of Photogrammetry, Remote Sensing and Spatial Information Sciences (APRS)","author":"Hackel","year":"2016"},{"key":"2026032615364252300_ref268","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.89","article-title":"Class specific 3D object shape priors using surface normals","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Haene","year":"2014"},{"key":"2026032615364252300_ref269","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.20","article-title":"Joint 3D scene reconstruction and class segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Haene","year":"2013"},{"key":"2026032615364252300_ref270","doi-asserted-by":"crossref","DOI":"10.1109\/3DIMPVT.2012.55","article-title":"A patch prior for dense 3D reconstruction in man-made environments","volume-title":"Proc. of the International Conf. on 3D Digital Imaging, Modeling, Data Processing, Visualization and Transmission (THREEDIM-PVT)","author":"Haene","year":"2012"},{"key":"2026032615364252300_ref271","article-title":"MatchNet: Unifying feature and metric learning for patch-based matching","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Han","year":"2015"},{"key":"2026032615364252300_ref272","article-title":"Real-time direct dense matching on fisheye images using planesweeping stereo","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"H\u00e4ne","year":"2014"},{"key":"2026032615364252300_ref273","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2015.7354095","article-title":"Obstacle detection for self-driving cars using only monocular cameras and wheel odometry","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"H\u00e4ne","year":"2015"},{"key":"2026032615364252300_ref274","volume-title":"Computer Vision Systems","author":"Hanson","year":"1978"},{"key":"2026032615364252300_ref275","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10584-0_20","article-title":"Simultaneous detection and segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Hariharan","year":"2014"},{"key":"2026032615364252300_ref276","first-page":"447","article-title":"Hypercolumns for object segmentation and fine-grained localization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hariharan","year":"2015"},{"key":"2026032615364252300_ref277","first-page":"991","article-title":"Semantic contours from inverse detectors","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Hariharan","year":"2011"},{"key":"2026032615364252300_ref278","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.1995.466816","article-title":"In defence of the 8-point algorithm","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Hartley","year":"1995"},{"key":"2026032615364252300_ref279","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511811685","volume-title":"Multiple View Geometry in Computer Vision","author":"Hartley","year":"2004"},{"key":"2026032615364252300_ref280","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.176","article-title":"Learned multi-patch similarity","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Hartmann","year":"2017"},{"key":"2026032615364252300_ref281","first-page":"587","article-title":"Boundary-aware instance segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hayder","year":"2017"},{"key":"2026032615364252300_ref282","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587784","article-title":"im2gps: Estimating geographic information from a single image","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hays","year":"2008"},{"key":"2026032615364252300_ref283","article-title":"FuseNet: Incorporating depth into semantic segmentation via fusion-based CNN architecture","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Hazirbas","year":"2016"},{"key":"2026032615364252300_ref284","first-page":"2980","article-title":"Mask R-CNN","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"He","year":"2017"},{"key":"2026032615364252300_ref285","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10578-9_23","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"He","year":"2014"},{"key":"2026032615364252300_ref286","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.90","article-title":"Deep residual learning for image recognition","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"He","year":"2016"},{"key":"2026032615364252300_ref287","first-page":"630","article-title":"Identity mappings in deep residual networks","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"He","year":"2016"},{"key":"2026032615364252300_ref288","article-title":"Multiscale conditional random fields for image labeling","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"He","year":"2004"},{"key":"2026032615364252300_ref289","doi-asserted-by":"crossref","DOI":"10.1007\/11744023_27","article-title":"Learning and incorporating top-down cues in image segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"He","year":"2006"},{"issue":"4","key":"2026032615364252300_ref290","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1007\/BF00133568","article-title":"Optical flow using spatiotemporal filters","volume":"1","author":"Heeger","year":"1988","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref291","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.1997.609468","article-title":"A four-step camera calibration procedure with implicit image correction","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Heikkila","year":"1997"},{"key":"2026032615364252300_ref292","unstructured":"Henaff, M., A.Canziani, and Y.LeCun (2019). \u201cModel-predictive policy learning with uncertainty regularization for driving in dense traffic\u201d. arXiv: 1901.02705[cs.LG]."},{"issue":"5","key":"2026032615364252300_ref293","doi-asserted-by":"crossref","first-page":"775","DOI":"10.1002\/rob.21540","article-title":"Leveraging image-based localization for infrastructure-based calibration of a multi-camera rig","volume":"32","author":"Heng","year":"2015","journal-title":"Journal of Field Robotics (JFR)"},{"key":"2026032615364252300_ref294","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2013.6696592","article-title":"CamOdoCal: Automatic intrinsic and extrinsic calibration of a rig with multiple generic cameras and odometry","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Heng","year":"2013"},{"key":"2026032615364252300_ref295","first-page":"1428","article-title":"Fusion of head and full-body detectors for multi-object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Henschel","year":"2018"},{"key":"2026032615364252300_ref296","article-title":"Fusion of head and full-body detectors for multi-object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Henschel","year":"2019"},{"key":"2026032615364252300_ref297","first-page":"1","article-title":"Evaluation of cost functions for stereo matching","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hirschm\u00fcller","year":"2007"},{"issue":"2","key":"2026032615364252300_ref298","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1109\/TPAMI.2007.1166","article-title":"Stereo processing by semiglobal matching and mutual information","volume":"30","author":"Hirschm\u00fcller","year":"2008","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"8","key":"2026032615364252300_ref299","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"key":"2026032615364252300_ref300","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/s11263-008-0137-5","article-title":"Putting objects in perspective","volume":"80","author":"Hoiem","year":"2008","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"1","key":"2026032615364252300_ref301","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1007\/s11263-006-0031-y","article-title":"Recovering surface layout from an image","volume":"75","author":"Hoiem","year":"2007","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"1\u20133","key":"2026032615364252300_ref302","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1016\/0004-3702(81)90024-2","article-title":"Determining optical flow","volume":"17","author":"Horn","year":"1981","journal-title":"Artificial Intelligence (AI)"},{"key":"2026032615364252300_ref303","first-page":"1","article-title":"Detection of traffic signs in real-world images: The German traffic sign detection benchmark","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Houben","year":"2013"},{"key":"2026032615364252300_ref304","first-page":"2297","article-title":"Efficient 3-D scene analysis from streaming data","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Hu","year":"2013"},{"key":"2026032615364252300_ref305","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2019.00549","article-title":"Joint monocular 3D vehicle detection and tracking","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Hu","year":"2019"},{"key":"2026032615364252300_ref306","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.153","article-title":"Deep 360 pilot: Learning a deep agent for piloting through 360\u00b0 sports video","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hu","year":"2017"},{"issue":"3","key":"2026032615364252300_ref307","doi-asserted-by":"crossref","first-page":"1010","DOI":"10.1109\/TITS.2018.2838132","article-title":"SINet: A scale-insensitive convolutional neural network for fast vehicle detection","volume":"20","author":"Hu","year":"2018","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref308","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-88688-4_58","article-title":"Robust object tracking by hierarchical association of detection responses","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Huang","year":"2008"},{"key":"2026032615364252300_ref309","first-page":"2261","article-title":"Densely connected convolutional networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Huang","year":"2017"},{"key":"2026032615364252300_ref310","unstructured":"Huang, G., Z.Liu, L.van der Maaten, and K. Q.Weinberger (2017b). \u201cDensely connected convolutional networks\u201d. https:\/\/www.youtube.com\/watch?v=-W6y8xnd--U. Online: accessed 18-February-2020."},{"key":"2026032615364252300_ref311","doi-asserted-by":"crossref","DOI":"10.1109\/ICPR.2016.7900038","article-title":"Point cloud labeling using 3D convolutional neural network","volume-title":"Proc. of the International Conf. on Pattern Recognition (ICPR)","author":"Huang","year":"2016"},{"key":"2026032615364252300_ref312","first-page":"2821","article-title":"DeepMVS: Learning multi-view stereopsis","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Huang","year":"2018"},{"key":"2026032615364252300_ref313","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2018.00141","article-title":"The apolloscape dataset for autonomous driving","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Huang","year":"2018"},{"key":"2026032615364252300_ref314","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2007.4409000","article-title":"A variational method for scene flow estimation from stereo sequences","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Huguet","year":"2007"},{"key":"2026032615364252300_ref315","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00936","article-title":"LiteFlowNet: A lightweight convolutional neural network for optical flow estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hui","year":"2018"},{"key":"2026032615364252300_ref316","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.42","article-title":"MirrorFlow: Exploiting symmetries in joint optical flow and occlusion estimation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Hur","year":"2017"},{"key":"2026032615364252300_ref317","first-page":"677","article-title":"Uncertainty estimates and multi-hypotheses networks for optical flow","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV), Part VII","author":"Ilg","year":"2018"},{"key":"2026032615364252300_ref318","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.179","article-title":"FlowNet 2.0: Evolution of optical flow estimation with deep networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ilg","year":"2017"},{"key":"2026032615364252300_ref319","article-title":"Batch normalization: Accelerating deep network training by reducing internal covariate shift","volume-title":"Proc. of the International Conf. on Machine Learning (ICML)","author":"Ioffe","year":"2015"},{"key":"2026032615364252300_ref320","first-page":"2599","article-title":"From structure-from-motion point clouds to fast location recognition","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Irschara","year":"2009"},{"key":"2026032615364252300_ref321","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.482","article-title":"Learning sparse high dimensional filters: Image filtering, dense CRFs and bilateral neural networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Jampani","year":"2016"},{"key":"2026032615364252300_ref322","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01270-0_42","article-title":"Unsupervised learning of multi-frame optical flow with occlusions","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Janai","year":"2018"},{"key":"2026032615364252300_ref323","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.154","article-title":"Slow flow: Exploiting high-speed cameras for accurate and diverse optical flow reference data","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Janai","year":"2017"},{"key":"2026032615364252300_ref324","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995693","article-title":"Multi-view reconstruction preserving weakly-supported surfaces","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Jancosek","year":"2011"},{"key":"2026032615364252300_ref325","first-page":"1175","article-title":"The one hundred layers Tiramisu: Fully convolutional DenseNets for semantic segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"J\u00e9gou","year":"2017"},{"key":"2026032615364252300_ref326","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.59","article-title":"Large scale multi-view stereopsis evaluation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Jensen","year":"2014"},{"key":"2026032615364252300_ref327","first-page":"650","article-title":"CPFG-SLAM: A robust simultaneous localization and mapping based on LIDAR in off-road environment","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Ji","year":"2018"},{"key":"2026032615364252300_ref328","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.253","article-title":"SurfaceNet: An end-to-end 3D neural network for multiview stereopsis","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Ji","year":"2017"},{"key":"2026032615364252300_ref329","doi-asserted-by":"crossref","DOI":"10.1145\/2647868.2654889","article-title":"Caffe: Convolutional architecture for fast feature embedding","volume-title":"Proc. of the International Conf. on Multimedia (ICM)","author":"Jia","year":"2014"},{"key":"2026032615364252300_ref330","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2007.383180","article-title":"A linear programming approach for multiple object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Jiang","year":"2007"},{"issue":"5","key":"2026032615364252300_ref331","doi-asserted-by":"crossref","first-page":"1991","DOI":"10.1109\/TITS.2014.2308281","article-title":"Traffic sign recognition with hinge loss trained convolutional neural networks","volume":"15","author":"Jin","year":"2014","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref332","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.2001.933280","article-title":"A counter example to the theory of simultaneous localization and map building","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Julier","year":"2001"},{"key":"2026032615364252300_ref333","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1016\/j.patrec.2017.02.018","article-title":"Deep network aided by guiding network for pedestrian detection","volume":"90","author":"Jung","year":"2017","journal-title":"Pattern Recognition Letters"},{"issue":"2","key":"2026032615364252300_ref334","first-page":"217","article-title":"iSAM2: Incremental smoothing and mapping using the Bayes tree","volume":"31","author":"Kaess","year":"2012","journal-title":"International Journal of Robotics Research (IJRR)"},{"issue":"6","key":"2026032615364252300_ref335","doi-asserted-by":"crossref","first-page":"1365","DOI":"10.1109\/TRO.2008.2006706","article-title":"iSAM: Incremental smoothing and mapping","volume":"24","author":"Kaess","year":"2008","journal-title":"IEEE Trans. on Robotics"},{"key":"2026032615364252300_ref336","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2009.5204180","article-title":"Alignment of 3D point clouds to overhead images","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Kaminsky","year":"2009"},{"key":"2026032615364252300_ref337","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.101","article-title":"Object detection in videos with tubelet proposal networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kang","year":"2017"},{"key":"2026032615364252300_ref338","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.95","article-title":"Object detection from video tubelets with convolutional neural networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kang","year":"2016"},{"issue":"2","key":"2026032615364252300_ref339","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TIV.2018.2886678","article-title":"Test your self-driving algorithm: An overview of publicly available driving datasets and virtual testing environments","volume":"4","author":"Kang","year":"2019","journal-title":"Proc. IEEE Transactions on Intelligent Vehicles (T-IV)"},{"key":"2026032615364252300_ref340","article-title":"Learning a multi-view stereo machine","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Kar","year":"2017"},{"key":"2026032615364252300_ref341","article-title":"Reliable automatic cameralaser calibration","volume-title":"Proc. IEEE Australasian Conf. on Robotics and Automation (ACRA)","author":"Kassir","year":"2010"},{"issue":"4","key":"2026032615364252300_ref342","doi-asserted-by":"crossref","first-page":"1096","DOI":"10.1109\/TITS.2011.2143410","article-title":"The benefits of dense stereo for pedestrian detection","volume":"12","author":"Keller","year":"2011","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref343","first-page":"7482","article-title":"Multi-task learning using uncertainty to weigh losses for scene geometry and semantics","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kendall","year":"2018"},{"key":"2026032615364252300_ref344","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.336","article-title":"PoseNet: A convolutional network for real-time 6-DOF camera relocalization","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Kendall","year":"2015"},{"key":"2026032615364252300_ref345","unstructured":"Kendall, A., J.Hawke, D.Janz, P.Mazur, D.Reda, J. M.Allen, V. D.Lam, A.Bewley, and A.Shah (2018b). \u201cLearning to drive in a day\u201d. arXiv: 1807.00412[cs.LG]."},{"key":"2026032615364252300_ref346","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.17","article-title":"End-to-end learning of geometry and context for deep stereo regression","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Kendall","year":"2017"},{"key":"2026032615364252300_ref347","first-page":"3748","article-title":"Robust odometry estimation for RGB-D cameras","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Kerl","year":"2013"},{"key":"2026032615364252300_ref348","unstructured":"Kesten, R., M.Usman, J.Houston, T.Pandya, K.Nadhamuni, A.Ferreira, M.Yuan, B.Low, A.Jain, P.Ondruska, S.Omari, S.Shah, A.Kulkarni, A.Kazakova, C.Tao, L.Platinsky, W.Jiang, and V.Shet (2019). \u201cLyft level 5 AV dataset 2019\u201d. https:\/\/level5.lyft.com\/dataset\/. Online: accessed 25-March-2020."},{"issue":"1","key":"2026032615364252300_ref349","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1109\/TPAMI.2018.2876253","article-title":"Motion segmentation & multiple object tracking by correlation co-clustering","volume":"42","author":"Keuper","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref350","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.533","article-title":"Multiple hypothesis tracking revisited","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Kim","year":"2015"},{"key":"2026032615364252300_ref351","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01237-3_13","article-title":"Multi-object tracking with neural gating using bilinear LSTM","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Kim","year":"2018"},{"key":"2026032615364252300_ref352","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00656","article-title":"Panoptic feature pyramid networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kirillov","year":"2019"},{"key":"2026032615364252300_ref353","first-page":"9404","article-title":"Panoptic segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kirillov","year":"2019"},{"key":"2026032615364252300_ref354","first-page":"7322","article-title":"InstanceCut: From edges to instances with MultiCut","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kirillov","year":"2017"},{"key":"2026032615364252300_ref355","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2010.5548123","article-title":"Visual odometry based on stereo image sequences with RANSAC-based outlier rejection scheme","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Kitt","year":"2010"},{"key":"2026032615364252300_ref356","first-page":"225","article-title":"Parallel tracking and mapping for small AR workspaces","volume-title":"Proc. of the International Symposium on Mixed and Augmented Reality (ISMAR)","author":"Klein","year":"2007"},{"key":"2026032615364252300_ref357","volume-title":"Tech. rep","author":"Klette","year":"2015"},{"key":"2026032615364252300_ref358","first-page":"953","article-title":"Street view motion-from-structure-from-motion","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Klingner","year":"2013"},{"issue":"4","key":"2026032615364252300_ref359","doi-asserted-by":"crossref","DOI":"10.1145\/3072959.3073599","article-title":"Tanks and temples: Benchmarking large-scale scene reconstruction","volume":"36","author":"Knapitsch","year":"2017","journal-title":"ACM Trans. on Graphics"},{"issue":"3","key":"2026032615364252300_ref360","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1007\/s11263-008-0202-0","article-title":"Robust higher order potentials for enforcing label consistency","volume":"82","author":"Kohli","year":"2009","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"10","key":"2026032615364252300_ref361","doi-asserted-by":"crossref","first-page":"1568","DOI":"10.1109\/TPAMI.2006.200","article-title":"Convergent tree-reweighted message passing for energy minimization","volume":"28","author":"Kolmogorov","year":"2006","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref362","first-page":"1008","article-title":"Actor-critic algorithms","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Konda","year":"1999"},{"key":"2026032615364252300_ref363","first-page":"132","article-title":"An adaptive confidence measure for optical flows based on linear subspace projections","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Kondermann","year":"2007"},{"key":"2026032615364252300_ref364","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-88690-7_22","article-title":"A statistical confidence measure for optical flows","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Kondermann","year":"2008"},{"key":"2026032615364252300_ref365","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2016.10","article-title":"The HCI benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Kondermann","year":"2016"},{"key":"2026032615364252300_ref366","article-title":"Efficient inference in fully connected CRFs with Gaussian edge potentials","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Kr\u00e4henb\u00fchl","year":"2011"},{"key":"2026032615364252300_ref367","doi-asserted-by":"crossref","DOI":"10.5220\/0005316103470356","article-title":"Improving the egomotion estimation by correcting the calibration bias","volume-title":"Proc. of the Conf. on Computer Vision Theory and Applications (VISAPP)","author":"Kre\u0161o","year":"2015"},{"key":"2026032615364252300_ref368","article-title":"ImageNet classification with deep convolutional neural networks","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Krizhevsky","year":"2012"},{"key":"2026032615364252300_ref369","first-page":"471","article-title":"Fast optical flow using dense inverse search","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Kroeger","year":"2016"},{"key":"2026032615364252300_ref370","first-page":"1","article-title":"Joint 3D proposal generation and object detection from view aggregation","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Ku","year":"2018"},{"key":"2026032615364252300_ref371","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.01214","article-title":"Monocular 3D object detection leveraging accurate proposals and shape reconstruction","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ku","year":"2019"},{"key":"2026032615364252300_ref372","doi-asserted-by":"crossref","DOI":"10.1109\/ITSC.2012.6338740","article-title":"Spatial ray features for real-time ego-lane extraction","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Kuehnl","year":"2012"},{"key":"2026032615364252300_ref373","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5539869","article-title":"What\u2019s going on?: Discovering spatio-temporal dependencies in dynamic scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kuettel","year":"2010"},{"issue":"1","key":"2026032615364252300_ref374","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1007\/s11263-016-0946-x","article-title":"A TV prior for high-quality scalable multi-view stereo reconstruction","volume":"124","author":"Kuhn","year":"2017","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref375","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2005.9","article-title":"A hierarchical field framework for unified context-based classification","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Kumar","year":"2005"},{"key":"2026032615364252300_ref376","first-page":"3607","article-title":"G2o: A general framework for graph optimization","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"K\u00fcmmerle","year":"2011"},{"key":"2026032615364252300_ref377","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10599-4_45","article-title":"Joint semantic segmentation and 3D reconstruction from monocular video","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Kundu","year":"2014"},{"key":"2026032615364252300_ref378","first-page":"3559","article-title":"3D-RCNN: Instancelevel 3D object reconstruction via render-and-compare","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kundu","year":"2018"},{"key":"2026032615364252300_ref379","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.345","article-title":"Feature space optimization for semantic video segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Kundu","year":"2016"},{"key":"2026032615364252300_ref380","article-title":"Fast and accurate largescale stereo reconstruction using variational methods","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV) Workshops","author":"Kuschk","year":"2013"},{"issue":"3","key":"2026032615364252300_ref381","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1023\/A:1008191222954","article-title":"A theory of shape by space carving","volume":"38","author":"Kutulakos","year":"2000","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"10","key":"2026032615364252300_ref382","doi-asserted-by":"crossref","first-page":"1449","DOI":"10.1016\/j.cviu.2011.06.008","article-title":"Bootstrap optical flow confidence and uncertainty measure","volume":"115","author":"Kybic","year":"2011","journal-title":"Computer Vision and Image Understanding (CVIU)"},{"key":"2026032615364252300_ref383","first-page":"1","article-title":"Efficient multiview reconstruction of large-scale scenes using interest points, delaunay triangulation and graph cuts","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Labatut","year":"2007"},{"key":"2026032615364252300_ref384","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2016.7535374","article-title":"Map-supervised road detection","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Laddha","year":"2016"},{"key":"2026032615364252300_ref385","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2009.5459248","article-title":"Associative hierarchical CRFs for object class image segmentation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Ladicky","year":"2009"},{"key":"2026032615364252300_ref386","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15555-0_18","article-title":"Graph cut based inference with co-occurrence statistics","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Ladicky","year":"2010"},{"issue":"6","key":"2026032615364252300_ref387","doi-asserted-by":"crossref","first-page":"1056","DOI":"10.1109\/TPAMI.2013.165","article-title":"Associative hierarchical random fields","volume":"36","author":"Ladicky","year":"2014","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"1","key":"2026032615364252300_ref388","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1109\/TPAMI.2008.281","article-title":"Structural approach for building reconstruction from a single DSM","volume":"32","author":"Lafarge","year":"2010","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"1","key":"2026032615364252300_ref389","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1109\/TPAMI.2012.84","article-title":"A hybrid multiview stereo algorithm for modeling urban scenes","volume":"35","author":"Lafarge","year":"2013","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"1","key":"2026032615364252300_ref390","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1007\/s11263-012-0517-8","article-title":"Creating large-scale city models from 3D-point clouds: A robust approach with hybrid representation","volume":"99","author":"Lafarge","year":"2012","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref391","first-page":"143","article-title":"DART: Noise injection for robust imitation learning","volume-title":"Proc. Conf. on Robot Learning (CoRL)","author":"Laskey","year":"2017"},{"key":"2026032615364252300_ref392","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2011.5979711","article-title":"Visual SLAM for autonomous ground vehicles","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Lategahn","year":"2011"},{"issue":"3","key":"2026032615364252300_ref393","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1109\/MITS.2011.942107","article-title":"Grand cooperative driving challenge 2011","volume":"3","author":"Lauer","year":"2011","journal-title":"Proc. IEEE Intelligent Transportation Systems Magazine (ITSM)"},{"key":"2026032615364252300_ref394","first-page":"765","article-title":"CornerNet: Detecting objects as paired keypoints","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Law","year":"2018"},{"key":"2026032615364252300_ref395","article-title":"Long-term timesensitive costs for CRF-based tracking by detection","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV) Workshops","author":"Le","year":"2016"},{"key":"2026032615364252300_ref396","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2016.59","article-title":"Learning by tracking: Siamese CNN for robust target Association","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Leal-Taix\u00e9","year":"2016"},{"key":"2026032615364252300_ref397","unstructured":"Leal-Taix\u00e9, L., A.Milan, I. D.Reid, S.Roth, and K.Schindler (2015). \u201cMOTChallenge 2015: Towards a benchmark for multitarget tracking\u201d. arXiv: 1504.01942[cs.CV]."},{"key":"2026032615364252300_ref398","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.354","article-title":"Motion estimation for self-driving cars with a generalized camera","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Lee","year":"2013"},{"key":"2026032615364252300_ref399","article-title":"Structureless pose-graph loop-closure with a multi-camera system on a self-driving car","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Lee","year":"2013"},{"key":"2026032615364252300_ref400","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.76","article-title":"Relative pose estimation for a multi-camera system with known vertical direction","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Lee","year":"2014"},{"key":"2026032615364252300_ref401","first-page":"1965","article-title":"VPGNet: Vanishing point guided network for lane and road marking detection and recognition","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Lee","year":"2017"},{"key":"2026032615364252300_ref402","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2007.383146","article-title":"Dynamic 3D scene analysis from a moving vehicle","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Leibe","year":"2007"},{"issue":"1-3","key":"2026032615364252300_ref403","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1007\/s11263-007-0095-3","article-title":"Robust object detection with interleaved categorization and segmentation","volume":"77","author":"Leibe","year":"2008","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"10","key":"2026032615364252300_ref404","doi-asserted-by":"crossref","first-page":"1683","DOI":"10.1109\/TPAMI.2008.170","article-title":"Coupled detection and tracking from static cameras and moving vehicles","volume":"30","author":"Leibe","year":"2008","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"6","key":"2026032615364252300_ref405","doi-asserted-by":"crossref","first-page":"585","DOI":"10.1177\/0278364918767756","article-title":"Exactly sparse delayed state filter on Lie groups for long-term pose graph SLAM","volume":"37","author":"Lenac","year":"2018","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref406","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.496","article-title":"FollowMe: Efficient online min-cost flow tracking with bounded memory and computation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Lenz","year":"2015"},{"key":"2026032615364252300_ref407","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2013.IX.037","article-title":"Keyframe-based visual-inertial SLAM using nonlinear optimization","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Leutenegger","year":"2013"},{"key":"2026032615364252300_ref408","doi-asserted-by":"crossref","DOI":"10.5244\/C.29.109","article-title":"StixelNet: A deep convolutional network for obstacle detection and road segmentation","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Levi","year":"2015"},{"key":"2026032615364252300_ref409","first-page":"1904","article-title":"Joint graph decomposition & node labeling: Problem, algorithms, Applications","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Levinkov","year":"2017"},{"key":"2026032615364252300_ref410","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2007.III.016","article-title":"Map-based precision vehicle localization in urban environments","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Levinson","year":"2007"},{"key":"2026032615364252300_ref411","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.2010.5509700","article-title":"Robust vehicle localization in urban environments using probabilistic maps","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Levinson","year":"2010"},{"key":"2026032615364252300_ref412","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2016.XII.042","article-title":"Vehicle detection from 3D lidar using fully convolutional network","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Li","year":"2016"},{"key":"2026032615364252300_ref413","doi-asserted-by":"crossref","unstructured":"Li, G., M.Mueller, V.Casser, N.Smith, D. L.Michels, and B.Ghanem (2018a). \u201cTeaching UAVs to race with observational imitation learning\u201d. arXiv: 1803.01129[cs.CV].","DOI":"10.15607\/RSS.2019.XV.005"},{"issue":"3","key":"2026032615364252300_ref414","doi-asserted-by":"crossref","first-page":"690","DOI":"10.1109\/TNNLS.2016.2522428","article-title":"Deep neural network for structural prediction and lane detection in traffic scene","volume":"28","author":"Li","year":"2017","journal-title":"IEEE Trans. on Neural Networks and Learning Systems"},{"key":"2026032615364252300_ref415","first-page":"106","article-title":"Weakly- and semisupervised panoptic segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Li","year":"2018"},{"key":"2026032615364252300_ref416","unstructured":"Li, X., H.Zhao, L.Han, Y.Tong, and K.Yang (2019). \u201cGFF: Gated fully fusion for semantic segmentation\u201d. arXiv: 1904.01803[cs.CV]."},{"key":"2026032615364252300_ref417","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15552-9_57","article-title":"Location recognition using prioritized feature matching","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Li","year":"2010"},{"key":"2026032615364252300_ref418","first-page":"4438","article-title":"Fully convolutional instance-aware semantic segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Li","year":"2017"},{"key":"2026032615364252300_ref419","article-title":"Mean-field networks","volume-title":"Proc. of the International Conf. on Machine Learning (ICML) Workshops","author":"Li","year":"2014"},{"key":"2026032615364252300_ref420","article-title":"Landmark classification in large-scale image collections","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Li","year":"2009"},{"key":"2026032615364252300_ref421","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33718-5_2","article-title":"Worldwide pose estimation using 3D point clouds","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Li","year":"2012"},{"key":"2026032615364252300_ref422","doi-asserted-by":"crossref","unstructured":"Liang, X., T.Wang, L.Yang, and E.Xing (2018). \u201cCIRL: Controllable imitative reinforcement learning for vision-based selfdriving\u201d. arXiv: 1807.03776[cs.CV].","DOI":"10.1007\/978-3-030-01234-2_36"},{"key":"2026032615364252300_ref423","first-page":"891","article-title":"Cross-view image geolocalization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Lin","year":"2013"},{"key":"2026032615364252300_ref424","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7299135","article-title":"Learning deep representations for ground-to-aerial geolocalization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Lin","year":"2015"},{"key":"2026032615364252300_ref425","first-page":"2999","article-title":"Focal loss for dense object detection","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Lin","year":"2017"},{"key":"2026032615364252300_ref426","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common objects in context","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Lin","year":"2014"},{"key":"2026032615364252300_ref427","first-page":"2391","article-title":"Efficient Global 2D-3D matching for camera localization in a large-scale 3D map","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Liu","year":"2017"},{"issue":"3","key":"2026032615364252300_ref428","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1016\/j.conb.2010.03.007","article-title":"Neuromorphic sensory systems","volume":"20","author":"Liu","year":"2010","journal-title":"Current Opinion in Neurobiology"},{"key":"2026032615364252300_ref429","first-page":"3516","article-title":"SGN: Sequential grouping networks for instance segmentation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Liu","year":"2017"},{"key":"2026032615364252300_ref430","first-page":"8759","article-title":"Path aggregation network for instance segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Liu","year":"2018"},{"key":"2026032615364252300_ref431","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46448-0_2","article-title":"SSD: Single shot MultiBox detector","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Liu","year":"2016"},{"key":"2026032615364252300_ref432","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298965","article-title":"Fully convolutional networks for semantic segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Long","year":"2015"},{"key":"2026032615364252300_ref433","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1038\/293133a0","article-title":"A computer algorithm for reconstructing a scene from two projections","volume":"293","author":"Longuet-Higgins","year":"1981","journal-title":"Nature"},{"key":"2026032615364252300_ref434","doi-asserted-by":"crossref","DOI":"10.1145\/2816795.2818013","article-title":"SMPL: A skinned multi-person linear model","volume-title":"ACM Trans. on Graphics","author":"Loper","year":"2015"},{"key":"2026032615364252300_ref435","first-page":"163","article-title":"Marching cubes: A high resolution 3D surface construction algorithm","volume-title":"ACM Trans. on Graphics","author":"Lorensen","year":"1987"},{"issue":"2","key":"2026032615364252300_ref436","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scaleinvariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"1","key":"2026032615364252300_ref437","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TRO.2015.2496823","article-title":"Visual place recognition: A survey","volume":"32","author":"Lowry","year":"2016","journal-title":"IEEE Trans. on Robotics"},{"issue":"11","key":"2026032615364252300_ref438","article-title":"Sparse cost volume for efficient stereo matching","volume":"10","author":"Lu","year":"2018","journal-title":"Remote Sensing (RS)"},{"key":"2026032615364252300_ref439","unstructured":"Luiten, J., T.Fischer, and B.Leibe (2019). \u201cTrack to reconstruct and reconstruct to track\u201d. arXiv: 1910.00130[cs.CV]."},{"key":"2026032615364252300_ref440","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.614","article-title":"Efficient deep learning for stereo matching","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Luo","year":"2016"},{"key":"2026032615364252300_ref441","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46484-8_46","article-title":"A continuous optimization approach for efficient and accurate scene flow","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Lv","year":"2016"},{"key":"2026032615364252300_ref442","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2015.XI.037","article-title":"Get out of my lab: Large-scale, real-time visual-inertial localization","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Lynen","year":"2015"},{"key":"2026032615364252300_ref443","article-title":"Customized multi-person tracker","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Ma","year":"2018"},{"key":"2026032615364252300_ref444","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00373","article-title":"Deep rigid instance scene flow","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ma","year":"2019"},{"issue":"5","key":"2026032615364252300_ref445","doi-asserted-by":"crossref","first-page":"1107","DOI":"10.1109\/TPAMI.2012.171","article-title":"Learning a confidence measure for optical flow","volume":"35","author":"Mac Aodha","year":"2013","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"1","key":"2026032615364252300_ref446","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1177\/0278364916679498","article-title":"1 year, 1000 km: The Oxford RobotCar dataset","volume":"36","author":"Maddern","year":"2016","journal-title":"International Journal of Robotics Research (IJRR)"},{"issue":"2","key":"2026032615364252300_ref447","doi-asserted-by":"crossref","first-page":"264","DOI":"10.1109\/TITS.2007.895311","article-title":"Road-sign detection and recognition based on support vector machines","volume":"8","author":"Maldonado-Basc\u00f3n","year":"2007","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref448","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00217","article-title":"ROI-10D: Monocular lifting of 2D detection to 6D pose and metric shape","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Manhardt","year":"2019"},{"key":"2026032615364252300_ref449","article-title":"Approximate Bayesian image interpretation using generative probabilistic graphics programs","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Mansinghka","year":"2013"},{"key":"2026032615364252300_ref450","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00568","article-title":"Event-based vision meets deep learning on steering prediction for self-driving cars","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Maqueda","year":"2018"},{"key":"2026032615364252300_ref451","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7299075","article-title":"3D all the way: Semantic segmentation of urban scenes from start to end in 3D","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Martinovi\u0107","year":"2015"},{"issue":"1","key":"2026032615364252300_ref452","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1007\/s11263-015-0868-z","article-title":"ATLAS: A three-layered approach to facade parsing","volume":"118","author":"Mathias","year":"2016","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref453","first-page":"1","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Mathias","year":"2013"},{"key":"2026032615364252300_ref454","first-page":"366","article-title":"Incremental estimation of dense depth maps from image sequences","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Matthies","year":"1988"},{"key":"2026032615364252300_ref455","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.197","article-title":"Enhancing road maps by parsing aerial images around the world","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Mattyus","year":"2015"},{"key":"2026032615364252300_ref456","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.393","article-title":"HD maps: Fine-grained road segmentation by parsing ground and aerial images","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Mattyus","year":"2016"},{"key":"2026032615364252300_ref457","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2015.7353481","article-title":"VoxNet: A 3D convolutional neural network for real-time object recognition","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Maturana","year":"2015"},{"key":"2026032615364252300_ref458","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.438","article-title":"A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Mayer","year":"2016"},{"key":"2026032615364252300_ref459","first-page":"4628","article-title":"SemanticFusion: Dense 3D semantic mapping with convolutional neural networks","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"McCormac","year":"2017"},{"key":"2026032615364252300_ref460","doi-asserted-by":"crossref","unstructured":"Mehta, A., A.Subramanian, and A.Subramanian (2018). \u201cLearning end-to-end autonomous driving using guided auxiliary supervision\u201d. arXiv: 1808.10393[cs.LG].","DOI":"10.1145\/3293353.3293364"},{"key":"2026032615364252300_ref461","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.2007.364084","article-title":"Single view point omnidirectional camera calibration from planar grids","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Mei","year":"2007"},{"key":"2026032615364252300_ref462","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.12276","article-title":"UnFlow: Unsupervised learning of optical flow with a bidirectional census loss","volume-title":"Proc. of the Conf. on Artificial Intelligence (AAAI)","author":"Meister","year":"2018"},{"key":"2026032615364252300_ref463","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298925","article-title":"Object scene flow for autonomous vehicles","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Menze","year":"2015"},{"key":"2026032615364252300_ref464","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-24947-6_2","article-title":"Discrete optimization for optical flow","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Menze","year":"2015"},{"key":"2026032615364252300_ref465","doi-asserted-by":"crossref","DOI":"10.5194\/isprsannals-II-3-W5-427-2015","article-title":"Joint 3D estimation of vehicles and scene flow","volume-title":"Proc. of the ISPRS Workshop on Image Sequence Analysis (ISA)","author":"Menze","year":"2015"},{"key":"2026032615364252300_ref466","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1016\/j.isprsjprs.2017.09.013","article-title":"Object scene flow","volume":"140","author":"Menze","year":"2018","journal-title":"ISPRS Journal of Photogrammetry and Remote Sensing (JPRS)"},{"key":"2026032615364252300_ref467","first-page":"1","article-title":"Real-time visibilitybased fusion of depth maps","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Merrell","year":"2007"},{"key":"2026032615364252300_ref468","unstructured":"Metz, C.\n           (2018). \u201cA Toaster on wheels to deliver groceries? Selfdriving tech tests practical uses\u201d. https:\/\/www.nytimes.com\/2018\/12\/18\/technology\/driverless-mini-car-deliver-groceries.html. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref469","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206535","article-title":"Piecewise planar city 3D modeling from street view panoramic sequences","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Micusik","year":"2009"},{"issue":"1","key":"2026032615364252300_ref470","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1109\/TPAMI.2013.103","article-title":"Continuous energy minimization for multitarget tracking","volume":"36","author":"Milan","year":"2014","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref471","unstructured":"Milan, A., L.Leal-Taix\u00e9, I. D.Reid, S.Roth, and K.Schindler (2016). \u201cMOT16: A benchmark for multi-object tracking\u201d. arXiv: 1603.00831[cs.CV]."},{"key":"2026032615364252300_ref472","first-page":"4225","article-title":"Online multi-target tracking using recurrent neural networks","volume-title":"Proc. of the Conf. on Artificial Intelligence (AAAI)","author":"Milan","year":"2017"},{"key":"2026032615364252300_ref473","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.472","article-title":"Detection- and trajectory-level exclusion in multiple object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Milan","year":"2013"},{"key":"2026032615364252300_ref474","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-11752-2_45","article-title":"On the second order statistics of essential matrix elements","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Mirabdollah","year":"2014"},{"key":"2026032615364252300_ref475","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-24947-6_24","article-title":"Fast techniques for monocular visual odometry","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Mirabdollah","year":"2015"},{"key":"2026032615364252300_ref476","unstructured":"Mirowski, P., A.Banki-Horvath, K.Anderson, D.Teplyashin, K. M.Hermann, M.Malinowski, M. K.Grimes, K.Simonyan, K.Kavukcuoglu, A.Zisserman, and R.Hadsell (2019). \u201cThe streetlearn environment and dataset\u201d. arXiv: 1903.01292[cs.AI]."},{"key":"2026032615364252300_ref477","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33715-4_41","article-title":"Taking mobile multi-object tracking to the next level: People, unknown objects, and carried items","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Mitzel","year":"2012"},{"key":"2026032615364252300_ref478","article-title":"Deep deconvolutional networks for scene parsing","volume-title":"arXiv.org. 1411.4101","author":"Mohan","year":"2014"},{"key":"2026032615364252300_ref479","first-page":"593","article-title":"FastSLAM: A factored solution to the simultaneous localization and mapping problem","volume-title":"Artificial Intelligence (AI)","author":"Montemerlo","year":"2002"},{"key":"2026032615364252300_ref480","article-title":"Semantic segmentation of aerial images in urban areas with classspecific higher-order cliques","volume-title":"ISPRS Conf. Photogrammetric Image Analysis (PIA)","author":"Montoya","year":"2015"},{"key":"2026032615364252300_ref481","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2013.6630716","article-title":"Joint self-localization and tracking of generic objects in 3D range data","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Moosmann","year":"2013"},{"key":"2026032615364252300_ref482","unstructured":"Moulon, P., P.Monasse, R.Marlet, et al. (2012). \u201cOpenMVG. An open multiple view geometry library\u201d. https:\/\/github.com\/openMVG\/openMVG. Online: accessed 23-April-2019."},{"key":"2026032615364252300_ref483","first-page":"5632","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Mousavian","year":"2017"},{"key":"2026032615364252300_ref484","article-title":"Continuoustime trajectory estimation for event-based vision sensors","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Mueggler","year":"2015"},{"key":"2026032615364252300_ref485","doi-asserted-by":"crossref","DOI":"10.1177\/0278364917691115","article-title":"The event-camera dataset and simulator: Eventbased data for pose estimation, visual odometry, and SLAM","volume-title":"International Journal of Robotics Research (IJRR)","author":"Mueggler","year":"2017"},{"key":"2026032615364252300_ref486","article-title":"Driving policy transfer via modularity and abstraction","volume-title":"arXiv. org. abs\/1804.09364","author":"M\u00fcller","year":"2018"},{"key":"2026032615364252300_ref487","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15567-3_5","article-title":"Stacked hierarchical labeling","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Munoz","year":"2010"},{"key":"2026032615364252300_ref488","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206590","article-title":"Contextual classification with functional max-margin Markov networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Munoz","year":"2009"},{"issue":"5","key":"2026032615364252300_ref489","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1109\/TRO.2015.2463671","article-title":"ORBSLAM: A versatile and accurate monocular SLAM system","volume":"31","author":"Mur-Artal","year":"2015","journal-title":"IEEE Trans. on Robotics"},{"issue":"6","key":"2026032615364252300_ref490","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1111\/cgf.12077","article-title":"A survey of urban reconstruction","volume":"32","author":"Musialski","year":"2013","journal-title":"Computer Graphics Forum"},{"key":"2026032615364252300_ref491","article-title":"Learning object relationships via graph-based context model","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Myeong","year":"2012"},{"key":"2026032615364252300_ref492","article-title":"Object scene flow with temporal consistency","volume-title":"Proc. of the Computer Vision Winter Workshop (CVWW)","author":"Neoral","year":"2017"},{"key":"2026032615364252300_ref493","article-title":"Continual occlusions and optical flow estimation","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Neoral","year":"2018"},{"key":"2026032615364252300_ref494","first-page":"60","article-title":"MC2-SLAM: Real-time inertial lidar odometry using two-scan motion compensation","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Neuhaus","year":"2018"},{"key":"2026032615364252300_ref495","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.534","article-title":"The mapillary vistas dataset for semantic understanding of street scenes","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Neuhold","year":"2017"},{"key":"2026032615364252300_ref496","doi-asserted-by":"crossref","DOI":"10.1109\/ISMAR.2011.6092378","article-title":"KinectFusion: Real-time dense surface mapping and tracking","volume-title":"Proc. of the International Symposium on Mixed and Augmented Reality (ISMAR)","author":"Newcombe","year":"2011"},{"key":"2026032615364252300_ref497","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2011.6126513","article-title":"DTAM: Dense tracking and mapping in real-time","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Newcombe","year":"2011"},{"key":"2026032615364252300_ref498","doi-asserted-by":"crossref","DOI":"10.1145\/2508363.2508374","article-title":"Real-time 3D reconstruction at scale using voxel hashing","volume-title":"ACM Trans. on Graphics","author":"Nie\u00dfner","year":"2013"},{"issue":"6","key":"2026032615364252300_ref499","doi-asserted-by":"crossref","first-page":"756","DOI":"10.1109\/TPAMI.2004.17","article-title":"An efficient solution to the five-point relative pose problem","volume":"26","author":"Nist\u00e9r","year":"2004","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref500","article-title":"Mapbased priors for localization","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Oh","year":"2004"},{"key":"2026032615364252300_ref501","volume-title":"Knowledge-Based Interpretation of Outdoor Natural Color Scenes","author":"Ohta","year":"1985"},{"key":"2026032615364252300_ref502","article-title":"Efficient deep methods for monocular road segmentation","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Oliveira","year":"2016"},{"key":"2026032615364252300_ref503","doi-asserted-by":"crossref","first-page":"312","DOI":"10.1016\/j.robot.2016.05.011","article-title":"Incremental scenario representations for autonomous driving using geometric polygonal primitives","volume":"83","author":"Oliveira","year":"2016","journal-title":"Robotics and Autonomous Systems (RAS)"},{"key":"2026032615364252300_ref504","first-page":"371","article-title":"Foveal vision for instance segmentation of road images","volume-title":"Proc. of the International Joint Conf. on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP)","author":"Ortelt","year":"2018"},{"key":"2026032615364252300_ref505","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2017.7989230","article-title":"Combined image- and world-space tracking in traffic scenes","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Osep","year":"2017"},{"key":"2026032615364252300_ref506","doi-asserted-by":"crossref","DOI":"10.5244\/C.31.11","article-title":"Virtual to real reinforcement learning for autonomous driving","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Pan","year":"2017"},{"key":"2026032615364252300_ref507","doi-asserted-by":"crossref","first-page":"1543","DOI":"10.1177\/0278364911400640","article-title":"Ford campus vision and lidar data set","volume":"30","author":"Pandey","year":"2011","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref508","first-page":"878","article-title":"Cascade residual learning: A two-stage convolutional neural network for stereo matching","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Pang","year":"2017"},{"issue":"1","key":"2026032615364252300_ref509","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1023\/A:1008162616689","article-title":"A trainable system for object detection","volume":"38","author":"Papageorgiou","year":"2000","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref510","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00410","article-title":"RayNet: Learning volumetric 3D reconstruction with ray potentials","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Paschalidou","year":"2018"},{"key":"2026032615364252300_ref511","first-page":"1157","article-title":"Thin junction tree filters for simultaneous localization and mapping","volume-title":"Proc. of the International Joint Conf. on Artificial Intelligence (IJCAI)","author":"Paskin","year":"2003"},{"key":"2026032615364252300_ref512","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.2010.5509587","article-title":"FAB-MAP 3D: Topological mapping with spatial and visual appearance","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Paul","year":"2010"},{"key":"2026032615364252300_ref513","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2009.5459260","article-title":"You\u2019ll never walk alone: Modeling social behavior for multitarget tracking","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Pellegrini","year":"2009"},{"issue":"11","key":"2026032615364252300_ref514","doi-asserted-by":"crossref","first-page":"2232","DOI":"10.1109\/TPAMI.2015.2408347","article-title":"Multiview and 3D deformable part models","volume":"37","author":"Pepik","year":"2015","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref515","first-page":"61","article-title":"Backward motion for estimation enhancement in sparse visual odometry","volume-title":"Workshop of Computer Vision (WVC)","author":"Pereira","year":"2017"},{"key":"2026032615364252300_ref516","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2006.195","article-title":"Multi-object tracking through simultaneous long occlusions and split-merge conditions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Perera","year":"2006"},{"key":"2026032615364252300_ref517","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2015.7225764","article-title":"Robust stereo visual odometry from monocular techniques","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Persson","year":"2015"},{"key":"2026032615364252300_ref518","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2010.5548114","article-title":"Efficient representation of traffic scenes by means of dynamic stixels","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Pfeiffer","year":"2010"},{"key":"2026032615364252300_ref519","doi-asserted-by":"crossref","DOI":"10.5244\/C.25.51","article-title":"Towards a global optimal multi-layer stixel representation of dense 3D data","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Pfeiffer","year":"2011"},{"key":"2026032615364252300_ref520","first-page":"110","article-title":"Robust object proposals re-ranking for object detection in autonomous driving using convolutional neural networks","volume":"53","author":"Pham","year":"2017","journal-title":"Signal Processing: Image Communication (SPIC)"},{"key":"2026032615364252300_ref521","first-page":"9250","article-title":"SuperDepth: Selfsupervised, super-resolved monocular depth estimation","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Pillai","year":"2019"},{"key":"2026032615364252300_ref522","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2015.7353537","article-title":"High-performance long range obstacle detection using stereo vision","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Pinggera","year":"2015"},{"key":"2026032615364252300_ref523","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2016.7759186","article-title":"Lost and found: Detecting small road hazards for self-driving vehicles","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Pinggera","year":"2016"},{"key":"2026032615364252300_ref524","first-page":"1990","article-title":"Learning to segment object candidates","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Pinheiro","year":"2015"},{"key":"2026032615364252300_ref525","first-page":"75","article-title":"Learning to refine object segments","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Pinheiro","year":"2016"},{"key":"2026032615364252300_ref526","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995604","article-title":"Globally-optimal greedy algorithms for tracking a variable number of objects","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pirsiavash","year":"2011"},{"key":"2026032615364252300_ref527","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.533","article-title":"DeepCut: Joint subset partition and labeling for multi person pose estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pishchulin","year":"2016"},{"key":"2026032615364252300_ref528","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2012.6248052","article-title":"Articulated people detection and pose estimation: Reshaping the future","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pishchulin","year":"2012"},{"key":"2026032615364252300_ref529","article-title":"PyDriver: Entwicklung eines Frameworks f\u00fcr r\u00e4umliche Detektion und Klassifikation von Objekten in Fahrzeugumgebung","volume-title":"MA thesis","author":"Plotkin","year":"2015"},{"key":"2026032615364252300_ref530","first-page":"3309","article-title":"Fullresolution residual networks for semantic segmentation in street scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pohlen","year":"2017"},{"key":"2026032615364252300_ref531","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2007.383073","article-title":"Change detection in a 3-d world","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Pollard","year":"2007"},{"issue":"2-3","key":"2026032615364252300_ref532","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1007\/s11263-007-0086-4","article-title":"Detailed real-time urban 3D reconstruction from video","volume":"78","author":"Pollefeys","year":"2008","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"2","key":"2026032615364252300_ref533","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1109\/64.491277","article-title":"Rapidly adapting machine vision for automated vehicle steering","volume":"11","author":"Pomerleau","year":"1996","journal-title":"IEEE Expert"},{"key":"2026032615364252300_ref534","first-page":"305","article-title":"ALVINN: An autonomous land vehicle in a neural network","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Pomerleau","year":"1988"},{"key":"2026032615364252300_ref535","unstructured":"Pomerleau, D. and T.Jochem (2015). \u201cLook, Ma, No Hands\u201d. https:\/\/www.cmu.edu\/news\/stories\/archives\/2015\/july\/look-ma-no-hands.html. Online: accessed 18-June-2019."},{"issue":"1","key":"2026032615364252300_ref536","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1109\/TPAMI.2016.2537320","article-title":"Multiscale combinatorial grouping for image segmentation and object proposal generation","volume":"39","author":"Pont-Tuset","year":"2017","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref537","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00102","article-title":"Frustum pointnets for 3D object detection from RGB-D data","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Qi","year":"2018"},{"key":"2026032615364252300_ref538","article-title":"Improving multi-target tracking via social grouping","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Qin","year":"2012"},{"key":"2026032615364252300_ref539","first-page":"1221","article-title":"UnrealCV: Virtual worlds for computer vision","volume-title":"ACM Multimedia Open Source Software Competition","author":"Qiu","year":"2017"},{"key":"2026032615364252300_ref540","first-page":"149","article-title":"Hierarchical warp stereo","volume-title":"DARPA Image Understanding Workshop","author":"Quam","year":"1984"},{"key":"2026032615364252300_ref541","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15561-1_42","article-title":"Dense, robust, and accurate motion field estimation from stereo image sequences in real-time","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Rabe","year":"2010"},{"key":"2026032615364252300_ref542","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2007.4408986","article-title":"Objects in context","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Rabinovich","year":"2007"},{"issue":"4","key":"2026032615364252300_ref543","doi-asserted-by":"crossref","first-page":"4407","DOI":"10.1109\/LRA.2018.2869640","article-title":"VLocNet++: Deep multitask learning for semantic visual localization and odometry","volume":"3","author":"Radwan","year":"2018","journal-title":"IEEE Robotics and Automation Letters (RA-L)"},{"key":"2026032615364252300_ref544","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10590-1_29","article-title":"Non-local total generalized variation for optical flow estimation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Ranftl","year":"2014"},{"key":"2026032615364252300_ref545","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-38267-3_24","article-title":"Minimizing TGV-based variational models with non-convex data terms","volume-title":"Proc. of the International Conf. on Scale Space and Variational Methods in Computer Vision (SSVM)","author":"Ranftl","year":"2013"},{"key":"2026032615364252300_ref546","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.291","article-title":"Optical flow estimation using a spatial pyramid network","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ranjan","year":"2017"},{"key":"2026032615364252300_ref547","article-title":"EVO: A geometric approach to event-based 6-DOF parallel tracking and mapping in real-time","volume-title":"IEEE Robotics and Automation Letters (RA-L)","author":"Rebecq","year":"2016"},{"key":"2026032615364252300_ref548","first-page":"779","article-title":"You only look once: Unified, real-time object detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Redmon","year":"2016"},{"key":"2026032615364252300_ref549","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.690","article-title":"YOLO9000: Better, faster, stronger","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Redmon","year":"2017"},{"key":"2026032615364252300_ref550","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15558-1_14","article-title":"Detection and tracking of large number of targets in wide area surveillance","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Reilly","year":"2010"},{"key":"2026032615364252300_ref551","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.87","article-title":"Accurate single stage detector using recurrent rolling convolution","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ren","year":"2017"},{"key":"2026032615364252300_ref552","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Ren","year":"2015"},{"issue":"6","key":"2026032615364252300_ref553","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"7","key":"2026032615364252300_ref554","doi-asserted-by":"crossref","first-page":"1476","DOI":"10.1109\/TPAMI.2016.2601099","article-title":"Object detection networks on convolutional feature maps","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref555","first-page":"225","article-title":"Cascaded scene flow prediction using semantic segmentation","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Ren","year":"2017"},{"key":"2026032615364252300_ref556","doi-asserted-by":"crossref","DOI":"10.1016\/j.isprsjprs.2014.09.010","article-title":"Evaluation of feature-based 3-D registration of probabilistic volumetric scenes","volume-title":"ISPRS Journal of Photogrammetry and Remote Sensing (JPRS)","author":"Restrepo","year":"2014"},{"key":"2026032615364252300_ref557","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298720","article-title":"EpicFlow: Edge-preserving interpolation of correspondences for optical flow","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Revaud","year":"2015"},{"key":"2026032615364252300_ref558","doi-asserted-by":"crossref","DOI":"10.1109\/3DV.2016.36","article-title":"Dense wide-baseline scene flow from two handheld video cameras","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Richardt","year":"2016"},{"key":"2026032615364252300_ref559","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.243","article-title":"Playing for benchmarks","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Richter","year":"2017"},{"key":"2026032615364252300_ref560","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46475-6_7","article-title":"Playing for data: Ground truth from computer games","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Richter","year":"2016"},{"key":"2026032615364252300_ref561","doi-asserted-by":"crossref","unstructured":"Richter, S. R., V.Vineet, S.Roth, and V.Koltun (2016b). \u201cPlaying for data: Ground truth from computer games\u201d. https:\/\/download.visinf.tu-darmstadt.de\/data\/from_games\/. Online: accessed 18-October-2019.","DOI":"10.1007\/978-3-319-46475-6_7"},{"key":"2026032615364252300_ref562","doi-asserted-by":"crossref","DOI":"10.1109\/3DV.2017.00017","article-title":"OctNetFusion: Learning depth fusion from data","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Riegler","year":"2017"},{"key":"2026032615364252300_ref563","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.701","article-title":"OctNet: Learning deep 3D representations at high resolutions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Riegler","year":"2017"},{"key":"2026032615364252300_ref564","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10602-1_34","article-title":"Learning where to classify in multi-view semantic segmentation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Riemenschneider","year":"2014"},{"key":"2026032615364252300_ref565","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-48881-3_2","article-title":"Performance measures and a data set for multi-target, multi-camera tracking","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV) Workshops","author":"Ristani","year":"2016"},{"key":"2026032615364252300_ref566","article-title":"Machine perception of three-dimensional solids","volume-title":"PhD thesis","author":"Roberts","year":"1963"},{"key":"2026032615364252300_ref567","first-page":"234","volume-title":"Medical Image Computing and Computer-Assisted Intervention (MICCAI)","author":"Ronneberger","year":"2015"},{"key":"2026032615364252300_ref568","unstructured":"Ros, G., L.Sellart, J.Materzynska, D.Vazquez, and A.Lopez (2016a). \u201cSYNTHIA dataset\u201d. http:\/\/synthia-dataset.net\/. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref569","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.352","article-title":"The SYNTHIA dataset: A large collection of synthetic images for semantic segmentation of urban scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ros","year":"2016"},{"key":"2026032615364252300_ref570","first-page":"661","article-title":"Efficient reductions for imitation learning","volume":"9","author":"Ross","year":"2010","journal-title":"Conference on Artificial Intelligence and Statistics AISTATS"},{"issue":"3","key":"2026032615364252300_ref571","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/1015706.1015720","article-title":"GrabCut: Interactive foreground extraction using iterated graph cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM Trans. on Graphics"},{"key":"2026032615364252300_ref572","first-page":"2564","article-title":"ORB: An efficient alternative to SIFT or SURF","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Rublee","year":"2011"},{"issue":"1\u20134","key":"2026032615364252300_ref573","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1016\/0167-2789(92)90242-F","article-title":"Nonlinear total variation based noise removal algorithms","volume":"60","author":"Rudin","year":"1992","journal-title":"Physica D: Nonlinear Phenomena"},{"key":"2026032615364252300_ref574","doi-asserted-by":"crossref","first-page":"533","DOI":"10.1038\/323533a0","article-title":"Learning representations by back-propagating errors","volume":"323","author":"Rumelhart","year":"1986","journal-title":"Nature"},{"issue":"3","key":"2026032615364252300_ref575","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref576","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.41","article-title":"Tracking the untrackable: Learning to track multiple cues with long-term dependencies","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Sadeghian","year":"2017"},{"key":"2026032615364252300_ref577","first-page":"164","article-title":"Improved visual relocalization by discovering anchor points","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Saha","year":"2018"},{"key":"2026032615364252300_ref578","unstructured":"Santana, E. and G.Hotz (2016). \u201cLearning a driving simulator\u201d. arXiv: 1608.01230[cs.LG]."},{"key":"2026032615364252300_ref579","article-title":"From coarse to fine: Robust hierarchical localization at large scale","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sarlin","year":"2018"},{"issue":"9","key":"2026032615364252300_ref580","doi-asserted-by":"crossref","first-page":"1744","DOI":"10.1109\/TPAMI.2016.2611662","article-title":"Efficient effective prioritized matching for large-scale image-based localization","volume":"39","author":"Sattler","year":"2016","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref581","article-title":"Hyperpoints and fine vocabularies for largescale location recognition","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Sattler","year":"2015"},{"key":"2026032615364252300_ref582","first-page":"8601","article-title":"Benchmarking 6DOF outdoor visual localization in changing conditions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sattler","year":"2018"},{"key":"2026032615364252300_ref583","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00342","article-title":"Understanding the limitations of CNN-based absolute camera pose regression","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sattler","year":"2019"},{"key":"2026032615364252300_ref584","article-title":"Conditional affordance learning for driving in urban environments","volume-title":"Proc. Conf. on Robot Learning (CoRL)","author":"Sauer","year":"2018"},{"key":"2026032615364252300_ref585","doi-asserted-by":"crossref","DOI":"10.1109\/CVPRW.2015.7301289","article-title":"Semantically-enriched 3D models for common-sense knowledge","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops","author":"Savva","year":"2015"},{"key":"2026032615364252300_ref586","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2019.8814146","article-title":"PWOC-3D: Deep occlusion-aware end-to-end scene flow estimation","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Saxena","year":"2019"},{"issue":"5","key":"2026032615364252300_ref587","doi-asserted-by":"crossref","first-page":"1015","DOI":"10.1109\/TRO.2008.2004490","article-title":"Appearance-guided monocular omnidirectional visual odometry for outdoor ground vehicles","volume":"24","author":"Scaramuzza","year":"2008","journal-title":"IEEE Trans. on Robotics"},{"issue":"4","key":"2026032615364252300_ref588","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1109\/MRA.2011.943233","article-title":"Visual odometry [tutorial]","volume":"18","author":"Scaramuzza","year":"2011","journal-title":"Robotics and Automation Magazine (RAM)"},{"key":"2026032615364252300_ref589","article-title":"Realtime monocular visual odometry for on-road vehicles with 1-point RANSAC","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Scaramuzza","year":"2009"},{"key":"2026032615364252300_ref590","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2006.282372","article-title":"A toolbox for easily calibrating omnidirectional cameras","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Scaramuzza","year":"2006"},{"key":"2026032615364252300_ref591","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-11752-2_3","article-title":"High-resolution stereo datasets with subpixel-accurate ground truth","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Scharstein","year":"2014"},{"key":"2026032615364252300_ref592","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein","year":"2002","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref593","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2003.1211354","article-title":"High-accuracy stereo depth maps using structured light","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Scharstein","year":"2003"},{"key":"2026032615364252300_ref594","first-page":"5171","article-title":"A new feature detector and stereo matching method for accurate high-performance sparse stereo matching","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Schauwecker","year":"2012"},{"key":"2026032615364252300_ref595","first-page":"433","article-title":"Mono-camera 3D multi-object tracking using deep learning detections and PMBM filtering","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Scheidegger","year":"2018"},{"key":"2026032615364252300_ref596","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2016.7535373","article-title":"Semantic stixels: Depth is not enough","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Schneider","year":"2016"},{"key":"2026032615364252300_ref597","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2014.6942637","article-title":"Omnidirectional 3D reconstruction in augmented manhattan worlds","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Sch\u00f6nbein","year":"2014"},{"key":"2026032615364252300_ref598","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2014.6907507","article-title":"Calibrating and centering quasi-central catadioptric cameras","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Sch\u00f6nbein","year":"2014"},{"key":"2026032615364252300_ref599","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.445","article-title":"Structure-from-motion revisited","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sch\u00f6nberger","year":"2016"},{"key":"2026032615364252300_ref600","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-46487-9_31","article-title":"Pixelwise view selection for unstructured multi-view stereo","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Sch\u00f6nberger","year":"2016"},{"key":"2026032615364252300_ref601","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00721","article-title":"Semantic visual localization","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sch\u00f6nberger","year":"2018"},{"key":"2026032615364252300_ref602","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.272","article-title":"A multi-view stereo benchmark with high-resolution images and multi-camera videos","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sch\u00f6ps","year":"2017"},{"key":"2026032615364252300_ref603","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2013.6629509","article-title":"LaneLoc: Lane marking based localization using highly accurate maps","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Schreiber","year":"2013"},{"key":"2026032615364252300_ref604","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.292","article-title":"Deep network flow for multi-object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Schulter","year":"2017"},{"key":"2026032615364252300_ref605","unstructured":"Scott, S.\n           (2019). \u201cMeet Scout\u201d. https:\/\/blog.aboutamazon.com\/transportation\/meet-scout. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref606","unstructured":"Seff, A. and J.Xiao (2016). \u201cLearning from maps: Visual common sense for autonomous driving\u201d. arXiv: 1611.08583[cs.CV]."},{"key":"2026032615364252300_ref607","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2006.19","article-title":"A comparison and evaluation of multi-view stereo reconstruction algorithms","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Seitz","year":"2006"},{"key":"2026032615364252300_ref608","doi-asserted-by":"crossref","DOI":"10.5244\/C.30.23","article-title":"Patch based confidence prediction for dense disparity map","volume-title":"Proc. of the British Machine Vision Conf. (BMVC)","author":"Seki","year":"2016"},{"key":"2026032615364252300_ref609","first-page":"618","article-title":"Grad-CAM: Visual explanations from deep networks via gradient-based localization","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Selvaraju","year":"2017"},{"key":"2026032615364252300_ref610","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2013.6630632","article-title":"Urban 3D semantic modelling using stereo vision","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Sengupta","year":"2013"},{"key":"2026032615364252300_ref611","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2012.6385958","article-title":"Automatic dense visual semantic mapping from street-level imagery","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Sengupta","year":"2012"},{"key":"2026032615364252300_ref612","article-title":"OverFeat: Integrated recognition, localization and detection using convolutional networks","volume-title":"Proc. of the International Conf. on Learning Representations (ICLR)","author":"Sermanet","year":"2014"},{"key":"2026032615364252300_ref613","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.465","article-title":"Pedestrian detection with unsupervised multi-stage feature learning","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sermanet","year":"2013"},{"key":"2026032615364252300_ref614","first-page":"2809","article-title":"Traffic sign recognition with multi-scale convolutional networks","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Sermanet","year":"2011"},{"key":"2026032615364252300_ref615","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.422","article-title":"Optical flow with semantic segmentation and localized layers","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sevilla-Lara","year":"2016"},{"key":"2026032615364252300_ref616","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587577","article-title":"A rank constrained continuous formulation of multi-frame multi-target tracking problem","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Shafique","year":"2008"},{"key":"2026032615364252300_ref617","article-title":"Accurate geo-registration by ground-to-aerial image matching","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Shan","year":"2014"},{"key":"2026032615364252300_ref618","article-title":"Learning to drive using inverse reinforcement learning and deep Q-networks","volume-title":"Advances in Neural Information Processing Systems (NeurIPS) Workshops","author":"Sharifzadeh","year":"2016"},{"key":"2026032615364252300_ref619","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2018.8461018","article-title":"Beyond pixels: Leveraging geometry and shape cues for online multi-object tracking","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Sharma","year":"2018"},{"key":"2026032615364252300_ref620","doi-asserted-by":"crossref","DOI":"10.1109\/IVS.2004.1336346","article-title":"Pedestrian detection for driving assistance systems: Single-frame classification and system level performance","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Shashua","year":"2004"},{"key":"2026032615364252300_ref621","unstructured":"Shen, H., L.Huang, C.Huang, and W.Xu (2018). \u201cTracklet association tracker: An end-to-end learning-based association approach for multi-object tracking\u201d. arXiv: 1808.01562[cs.CV]."},{"issue":"11","key":"2026032615364252300_ref622","doi-asserted-by":"crossref","first-page":"3269","DOI":"10.1109\/TCSVT.2018.2882192","article-title":"Heterogeneous association graph fusion for target association in multiple object tracking","volume":"29","author":"Sheng","year":"2019","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026032615364252300_ref623","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1007\/s11263-007-0109-1","article-title":"TextonBoost for image understanding: Multi-class object recognition and segmentation by jointly modeling texture, layout, and context","volume":"81","author":"Shotton","year":"2009","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref624","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.377","article-title":"Scene coordinate regression forests for camera relocalization in RGB-D images","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Shotton","year":"2013"},{"key":"2026032615364252300_ref625","article-title":"Part-based multiple-person tracking with partial occlusion handling","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Shu","year":"2012"},{"key":"2026032615364252300_ref626","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.1991.139707","article-title":"Probability distributions of optical flow","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Simoncelli","year":"1991"},{"key":"2026032615364252300_ref627","article-title":"Very deep convolutional networks for large-scale image recognition","volume-title":"Proc. of the International Conf. on Learning Representations (ICLR)","author":"Simonyan","year":"2015"},{"key":"2026032615364252300_ref628","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1177\/0278364909103911","article-title":"The new college vision and laser data set","volume":"28","author":"Smith","year":"2009","journal-title":"International Journal of Robotics Research (IJRR)"},{"key":"2026032615364252300_ref629","doi-asserted-by":"crossref","DOI":"10.1109\/ROBOT.1987.1087846","article-title":"Estimating uncertain spatial relationships in robotics","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Smith","year":"1987"},{"key":"2026032615364252300_ref630","first-page":"835","article-title":"Photo tourism: Exploring photo collections in 3D","volume-title":"ACM Trans. on Graphics","author":"Snavely","year":"2006"},{"issue":"2","key":"2026032615364252300_ref631","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1007\/s11263-007-0107-3","article-title":"Modeling the world from internet photo collections","volume":"80","author":"Snavely","year":"2008","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref632","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2014.203","article-title":"Robust scale estimation in real-time monocular SFM for autonomous driving","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Song","year":"2014"},{"key":"2026032615364252300_ref633","article-title":"EdgeStereo: A context integrated residual pyramid network for stereo matching","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Song","year":"2018"},{"key":"2026032615364252300_ref634","first-page":"1453","article-title":"The German traffic sign recognition benchmark: A multi-class classification competition","volume-title":"International Joint Conference on Neural Networks (IJCNN)","author":"Stallkamp","year":"2011"},{"key":"2026032615364252300_ref635","first-page":"1","volume-title":"CLEAR","author":"Stiefelhagen","year":"2007"},{"key":"2026032615364252300_ref636","first-page":"2352","article-title":"Double window optimisation for constant time visual SLAM","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Strasdat","year":"2011"},{"key":"2026032615364252300_ref637","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2010.VI.010","article-title":"Scale drift-aware large scale monocular SLAM","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Strasdat","year":"2010"},{"key":"2026032615364252300_ref638","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587706","article-title":"On benchmarking camera calibration and multi-view stereo for high resolution imagery","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Strecha","year":"2008"},{"key":"2026032615364252300_ref639","first-page":"573","article-title":"A benchmark for the evaluation of RGB-D SLAM systems","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Sturm","year":"2012"},{"key":"2026032615364252300_ref640","first-page":"206","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Suard","year":"2006"},{"key":"2026032615364252300_ref641","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2016.7759533","article-title":"The path less taken: A fast variational approach for scene segmentation used for closed loop control","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Suleymanov","year":"2016"},{"issue":"2","key":"2026032615364252300_ref642","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1007\/s11263-013-0644-x","article-title":"A quantitative analysis of current practices in optical flow estimation and the principles behind them","volume":"106","author":"Sun","year":"2014","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref643","unstructured":"Sun, D., X.Yang, M.Liu, and J.Kautz (2018a). \u201cModels matter, so does training: An empirical study of CNNs for optical flow estimation\u201d. arXiv: 1809.05571[cs.CV]."},{"key":"2026032615364252300_ref644","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00931","article-title":"PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sun","year":"2018"},{"key":"2026032615364252300_ref645","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2011.6126309","article-title":"Articulated part-based model for joint object detection and pose estimation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Sun","year":"2011"},{"key":"2026032615364252300_ref646","doi-asserted-by":"crossref","unstructured":"Sun, P., H.Kretzschmar, X.Dotiwalla, A.Chouard, V.Patnaik, P.Tsui, J.Guo, Y.Zhou, Y.Chai, B.Caine, V.Vasudevan, W.Han, J.Ngiam, H.Zhao, A.Timofeev, S.Ettinger, M.Krivokon, A.Gao, A.Joshi, Y.Zhang, J.Shlens, Z.Chen, and D.Anguelov (2019). \u201cScalability in perception for autonomous driving: Waymo open dataset\u201d. arXiv: 1912.04838[cs.CV].","DOI":"10.1109\/CVPR42600.2020.00252"},{"issue":"7","key":"2026032615364252300_ref647","doi-asserted-by":"crossref","first-page":"1455","DOI":"10.1109\/TPAMI.2016.2598331","article-title":"Cityscale localization for cameras with known vertical direction","volume":"39","author":"Svarm","year":"2017","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref648","first-page":"532","article-title":"Accurate localization and pose estimation for large 3D models","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Sv\u00e4rm","year":"2014"},{"issue":"1","key":"2026032615364252300_ref649","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1023\/A:1019869530073","article-title":"Epipolar geometry for central catadioptric cameras","volume":"49","author":"Svoboda","year":"2002","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref650","unstructured":"Sweeney, C.\n           (2016). \u201cTheia multiview geometry library: Tutorial & reference\u201d. http:\/\/theia-sfm.org. Online: accessed 23-April-2019."},{"key":"2026032615364252300_ref651","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298594","article-title":"Going deeper with convolutions","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Szegedy","year":"2015"},{"key":"2026032615364252300_ref652","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-84882-935-0","volume-title":"Computer Vision \u2013 Algorithms and Applications","author":"Szeliski","year":"2011"},{"issue":"5","key":"2026032615364252300_ref653","doi-asserted-by":"crossref","first-page":"1272","DOI":"10.1109\/TPAMI.2019.2910529","article-title":"Object detection in videos by high quality object linking","volume":"42","author":"Tang","year":"2019","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref654","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7299138","article-title":"Subgraph decomposition for multi-target tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Tang","year":"2015"},{"key":"2026032615364252300_ref655","article-title":"Multiperson tracking by multicut and deep matching","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV) Workshops","author":"Tang","year":"2016"},{"key":"2026032615364252300_ref656","first-page":"3701","article-title":"Multiple people tracking by lifted multicut and person re-identification","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Tang","year":"2017"},{"issue":"2","key":"2026032615364252300_ref657","doi-asserted-by":"crossref","first-page":"594","DOI":"10.1109\/LRA.2019.2891492","article-title":"A whitenoise-on-jerk motion prior for continuous-time trajectory estimation on SE(3)","volume":"4","author":"Tang","year":"2019","journal-title":"IEEE Robotics and Automation Letters (RA-L)"},{"key":"2026032615364252300_ref658","first-page":"6891","article-title":"Fast multi-frame stereo scene flow with motion segmentation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Taniai","year":"2017"},{"key":"2026032615364252300_ref659","unstructured":"Tesla\n           (2014). \u201cTesla autopilot\u201d. https:\/\/www.tesla.com\/autopilot. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref660","unstructured":"Tesla\n           (2018). \u201cIntroducing software version 9.0\u201d. https:\/\/www.tesla.com\/blog\/introducing-software-version-9?redirect=no. Online: accessed 8-June-2019."},{"issue":"3","key":"2026032615364252300_ref661","doi-asserted-by":"crossref","first-page":"362","DOI":"10.1109\/34.3900","article-title":"Vision and navigation for the Carnegie-Mellon Navlab","volume":"10","author":"Thorpe","year":"1988","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref662","volume-title":"Probabilistic Robotics","author":"Thrun","year":"2005"},{"issue":"1","key":"2026032615364252300_ref663","doi-asserted-by":"crossref","first-page":"374","DOI":"10.1109\/TITS.2019.2892413","article-title":"Online multi-object tracking using joint domain information in traffic scenarios","volume":"21","author":"Tian","year":"2019","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"issue":"9","key":"2026032615364252300_ref664","doi-asserted-by":"crossref","first-page":"2146","DOI":"10.1109\/TPAMI.2018.2849374","article-title":"On detection, data association and segmentation for multi-target tracking","volume":"41","author":"Tian","year":"2018","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref665","unstructured":"TIME USA\n           (1925). \u201cScience: Radio auto\u201d. http:\/\/content.time.com\/time\/magazine\/article\/0,9171,720720,00.html. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref666","doi-asserted-by":"crossref","DOI":"10.1109\/WACV.2015.151","article-title":"Sparse flow: Sparse matching for small to large displacement optical flow","volume-title":"Proc. of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Timofte","year":"2015"},{"key":"2026032615364252300_ref667","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15558-1_26","article-title":"Unique signatures of histograms for local surface description","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Tombari","year":"2010"},{"issue":"1","key":"2026032615364252300_ref668","doi-asserted-by":"crossref","first-page":"441","DOI":"10.1109\/TITS.2014.2354243","article-title":"Efficient road scene understanding for intelligent vehicles using compositional hierarchical models","volume":"16","author":"Topfer","year":"2015","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref669","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298790","article-title":"24\/7 place recognition by view synthesis","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Torii","year":"2015"},{"issue":"2","key":"2026032615364252300_ref670","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1109\/TMI.2002.808355","article-title":"A shape-based approach to the segmentation of medical imagery using level sets","volume":"22","author":"Tsai","year":"2003","journal-title":"Medical Imaging"},{"issue":"10","key":"2026032615364252300_ref671","doi-asserted-by":"crossref","first-page":"1744","DOI":"10.1109\/TPAMI.2009.186","article-title":"Auto-context and its application to high-level vision tasks and 3D brain image segmentation","volume":"32","author":"Tu","year":"2010","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref672","first-page":"5875","article-title":"Practical deep stereo (PDS): Toward applications-friendly deep stereo matching","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Tulyakov","year":"2018"},{"key":"2026032615364252300_ref673","unstructured":"Uber\n           (2015). \u201cAdvanced technologies group\u201d. https:\/\/www.uber.com\/de\/de\/atg\/. Online: accessed 18-October-2019."},{"key":"2026032615364252300_ref674","first-page":"14","article-title":"Pixel-level encoding and depth layering for instance-level semantic labeling","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Uhrig","year":"2016"},{"issue":"2","key":"2026032615364252300_ref675","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref676","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.482","article-title":"Semantic multi-view stereo: Jointly estimating objects and voxels","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Ulusoy","year":"2017"},{"key":"2026032615364252300_ref677","doi-asserted-by":"crossref","DOI":"10.1109\/3DV.2015.9","article-title":"Towards probabilistic volumetric reconstruction using ray potentials","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Ulusoy","year":"2015"},{"issue":"2","key":"2026032615364252300_ref678","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1007\/BF00202895","article-title":"A computational approach to motion perception","volume":"60","author":"Uras","year":"1988","journal-title":"Biological Cybernetics"},{"key":"2026032615364252300_ref679","first-page":"1885","article-title":"Direct visual-inertial odometry with stereo cameras","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Usenko","year":"2016"},{"key":"2026032615364252300_ref680","unstructured":"Valada, A., R.Mohan, and W.Burgard (2018). \u201cSelf-supervised model adaptation for multimodal semantic segmentation\u201d. arXiv: 1808.03833[cs.CV]."},{"key":"2026032615364252300_ref681","first-page":"4644","article-title":"Adap-Net: Adaptive semantic segmentation in adverse environmental conditions","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Valada","year":"2017"},{"key":"2026032615364252300_ref682","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.269","article-title":"Mesh based semantic modelling for indoor and outdoor scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Valentin","year":"2013"},{"key":"2026032615364252300_ref683","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.1999.790293","article-title":"Three-dimensional scene flow","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Vedula","year":"1999"},{"key":"2026032615364252300_ref684","first-page":"1553","article-title":"Scene segmentation with CRFs learned from partially labeled images","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Verbeek","year":"2007"},{"issue":"1","key":"2026032615364252300_ref685","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1007\/s11263-013-0641-0","article-title":"Detecting parametric objects in large scenes by Monte Carlo sampling","volume":"106","author":"Verdie","year":"2014","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref686","doi-asserted-by":"crossref","DOI":"10.1109\/LRA.2018.2793357","article-title":"Ultimate SLAM? Combining events, images, and IMU for robust visual SLAM in HDR and high-speed scenarios","volume-title":"IEEE Robotics and Automation Letters (RA-L)","author":"Vidal","year":"2018"},{"key":"2026032615364252300_ref687","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33715-4_36","article-title":"Active frame selection for label propagation in videos","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Vijayanarasimhan","year":"2012"},{"key":"2026032615364252300_ref688","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2015.7138983","article-title":"Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Vineet","year":"2015"},{"issue":"2","key":"2026032615364252300_ref689","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1007\/s11263-005-6644-8","article-title":"Detecting pedestrians using patterns of motion and appearance","volume":"63","author":"Viola","year":"2005","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"2","key":"2026032615364252300_ref690","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1023\/B:VISI.0000013087.49260.fb","article-title":"Robust real-time face detection","volume":"57","author":"Viola","year":"2004","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref691","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-40602-7_37","article-title":"An evaluation of data costs for optical flow","volume-title":"Proc. of the German Conference on Pattern Recognition (GCPR)","author":"Vogel","year":"2013"},{"issue":"1","key":"2026032615364252300_ref692","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11263-015-0806-0","article-title":"3D scene flow estimation with a piecewise rigid scene model","volume":"115","author":"Vogel","year":"2015","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"3","key":"2026032615364252300_ref693","first-page":"37","article-title":"3D building model reconstruction from point clouds and ground plans","volume":"34","author":"Vosselman","year":"2001","journal-title":"Proc. of the ISPRS Workshop Land Surface Mapping and Characterization Using Laser Altimetry"},{"key":"2026032615364252300_ref694","first-page":"627","article-title":"Image-based localization using LSTMs for structured feature correlation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Walch","year":"2017"},{"key":"2026032615364252300_ref695","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2010.5540102","article-title":"New features and insights for pedestrian detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Walk","year":"2010"},{"key":"2026032615364252300_ref696","doi-asserted-by":"crossref","unstructured":"Wang, D., C.Devin, Q.Cai, P.Kr\u00e4henb\u00fchl, and T.Darrell (2019a). \u201cMonocular plan view networks for autonomous driving\u201d. arXiv: 1905.06937[cs.CV].","DOI":"10.1109\/IROS40897.2019.8967897"},{"key":"2026032615364252300_ref697","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2015.XI.035","article-title":"Voting for voting in online point cloud object detection","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Wang","year":"2015"},{"key":"2026032615364252300_ref698","unstructured":"Wang, G., Y.Wang, H.Zhang, R.Gu, and J.Hwang (2018a). \u201cExploit the connectivity: Multi-object tracking with Tracklet-Net\u201d. arXiv: 1811.07258[cs.CV]."},{"key":"2026032615364252300_ref699","first-page":"3923","article-title":"Stereo DSO: Large-scale direct sparse visual odometry with stereo cameras","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Wang","year":"2017"},{"key":"2026032615364252300_ref700","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00274","article-title":"Deep parametric continuous convolutional neural networks","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wang","year":"2018"},{"key":"2026032615364252300_ref701","article-title":"Fully motionaware network for video object detection","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Wang","year":"2018"},{"issue":"10","key":"2026032615364252300_ref702","doi-asserted-by":"crossref","first-page":"2071","DOI":"10.1109\/TPAMI.2015.2389830","article-title":"Regionlets for generic object detection","volume":"37","author":"Wang","year":"2015","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref703","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00513","article-title":"Occlusion aware unsupervised learning of optical flow","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wang","year":"2018"},{"key":"2026032615364252300_ref704","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.01057","article-title":"A parametric top-view representation of complex road scenes","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wang","year":"2019"},{"key":"2026032615364252300_ref705","first-page":"1182","article-title":"ProbFlow: Joint optical flow and uncertainty estimation","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Wannenwetsch","year":"2017"},{"key":"2026032615364252300_ref706","unstructured":"Waymo\n           (2019). \u201cBe an early rider\u201d. https:\/\/waymo.com\/apply. Online: accessed 18-October-2019."},{"issue":"1","key":"2026032615364252300_ref707","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1007\/s11263-010-0404-0","article-title":"Stereoscopic scene flow computation for 3D motion understanding","volume":"95","author":"Wedel","year":"2011","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"4","key":"2026032615364252300_ref708","doi-asserted-by":"crossref","first-page":"572","DOI":"10.1109\/TITS.2009.2027223","article-title":"B-spline modeling of road surfaces with an application to free space estimation","volume":"10","author":"Wedel","year":"2009","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref709","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-88682-2_56","article-title":"Efficient dense scene flow from sparse or dense stereo data","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Wedel","year":"2008"},{"key":"2026032615364252300_ref710","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.647","article-title":"Cataloging public objects using aerial and street-level images - Urban trees","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wegner","year":"2016"},{"key":"2026032615364252300_ref711","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.222","article-title":"A higher-order CRF model for road network extraction","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wegner","year":"2013"},{"key":"2026032615364252300_ref712","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1016\/j.isprsjprs.2015.07.002","article-title":"Road networks as collections of minimum cost paths","volume":"108","author":"Wegner","year":"2015","journal-title":"ISPRS Journal of Photogrammetry and Remote Sensing (JPRS)"},{"key":"2026032615364252300_ref713","article-title":"A data-driven regularization model for stereo and flow","volume-title":"Proc. of the International Conf. on 3D Vision (3DV)","author":"Wei","year":"2014"},{"key":"2026032615364252300_ref714","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2013.175","article-title":"DeepFlow: Large displacement optical flow with deep matching","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Weinzaepfel","year":"2013"},{"key":"2026032615364252300_ref715","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2015.XI.001","article-title":"ElasticFusion: Dense SLAM without A pose graph","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Whelan","year":"2015"},{"key":"2026032615364252300_ref716","volume-title":"Handbook of Driver Assistance Systems","author":"Winner","year":"2015"},{"key":"2026032615364252300_ref717","doi-asserted-by":"crossref","DOI":"10.1016\/S0262-8856(98)00108-5","article-title":"A time delay neural network algorithm for estimating image-pattern shape and motion","volume-title":"Image and Vision Computing (IVC)","author":"W\u00f6hler","year":"1999"},{"key":"2026032615364252300_ref718","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15561-1_34","article-title":"Monocular 3D scene modeling and inference: Understanding multi-object traffic scenes","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Wojek","year":"2010"},{"key":"2026032615364252300_ref719","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206638","article-title":"Multi-cue onboard pedestrian detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wojek","year":"2009"},{"key":"2026032615364252300_ref720","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-88693-8_54","article-title":"A dynamic conditional random field model for joint labeling of object and scene classes","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Wojek","year":"2008"},{"key":"2026032615364252300_ref721","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-69321-5_9","article-title":"A performance evaluation of single and multi-feature people detection","volume-title":"Proc. of the DAGM Symposium on Pattern Recognition (DAGM)","author":"Wojek","year":"2008"},{"key":"2026032615364252300_ref722","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995547","article-title":"Monocular 3D scene understanding with explicit occlusion reasoning","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wojek","year":"2011"},{"issue":"4","key":"2026032615364252300_ref723","doi-asserted-by":"crossref","first-page":"882","DOI":"10.1109\/TPAMI.2012.174","article-title":"Monocular visual scene understanding: Understanding multiobject traffic scenes","volume":"35","author":"Wojek","year":"2013","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref724","first-page":"3645","article-title":"Simple online and realtime tracking with a deep association metric","volume-title":"Proc. IEEE International Conf. on Image Processing (ICIP)","author":"Wojke","year":"2017"},{"key":"2026032615364252300_ref725","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.176","article-title":"Regularity-driven facade matching between aerial and street views","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wolff","year":"2016"},{"key":"2026032615364252300_ref726","doi-asserted-by":"crossref","first-page":"2115","DOI":"10.1109\/TPAMI.2009.131","article-title":"Global stereo reconstruction under second-order smoothness priors","volume":"31","author":"Woodford","year":"2009","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref727","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.451","article-title":"Wide-area image geolocalization with aerial reference imagery","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Workman","year":"2015"},{"issue":"2","key":"2026032615364252300_ref728","doi-asserted-by":"crossref","first-page":"247","DOI":"10.1007\/s11263-006-0027-7","article-title":"Detection and tracking of multiple, partially occluded humans by bayesian combination of edgelet part detectors","volume":"75","author":"Wu","year":"2007","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2026032615364252300_ref729","article-title":"VisualSFM: A visual structure from motion system","author":"Wu","year":"2011"},{"key":"2026032615364252300_ref730","first-page":"25","article-title":"A practical system for road marking detection and recognition","volume-title":"Proc. IEEE Intelligent Vehicles Symposium (IV)","author":"Wu","year":"2012"},{"issue":"9","key":"2026032615364252300_ref731","doi-asserted-by":"crossref","first-page":"1829","DOI":"10.1109\/TPAMI.2015.2497699","article-title":"Learning and-or model to represent context and occlusion for car detection and viewpoint estimation","volume":"38","author":"Wu","year":"2016","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"key":"2026032615364252300_ref732","first-page":"1185","article-title":"Efficient track linking methods for track graphs using network-flow and setcover techniques","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wu","year":"2011"},{"key":"2026032615364252300_ref733","article-title":"Coupling detection and data association for multiple object tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wu","year":"2012"},{"key":"2026032615364252300_ref734","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1016\/j.patcog.2019.01.006","article-title":"Wider or deeper: Revisiting the ResNet model for visual recognition","volume":"90","author":"Wu","year":"2019","journal-title":"Pattern Recognition"},{"key":"2026032615364252300_ref735","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298607","article-title":"Efficient sparse-to-dense optical flow estimation using a learned basis and layers","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Wulff","year":"2015"},{"key":"2026032615364252300_ref736","article-title":"TORCS: The open racing car simulator","author":"Wymann","year":"2015"},{"key":"2026032615364252300_ref737","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.534","article-title":"Learning to track: Online multi-object tracking by decision making","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Xiang","year":"2015"},{"key":"2026032615364252300_ref738","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2015.7298800","article-title":"Data-driven 3D voxel patterns for object category recognition","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xiang","year":"2015"},{"key":"2026032615364252300_ref739","first-page":"924","article-title":"Subcategory-aware convolutional neural networks for object proposals and detection","volume-title":"Proc. of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Xiang","year":"2017"},{"issue":"5","key":"2026032615364252300_ref740","doi-asserted-by":"crossref","first-page":"114:1","DOI":"10.1145\/1618452.1618460","article-title":"Image-based street-side city modeling","volume":"28","author":"Xiao","year":"2009","journal-title":"ACM Trans. on Graphics"},{"key":"2026032615364252300_ref741","article-title":"Multiple view semantic segmentation for street view images","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Xiao","year":"2009"},{"key":"2026032615364252300_ref742","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.401","article-title":"Semantic instance annotation of street scenes by 3D to 2D label transfer","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xie","year":"2016"},{"key":"2026032615364252300_ref743","article-title":"Multi-object tracking through occlusions by local tracklets filtering and global tracklets association with detection responses","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xing","year":"2009"},{"key":"2026032615364252300_ref744","first-page":"2609","article-title":"3-D scene analysis via sequenced predictions over points and regions","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Xiong","year":"2011"},{"key":"2026032615364252300_ref745","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00902","article-title":"UPSNet: A unified panoptic segmentation network","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xiong","year":"2019"},{"key":"2026032615364252300_ref746","first-page":"3530","article-title":"End-to-end learning of driving models from large-scale video datasets","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xu","year":"2017"},{"key":"2026032615364252300_ref747","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00563","article-title":"Multi-scale geometric consistency guided multi-view stereo","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Xu","year":"2019"},{"key":"2026032615364252300_ref748","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2013.243","article-title":"Robust monocular epipolar flow estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yamaguchi","year":"2013"},{"key":"2026032615364252300_ref749","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33715-4_4","article-title":"Continuous Markov random fields for robust stereo estimation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yamaguchi","year":"2012"},{"key":"2026032615364252300_ref750","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10602-1_49","article-title":"Efficient joint segmentation, occlusion labeling, stereo and flow estimation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yamaguchi","year":"2014"},{"issue":"10","key":"2026032615364252300_ref751","doi-asserted-by":"crossref","DOI":"10.3390\/s18103337","article-title":"SECOND: Sparsely embedded convolutional detection","volume":"18","author":"Yan","year":"2018","journal-title":"Sensors"},{"key":"2026032615364252300_ref752","first-page":"6043","article-title":"Craft objects from images","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yang","year":"2016"},{"key":"2026032615364252300_ref753","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2011.5995587","article-title":"Learning affinities and dependencies for multi-target tracking using a CRF model","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yang","year":"2011"},{"key":"2026032615364252300_ref754","article-title":"An online learned CRF model for multi-target tracking","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yang","year":"2012"},{"key":"2026032615364252300_ref755","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.234","article-title":"Exploit all the layers: Fast and accurate CNN object detector with scale dependent pooling and cascaded rejection classifiers","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yang","year":"2016"},{"key":"2026032615364252300_ref756","first-page":"660","article-title":"Seg-Stereo: Exploiting semantic information for disparity estimation","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yang","year":"2018"},{"key":"2026032615364252300_ref757","first-page":"835","article-title":"Deep virtual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yang","year":"2018"},{"key":"2026032615364252300_ref758","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1016\/j.comnet.2018.02.026","article-title":"Deep detection network for real-life traffic sign in vehicular networks","volume":"136","author":"Yang","year":"2018","journal-title":"Computer Networks"},{"key":"2026032615364252300_ref759","first-page":"3522","article-title":"Recognizing proxemics in personal photos","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yang","year":"2012"},{"key":"2026032615364252300_ref760","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01237-3_47","article-title":"MVSNet: Depth inference for unstructured multi-view stereo","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yao","year":"2018"},{"key":"2026032615364252300_ref761","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00567","article-title":"Recurrent MVSNet for high-resolution multi-view stereo depth inference","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yao","year":"2019"},{"key":"2026032615364252300_ref762","first-page":"6044","article-title":"Hierarchical discrete distribution decomposition for match density estimation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yin","year":"2019"},{"key":"2026032615364252300_ref763","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.155","article-title":"Online multi-object tracking via structural constraint event aggregation","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yoon","year":"2016"},{"key":"2026032615364252300_ref764","doi-asserted-by":"crossref","DOI":"10.1109\/WACV.2015.12","article-title":"Bayesian multi-object tracking using motion context from multiple objects","volume-title":"Proc. of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Yoon","year":"2015"},{"key":"2026032615364252300_ref765","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-48881-3_3","article-title":"POI: Multiple object tracking with high performance detection and appearance feature","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV) Workshops","author":"Yu","year":"2016"},{"key":"2026032615364252300_ref766","article-title":"Multi-scale context aggregation by dilated convolutions","volume-title":"Proc. of the International Conf. on Learning Representations (ICLR)","author":"Yu","year":"2016"},{"key":"2026032615364252300_ref767","unstructured":"Yu, F., W.Xian, Y.Chen, F.Liu, M.Liao, V.Madhavan, and T.Darrell (2018). \u201cBDD100K: A diverse driving video database with scalable annotation tooling\u201d. arXiv: 1805.04687[cs.CV]."},{"key":"2026032615364252300_ref768","unstructured":"Yu, F., W.Xian, Y.Chen, F.Liu, M.Liao, V.Madhavan, and T.Darrell (2019). \u201cBerkeley DeepDrive\u201d. https:\/\/bdd-data.berkeley.edu\/. Online: accessed 05-June-2019."},{"key":"2026032615364252300_ref769","first-page":"1722","article-title":"Semantic alignment of LiDAR data at city scale","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Yu","year":"2015"},{"key":"2026032615364252300_ref770","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-49409-8_1","article-title":"Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Yu","year":"2016"},{"key":"2026032615364252300_ref771","first-page":"214","article-title":"A duality based approach for realtime TV-L1 optical flow","volume-title":"Proc. of the DAGM Symposium on Pattern Recognition (DAGM)","author":"Zach","year":"2007"},{"key":"2026032615364252300_ref772","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1007\/978-3-540-74936-3_22","volume-title":"Pattern Recognition Letters","author":"Zach","year":"2007"},{"key":"2026032615364252300_ref773","article-title":"A globally optimal algorithm for robust TV-L1 range image integration","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zach","year":"2007"},{"key":"2026032615364252300_ref774","article-title":"GMCP-tracker: Global multi-object tracking using generalized minimum clique graphs","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Zamir","year":"2012"},{"issue":"65","key":"2026032615364252300_ref775","first-page":"1","article-title":"Stereo matching by training a convolutional neural network to compare image patches","volume":"17","author":"\u017dbontar","year":"2016","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026032615364252300_ref776","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2011.6126474","article-title":"Adaptive deconvolutional networks for mid and high level feature learning","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zeiler","year":"2011"},{"key":"2026032615364252300_ref777","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.310","article-title":"Camera pose voting for large-scale image-based localization","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zeisl","year":"2015"},{"key":"2026032615364252300_ref778","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2013.379","article-title":"Understanding high-level semantics by modeling traffic patterns","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zhang","year":"2013"},{"key":"2026032615364252300_ref779","doi-asserted-by":"crossref","DOI":"10.1109\/IROS.2014.6943269","article-title":"Real-time depth enhanced monocular odometry","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Zhang","year":"2014"},{"key":"2026032615364252300_ref780","doi-asserted-by":"crossref","DOI":"10.15607\/RSS.2014.X.007","article-title":"LOAM: Lidar odometry and mapping in real-time","volume-title":"Proc. Robotics: Science and Systems (RSS)","author":"Zhang","year":"2014"},{"key":"2026032615364252300_ref781","doi-asserted-by":"crossref","DOI":"10.1109\/ICRA.2015.7139486","article-title":"Visual-lidar odometry and mapping: Low-drift, robust, and fast","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Zhang","year":"2015"},{"key":"2026032615364252300_ref782","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2008.4587584","article-title":"Global data association for multi-object tracking using network flows","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang","year":"2008"},{"key":"2026032615364252300_ref783","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-319-10590-1_54","article-title":"Part-based R-CNNs for fine-grained category detection","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Zhang","year":"2014"},{"key":"2026032615364252300_ref784","article-title":"Extrinsic calibration of a camera and laser range finder","volume-title":"Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS)","author":"Zhang","year":"2004"},{"key":"2026032615364252300_ref785","first-page":"1850","article-title":"Sensor fusion for semantic segmentation of urban scenes","volume-title":"Proc. IEEE International Conf. on Robotics and Automation (ICRA)","author":"Zhang","year":"2015"},{"key":"2026032615364252300_ref786","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.141","article-title":"How far are we from solving pedestrian detection?","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang","year":"2016"},{"key":"2026032615364252300_ref787","first-page":"584","article-title":"Led: Localization-quality estimation embedded detector","volume-title":"Proc. IEEE International Conf. on Image Processing (ICIP)","author":"Zhang","year":"2018"},{"key":"2026032615364252300_ref788","article-title":"Efficient inference for fully-connected CRFs with stationarity","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang","year":"2012"},{"key":"2026032615364252300_ref789","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2016.79","article-title":"Instance-level segmentation for autonomous driving with deep densely connected MRFs","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang","year":"2016"},{"key":"2026032615364252300_ref790","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.300","article-title":"Monocular object instance segmentation and depth ordering with CNNs","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zhang","year":"2015"},{"key":"2026032615364252300_ref791","first-page":"6230","article-title":"Pyramid scene parsing network","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhao","year":"2017"},{"key":"2026032615364252300_ref792","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.179","article-title":"Conditional random fields as recurrent neural networks","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zheng","year":"2015"},{"key":"2026032615364252300_ref793","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2009.5206749","article-title":"Tour the world: Building a web-scale landmark recognition engine","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zheng","year":"2009"},{"key":"2026032615364252300_ref794","article-title":"Learning deep features for scene recognition using places database","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Zhou","year":"2014"},{"key":"2026032615364252300_ref795","doi-asserted-by":"crossref","unstructured":"Zhou, B., P.Kr\u00e4henb\u00fchl, and V.Koltun (2019). \u201cDoes computer vision matter for action?\u201d arXiv: 1905.12887[cs.CV].","DOI":"10.1126\/scirobotics.aaw6661"},{"key":"2026032615364252300_ref796","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2015.254","article-title":"Exploiting object similarity in 3D reconstruction","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zhou","year":"2015"},{"key":"2026032615364252300_ref797","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00472","article-title":"VoxelNet: End-to-end learning for point cloud based 3D object detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhou","year":"2018"},{"issue":"10","key":"2026032615364252300_ref798","doi-asserted-by":"crossref","first-page":"2584","DOI":"10.1109\/TITS.2017.2658662","article-title":"Overview of environment perception for intelligent vehicles","volume":"18","author":"Zhu","year":"2017","journal-title":"IEEE Trans. on Intelligent Transportation Systems (T-ITS)"},{"key":"2026032615364252300_ref799","first-page":"379","article-title":"Online multi-object tracking with dual matching attention networks","volume-title":"Proc. of the European Conf. on Computer Vision (ECCV)","author":"Zhu","year":"2018"},{"key":"2026032615364252300_ref800","first-page":"4558","article-title":"Image gradient-based joint direct visual odometry for stereo camera","volume-title":"Proc. of the International Joint Conf. on Artificial Intelligence (IJCAI)","author":"Zhu","year":"2017"},{"key":"2026032615364252300_ref801","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2018.00753","article-title":"Towards high performance video object detection","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu","year":"2018"},{"key":"2026032615364252300_ref802","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.52","article-title":"Flow-guided feature aggregation for video object detection","volume-title":"Proc. of the IEEE International Conf. on Computer Vision (ICCV)","author":"Zhu","year":"2017"},{"key":"2026032615364252300_ref803","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.441","article-title":"Deep feature flow for video recognition","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu","year":"2017"},{"key":"2026032615364252300_ref804","article-title":"Scaleadaptive deconvolutional regression network for pedestrian detection","volume-title":"Proc. of the Asian Conf. on Computer Vision (ACCV)","author":"Zhu","year":"2016"},{"key":"2026032615364252300_ref805","first-page":"2110","article-title":"Traffic-sign detection and classification in the wild","volume-title":"Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu","year":"2016"},{"key":"2026032615364252300_ref806","first-page":"1506","article-title":"RelationNet: Learning deep-aligned representation for semantic image segmentation","volume-title":"Proc. of the International Conf. on Pattern Recognition (ICPR)","author":"Zhuang","year":"2018"},{"key":"2026032615364252300_ref807","first-page":"3698","article-title":"Dense relation network: Learning consistent and context-aware representation for semantic image segmentation","volume-title":"Proc. IEEE International Conf. on Image Processing (ICIP)","author":"Zhuang","year":"2018"},{"issue":"11","key":"2026032615364252300_ref808","doi-asserted-by":"crossref","first-page":"2608","DOI":"10.1109\/TPAMI.2013.87","article-title":"Detailed 3D representations for object recognition and modeling","volume":"35","author":"Zia","year":"2013","journal-title":"IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI)"},{"issue":"2","key":"2026032615364252300_ref809","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1007\/s11263-014-0780-y","article-title":"Towards scene understanding with detailed 3D object representations","volume":"112","author":"Zia","year":"2015","journal-title":"International Journal of Computer Vision (IJCV)"},{"issue":"2","key":"2026032615364252300_ref810","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/MITS.2014.2306552","article-title":"Making bertha drive - An autonomous journey on a historic route","volume":"6","author":"Ziegler","year":"2014","journal-title":"Proc. IEEE Intelligent Transportation Systems Magazine (ITSM)"},{"issue":"3","key":"2026032615364252300_ref811","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1007\/s11263-011-0422-6","article-title":"Optic flow in harmony","volume":"93","author":"Zimmer","year":"2011","journal-title":"International Journal of Computer Vision (IJCV)"}],"container-title":["Foundations and Trends\u00ae in Computer Graphics and Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftcgv\/article-pdf\/12\/1-3\/1\/10865084\/0600000079en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftcgv\/article-pdf\/12\/1-3\/1\/10865084\/0600000079en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T14:09:11Z","timestamp":1777471751000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftcgv\/article\/12\/1-3\/1\/1319352\/Computer-Vision-for-Autonomous-VehiclesProblems"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,6]]},"references-count":811,"journal-issue":{"issue":"1-3","published-print":{"date-parts":[[2020,7,6]]}},"URL":"https:\/\/doi.org\/10.1561\/0600000079","relation":{},"ISSN":["1572-2740","1572-2759"],"issn-type":[{"value":"1572-2740","type":"print"},{"value":"1572-2759","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,6]]}}}