{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:10:22Z","timestamp":1760242222369,"version":"build-2065373602"},"reference-count":32,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2017,1,10]],"date-time":"2017-01-10T00:00:00Z","timestamp":1484006400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61403060","61332017","61602202","61402192","61603146"],"award-info":[{"award-number":["61403060","61332017","61602202","61402192","61603146"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"the six talent peaks project in Jiangsu Province","award":["XYDXXJS-011","XYDXXJS-012"],"award-info":[{"award-number":["XYDXXJS-011","XYDXXJS-012"]}]},{"name":"National Science and technology support project","award":["2015BAK04B05"],"award-info":[{"award-number":["2015BAK04B05"]}]},{"name":"National Natural Science Foundation of Jiangsu Province","award":["BK20160427","BK20160428"],"award-info":[{"award-number":["BK20160427","BK20160428"]}]},{"name":"Dalian Science and Technology Planning Project","award":["2015A11GX021"],"award-info":[{"award-number":["2015A11GX021"]}]},{"DOI":"10.13039\/501100010014","name":"Six talent peaks project in Jiangsu Province","doi-asserted-by":"publisher","award":["XYDXXJS-012"],"award-info":[{"award-number":["XYDXXJS-012"]}],"id":[{"id":"10.13039\/501100010014","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Visual object tracking technology is one of the key issues in computer vision. In this paper, we propose a visual object tracking algorithm based on cross-modality featuredeep learning using Gaussian-Bernoulli deep Boltzmann machines (DBM) with RGB-D sensors. First, a cross-modality featurelearning network based on aGaussian-Bernoulli DBM is constructed, which can extract cross-modality features of the samples in RGB-D video data. Second, the cross-modality features of the samples are input into the logistic regression classifier, andthe observation likelihood model is established according to the confidence score of the classifier. Finally, the object tracking results over RGB-D data are obtained using aBayesian maximum a posteriori (MAP) probability estimation algorithm. The experimental results show that the proposed method has strong robustness to abnormal changes (e.g., occlusion, rotation, illumination change, etc.). The algorithm can steadily track multiple targets and has higher accuracy.<\/jats:p>","DOI":"10.3390\/s17010121","type":"journal-article","created":{"date-parts":[[2017,1,10]],"date-time":"2017-01-10T10:16:18Z","timestamp":1484043378000},"page":"121","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Visual Object Tracking Based on Cross-Modality Gaussian-Bernoulli Deep Boltzmann Machines with RGB-D Sensors"],"prefix":"10.3390","volume":"17","author":[{"given":"Mingxin","family":"Jiang","sequence":"first","affiliation":[{"name":"Faculty of Computer and Software Engineering, Huaiyin Institute of Technology, Huai\u2019an 223003, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhigeng","family":"Pan","sequence":"additional","affiliation":[{"name":"Digital Media &amp;Interaction Research Center, Hangzhou Normal University, Hangzhou 310012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenzhou","family":"Tang","sequence":"additional","affiliation":[{"name":"College of Physics and Electronic Information Engineering, Wenzhou University, Wenzhou 325035, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2017,1,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"762","DOI":"10.1109\/TCSVT.2015.2416555","article-title":"Transductive People Tracking in Unconstrained Surveillance","volume":"26","author":"Coppi","year":"2016","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1408","DOI":"10.1109\/ACCESS.2015.2471935","article-title":"Track Detection of Low Observable Targets Using a Motion Model","volume":"3","author":"Daniel","year":"2015","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1007\/s11042-009-0368-7","article-title":"Dynamic tracking re-adjustment: A method for automatic tracking recovery in complex visual environments","volume":"50","author":"Doulamis","year":"2010","journal-title":"Multimed. Tools Appl."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"26877","DOI":"10.3390\/s151026877","article-title":"Visual tracking based on extreme learning machine and sparse representation","volume":"15","author":"Wang","year":"2015","journal-title":"Sensors"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2144","DOI":"10.1109\/TCYB.2015.2466437","article-title":"Visual Tracking via Random Walks on Graph Model","volume":"46","author":"Li","year":"2016","journal-title":"IEEE Trans. Cybern."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Munaro, M., Basso, F., and Menegatti, E. (2012, January 7\u201311). Tracking People within Groups with RGB-D Data. Proceedings of the IEEE International Conference on Intelligent Robots and Systems(IROS), Vilamoura, Portugal.","DOI":"10.1109\/IROS.2012.6385772"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Spinello, L., Luber, M., and Arras, K.O. (2011, January 9\u201313). Tracking people in 3D using a bottom-up top-down people detector. Proceedings of the International Conference on Robotics and Automation (ICRA), Shanghai, China.","DOI":"10.1109\/ICRA.2011.5980085"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Spinello, L., and Arras, K.O. (2011, January 25\u201330). People Detection in RGB-D Data. Proceedings of the IEEE International Conference on Intelligent Robots and Systems(IROS), San Francisco, CA, USA.","DOI":"10.1109\/IROS.2011.6095074"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1023\/B:VISI.0000013087.49260.fb","article-title":"Robust real-time face detection","volume":"57","author":"Viola","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","unstructured":"Navneet, D., and Triggs, B. (2005, January 20\u201326). Histograms of oriented gradients for human detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA."},{"key":"ref_11","unstructured":"Wang, X.Y., Han, T.X., and Yan, S.C. (October, January 29). An HOG-LBP human detector with partial occlusion handling. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Kyoto, Japan."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1834","DOI":"10.1109\/TIP.2015.2510583","article-title":"DeepTrack: Learning Discriminative Feature Representations Online for Robust Visual Tracking","volume":"25","author":"Li","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object Detection with Discriminatively Trained Part Based Models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lu, J.W., Wang, G., Deng, W.H., Moulin, P., and Zhou, J. (2015, January 7\u201312). Multi-manifold deep metric learning for image set classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298717"},{"key":"ref_15","unstructured":"Wang, N.Y., and Yeung, D.Y. (2013, January 5\u201310). Learning a deep compact image representation for visual tracking. Proceedings of the Conference and Workshop on Neural Information Processing Systems (NIPS), South Lake Tahoe, NV, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1424","DOI":"10.1109\/TIP.2015.2403231","article-title":"Video Tracking Using Learned Hierarchical Features","volume":"24","author":"Wang","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1109\/TCSVT.2015.2406231","article-title":"Severely Blurred Object Tracking by Learning Deep Image Representations","volume":"26","author":"Ding","year":"2016","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1409","DOI":"10.1109\/TPAMI.2011.239","article-title":"Tracking-learning-detection","volume":"34","author":"Kalal","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Babenko, B., Yang, M.H., and Belongie, S. (2009, January 20\u201321). Visual tracking with online multiple instance learning. Proceedings of the IEEE International Conference on Computer Vision (CVPR), Miami Beach, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206737"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kwon, J., and Lee, K.M. (2010, January 13\u201318). Visual tracking decomposition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539821"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Song, S., and Xiao, J.X. (2013, January 3\u20136). Tracking revisited using RGBD camera: Unified benchmark and baselines. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Darling Harbour, Sydney.","DOI":"10.1109\/ICCV.2013.36"},{"key":"ref_22","unstructured":"Hinton, G.E., and Sejnowski, T.J. (1983, January 8\u201310). Optimal perceptual inference. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), Los Alamitos, CA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Keronen, S., Cho, K., Raiko, T., Ilin, A., and Palom\u00e4ki, K.J. (2013, January 26\u201330). Gaussian-Bernoulli restricted Boltzmann machines and automatic feature extraction for noise robust missing data mask estimation. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing(ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638964"},{"key":"ref_24","unstructured":"Salakhutdinov, R., and Hinton, G.E. (2009, January 16\u201318). Deep Boltzmann Machines. Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), Clearwater Beach, FL, USA."},{"key":"ref_25","unstructured":"Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A.Y. (July, January 28). Multimodal Deep Learning. Proceedings of the International Conference on Machine Learning (ICML), Bellevue, WA, USA."},{"key":"ref_26","unstructured":"Srivastava, N., and Salakhutdinov, R. (2012, January 3\u20138). Multimodal Learning with Deep Boltzmann Machines. Proceedings of the International Conference and Workshop on Neural Information Processing Systems (NIPS), Lake Tahoe, NV, USA."},{"key":"ref_27","first-page":"404978","article-title":"Visual Object Tracking Based on 2DPCA and ML","volume":"2013","author":"Jiang","year":"2013","journal-title":"Math. Probl. Eng."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Luber, M., Spinello, L., and Arras, K.O. (2011, January 25\u201328). People tracking in RGB-D Data with on-line boosted target models. Proceedings of the IEEE International Conference on Intelligent Robots and Systems (IROS), San Francisco, CA, USA.","DOI":"10.1109\/IROS.2011.6048836"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hare, S., Saffari, A., and Torr, P.H.S. (2011, January 6\u201313). Struck: Structured output tracking with kernels. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126251"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, K., Zhang, L., and Yang, M.-H. (2012, January 7\u201313). Real-Time Compressive Tracking. Proceedings of the European Conference on Computer Vision (ECCV), Firenze, Italy.","DOI":"10.1007\/978-3-642-33712-3_62"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ruan, Y., and Wei, Z. (2016). Real-Time Visual Tracking through Fusion Features. Sensors, 16.","DOI":"10.3390\/s16070949"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"2964","DOI":"10.1016\/j.patcog.2015.02.012","article-title":"Self-taught learning of a deep invariant representation for visual tracking via temporal slowness principle","volume":"48","author":"Kuen","year":"2015","journal-title":"Pattern Recognit."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/17\/1\/121\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:25:47Z","timestamp":1760207147000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/17\/1\/121"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,1,10]]},"references-count":32,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2017,1]]}},"alternative-id":["s17010121"],"URL":"https:\/\/doi.org\/10.3390\/s17010121","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2017,1,10]]}}}