{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:35:00Z","timestamp":1760240100753,"version":"build-2065373602"},"reference-count":32,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2019,2,15]],"date-time":"2019-02-15T00:00:00Z","timestamp":1550188800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Radio-frequency (RF) tomographic imaging is a promising technique for inferring multi-dimensional physical space by processing RF signals traversed across a region of interest. Tensor-based approaches for tomographic imaging are superior at detecting the objects within higher dimensional spaces. The recently-proposed tensor sensing approach based on the transform tensor model achieves a lower error rate and faster speed than the previous tensor-based compress sensing approach. However, the running time of the tensor sensing approach increases exponentially with the dimension of tensors, thus not being very practical for big tensors. In this paper, we address this problem by exploiting massively-parallel GPUs. We design, implement, and optimize the tensor sensing approach on an NVIDIA Tesla GPU and evaluate the performance in terms of the running time and recovery error rate. Experimental results show that our GPU tensor sensing is as accurate as the CPU counterpart with an average of     44.79 \u00d7     and up to     84.70 \u00d7     speedups for varying-sized synthetic tensor data. For IKEA Model 3D model data of a smaller size, our GPU algorithm achieved 15.374\u00d7 speedup over the CPU tensor sensing. We further encapsulate the GPU algorithm into an open-source library, called cuTensorSensing (CUDA Tensor Sensing), which can be used for efficient RF tomographic imaging.<\/jats:p>","DOI":"10.3390\/fi11020046","type":"journal-article","created":{"date-parts":[[2019,2,17]],"date-time":"2019-02-17T22:11:50Z","timestamp":1550441510000},"page":"46","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Efficient Tensor Sensing for RF Tomographic Imaging on GPUs"],"prefix":"10.3390","volume":"11","author":[{"given":"Da","family":"Xu","sequence":"first","affiliation":[{"name":"Department of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8440-9924","authenticated-orcid":false,"given":"Tao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Institute for Advanced Communication and Data Science, Shanghai 200444, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3361","DOI":"10.1007\/s11277-017-4061-2","article-title":"Multi-dimensional wireless tomography using tensor-based compressed sensing","volume":"96","author":"Matsuda","year":"2017","journal-title":"Wirel. Pers. Commun."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"621","DOI":"10.1109\/TMC.2009.174","article-title":"Radio tomographic imaging with wireless networks","volume":"9","author":"Wilson","year":"2010","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"6474","DOI":"10.1109\/TWC.2016.2585141","article-title":"Ultrawideband Tomographic Imaging in Uncalibrated Networks","volume":"15","author":"Beck","year":"2016","journal-title":"IEEE Trans. Wirel. Commun."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Deng, T., Qian, F., Liu, X.Y., Zhang, M., and Walid, A. (2018, January 10\u201312). Tensor Sensing for Rf Tomographic Imaging. Proceedings of the 2018 IEEE International Conference on Multimedia and Expo (ICME), Miami, FL, USA.","DOI":"10.1109\/ICME.2018.8486609"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Cui, H., Zhang, H., Ganger, G.R., Gibbons, P.B., and Xing, E.P. (2016, January 18\u201321). GeePS: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server. Proceedings of the Eleventh European Conference on Computer Systems, London, UK.","DOI":"10.1145\/2901318.2901323"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Brito, R., Fong, S., Song, W., Cho, K., Bhatt, C., and Korzun, D. (2017). Detecting Unusual Human Activities Using GPU-Enabled Neural Network and Kinect Sensors. Internet of Things and Big Data Technologies for Next Generation Healthcare, Springer.","DOI":"10.1007\/978-3-319-49736-5_15"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Campos, V., Sastre, F., Yag\u00fces, M., Torres, J., and Gir\u00f3-i Nieto, X. (2017, January 14\u201317). Scaling a Convolutional Neural Network for Classification of Adjective Noun Pairs with TensorFlow on GPU Clusters. Proceedings of the 17th IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing, Madrid, Spain.","DOI":"10.1109\/CCGRID.2017.110"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1109\/TKDE.2017.2745562","article-title":"Frog: Asynchronous graph processing on GPU with hybrid coloring model","volume":"30","author":"Shi","year":"2018","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1149","DOI":"10.1109\/TPDS.2016.2611659","article-title":"Optimizing Graph Processing on GPUs","volume":"28","author":"Zhong","year":"2017","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Pan, Y., Wang, Y., Wu, Y., Yang, C., and Owens, J.D. (June, January 29). Multi-GPU graph analytics. Proceedings of the 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS), Orlando, FL, USA.","DOI":"10.1109\/IPDPS.2017.117"},{"key":"ref_11","first-page":"1","article-title":"SMOTE-GPU: Big Data preprocessing on commodity hardware for imbalanced classification","volume":"6","author":"Lastra","year":"2017","journal-title":"Prog. Artif. Intell."},{"key":"ref_12","first-page":"1","article-title":"Real-time big data stream processing using GPU with spark over hadoop ecosystem","volume":"46","author":"Rathore","year":"2017","journal-title":"Int. J. Parallel Program."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"5096","DOI":"10.1109\/TMTT.2017.2766060","article-title":"GPU-Accelerated Enhanced Resolution 3-D SAR Imaging With Dynamic Metamaterial Antennas","volume":"65","author":"Devadithya","year":"2017","journal-title":"IEEE Trans. Microw. Theory Tech."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Verma, K., Szewc, K., and Wille, R. (2017, January 12\u201314). Advanced load balancing for SPH simulations on multi-GPU architectures. Proceedings of the High Performance Extreme Computing Conference (HPEC), Waltham, MA, USA.","DOI":"10.1109\/HPEC.2017.8091093"},{"key":"ref_15","unstructured":"(2019, February 15). Intelligent Information Processing (IIP) Lab. Available online: http:\/\/www.findai.com."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Kanso, M.A., and Rabbat, M.G. (2009). Compressed RF tomography for wireless sensor networks: Centralized and decentralized approaches. International Conference on Distributed Computing in Sensor Systems, Springer.","DOI":"10.1007\/978-3-642-02085-8_13"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1769","DOI":"10.1109\/TMC.2011.31","article-title":"Compressive cooperative sensing and mapping in mobile networks","volume":"10","author":"Mostofi","year":"2011","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, Q., Schonfeld, D., and Friedland, S. (2013, January 15\u201319). Generalized tensor compressive sensing. Proceedings of the 2013 IEEE International Conference on Multimedia and Expo (ICME), San Jose, CA, USA.","DOI":"10.1109\/ICME.2013.6607560"},{"key":"ref_19","unstructured":"Liu, X.Y., and Wang, X. (arXiv, 2017). Fourth-order tensors with multidimensional discrete transforms, arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"122","DOI":"10.1109\/TC.2015.2417545","article-title":"Energy-efficient eDRAM-based on-chip storage architecture for GPGPUs","volume":"65","author":"Jing","year":"2016","journal-title":"IEEE Trans. Comput."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2744202","article-title":"Buddy SM: Sharing Pipeline Front-End for Improved Energy Efficiency in GPGPUs","volume":"12","author":"Zhang","year":"2015","journal-title":"ACM Trans. Archit. Code Optim. (TACO)"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1563","DOI":"10.1007\/s11227-015-1378-z","article-title":"Efficient graph computation on hybrid CPU and GPU systems","volume":"71","author":"Zhang","year":"2015","journal-title":"J. Supercomput."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2951","DOI":"10.1016\/j.jpdc.2014.07.004","article-title":"CUIRRE: An open-source library for load balancing and characterizing irregular applications on GPUs","volume":"74","author":"Zhang","year":"2014","journal-title":"J. Parallel Distrib. Comput."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, T., Tong, W., Shen, W., Peng, J., and Niu, Z. (2016). Efficient Graph Mining on Heterogeneous Platforms in the Cloud. Cloud Computing, Security, Privacy in New Computing Environments, Springer.","DOI":"10.1007\/978-3-319-69605-8_2"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Nelson, T., Rivera, A., Balaprakash, P., Hall, M., Hovland, P.D., Jessup, E., and Norris, B. (2015, January 1\u20134). Generating efficient tensor contractions for gpus. Proceedings of the 44th International Conference on Parallel Processing (ICPP), Beijing, China.","DOI":"10.1109\/ICPP.2015.106"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Shi, Y., Niranjan, U., Anandkumar, A., and Cecka, C. (2016, January 16\u201319). Tensor contractions with extended BLAS kernels on CPU and GPU. Proceedings of the IEEE 23rd International Conference on High Performance Computing (HiPC), Kochi, India.","DOI":"10.1109\/HiPC.2016.031"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1135","DOI":"10.1109\/TPDS.2010.194","article-title":"Nonnegative tensor factorization accelerated using GPGPU","volume":"22","author":"Antikainen","year":"2011","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1016\/j.cpc.2014.12.013","article-title":"An efficient tensor transpose algorithm for multicore CPU, Intel Xeon Phi, and NVidia Tesla GPU","volume":"189","author":"Lyakh","year":"2015","journal-title":"Comput. Phys. Commun."},{"key":"ref_29","unstructured":"Hynninen, A.P., and Lyakh, D.I. (arXiv, 2017). cuTT: A high-performance tensor transpose library for CUDA compatible GPUs, arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Rogers, D.M. (2016, January 17\u201321). Efficient primitives for standard tensor linear algebra. Proceedings of the XSEDE16 Conference on Diversity, Big Data, and Science at Scale, ACM, Miami, FL, USA.","DOI":"10.1145\/2949550.2949580"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1016\/j.ins.2014.12.004","article-title":"GPUTENSOR: Efficient tensor factorization for context-aware recommendations","volume":"299","author":"Zou","year":"2015","journal-title":"Inf. Sci."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, J., Ma, Y., Yan, C., and Vuduc, R. (2016, January 13\u201318). Optimizing sparse tensor times matrix on multi-core and many-core architectures. Proceedings of the IEEE Workshop on Irregular Applications: Architecture and Algorithms (IA3), Salt Lake City, UT, USA.","DOI":"10.1109\/IA3.2016.010"}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/11\/2\/46\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:32:27Z","timestamp":1760185947000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/11\/2\/46"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,15]]},"references-count":32,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2019,2]]}},"alternative-id":["fi11020046"],"URL":"https:\/\/doi.org\/10.3390\/fi11020046","relation":{},"ISSN":["1999-5903"],"issn-type":[{"type":"electronic","value":"1999-5903"}],"subject":[],"published":{"date-parts":[[2019,2,15]]}}}