{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T14:31:38Z","timestamp":1778509898732,"version":"3.51.4"},"reference-count":39,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T00:00:00Z","timestamp":1778457600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Point cloud upsampling is a fundamental task in 3D vision, yet most existing methods adopt a global and uniform strategy, which is computationally inefficient and fails to address the need for region-specific refinement. To address this challenge, we propose PartSPUNet, a novel self-supervised, text-driven point cloud upsampling framework designed to enhance robotic perception through task-oriented local refinement. Inspired by the human cognitive process where high-level language instructions guide visual attention to specific regions of interest, our method allows an operator to use intuitive natural language prompts to direct the upsampling process. Specifically, PartSPUNet leverages a pretrained vision\u2013language model to zero-shot localize the user-specified semantic part within a sparse point cloud. It then performs geometry-aware densification exclusively on this target region, recovering rich geometric details while preserving the global structure. Experimental results demonstrate that our approach significantly outperforms existing methods in reconstructing specified areas, offering a powerful and intuitive tool for enhancing the 3D perception pipeline in intelligent robotic systems.<\/jats:p>","DOI":"10.3390\/jimaging12050204","type":"journal-article","created":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T13:10:42Z","timestamp":1778505042000},"page":"204","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Self-Supervised Text-Driven Point Cloud Upsampling via Semantic Text Guidance"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8632-5540","authenticated-orcid":false,"given":"Zhiyong","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meiling","family":"Qiu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuo","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2130-9122","authenticated-orcid":false,"given":"Ruyu","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Hangzhou Normal University, Hangzhou 310000, China"},{"name":"Department of Technology, Management and Economics, Technical University of Denmark, 2800 Lyngby, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7844-6035","authenticated-orcid":false,"given":"Jianhua","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengyong","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Meng, S., Wang, S., and Ren, Y. (2026). Accelerating Point Cloud Computation via Memory in Embedded Structured Light Cameras. J. Imaging, 12.","DOI":"10.3390\/jimaging12020091"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bao, J., Kong, L., and Wang, W. (2026). SFD-ADNet: Spatial\u2013Frequency Dual-Domain Adaptive Deformation for Point Cloud Data Augmentation. J. Imaging, 12.","DOI":"10.3390\/jimaging12020058"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1944","DOI":"10.1109\/LRA.2025.3528229","article-title":"Funabot-sleeve: A wearable device employing McKibben artificial muscles for haptic sensation in the forearm","volume":"10","author":"Peng","year":"2025","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, T., and Wang, H. (2025). Pov9D: Point Cloud-Based Open-Vocabulary 9D Object Pose Estimation. J. Imaging, 11.","DOI":"10.3390\/jimaging11110380"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhu, X., Zhang, R., and Gao, P. (2025). Adapting CLIP for 3D Understanding. Large Vision-Language Models: Pre-training, Prompting, and Applications, Springer.","DOI":"10.1007\/978-3-031-94969-2_12"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Nguyen, K., Hassan, G.M., and Mian, A. (2025). Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition. Proceedings of the Computer Vision and Pattern Recognition Conference, IEEE.","DOI":"10.1109\/CVPR52734.2025.01581"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3432","DOI":"10.1109\/TMM.2022.3160604","article-title":"Semantic point cloud upsampling","volume":"25","author":"Li","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Qian, Y., Hou, J., Kwong, S., and He, Y. (2020). PUGeo-Net: A geometry-centric network for 3D point cloud upsampling. Proceedings of the European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58529-7_44"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yu, L., Li, X., Fu, C.W., Cohen-Or, D., and Heng, P.A. (2018). Pu-net: Point cloud upsampling network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR.2018.00295"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, R., Li, X., Fu, C.W., Cohen-Or, D., and Heng, P.A. (2019). Pu-gan: A point cloud upsampling adversarial network. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV.2019.00730"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Qian, G., Abualshour, A., Li, G., Thabet, A., and Ghanem, B. (2021). Pu-gcn: Point cloud upsampling using graph convolutional networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR46437.2021.01151"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"5472","DOI":"10.1609\/aaai.v38i6.28356","article-title":"Pointattn: You only need attention for point cloud completion","volume":"Volume 38","author":"Wang","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, J., Chen, J., Shi, Y., Ling, N., and Yin, B. (2023). Sspu-net: A structure sensitive point cloud upsampling network with multi-scale spatial refinement. Proceedings of the 31st ACM International Conference on Multimedia, Association for Computing Machinery.","DOI":"10.1145\/3581783.3613807"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Feng, W., Li, J., Cai, H., Luo, X., and Zhang, J. (2022). Neural points: Point cloud representation with neural fields for arbitrary upsampling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.01808"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"7458","DOI":"10.1609\/aaai.v40i9.37685","article-title":"PUFM: Efficient Point Cloud Upsampling via Flow Matching","volume":"Volume 40","author":"Liu","year":"2026","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1109\/TIP.2025.3571680","article-title":"SPU+: Dimension Folding for Semantic Point Cloud Upsampling","volume":"34","author":"Li","year":"2025","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems, Neural Information Processing Systems Foundation, Inc."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"7389","DOI":"10.1109\/TIP.2022.3222918","article-title":"PUFA-GAN: A frequency-aware generative adversarial network for 3D point cloud upsampling","volume":"31","author":"Liu","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3206","DOI":"10.1109\/TVCG.2021.3058311","article-title":"Meta-PU: An arbitrary-scale upsampling network for point cloud","volume":"28","author":"Ye","year":"2021","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Du, H., Yan, X., Wang, J., Xie, D., and Pu, S. (2022). Point cloud upsampling via cascaded refinement network. Proceedings of the Asian Conference on Computer Vision, ACCV.","DOI":"10.1007\/978-3-031-26319-4_7"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Qiu, S., Anwar, S., and Barnes, N. (2022). Pu-transformer: Point cloud upsampling transformer. Proceedings of the Asian Conference on Computer Vision, ACCV.","DOI":"10.1007\/978-3-031-26319-4_20"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1626","DOI":"10.1609\/aaai.v38i2.27929","article-title":"Arbitrary-scale point cloud upsampling by voxel-based network with latent geometric-consistent learning","volume":"Volume 38","author":"Du","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"6489","DOI":"10.1109\/TCSVT.2024.3370001","article-title":"Pu-mask: 3d point cloud upsampling via an implicit virtual mask","volume":"34","author":"Liu","year":"2024","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Liu, R., Wang, C., Zhang, X., Zhang, J., and Liu, X. (2025). SPRGAN: Streamlined Progressive Refinement for Adversarial Point Cloud Video Upsampling. Proceedings of the ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE.","DOI":"10.1109\/ICASSP49660.2025.10888638"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"4062","DOI":"10.1109\/TIP.2022.3166627","article-title":"VPU: A video-based point cloud upsampling framework","volume":"31","author":"Wang","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"4686","DOI":"10.1109\/TCSVT.2021.3104304","article-title":"Sequential point cloud upsampling by exploiting multi-scale temporal dependency","volume":"31","author":"Wang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, B., Yang, S., Chen, H., Yang, C., Jia, J., and Jiang, G. (2025). Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise Suppression. Proceedings of the Computer Vision and Pattern Recognition Conference, IEEE.","DOI":"10.1109\/CVPR52734.2025.01583"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Du, Y., Zhao, Z., Su, S., Golluri, S., Zheng, H., Yao, R., and Wang, C. (2025). SuperPC: A single diffusion model for point cloud completion, upsampling, denoising, and colorization. Proceedings of the Computer Vision and Pattern Recognition Conference, IEEE.","DOI":"10.1109\/CVPR52734.2025.01580"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Qu, W., Shao, Y., Meng, L., Huang, X., and Xiao, L. (2024). A conditional denoising diffusion probabilistic model for point cloud upsampling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.01964"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"4964","DOI":"10.1109\/TVCG.2022.3196334","article-title":"Pu-flow: A point cloud upsampling network with normalizing flows","volume":"29","author":"Mao","year":"2022","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhang, R., Guo, Z., Zhang, W., Li, K., Miao, X., Cui, B., Qiao, Y., Gao, P., and Li, H. (2022). Pointclip: Point cloud understanding by clip. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.00836"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Huang, T., Dong, B., Yang, Y., Huang, X., Lau, R.W., Ouyang, W., and Zuo, W. (2023). Clip2point: Transfer clip to point cloud classification with image-depth pre-training. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV51070.2023.02025"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"5694","DOI":"10.1609\/aaai.v39i6.32607","article-title":"CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment","volume":"Volume 39","author":"Liu","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Jiao, S., Dong, H., Yin, Y., Jie, Z., Qian, Y., Zhao, Y., Shi, H., and Wei, Y. (2025). CLIP-GS: Unifying vision-language representation with 3D Gaussian splatting. Proceedings of the IEEE\/CVF International Conference on Computer Vision, IEEE.","DOI":"10.1109\/ICCV51701.2025.00444"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"8767","DOI":"10.1109\/TCSVT.2025.3551084","article-title":"MV-CLIP: Multi-view CLIP for zero-shot 3D shape recognition","volume":"35","author":"Song","year":"2025","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., and Lu, J. (2022). Point-bert: Pre-training 3d point cloud transformers with masked point modeling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52688.2022.01871"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Xue, L., Yu, N., Zhang, S., Panagopoulou, A., Li, J., Mart\u00edn-Mart\u00edn, R., Wu, J., Xiong, C., Xu, R., and Niebles, J.C. (2024). Ulip-2: Towards scalable multimodal pre-training for 3d understanding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE.","DOI":"10.1109\/CVPR52733.2024.02558"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"10734","DOI":"10.1609\/aaai.v39i10.33166","article-title":"Position-aware Guided Point Cloud Completion with CLIP Model","volume":"Volume 39","author":"Zhou","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Xu, R., Yang, S., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D. (2025). Pointllm-v2: Empowering large language models to better understand point clouds. IEEE Trans. Pattern Anal. Mach. Intell., 1\u201315.","DOI":"10.1109\/TPAMI.2025.3590784"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/5\/204\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T13:57:21Z","timestamp":1778507841000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/5\/204"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,11]]},"references-count":39,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["jimaging12050204"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12050204","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,11]]}}}