{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T10:00:34Z","timestamp":1772445634085,"version":"3.50.1"},"reference-count":37,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2026,3,1]],"date-time":"2026-03-01T00:00:00Z","timestamp":1772323200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42371457"],"award-info":[{"award-number":["42371457"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42301468"],"award-info":[{"award-number":["42301468"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42371343"],"award-info":[{"award-number":["42371343"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2022J02045"],"award-info":[{"award-number":["2022J02045"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2022J01337"],"award-info":[{"award-number":["2022J01337"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2022J01819"],"award-info":[{"award-number":["2022J01819"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2023J01801"],"award-info":[{"award-number":["2023J01801"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2023J01799"],"award-info":[{"award-number":["2023J01799"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2022J05157"],"award-info":[{"award-number":["2022J05157"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003392","name":"Natural Science Foundation of Fujian Province, China","doi-asserted-by":"publisher","award":["2022J011394"],"award-info":[{"award-number":["2022J011394"]}],"id":[{"id":"10.13039\/501100003392","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Xiamen City\u2019s Leading Project","award":["3502Z20231038"],"award-info":[{"award-number":["3502Z20231038"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["3502Z202373036"],"award-info":[{"award-number":["3502Z202373036"]}]},{"name":"Natural Science Foundation of Xiamen, China","award":["3502Z202371019"],"award-info":[{"award-number":["3502Z202371019"]}]},{"name":"Open Competition for Innovative Projects of Xiamen","award":["3502Z20251012"],"award-info":[{"award-number":["3502Z20251012"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>3D visual grounding is a fundamental task for human\u2013machine interaction, aiming to localize specific objects in complex 3D point clouds based on natural language descriptions. Despite recent advancements, existing Transformer-based architectures often rely on absolute position embeddings and heuristic query initialization, which lack the capacity to capture fine-grained relative spatial dependencies and fail to effectively filter out scene clutter. In this paper, we propose SESQ, a novel framework that synergizes Spatially Aware Encoding and Semantically Guided Querying for 3D grounding. Our approach introduces two key innovations. First, we propose the Rotary Spatially Aware Encoder (RSAE), which incorporates Rotary Position Embeddings (RoPE) into the self-attention layers. By transforming 3D coordinates into a rotary representation, RSAE enables the model to inherently capture relative spatial distances and maintains geometric consistency throughout the encoding stage. Second, a Semantic Query Initialization (SQI) module is designed to initialize object queries by explicitly computing the cross-modal similarity between textual embeddings and visual point cloud features. By replacing traditional heuristic-based selection with semantic-aware alignment, SQI ensures that the decoding process originates from contextually relevant object candidates, significantly reducing the impact of task-irrelevant distractors. Extensive experiments on ScanRefer and ReferIt3D (Nr3D\/Sr3D) benchmarks demonstrate the effectiveness of our framework. Compared to the baseline EDA, our method achieves a significant performance gain of 2.68% in overall Acc@0.5 on ScanRefer, a 4.9% improvement on the challenging Nr3D \u201cHard\u201d subset, and a 1.1% increase in overall Acc@0.25 on Sr3D.<\/jats:p>","DOI":"10.3390\/computers15030145","type":"journal-article","created":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T08:51:57Z","timestamp":1772441517000},"page":"145","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SESQ: Spatially Aware Encoding and Semantically Guided Querying for 3D Grounding"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-1875-7712","authenticated-orcid":false,"given":"Jinyuan","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yundong","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tiancai","family":"Huang","sequence":"additional","affiliation":[{"name":"Xiamen Taqu Information Technology Co., Ltd., Xiamen 361000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mengyun","family":"Cao","sequence":"additional","affiliation":[{"name":"School of Computer Engineering, Jimei University, Xiamen 361000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,3,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhan, Y., Yuan, Y., and Xiong, Z. (2024, January 26\u201327). Mono3dvg: 3d visual grounding in monocular images. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i7.28525"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Yang, Z., Chen, L., Sun, Y., and Li, H. (2024, January 16\u201322). Visual point cloud forecasting enables scalable autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01390"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S.L., Liu, Y., and Li, H. (2024, January 16\u201322). Lmdrive: Closed-loop end-to-end driving with large language models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01432"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1016\/j.neucom.2019.07.081","article-title":"Sequential learning unification controller from human demonstrations for robotic compliant manipulation","volume":"366","author":"Duan","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_5","first-page":"110972","article-title":"Rg-san: Rule-guided spatial awareness network for end-to-end 3d referring expression segmentation","volume":"37","author":"Wu","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Mane, A.M., Weerakoon, D., Subbaraju, V., Sen, S., Sarma, S.E., and Misra, A. (2025, January 11\u201315). Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding. Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.00843"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, D.Z., Chang, A.X., and Nie\u00dfner, M. (2020). Scanrefer: 3d object localization in rgb-d scans using natural language. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-030-58565-5_13"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., and Guibas, L. (2020). Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-030-58452-8_25"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Huang, P.H., Lee, H.H., Chen, H.T., and Liu, T.L. (2021, January 2\u20139). Text-guided graph neural networks for referring 3d instance segmentation. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v35i2.16253"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yuan, Z., Yan, X., Liao, Y., Zhang, R., Wang, S., Li, Z., and Cui, S. (2021, January 10\u201317). Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00181"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yang, Z., Zhang, S., Wang, L., and Luo, J. (2021, January 10\u201317). Sat: 2d semantics assisted training for 3d visual grounding. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00187"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Luo, J., Fu, J., Kong, X., Gao, C., Ren, H., Shen, H., Xia, H., and Liu, S. (2022, January 18\u201324). 3d-sps: Single-stage 3d visual grounding via referred point progressive selection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01596"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Jain, A., Gkanatsios, N., Mediratta, I., and Fragkiadaki, K. (2022). Bottom up top down detection transformers for language grounding in images and point clouds. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-031-20059-5_24"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wu, Y., Cheng, X., Zhang, R., Cheng, Z., and Zhang, J. (2023, January 17\u201324). Eda: Explicit text-decoupling and dense alignment for 3d visual grounding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01843"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"128195","DOI":"10.1016\/j.neucom.2024.128195","article-title":"Revisiting 3D visual grounding with Context-aware Feature Aggregation","volume":"601","author":"Guo","year":"2024","journal-title":"Neurocomputing"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Gao, C., Chen, J., Liu, S., Wang, L., Zhang, Q., and Wu, Q. (2021, January 20\u201325). Room-and-object aware knowledge reasoning for remote embodied referring expression. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00308"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., and Darrell, T. (2016, January 27\u201330). Natural language object retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.493"},{"key":"ref_19","unstructured":"Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., and Berg, T.L. (2016, January 27\u201330). Mattnet: Modular attention network for referring expression comprehension. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"684","DOI":"10.1109\/TPAMI.2019.2911066","article-title":"Learning to compose and reason with language tree structures for visual grounding","volume":"44","author":"Hong","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","unstructured":"Liu, D., Zhang, H., Wu, F., and Zha, Z.J. (November, January 27). Learning to assemble neural module tree networks for visual grounding. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, H., Niu, Y., and Chang, S.F. (2018, January 18\u201323). Grounding referring expressions in images by variational context. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00437"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Deng, J., Yang, Z., Chen, T., Zhou, W., and Li, H. (2021, January 10\u201317). Transvg: End-to-end visual grounding with transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00179"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Liao, Y., Liu, S., Li, G., Wang, F., Chen, Y., Qian, C., and Li, B. (2020, January 13\u201319). A real-time cross-modality correlation filtering method for referring expression comprehension. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01089"},{"key":"ref_25","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_26","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems (NeurIPS), Curran Associates, Inc."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, H., Jiang, L., Jia, J., Torr, P.H., and Koltun, V. (2021, January 10\u201317). Point transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01595"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wu, X., Jiang, L., Wang, P.S., Liu, Z., Liu, X., Qiao, Y., Ouyang, W., He, T., and Zhao, H. (2024, January 16\u201322). Point Transformer V3: Simpler Faster Stronger. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00463"},{"key":"ref_29","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_30","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., and Nie\u00dfner, M. (2017, January 22\u201325). Scannet: Richly-annotated 3d reconstructions of indoor scenes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.261"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., and Carion, N. (2021, January 10\u201317). Mdetr-modulated detection for end-to-end multi-modal understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Qian, Z., Ma, Y., Lin, Z., Ji, J., Zheng, X., Sun, X., and Ji, R. (2024). Multi-branch Collaborative Learning Network for 3D Visual Grounding. Proceedings of the European Conference on Computer Vision (ECCV), Springer.","DOI":"10.1007\/978-3-031-72952-2_22"},{"key":"ref_34","unstructured":"Wang, X., Zhao, N., Han, Z., Guo, D., and Yang, X. (March, January 25). Augrefer: Advancing 3d visual grounding via cross-modal augmentation and spatial relation-based referring. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_35","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"113704","DOI":"10.1016\/j.asoc.2025.113704","article-title":"A focal quotient gradient system method for deep neural network training","volume":"184","author":"Lv","year":"2025","journal-title":"Appl. Soft Comput."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"He, Q. (2026). A Unified Metric Architecture for AI Infrastructure: A Cross-Layer Taxonomy Integrating Performance, Efficiency, and Cost. arXiv.","DOI":"10.2139\/ssrn.5808163"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/3\/145\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T09:06:27Z","timestamp":1772442387000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/3\/145"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,1]]},"references-count":37,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["computers15030145"],"URL":"https:\/\/doi.org\/10.3390\/computers15030145","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,1]]}}}