{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T16:17:36Z","timestamp":1777047456375,"version":"3.51.4"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62471013"],"award-info":[{"award-number":["62471013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Short videos follow the trend of creation, leading to a proliferation of homogenized video content. Textual overlays such as titles in short videos often reflect semantic homogeneity. This phenomenon manifests not only in the syntactic structure and overall thematic expression, but also in local semantic elements such as words and phrases. Based on information retrieval techniques, we propose a text semantic homogenization recognition network (TSHR-Net) for short video title overlays. The framework comprises key components: (1) a dynamic semantic representation that incorporates contextual information of title overlays using RoBERTa pre-trained word embeddings; (2) a dual-path semantic parser that integrates global semantics via BiLSTM-Attention and local semantics via multi-scale TextCNN; (3) a ranking loss optimization is designed to measure cosine similarity between semantic features, thereby improving homogenization recognition accuracy. Experimental results show that our TSHR-Net achieves the competitive performance in Chinese text semantic homogenization recognition, with \u03c1 and \u03c1\n                    <jats:italic toggle=\"yes\">\n                      <jats:sub>X,Y<\/jats:sub>\n                    <\/jats:italic>\n                    reaching 80.16% and 78.23% on LCQMC, 81.05% and 80.19% on STS-B(ZH), and 92.47% and 92.19% on our self-built BJUT-HCD. The model also exhibits generalization ability in English, attaining 79.68% and 79.89% on STS-B(EN), and 73.56% and 72.21% on SICK dataset, respectively.\n                  <\/jats:p>","DOI":"10.1145\/3805801","type":"journal-article","created":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T21:12:55Z","timestamp":1774645975000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["TSHR-Net: Text Semantic Homogenization Recognition Network for Short Video Title Overlays"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-8775-0147","authenticated-orcid":false,"given":"Yufei","family":"Feng","sequence":"first","affiliation":[{"name":"School of Information and Science Technology, Beijing University of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1290-0738","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information and Science Technology, Beijing University of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6188-9807","authenticated-orcid":false,"given":"Shuying","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9937-2669","authenticated-orcid":false,"given":"Li","family":"Zhuo","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,24]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Why Short-Form Video Is Popular: A Complete Guide. 2024-3-8. Retrieved from https:\/\/kajabi.com\/blog\/why-short-form-video-is-popular"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","unstructured":"N. Reimers and I. Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and International Joint Conference on Natural Language Processing Hong Kong China 3980--3990.","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_1_4_2","unstructured":"Y. Liu M. Ott N. Goyal J. Du M. Joshi D. Chen and O. Levy. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_1_5_2","unstructured":"S. Zhang D. Zheng X. Hu and M. Yang. 2015. Bidirectional long short-term memory networks for relation classification. In Proceedings of the Pacific Asia Conference on Language Information and Computation Shanghai China 73\u201378."},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"P. Zhou W. Shi J. Tian Z. Qi B. Li H. Hao and B. Xu. 2016. Attention-based bidirectional long short-term memory networks for relation classification. Annual Meeting of the Association for Computational Linguistics Berlin Germany 207\u2013212.","DOI":"10.18653\/v1\/P16-2034"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","unstructured":"Y. Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the Conference on Empirical Methods in Natural Language Processing Doha Qatar 1746\u20131751.","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"J. Devlin M W. Chang K. Lee and K. Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2018) 4171--4186.","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Y. Cui W. Che T. Liu B. Qin and Z. Yang. 2021. Pre-training with whole word masking for Chinese BERT. In Proceedings of the IEEE\/ACM Transactions on Audio Speech and Language Processing 29 (2021) 3504\u20133514.","DOI":"10.1109\/TASLP.2021.3124365"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","unstructured":"S. Li R. Pan H. Luo X. Liu and G. Zhao. 2021. Adaptive cross-contextual word embedding for word polysemy with unsupervised topic modeling. Knowledge-Based Systems 218 (2021) 106827.","DOI":"10.1016\/j.knosys.2021.106827"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"J. Chen Y. Wang S. Zhao P. Zhou and Y. Zhang. 2025. Contextualized quaternion embedding towards polysemy in knowledge graph for link prediction. ACM Transactions on Asian and Low-Resource Language Information Processing 24 4 (2025) 31.","DOI":"10.1145\/3714411"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Y. Chu H. Cao Y. Diao and H. Lin. 2023. Refined SBERT: Representing sentence BERT in manifold space. Neurocomputing 555 C (2023) 126453.","DOI":"10.1016\/j.neucom.2023.126453"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"X. Yuan K. Liu and Y. Wang. 2024. Contrastive language-knowledge graph pre-training. ACM Transactions on Asian and Low-Resource Language Information Processing 23 4 (2024) 51.","DOI":"10.1145\/3644820"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"K. Zhou Y. Zhou W X. Zhao and JR. Wen. 2023. Learning to perturb for contrastive learning of unsupervised sentence representations. In Proceedings of the IEEE\/ACM Transactions on Audio Speech and Language Processing. 3935\u20133944.","DOI":"10.1109\/TASLP.2023.3304485"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"J. Xu L. Xiao A. Wu T. Ma D. Dong and L. He. 2025. Bidirectional directed acyclic graph neural network for aspect-level sentiment classification. ACM Transactions on Asian and Low-Resource Language Information Processing 24 4 (2025) 33.","DOI":"10.1145\/3716501"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"T. Gao X. Yao and D. Chen. 2021. SimCSE: Simple contrastive learning of sentence embedding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing Virtual. 6894\u20136910.","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Y. Yan R. Li S. Wang F. Zhang W. Wu and W. Xu. 2021. ConSERT: A contrastive framework for self-supervised sentence representation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics and International Joint Conference on Natural Language Processing Virtual. 5065\u20135075.","DOI":"10.18653\/v1\/2021.acl-long.393"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"T. Yao Y. Li Y. Li Y. Zhu G. Wang and J. Yue. 2023. Cross-modal semantically augmented network for image-text matching. ACM Transactions on Multimedia Computing Communications and Applications 20 4 (2023) 99.","DOI":"10.1145\/3631356"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"D. Viji and S. Revathy. 2022. A hybrid approach of weighted fine-tuned BERT extraction with deep Siamese Bi-LSTM model for semantic text similarity identification. Multimedia Tools and Applications 81 5 (2022) 6131\u20136157.","DOI":"10.1007\/s11042-021-11771-6"},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"S. Yang G. Huang and B. Ofoghi. 2022. Yearwood J. Short text similarity measurement using context-aware weighted biterms. Concurrency and Computation: Practice and Experience 34 8 (2022) e5765.","DOI":"10.1002\/cpe.5765"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Y. Zhao and L. Cui. 2023. Fusion matrix-based text similarity measures for clustering of retrieval results. Scientometrics 128 2 (2023) 1163\u20131186.","DOI":"10.1007\/s11192-022-04596-z"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"Y. Zhao T. Xia Y. Jiang and Y. Tian. 2024. Enhancing inter-sentence attention for semantic textual similarity. Information Processing & Management 61 1 (2024) 103535.","DOI":"10.1016\/j.ipm.2023.103535"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"C. Wei B. Wang and C. C. J. Kuo. 2023. Synwmd: Syntax-aware word mover's distance for sentence similarity evaluation. Pattern Recognition Letters 170 6 (2023) 48\u201355.","DOI":"10.1016\/j.patrec.2023.04.012"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"L. Xu H. Xie F. L. Wang X. Tao W. Wang and Q. Li. 2024. Contrastive sentence representation learning with adaptive false negative cancellation. Information Fusion 102 2 (2024) 102065.","DOI":"10.1016\/j.inffus.2023.102065"},{"key":"e_1_3_1_25_2","unstructured":"C. Li W. Liu R. Guo X. Yin K. Jiang Y. Du Y. Du L. Zhu B. Lai X. Hu D. Yu and Y. Ma. 2022. PP-OCRv3: More attempts for the improvement of ultra-lightweight OCR system. arXiv:2206.03001. Retrieved from https:\/\/arxiv.org\/abs\/2206.03001"},{"key":"e_1_3_1_26_2","unstructured":"J. Su. 2022. CoSENT(I): A more efficient sentence embedding scheme than sentence-BERT. Retrieved from https:\/\/spaces.ac.cn\/archives\/8847"},{"key":"e_1_3_1_27_2","unstructured":"X. Liu Q. Chen C. Deng H. Zeng J. Chen D. Li and B. Tang. 2018. LCQMC: A large-scale Chinese question matching corpus. In Proceedings of the International Conference on Computational Linguistics Santa Fe USA 1952\u20131962."},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"D. Cer M. Diab E. Agirre I. Lopez and L. Specia. 2017. Semeval-2017 Task 1: Semantic Textual Similarity-Multilingual and Cross-Lingual Focused Evaluation. Association for Computational Linguistics Vancouver Canada 115\u2013119.","DOI":"10.18653\/v1\/S17-2001"},{"key":"e_1_3_1_29_2","unstructured":"M. Marelli S. Menini M. Baroni L. Bentivogli R. Bernardi and R. Zamparelli. 2014. A SICK cure for the evaluation of compositional distributional semantic models. In Proceedings of the International Conference on Language Resources and Evaluation (LREC) Reykjavik Iceland 216\u2013223."},{"key":"e_1_3_1_30_2","unstructured":"X. Wu C. Gao Y. Su J. Han Z. Wang and S. Hu. 2022. Smoothed contrastive learning for unsupervised sentence embedding. In Proceedings of the International Conference on Computational Linguistics Gyeongju Korea 4902\u20134906."},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"T. Jiang J. Jiao S. Huang Z. Zhang D. Wang F. Zhuang F. Wei H. Huang D. Deng and Q. Zhang. 2022. PromptBERT: Improving BERT sentence embeddings with prompts. In Proceedings of the Conference on Empirical Methods in Natural Language Processing Abu Dhabi UAE 8826\u20138837.","DOI":"10.18653\/v1\/2022.emnlp-main.603"},{"key":"e_1_3_1_32_2","unstructured":"X. Wu C. Gao L. Zang J. Han Z. Wang and S. Hu. 2022. ESimCSE: Enhanced sample building method for contrastive learning of unsupervised sentence embedding. In Proceedings of the International Conference on Computational Linguistics Gyeongju Korea 3898\u20133907."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805801","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T15:33:22Z","timestamp":1777044802000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805801"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,24]]},"references-count":31,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3805801"],"URL":"https:\/\/doi.org\/10.1145\/3805801","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,24]]},"assertion":[{"value":"2025-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}