{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T08:48:29Z","timestamp":1775810909876,"version":"3.50.1"},"reference-count":52,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2022,3,30]],"date-time":"2022-03-30T00:00:00Z","timestamp":1648598400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Sign language (SL) translation constitutes an extremely challenging task when undertaken in a general unconstrained setup, especially in the absence of vast training datasets that enable the use of end-to-end solutions employing deep architectures. In such cases, the ability to incorporate prior information can yield a significant improvement in the translation results by greatly restricting the search space of the potential solutions. In this work, we treat the translation problem in the limited confines of psychiatric interviews involving doctor-patient diagnostic sessions for deaf and hard of hearing patients with mental health problems.To overcome the lack of extensive training data and be able to improve the obtained translation performance, we follow a domain-specific approach combining data-driven feature extraction with the incorporation of prior information drawn from the available domain knowledge. This knowledge enables us to model the context of the interviews by using an appropriately defined hierarchical ontology for the contained dialogue, allowing for the classification of the current state of the interview, based on the doctor\u2019s question. Utilizing this information, video transcription is treated as a sentence retrieval problem. The goal is predicting the patient\u2019s sentence that has been signed in the SL video based on the available pool of possible responses, given the context of the current exchange. Our experimental evaluation using simulated scenarios of psychiatric interviews demonstrate the significant gains of incorporating context awareness in the system\u2019s decisions.<\/jats:p>","DOI":"10.3390\/s22072656","type":"journal-article","created":{"date-parts":[[2022,3,30]],"date-time":"2022-03-30T21:28:39Z","timestamp":1648675719000},"page":"2656","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Context-Aware Automatic Sign Language Video Transcription in Psychiatric Interviews"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2337-5142","authenticated-orcid":false,"given":"Erion-Vasilis","family":"Pikoulis","sequence":"first","affiliation":[{"name":"Computer Engineering and Informatics Department, University of Patras, 26504 Patras, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0246-1209","authenticated-orcid":false,"given":"Aristeidis","family":"Bifis","sequence":"additional","affiliation":[{"name":"Computer Engineering and Informatics Department, University of Patras, 26504 Patras, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7793-0407","authenticated-orcid":false,"given":"Maria","family":"Trigka","sequence":"additional","affiliation":[{"name":"Computer Engineering and Informatics Department, University of Patras, 26504 Patras, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Constantinos","family":"Constantinopoulos","sequence":"additional","affiliation":[{"name":"Computer Engineering and Informatics Department, University of Patras, 26504 Patras, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dimitrios","family":"Kosmopoulos","sequence":"additional","affiliation":[{"name":"Computer Engineering and Informatics Department, University of Patras, 26504 Patras, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,30]]},"reference":[{"key":"ref_1","unstructured":"(2022, February 22). World Federation of the Deaf. Available online: https:\/\/wfdeaf.org\/our-work\/."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1016\/j.linged.2010.10.001","article-title":"Interpreted writing center tutorials with college-level deaf students","volume":"22","author":"Babcock","year":"2011","journal-title":"Linguist. Educ."},{"key":"ref_3","unstructured":"Wheatley, M., and Pabsch, A. (2010, January 17\u201323). Sign Language in Europe. Proceedings of the 4th LREC Workshop on the Representation and Processing of Sign Languages: Corpora and Sign Language Technologies (CSLT), Valletta, Malta."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bragg, D., Koller, O., Bellard, M., Berke, L., Boudreault, P., Braffort, A., Caselli, N., Huenerfauth, M., Kacorri, H., and Verhoef, T. (2019, January 28\u201330). Sign language recognition, generation, and translation: An interdisciplinary perspective. Proceedings of the 21st International ACM SIGACCESS Conference on Computers and Accessibility, Pittsburgh, PA, USA.","DOI":"10.1145\/3308561.3353774"},{"key":"ref_5","unstructured":"Wilbur, R.B. (2013). Phonological and prosodic layering of nonmanuals in American Sign Language. The Signs of Language Revisited, Psychology Press."},{"key":"ref_6","unstructured":"Dudis, P.G. (2004). Depiction of Events in ASL: Conceptual Integration of Temporal Components, University of California."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Papastratis, I., Chatzikonstantinou, C., Konstantinidis, D., Dimitropoulos, K., and Daras, P. (2021). Artificial Intelligence Technologies for Sign Language. Sensors, 21.","DOI":"10.3390\/s21175843"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1016\/j.cviu.2015.09.013","article-title":"Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers","volume":"141","author":"Koller","year":"2015","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_9","unstructured":"Camgoz, N.C., Koller, O., Hadfield, S., and Bowden, R. (2020, January 14\u201319). Sign language transformers: Joint end-to-end sign language recognition and translation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Voskou, A., Panousis, K.P., Kosmopoulos, D., Metaxas, D.N., and Chatzis, S. (2021, January 11\u201317). Stochastic transformer networks with linear competing units: Application to end-to-end sl translation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01173"},{"key":"ref_11","unstructured":"Forster, J., Schmidt, C., Koller, O., Bellgardt, M., and Ney, H. (2014, January 26\u201331). Extensions of the Sign Language Recognition and Translation Corpus RWTH-PHOENIX-Weather. Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC\u201914), Reykjavik, Iceland."},{"key":"ref_12","unstructured":"Von Agris, U., and Kraiss, K.F. (2007, January 23\u201325). Towards a video corpus for signer-independent continuous sign language recognition. Proceedings of the Gesture in Human-Computer Interaction and Simulation: 7th International Gesture Workshop, Lisbon, Portugal."},{"key":"ref_13","unstructured":"Lugaresi, C., Tang, J., Nash, H., McClanahan, C., Uboweja, E., Hays, M., Zhang, F., Chang, C.L., Yong, M.G., and Lee, J. (2019). MediaPipe: A Framework for Building Perception Pipelines. arXiv."},{"key":"ref_14","unstructured":"Bifis, A., Trigka, M., Dedegkika, S., Goula, P., Constantinopoulos, C., and Kosmopoulos, D. (July, January 29). A Hierarchical Ontology for Dialogue Acts in Psychiatric Interviews. Proceedings of the 14th PErvasive Technologies Related to Assistive Environments Conference, Corfu, Greece."},{"key":"ref_15","unstructured":"Koller, O. (2020). Quantitative Survey of the State of the Art in Sign Language Recognition. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"113794","DOI":"10.1016\/j.eswa.2020.113794","article-title":"Sign language recognition: A deep survey","volume":"164","author":"Rastgoo","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1657","DOI":"10.1109\/TPAMI.2008.215","article-title":"Robust sequential data modeling using an outlier tolerant hidden Markov model","volume":"31","author":"Chatzis","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Vogler, C., and Metaxas, D. (2003). Handshapes and Movements: Multiple-Channel American Sign Language Recognition. International Gesture Workshop, Springer.","DOI":"10.1007\/978-3-540-24598-8_23"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lang, S., Block, M., and Rojas, R. (2012). Sign Language Recognition Using Kinect. Artificial Intelligence and Soft Computing, Springer.","DOI":"10.1007\/978-3-642-29347-4_46"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1685","DOI":"10.1109\/TPAMI.2008.203","article-title":"A unified framework for gesture recognition and spatiotemporal gesture segmentation","volume":"31","author":"Alon","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2040","DOI":"10.1109\/TPAMI.2008.123","article-title":"Sign language recognition by combining statistical DTW and independent classification","volume":"30","author":"Lichtenauer","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","unstructured":"Yang, R., and Sarkar, S. (2006, January 20\u201324). Detecting Coarticulation in Sign Language using Conditional Random Fields. Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906), Hong Kong, China."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yang, H., and Lee, S. (2010, January 23\u201326). Robust Sign Language Recognition with Hierarchical Conditional Random Fields. Proceedings of the International Conference on Pattern Recognition, Istanbul, Turkey.","DOI":"10.1109\/ICPR.2010.539"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Agapito, L., Bronstein, M.M., and Rother, C. (2015). Sign Language Recognition Using Convolutional Neural Networks. Computer Vision-ECCV 2014 Workshops, Springer International Publishing.","DOI":"10.1007\/978-3-319-16220-1"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1692","DOI":"10.1109\/TPAMI.2015.2461544","article-title":"ModDrop: Adaptive Multi-Modal Gesture Recognition","volume":"38","author":"Neverova","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Koller, O., Zargaran, S., and Ney, H. (2017, January 21\u201326). Re-Sign: Re-Aligned End-to-End Sequence Modelling with Deep Recurrent CNN-HMMs. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.364"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"22177","DOI":"10.1007\/s11042-020-08961-z","article-title":"Understanding vision-based continuous sign language recognition","volume":"79","author":"Aloysius","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2306","DOI":"10.1109\/TPAMI.2019.2911077","article-title":"Weakly Supervised Learning with Multi-Stream CNN-LSTM-HMMs to Discover Sequential Parallelism in Sign Language Videos","volume":"42","author":"Koller","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Koishybay, K., Mukushev, M., and Sandygulova, A. (2021, January 10\u201315). Continuous Sign Language Recognition with Iterative Spatiotemporal Fine-tuning. Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy.","DOI":"10.1109\/ICPR48806.2021.9412364"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.M. (2020). Fully Convolutional Networks for Continuous Sign Language Recognition. Computer Vision\u2013ECCV 2020, Springer International Publishing.","DOI":"10.1007\/978-3-030-58598-3"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Cui, R., Liu, H., and Zhang, C. (2017, January 21\u201326). Recurrent Convolutional Neural Networks for Continuous Sign Language Recognition by Staged Optimization. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.175"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"70948","DOI":"10.1109\/ACCESS.2021.3078638","article-title":"Boundary-adaptive encoder with attention method for Chinese sign language recognition","volume":"9","author":"Huang","year":"2021","journal-title":"IEEE Access"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhou, H., Zhou, W., Zhou, Y., and Li, H. (2020, January 7\u201312). Spatial-Temporal Multi-Cue Network for Continuous Sign Language Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.7001"},{"key":"ref_34","unstructured":"Graves, A., Fern\u00e1ndez, S., Gomez, F., and Schmidhuber, J. (2009, January 25\u201329). Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. Proceedings of the 23rd International Conference on Machine Learning, ICML \u201906, Pittsburgh, PA, USA."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Papastratis, I., Dimitropoulos, K., and Daras, P. (2021). Continuous Sign Language Recognition through a Context-Aware Generative Adversarial Network. Sensors, 21.","DOI":"10.3390\/s21072437"},{"key":"ref_36","unstructured":"Papadimitirou, G.N., Liappas, J.A., and Likouras, E. (2013). Modern Psychiatry, BETA Medical Publications."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Gelder, M., Andreasen, N., Lopez-Ibor, J., and Geddes, J. (2012). New Oxford Textbook of Psychiatry, Oxford University Press.","DOI":"10.1093\/med\/9780199696758.001.0001"},{"key":"ref_38","unstructured":"Sadock, B.J., Sadock, V.A., and Ruiz, P. (2017). Kaplan & Sadock\u2019s Comprehensive Textbook of Psychiatry, Wolters Kluwer. [10th ed.]."},{"key":"ref_39","first-page":"3440","article-title":"Dialogue act sequence labeling using hierarchical encoder with CRF","volume":"Volume 32","author":"Kumar","year":"2018","journal-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Kosmopoulos, D., Oikonomidis, I., Constantinopoulos, C., Arvanitis, N., Antzakas, K., Bifis, A., Lydakis, G., Roussos, A., and Argyros, A. (2020, January 16\u201320). Towards a visual Sign Language dataset for home care services. Proceedings of the 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), Buenos Aires, Argentina.","DOI":"10.1109\/FG47880.2020.00099"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"339","DOI":"10.1162\/089120100561737","article-title":"Dialogue act modeling for automatic tagging and recognition of conversational speech","volume":"26","author":"Stolcke","year":"2000","journal-title":"Comput. Linguist."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"4","DOI":"10.5087\/dad.2016.301","article-title":"The dialog state tracking challenge series: A review","volume":"7","author":"Williams","year":"2016","journal-title":"Dialogue Discourse"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Liu, Y., Han, K., Tan, Z., and Lei, Y. (2017, January 7\u201311). Using context information for dialog act classification in DNN framework. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1231"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019, January 3\u20137). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, Hong Kong, China.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_45","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1007\/s10618-010-0175-9","article-title":"A survey of hierarchical classification across different application domains","volume":"22","author":"Silla","year":"2011","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_47","unstructured":"(2022, February 22). MediaPipe. Available online: https:\/\/google.github.io\/mediapipe\/."},{"key":"ref_48","unstructured":"Salton, G., and McGill, M.J. (1983). Introduction to Modern Information Retrieval, Mcgraw-Hill."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1080\/01621459.1951.10500769","article-title":"The Kolmogorov-Smirnov test for goodness of fit","volume":"46","author":"Massey","year":"1951","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Leskovec, J., Rajaraman, A., and Ullman, J.D. (2014). Mining of Massive Datasets, Cambridge University Press. [2nd ed.].","DOI":"10.1017\/CBO9781139924801"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9","article-title":"Indexing by latent semantic analysis","volume":"41","author":"Deerwester","year":"1990","journal-title":"J. Am. Soc. Inf. Sci."},{"key":"ref_52","unstructured":"Hofmann, T. (2013). Probabilistic latent semantic analysis. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/7\/2656\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:46:32Z","timestamp":1760136392000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/7\/2656"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,30]]},"references-count":52,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["s22072656"],"URL":"https:\/\/doi.org\/10.3390\/s22072656","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,30]]}}}