{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:28:40Z","timestamp":1750220920807,"version":"3.41.0"},"reference-count":22,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,6,11]],"date-time":"2019-06-11T00:00:00Z","timestamp":1560211200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Access. Comput."],"published-print":{"date-parts":[[2019,6,30]]},"abstract":"<jats:p>Video sharing sites have become keepers of de-facto digital libraries of sign language content, being used to store videos including the experiences, knowledge, and opinions of many in the deaf or hard of hearing community. Due to limitations of term-based search over metadata, these videos can be difficult to find, reducing their value to the community. Another result is that community members frequently engage in a push-style delivery of content (e.g., emailing or posting links to videos for others in the sign language community) rather than having access be based on the information needs of community members. In prior work, we have shown the potential to detect sign language content using features derived from the video content rather than relying on metadata. Our prior technique was developed with a focus on accuracy of results and are quite computationally expensive, making it unrealistic to apply them on a corpus the size of YouTube or other large video sharing sites. Here, we describe and examine the performance of optimizations that reduce the cost of face detection and the length of video segments processed. We show that optimizations can reduce the computation time required by 96%, while losing only 1% in F1 score. Further, a keyframe-based approach is examined that removes the need to process continuous video. This approach achieves comparable recall but lower precision than the above techniques. Merging the advantages of the optimizations, we also present a staged classifier, where the keyframe approach is used to reduce the number of non-sign language videos fully processed. An analysis of the staged classifier shows a further reduction in average computation time per video while achieving similar quality of results.<\/jats:p>","DOI":"10.1145\/3325863","type":"journal-article","created":{"date-parts":[[2019,6,11]],"date-time":"2019-06-11T13:28:16Z","timestamp":1560259696000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Tradeoffs in the Efficient Detection of Sign Language Content in Video Sharing Sites"],"prefix":"10.1145","volume":"12","author":[{"given":"Caio D. D.","family":"Monteiro","sequence":"first","affiliation":[{"name":"Texas A8M University, College Station, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Frank M.","family":"Shipman","sequence":"additional","affiliation":[{"name":"Texas A8M University, College Station, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Satyakiran","family":"Duggina","sequence":"additional","affiliation":[{"name":"Texas A8M University, College Station, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ricardo","family":"Gutierrez-Osuna","sequence":"additional","affiliation":[{"name":"Texas A8M University, College Station, TX"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,6,11]]},"reference":[{"volume-title":"Proceedings of the International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908)","author":"Caridakis G.","key":"e_1_2_1_1_1","unstructured":"G. Caridakis , O. Diamanti , K. Karpouzis , and P. Maragos . 2008. Automatic sign language recognition: Vision-based feature extraction and probabilistic recognition scheme from multiple cues . In Proceedings of the International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908) . 8. G. Caridakis, O. Diamanti, K. Karpouzis, and P. Maragos. 2008. Automatic sign language recognition: Vision-based feature extraction and probabilistic recognition scheme from multiple cues. In Proceedings of the International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908). 8."},{"volume-title":"Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 1--6.","author":"Cherniavsky N.","key":"e_1_2_1_2_1","unstructured":"N. Cherniavsky , R. Ladner , and E. Riskin . 2008. Activity detection in conversational sign language video for mobile telecommunication . In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 1--6. N. Cherniavsky, R. Ladner, and E. Riskin. 2008. Activity detection in conversational sign language video for mobile telecommunication. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. 1--6."},{"volume-title":"Proceedings of the International Conference on Computer Systems and Technologies (CompSysTech\u201907)","author":"Dimov D.","key":"e_1_2_1_3_1","unstructured":"D. Dimov , A. Marinov , and N. Zlateva . 2007. CBIR approach to the recognition of a sign language alphabet . In Proceedings of the International Conference on Computer Systems and Technologies (CompSysTech\u201907) . ACM, New York, NY, 9. D. Dimov, A. Marinov, and N. Zlateva. 2007. CBIR approach to the recognition of a sign language alphabet. In Proceedings of the International Conference on Computer Systems and Technologies (CompSysTech\u201907). ACM, New York, NY, 9."},{"volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics. 370--376","author":"Gebre G.","key":"e_1_2_1_4_1","unstructured":"G. Gebre , O. Crasborn , P. Wittenburg , S. Drude , and T. Heskes . 2014. Unsupervised feature learning for visual sign language identification . In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 370--376 . G. Gebre, O. Crasborn, P. Wittenburg, S. Drude, and T. Heskes. 2014. Unsupervised feature learning for visual sign language identification. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. 370--376."},{"volume-title":"Proceedings of the IEEE International Conference on Image Processing. 2626--2630","author":"Gebre B.","key":"e_1_2_1_5_1","unstructured":"B. Gebre , P. Wittenburg , and T. Heskes . 2013. Automatic sign language identification . In Proceedings of the IEEE International Conference on Image Processing. 2626--2630 . B. Gebre, P. Wittenburg, and T. Heskes. 2013. Automatic sign language identification. In Proceedings of the IEEE International Conference on Image Processing. 2626--2630."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1088463.1088512"},{"volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201914)","author":"Karappa V.","key":"e_1_2_1_7_1","unstructured":"V. Karappa , C. Monteiro , F. Shipman , and R. Gutierrez-Osuna . 2014. Detection of sign-language content in video through polar motion profiles . In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201914) . 1290--1294. V. Karappa, C. Monteiro, F. Shipman, and R. Gutierrez-Osuna. 2014. Detection of sign-language content in video through polar motion profiles. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201914). 1290--1294."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2007.19.10.2756"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555400.1555438"},{"volume-title":"Proceedings of the IEEE International Symposium on Multimedia. 287--290","author":"Monteiro C.","key":"e_1_2_1_10_1","unstructured":"C. Monteiro , C. Mathew , R. Gutierrez-Osuna , and F. Shipman . 2016. Detecting and identifying sign languages through visual features . In Proceedings of the IEEE International Symposium on Multimedia. 287--290 . C. Monteiro, C. Mathew, R. Gutierrez-Osuna, and F. Shipman. 2016. Detecting and identifying sign languages through visual features. In Proceedings of the IEEE International Symposium on Multimedia. 287--290."},{"volume-title":"Proceedings of the IEEE International Conference on Multimedia Information Processing and Retrieval. To appear.","author":"Monteiro C.","key":"e_1_2_1_11_1","unstructured":"C. Monteiro , F. Shipman , and R. Gutierrez-Osuna . 2018. Comparing visual, textual, and multimodal features for detecting sign language in video sharing sites . In Proceedings of the IEEE International Conference on Multimedia Information Processing and Retrieval. To appear. C. Monteiro, F. Shipman, and R. Gutierrez-Osuna. 2018. Comparing visual, textual, and multimodal features for detecting sign language in video sharing sites. In Proceedings of the IEEE International Conference on Multimedia Information Processing and Retrieval. To appear."},{"volume-title":"Proceedings of the ACM Conference on Computers and Accessibility. 191--198","author":"Monteiro C.","key":"e_1_2_1_12_1","unstructured":"C. Monteiro , R. Gutierrez-Osuna , and F. Shipman . 2012. Design and evaluation of classifier for identifying sign language videos in video sharing sites . In Proceedings of the ACM Conference on Computers and Accessibility. 191--198 . C. Monteiro, R. Gutierrez-Osuna, and F. Shipman. 2012. Design and evaluation of classifier for identifying sign language videos in video sharing sites. In Proceedings of the ACM Conference on Computers and Accessibility. 191--198."},{"volume-title":"Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers","author":"Platt J. C.","key":"e_1_2_1_13_1","unstructured":"J. C. Platt . 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers . MIT Press , 61--74. J. C. Platt. 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers. MIT Press, 61--74."},{"volume-title":"Proceedings of the 1st International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908)","author":"Potamias M.","key":"e_1_2_1_14_1","unstructured":"M. Potamias and V. Athitsos . 2008. Nearest-neighbor search methods for handshape recognition . In Proceedings of the 1st International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908) . 30. M. Potamias and V. Athitsos. 2008. Nearest-neighbor search methods for handshape recognition. In Proceedings of the 1st International Conference on Pervasive Technologies Related to Assistive Environments (PETRA\u201908). 30."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2579698"},{"volume-title":"Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS\u201917)","author":"Shipman F.","key":"e_1_2_1_16_1","unstructured":"F. Shipman , S. Duggina , C. Monteiro , and R. Gutierrez-Osuna . 2017. Speed-accuracy tradeoffs for detecting sign language content in video sharing sites . In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS\u201917) , 185--189. F. Shipman, S. Duggina, C. Monteiro, and R. Gutierrez-Osuna. 2017. Speed-accuracy tradeoffs for detecting sign language content in video sharing sites. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS\u201917), 185--189."},{"volume-title":"Proceedings of the 1st International Symposium on Information and Communication Technologies (ISICT\u201903)","author":"Somers G.","key":"e_1_2_1_17_1","unstructured":"G. Somers and R. N. Whyte . 2003. Hand posture matching for Irish sign language interpretation . In Proceedings of the 1st International Symposium on Information and Communication Technologies (ISICT\u201903) . 439--444. G. Somers and R. N. Whyte. 2003. Hand posture matching for Irish sign language interpretation. In Proceedings of the 1st International Symposium on Information and Communication Technologies (ISICT\u201903). 439--444."},{"volume-title":"Proceedings of the International Symposium on Computer Vision. Springer, 265--270","author":"Starner T.","key":"e_1_2_1_18_1","unstructured":"T. Starner and A. Pentland . 1995. Real-time American sign language recognition from video using hidden Markov models . In Proceedings of the International Symposium on Computer Vision. Springer, 265--270 . T. Starner and A. Pentland. 1995. Real-time American sign language recognition from video using hidden Markov models. In Proceedings of the International Symposium on Computer Vision. Springer, 265--270."},{"key":"e_1_2_1_19_1","unstructured":"P. Toman. 2016. Complexity Beyond the Trigram: Identifying Sign Languages from Video Using Neural Networks project report. Retrieved from http:\/\/cs231n.stanford.edu\/reports\/2016\/pdfs\/207_Report.pdf.  P. Toman. 2016. Complexity Beyond the Trigram: Identifying Sign Languages from Video Using Neural Networks project report. Retrieved from http:\/\/cs231n.stanford.edu\/reports\/2016\/pdfs\/207_Report.pdf."},{"volume-title":"Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201901)","author":"Viola P.","key":"e_1_2_1_20_1","unstructured":"P. Viola and M. Jones . 2001. Rapid object detection using a boosted cascade of simple features . In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201901) . I--511. P. Viola and M. Jones. 2001. Rapid object detection using a boosted cascade of simple features. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR\u201901). I--511."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 7th IEEE International Conference on Computer Vision","volume":"1","author":"Vogler C.","unstructured":"C. Vogler and D. Metaxas . 1999. Parallel hidden Markov models for American sign language recognition . In Proceedings of the 7th IEEE International Conference on Computer Vision , vol. 1 . 116--122. C. Vogler and D. Metaxas. 1999. Parallel hidden Markov models for American sign language recognition. In Proceedings of the 7th IEEE International Conference on Computer Vision, vol. 1. 116--122."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/1018428.1020644"}],"container-title":["ACM Transactions on Accessible Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3325863","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3325863","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:08Z","timestamp":1750204388000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3325863"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,11]]},"references-count":22,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,30]]}},"alternative-id":["10.1145\/3325863"],"URL":"https:\/\/doi.org\/10.1145\/3325863","relation":{},"ISSN":["1936-7228","1936-7236"],"issn-type":[{"type":"print","value":"1936-7228"},{"type":"electronic","value":"1936-7236"}],"subject":[],"published":{"date-parts":[[2019,6,11]]},"assertion":[{"value":"2018-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-06-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}