{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T18:06:28Z","timestamp":1772042788106,"version":"3.50.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2010,12,1]],"date-time":"2010-12-01T00:00:00Z","timestamp":1291161600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000180","name":"U.S. Department of Homeland Security","doi-asserted-by":"publisher","award":["FA8750-08-2-0147"],"award-info":[{"award-number":["FA8750-08-2-0147"]}],"id":[{"id":"10.13039\/100000180","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CNS-0546350"],"award-info":[{"award-number":["CNS-0546350"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst. Secur."],"published-print":{"date-parts":[[2010,12]]},"abstract":"<jats:p>Although Voice over IP (VoIP) is rapidly being adopted, its security implications are not yet fully understood. Since VoIP calls may traverse untrusted networks, packets should be encrypted to ensure confidentiality. However, we show that it is possible to<jats:italic>identify the phrases spoken within encrypted VoIP calls<\/jats:italic>when the audio is encoded using variable bit rate codecs. To do so, we train a hidden Markov model using only knowledge of the phonetic pronunciations of words, such as those provided by a dictionary, and search packet sequences for instances of specified phrases. Our approach does not require examples of the speaker\u2019s voice, or even example recordings of the words that make up the target phrase. We evaluate our techniques on a standard speech recognition corpus containing over 2,000 phonetically rich phrases spoken by 630 distinct speakers from across the continental United States. Our results indicate that we can identify phrases within encrypted calls with an average accuracy of 50%, and with accuracy greater than 90% for some phrases. Clearly, such an attack calls into question the efficacy of current VoIP encryption standards. In addition, we examine the impact of various features of the underlying audio on our performance and discuss methods for mitigation.<\/jats:p>","DOI":"10.1145\/1880022.1880029","type":"journal-article","created":{"date-parts":[[2010,12,29]],"date-time":"2010-12-29T14:32:48Z","timestamp":1293633168000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":29,"title":["Uncovering Spoken Phrases in Encrypted Voice over IP Conversations"],"prefix":"10.1145","volume":"13","author":[{"given":"Charles V.","family":"Wright","sequence":"first","affiliation":[{"name":"MIT Lincoln Laboratory"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lucas","family":"Ballard","sequence":"additional","affiliation":[{"name":"Google Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Scott E.","family":"Coull","sequence":"additional","affiliation":[{"name":"University of North Carolina, Chapel Hill"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fabian","family":"Monrose","sequence":"additional","affiliation":[{"name":"University of North Carolina, Chapel Hill"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gerald M.","family":"Masson","sequence":"additional","affiliation":[{"name":"Johns Hopkins University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2010,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo. 970--973","author":"Aggarwal C."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Baugher M. McGrew D. Naslund M. Carrara E. and Norrman K. 2004. The secure real-time transport protocol (SRTP). RFC 3711. Baugher M. McGrew D. Naslund M. Carrara E. and Norrman K. 2004. The secure real-time transport protocol (SRTP). RFC 3711.","DOI":"10.17487\/rfc3711"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177697196"},{"key":"e_1_2_1_4_1","volume-title":"Speech Coding Algorithms","author":"Chu W. C."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster A. P.","year":"1977","journal-title":"J. Royal Statist. Soc."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/178523.178524"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 13th USENIX Security Symposium. 303--320","author":"Dingledine R."},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Durbin R. Eddy S. R. Krogh A. and Mitchison G. 1999. Biological Sequence Analysis : Probabilistic Models of Proteins and Nucleic Acids. Cambridge University Press. Durbin R. Eddy S. R. Krogh A. and Mitchison G. 1999. Biological Sequence Analysis : Probabilistic Models of Proteins and Nucleic Acids . Cambridge University Press.","DOI":"10.1017\/CBO9780511790492"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 3rd International Conference on Intelligent Systems for Molecular Biology. 114--120","author":"Eddy S.","year":"1995"},{"key":"e_1_2_1_10_1","volume-title":"HMMER: Biosequence analysis using profile hidden markov models","author":"Eddy S.","year":"2009"},{"key":"e_1_2_1_11_1","unstructured":"ETSI\/GSM. 1991. GSM full rate speech transcoding. GSM Recommendation 06.10. ETSI\/GSM. 1991. GSM full rate speech transcoding. GSM Recommendation 06.10."},{"key":"e_1_2_1_12_1","volume-title":"QCELP: A variable bit rate speech coder for CDMA digital cellular. Speech Audio Coding Wirel. Netw. Appl., 85--92.","author":"Gardner W.","year":"1993"},{"key":"e_1_2_1_13_1","unstructured":"Garofolo J. S. Lamel L. F. Fisher W. M. Fiscus J. G. Pallett D. S. Dahlgren N. L. and Zue V. 1993. TIMIT acoustic-phonetic continuous speech corpus. Linguistic Data Consortium Philadelphia PA. Garofolo J. S. Lamel L. F. Fisher W. M. Fiscus J. G. Pallett D. S. Dahlgren N. L. and Zue V. 1993. TIMIT acoustic-phonetic continuous speech corpus. Linguistic Data Consortium Philadelphia PA."},{"key":"e_1_2_1_14_1","volume-title":"Statistical Methods for Speech Recognition","author":"Jelinek F."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/89.294354"},{"key":"e_1_2_1_16_1","volume-title":"CALLHOME american english lexicon (PRONLEX)","author":"Kingsbury P."},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Kirkpatrick S. Gelatt C. D. and Vecchi M. P. 1983. Optimization by simulated annealing. Sci. 220 4598 671--680. Kirkpatrick S. Gelatt C. D. and Vecchi M. P. 1983. Optimization by simulated annealing. Sci. 220 4598 671--680.","DOI":"10.1126\/science.220.4598.671"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1006\/jmbi.1994.1104"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1180405.1180437"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing.","author":"Okawa S."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/6046.923820"},{"key":"e_1_2_1_22_1","unstructured":"Provos N. 2004. Voice over misconfigured Internet telephones. http:\/\/vomit.xtdnet.nl. Provos N. 2004. Voice over misconfigured Internet telephones. http:\/\/vomit.xtdnet.nl."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/108235.108253"},{"key":"e_1_2_1_24_1","unstructured":"Rix A. W. Beerends J. G. Hollier M. P. and Hekstra A. P. 2001. Perceptual evaluation of speech quality (PESQ) an objective method for end-to-end speech quality assessment of narrowband telephone networks and speech codecs. ITU-T recommendation P.862. Rix A. W. Beerends J. G. Hollier M. P. and Hekstra A. P. 2001. Perceptual evaluation of speech quality (PESQ) an objective method for end-to-end speech quality assessment of narrowband telephone networks and speech codecs. ITU-T recommendation P.862."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing. 627--630","author":"Rohlicek J. R."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing.","author":"Rohlicek J. R."},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing. 129--132","author":"Rose R. C."},{"key":"e_1_2_1_28_1","volume-title":"SIP: Session initiation protocol. RFC 3261.","author":"Rosenberg J.","year":"2002"},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Saint-Andre P. 2004. Extensible messaging and presence protocol (XMPP): Core. RFC 3920. Saint-Andre P. 2004. Extensible messaging and presence protocol (XMPP): Core. RFC 3920.","DOI":"10.17487\/rfc3920"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 16th Annual USENIX Security Symposium. 55--70","author":"Saponas T. S."},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing.","volume":"10","author":"Schroeder M. R."},{"key":"e_1_2_1_32_1","volume-title":"RTP: A transport protocol for real-time applications. RFC","author":"Schulzrinne H.","year":"1996"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the IEEE Symposium on Security and Privacy. 117--128","author":"Simmons G. J."},{"key":"e_1_2_1_34_1","unstructured":"Skype. 2009. Skype. http:\/\/www.skype.com. Skype. 2009. Skype. http:\/\/www.skype.com."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 10th USENIX Security Symposium.","author":"Song D."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the IEEE Symposium on Security and Privacy. 19--30","author":"Sun Q."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing. 1255--1258","author":"Tibrewala S."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the Audio Engineering Society Convention. http:\/\/www.speex.org.","author":"Valin J.-M."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2006.72"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1967.1054010"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.3115\/993268.993313"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102120.1102133"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/29.103088"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2008.21"},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of the 16th Annual USENIX Security Symposium. 43--54","author":"Wright C. V."},{"key":"e_1_2_1_46_1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing.","volume":"2","author":"Zhang L."},{"key":"e_1_2_1_47_1","unstructured":"Zimmerman P. 2008. The Zfone project. http:\/\/zfoneproject.com\/faq.html. Zimmerman P. 2008. The Zfone project. http:\/\/zfoneproject.com\/faq.html."}],"container-title":["ACM Transactions on Information and System Security"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1880022.1880029","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1880022.1880029","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T10:52:15Z","timestamp":1750243935000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1880022.1880029"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,12]]},"references-count":47,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2010,12]]}},"alternative-id":["10.1145\/1880022.1880029"],"URL":"https:\/\/doi.org\/10.1145\/1880022.1880029","relation":{},"ISSN":["1094-9224","1557-7406"],"issn-type":[{"value":"1094-9224","type":"print"},{"value":"1557-7406","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,12]]},"assertion":[{"value":"2008-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}