{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T16:13:42Z","timestamp":1782317622286,"version":"3.54.5"},"reference-count":200,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,8,2]],"date-time":"2023-08-02T00:00:00Z","timestamp":1690934400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,8,2]],"date-time":"2023-08-02T00:00:00Z","timestamp":1690934400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Intelligent video coding (IVC), which dates back to the late 1980s with the concept of encoding videos with knowledge and semantics, includes visual content compact representation models and methods enabling structural, detailed descriptions of visual information at different granularity levels (i.e., block, mesh, region, and object) and in different areas. It aims to support and facilitate a wide range of applications, such as visual media coding, content broadcasting, and ubiquitous multimedia computing. We present a high-level overview of the IVC technology from model-based coding (MBC) to learning-based coding (LBC). MBC mainly adopts a manually designed coding scheme to explicitly decompose videos to be coded into blocks or semantic components. Thanks to emerging deep learning technologies such as neural networks and generative models, LBC has become a rising topic in the coding area. In this paper, we first review the classical MBC approaches, followed by the LBC approaches for image and video data. We also discuss and overview our recent attempts at neural coding approaches, which are inspiring for both academic research and industrial implementation. Some critical yet less studied issues are discussed at the end of this paper.<\/jats:p>","DOI":"10.1007\/s44267-023-00018-7","type":"journal-article","created":{"date-parts":[[2023,8,2]],"date-time":"2023-08-02T08:01:34Z","timestamp":1690963294000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":37,"title":["Overview of intelligent video coding: from model-based to learning-based approaches"],"prefix":"10.1007","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2731-5403","authenticated-orcid":false,"given":"Siwei","family":"Ma","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junlong","family":"Gao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruofan","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhui","family":"Chang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9362-6237","authenticated-orcid":false,"given":"Qi","family":"Mao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8026-9349","authenticated-orcid":false,"given":"Zhimeng","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7418-6245","authenticated-orcid":false,"given":"Chuanmin","family":"Jia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,8,2]]},"reference":[{"issue":"9","key":"18_CR1","doi-asserted-by":"publisher","first-page":"1098","DOI":"10.1109\/JRPROC.1952.273898","volume":"40","author":"D. A. Huffman","year":"1952","unstructured":"Huffman, D. A. (1952). A method for the construction of minimum-redundancy codes. Proceedings of the IRE, 40(9), 1098\u20131101.","journal-title":"Proceedings of the IRE"},{"issue":"3","key":"18_CR2","doi-asserted-by":"publisher","first-page":"399","DOI":"10.1109\/TIT.1966.1053907","volume":"12","author":"S. Golomb","year":"1966","unstructured":"Golomb, S. (1966). Run-length encodings. IEEE Transactions on Information Theory, 12(3), 399\u2013401.","journal-title":"IEEE Transactions on Information Theory"},{"key":"18_CR3","first-page":"677","volume-title":"Hawaii international conference on system sciences","author":"H. Andrews","year":"1968","unstructured":"Andrews, H., & Pratt, W. (1968). Fourier transform coding of images. In Hawaii international conference on system sciences (pp.\u00a0677\u2013679)."},{"issue":"1","key":"18_CR4","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1109\/PROC.1969.6869","volume":"57","author":"W. K. Pratt","year":"1969","unstructured":"Pratt, W. K., Kane, J., & Andrews, H. C. (1969). Hadamard transform image coding. Proceedings of the IEEE, 57(1), 58\u201368.","journal-title":"Proceedings of the IEEE"},{"issue":"1","key":"18_CR5","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1109\/T-C.1974.223784","volume":"100","author":"N. Ahmed","year":"1974","unstructured":"Ahmed, N., Natarajan, T., & Rao, K. R. (1974). Discrete cosine transform. IEEE Transactions on Computers, 100(1), 90\u201393.","journal-title":"IEEE Transactions on Computers"},{"issue":"4","key":"18_CR6","doi-asserted-by":"publisher","first-page":"764","DOI":"10.1002\/j.1538-7305.1952.tb01405.x","volume":"31","author":"C. W. Harrison","year":"1952","unstructured":"Harrison, C. W. (1952). Experiments with linear prediction in television. The Bell System Technical Journal, 31(4), 764\u2013783.","journal-title":"The Bell System Technical Journal"},{"issue":"16","key":"18_CR7","first-page":"676","volume":"109","author":"A. J. Seyler","year":"1962","unstructured":"Seyler, A. J. (1962). The coding of visual signals to reduce channel-capacity requirements. Proceedings of the IEE-Part C: Monographs, 109(16), 676\u2013684.","journal-title":"Proceedings of the IEE-Part C: Monographs"},{"key":"18_CR8","volume-title":"Proceedings of institute of electronics and communication engineers of Japan (IECE) annual convention","author":"Y. Taki","year":"1974","unstructured":"Taki, Y., Hatori, M., & Tanaka, S. (1974). Inter frame coding that follows the motion. In Proceedings of institute of electronics and communication engineers of Japan (IECE) annual convention (pp.\u00a01263). Tokyo: IEICE."},{"issue":"7","key":"18_CR9","doi-asserted-by":"publisher","first-page":"1703","DOI":"10.1002\/j.1538-7305.1979.tb02277.x","volume":"58","author":"A. N. Netravali","year":"1979","unstructured":"Netravali, A. N., & Stuller, J.\u00a0A. (1979). Motion-compensated transform coding. The Bell System Technical Journal, 58(7), 1703\u20131718.","journal-title":"The Bell System Technical Journal"},{"key":"18_CR10","unstructured":"Reader, C. History of video compression (draft). Retrieved March 30, 2023, from https:\/\/www.itu.int\/wftp3\/av-arch\/jVt-site\/2002_07_Klagenfurt\/JVT-D068.doc."},{"issue":"7","key":"18_CR11","doi-asserted-by":"publisher","first-page":"560","DOI":"10.1109\/TCSVT.2003.815165","volume":"13","author":"T. Wiegand","year":"2003","unstructured":"Wiegand, T., Sullivan, G. J., Bjontegaard, G., & Luthra, A. (2003). Overview of the h. 264\/avc video coding standard. IEEE Transactions on Circuits and Systems for Video Technology, 13(7), 560\u2013576.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR12","first-page":"423","volume-title":"IEEE international conference on multimedia and expo","author":"L. Fan","year":"2004","unstructured":"Fan, L., Ma, S., & Wu, F. (2004). Overview of AVS video standard. In IEEE international conference on multimedia and expo (Vol.\u00a01, pp.\u00a0423\u2013426). Los Alamitos: IEEE."},{"issue":"9","key":"18_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11432-021-3461-9","volume":"65","author":"S. Ma","year":"2022","unstructured":"Ma, S., Zhang, L., Wang, S., Jia, C., Wang, S., Huang, T., et al. (2022). Evolution of AVS video coding standards: twenty years of innovation and development. Science China. Information Sciences, 65(9), 1\u201324.","journal-title":"Science China. Information Sciences"},{"key":"18_CR14","first-page":"1","volume-title":"IEEE picture coding symposium","author":"J. Zhang","year":"2019","unstructured":"Zhang, J., Jia, C., Lei, M., Wang, S., Ma, S., & Gao, W. (2019). Recent development of AVS video coding standard: AVS3. In IEEE picture coding symposium (pp.\u00a01\u20135). Los Alamitos: IEEE."},{"issue":"2","key":"18_CR15","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1109\/MSP.2014.2371951","volume":"32","author":"S. Ma","year":"2015","unstructured":"Ma, S., Huang, T., Reader, C., & Gao, W. (2015). AVS2? Making video coding smarter [standards in a nutshell]. IEEE Signal Processing Magazine, 32(2), 172\u2013183.","journal-title":"IEEE Signal Processing Magazine"},{"issue":"12","key":"18_CR16","doi-asserted-by":"publisher","first-page":"1649","DOI":"10.1109\/TCSVT.2012.2221191","volume":"22","author":"G. J. Sullivan","year":"2012","unstructured":"Sullivan, G. J., Ohm, J.-R., Han, W.-J., & Wiegand, T. (2012). Overview of the high efficiency video coding (hevc) standard. IEEE Transactions on Circuits and Systems for Video Technology, 22(12), 1649\u20131668.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"10","key":"18_CR17","doi-asserted-by":"publisher","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","volume":"31","author":"B. Bross","year":"2021","unstructured":"Bross, B., Wang, Y.-K., Ye, Y., Liu, S., Chen, J., Sullivan, G. J., & Ohm, J.-R. (2021). Overview of the versatile video coding (vvc) standard and its applications. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 3736\u20133764.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"11","key":"18_CR18","doi-asserted-by":"publisher","first-page":"1349","DOI":"10.1109\/TCOM.1977.1093774","volume":"25","author":"J. Limb","year":"1977","unstructured":"Limb, J., Rubinstein, C., & Thompson, J. (1977). Digital coding of color video signals-a review. IEEE Transactions on Communications, 25(11), 1349\u20131385.","journal-title":"IEEE Transactions on Communications"},{"issue":"3","key":"18_CR19","doi-asserted-by":"publisher","first-page":"366","DOI":"10.1109\/PROC.1980.11647","volume":"68","author":"A. N. Netravali","year":"1980","unstructured":"Netravali, A. N., & Limb, J.\u00a0O. (1980). Picture coding: a review. Proceedings of the IEEE, 68(3), 366\u2013406.","journal-title":"Proceedings of the IEEE"},{"issue":"3","key":"18_CR20","doi-asserted-by":"publisher","first-page":"349","DOI":"10.1109\/PROC.1981.11971","volume":"69","author":"A. K. Jain","year":"1981","unstructured":"Jain, A. K. (1981). Image data compression: a review. Proceedings of the IEEE, 69(3), 349\u2013389.","journal-title":"Proceedings of the IEEE"},{"issue":"4","key":"18_CR21","doi-asserted-by":"publisher","first-page":"523","DOI":"10.1109\/PROC.1985.13183","volume":"73","author":"H. G. Musmann","year":"1985","unstructured":"Musmann, H. G., Pirsch, P., & Grallert, H.-J. (1985). Advances in picture coding. Proceedings of the IEEE, 73(4), 523\u2013548.","journal-title":"Proceedings of the IEEE"},{"issue":"5","key":"18_CR22","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/79.618010","volume":"14","author":"T. Sikora","year":"1997","unstructured":"Sikora, T. (1997). MPEG digital video-coding standards. IEEE Signal Processing Magazine, 14(5), 82\u2013100.","journal-title":"IEEE Signal Processing Magazine"},{"issue":"7","key":"18_CR23","doi-asserted-by":"publisher","first-page":"814","DOI":"10.1109\/76.735379","volume":"8","author":"B. G. Haskell","year":"1998","unstructured":"Haskell, B. G., Howard, P. G., LeCun, Y. A., Puri, A., Ostermann, J., Civanlar, M. R., et al. (1998). Image and video coding-emerging standards and beyond. IEEE Transactions on Circuits and Systems for Video Technology, 8(7), 814\u2013837.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR24","doi-asserted-by":"publisher","first-page":"454","DOI":"10.1117\/12.564457","volume-title":"SPIE Conference on applications of digital image processing XXVII","author":"G. J. Sullivan","year":"2004","unstructured":"Sullivan, G. J., Topiwala, P. N., & Luthra, A. (2004). The h. 264\/AVC advanced video coding standard: overview and introduction to the fidelity range extensions. In SPIE Conference on applications of digital image processing XXVII (Vol.\u00a05558, pp.\u00a0454\u2013474). Bellingham, Washington, SPIE."},{"issue":"9","key":"18_CR25","doi-asserted-by":"publisher","first-page":"1103","DOI":"10.1109\/TCSVT.2007.905532","volume":"17","author":"H. Schwarz","year":"2007","unstructured":"Schwarz, H., Marpe, D., & Wiegand, T. (2007). Overview of the scalable video coding extension of the h. 264\/AVC standard. IEEE Transactions on Circuits and Systems for Video Technology, 17(9), 1103\u20131120.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"4","key":"18_CR26","doi-asserted-by":"publisher","first-page":"626","DOI":"10.1109\/JPROC.2010.2098830","volume":"99","author":"A. Vetro","year":"2011","unstructured":"Vetro, A., Wiegand, T., & Sullivan, G. J. (2011). Overview of the stereo and multiview video coding extensions of the h. 264\/MPEG-4 AVC standard. Proceedings of the IEEE, 99(4), 626\u2013642.","journal-title":"Proceedings of the IEEE"},{"issue":"4","key":"18_CR27","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1109\/PROC.1985.13184","volume":"73","author":"M. Kunt","year":"1985","unstructured":"Kunt, M., Ikonomopoulos, A., & Kocher, M. (1985). Second-generation image-coding techniques. Proceedings of the IEEE, 73(4), 549\u2013574.","journal-title":"Proceedings of the IEEE"},{"key":"18_CR28","doi-asserted-by":"publisher","first-page":"132","DOI":"10.1117\/12.937257","volume-title":"SPIE visual communications and image processing","author":"M. R. Civanlar","year":"1986","unstructured":"Civanlar, M. R., Rajala, S. A., & Lee, W. M. (1986). Second generation hybrid image-coding techniques. In SPIE visual communications and image processing (Vol.\u00a0707, pp.\u00a0132\u2013137). Bellingham, Washington, SPIE."},{"key":"18_CR29","volume-title":"Video coding: the second generation approach","author":"L. Torres","year":"2012","unstructured":"Torres, L., & Kunt, M. (2012). Video coding: the second generation approach. Berlin: Springer."},{"issue":"1","key":"18_CR30","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1145\/248621.248622","volume":"29","author":"M. M. Reid","year":"1997","unstructured":"Reid, M. M., Millar, R. J., & Black, N. D. (1997). Second-generation image coding: an overview. ACM Computing Surveys, 29(1), 3\u201329.","journal-title":"ACM Computing Surveys"},{"issue":"4\u20136","key":"18_CR31","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1016\/0923-5965(95)00010-5","volume":"7","author":"H. G. Musmann","year":"1995","unstructured":"Musmann, H. G. (1995). A layered coding system for very low bit rate video coding. Signal Processing. Image Communication, 7(4\u20136), 267\u2013278.","journal-title":"Signal Processing. Image Communication"},{"issue":"2","key":"18_CR32","doi-asserted-by":"publisher","first-page":"769","DOI":"10.1109\/TIP.2013.2294549","volume":"23","author":"X. Zhang","year":"2013","unstructured":"Zhang, X., Huang, T., Tian, Y., & Gao, W. (2013). Background-modeling-based adaptive prediction for surveillance video coding. IEEE Transactions on Image Processing, 23(2), 769\u2013784.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR33","doi-asserted-by":"publisher","first-page":"30","DOI":"10.1109\/TIP.2021.3126420","volume":"31","author":"X. Meng","year":"2021","unstructured":"Meng, X., Jia, C., Zhang, X., Wang, S., & Ma, S. (2021). Spatio-temporal correlation guided geometric partitioning for versatile video coding. IEEE Transactions on Image Processing, 31, 30\u201342.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR34","volume-title":"IEEE Transactions on Circuits and Systems for Video Technology","author":"S. Wang","year":"2022","unstructured":"Wang, S., Jia, C., Zhang, X., Wang, S., Ma, S., & Gao, W. (2022). A\u00a0pixel-level segmentation-synthesis framework for dynamic texture video compression. IEEE Transactions on Circuits and Systems for Video Technology, 32(10), 7077\u20137091."},{"issue":"5","key":"18_CR35","first-page":"452","volume":"72","author":"H. Harashima","year":"1989","unstructured":"Harashima, H., Aizawa, K., & Saito, T. (1989). Model-based analysis synthesis coding of videotelephone images\u2013conception and basic study of intelligent image coding. IEICE Transactions (1976-1990), 72(5), 452\u2013459.","journal-title":"IEICE Transactions (1976-1990)"},{"issue":"2","key":"18_CR36","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1016\/0895-6111(94)90019-1","volume":"18","author":"P. Cicconi","year":"1994","unstructured":"Cicconi, P., Reusens, E., Dufaux, F., Moccagatta, I., Rouchouze, B., Ebrahimi, T., & Kunt, M. (1994). New trends in image data compression. Computerized Medical Imaging and Graphics, 18(2), 107\u2013124.","journal-title":"Computerized Medical Imaging and Graphics"},{"issue":"5","key":"18_CR37","doi-asserted-by":"publisher","first-page":"589","DOI":"10.1109\/83.334983","volume":"3","author":"H. Li","year":"1994","unstructured":"Li, H., Lundmark, A., & Forchheimer, R. (1994). Image sequence coding at very low bit rates: a review. IEEE Transactions on Image Processing, 3(5), 589\u2013609.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"6","key":"18_CR38","doi-asserted-by":"publisher","first-page":"892","DOI":"10.1109\/5.387091","volume":"83","author":"D. E. Pearson","year":"1995","unstructured":"Pearson, D. E. (1995). Developments in model-based video coding. Proceedings of the IEEE, 83(6), 892\u2013906.","journal-title":"Proceedings of the IEEE"},{"issue":"2","key":"18_CR39","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1109\/5.364463","volume":"83","author":"K. Aizawa","year":"1995","unstructured":"Aizawa, K., & Huang, T. S. (1995). Model-based image coding advanced video coding techniques for very low bit-rate applications. Proceedings of the IEEE, 83(2), 259\u2013271.","journal-title":"Proceedings of the IEEE"},{"key":"18_CR40","first-page":"580","volume-title":"IEEE international symposium on circuits and systems","author":"B. Girod","year":"1996","unstructured":"Girod, B., Ben Younes, K., Bernstein, R., Eisert, P., Farber, N., Hartung, F., et al. (1996). Recent advances in video compression. In IEEE international symposium on circuits and systems (Vol.\u00a02, pp.\u00a0580\u2013583)."},{"issue":"8","key":"18_CR41","doi-asserted-by":"publisher","first-page":"1144","DOI":"10.1109\/TCSVT.1999.809152","volume":"9","author":"F. Pereira","year":"1999","unstructured":"Pereira, F., Chang, S.\u00a0f., Koenen, R., Puri, A., & Avaro, O. (1999). Introduction to the special issue on object-based video coding and description. IEEE Transactions on Circuits and Systems for Video Technology, 9(8), 1144\u20131146.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"1","key":"18_CR42","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1109\/JPROC.2004.839601","volume":"93","author":"T. Sikora","year":"2005","unstructured":"Sikora, T. (2005). Trends and perspectives in image and video coding. Proceedings of the IEEE, 93(1), 6\u201317.","journal-title":"Proceedings of the IEEE"},{"issue":"8","key":"18_CR43","first-page":"525","volume":"68","author":"W. F. Schreiber","year":"1959","unstructured":"Schreiber, W. F., Knapp, C. F., & Kay, N. D. (1959). Synthetic highs\u2014an experimental TV bandwidth reduction system. Journal of the Society of Motion Picture and Television Engineers, 68(8), 525\u2013537.","journal-title":"Journal of the Society of Motion Picture and Television Engineers"},{"issue":"9","key":"18_CR44","doi-asserted-by":"publisher","first-page":"1546","DOI":"10.1109\/JRPROC.1960.287668","volume":"48","author":"R. L. Carbrey","year":"1960","unstructured":"Carbrey, R. L. (1960). Video transmission over telephone cable pairs by pulse code modulation. Proceedings of the IRE, 48(9), 1546\u20131561.","journal-title":"Proceedings of the IRE"},{"issue":"11","key":"18_CR45","doi-asserted-by":"publisher","first-page":"2073","DOI":"10.1109\/26.61489","volume":"38","author":"I. Dinstein","year":"1990","unstructured":"Dinstein, I., Rose, K., & Heiman, A. (1990). Variable block-size transform image coder. IEEE Transactions on Communications, 38(11), 2073\u20132078.","journal-title":"IEEE Transactions on Communications"},{"key":"18_CR46","unstructured":"Wien, M., & Ohm, J.-R. (2001). Simplified adaptive block transforms. Document VCEG-O30. Retrieved March 30, 2023, from https:\/\/www.itu.int\/wftp3\/av-arch\/jvt-site\/2002_07_Klagenfurt\/JVT-D068.doc."},{"issue":"12","key":"18_CR47","doi-asserted-by":"publisher","first-page":"1697","DOI":"10.1109\/TCSVT.2012.2223011","volume":"22","author":"I.-K. Kim","year":"2012","unstructured":"Kim, I.-K., Min, J., Lee, T., Han, W.-J., & Park, J. (2012). Block partitioning structure in the HEVC standard. IEEE Transactions on Circuits and Systems for Video Technology, 22(12), 1697\u20131706.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"10","key":"18_CR48","doi-asserted-by":"publisher","first-page":"3818","DOI":"10.1109\/TCSVT.2021.3088134","volume":"31","author":"Y.-W. Huang","year":"2021","unstructured":"Huang, Y.-W., An, J., Huang, H., Li, X., Hsiang, S.-T., Zhang, K., et al. (2021). Block partitioning structure in the VVC standard. IEEE Transactions on Circuits and Systems for Video Technology, 31(10), 3818\u20133833.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"7","key":"18_CR49","doi-asserted-by":"publisher","first-page":"1464","DOI":"10.1117\/12.138613","volume":"32","author":"V. E. Seferidis","year":"1993","unstructured":"Seferidis, V. E., & Ghanbari, M. (1993). General approach to block-matching motion estimation. Optical Engineering, 32(7), 1464\u20131474.","journal-title":"Optical Engineering"},{"key":"18_CR50","first-page":"2713","volume-title":"IEEE international conference on acoustics, speech, and signal processing","author":"G. J. Sullivan","year":"1991","unstructured":"Sullivan, G. J., & Baker, R. L. (1991). Motion compensation for video compression using control grid interpolation. In IEEE international conference on acoustics, speech, and signal processing (pp.\u00a02713\u20132714). Los Alamitos: IEEE."},{"key":"18_CR51","first-page":"261","volume-title":"IEEE international conference on image processing","author":"K. Denis","year":"2006","unstructured":"Denis, K., & Guillemot, C. (2006). Mesh-based motion-compensated interpolation for side information extraction in distributed video coding. In IEEE international conference on image processing (pp.\u00a0261\u2013264). Los Alamitos: IEEE."},{"key":"18_CR52","unstructured":"Brusewitz, H. (1990). Motion compensation with triangles. [Paper presentation]. Proceedings of the 3rd international workshop on 64k bits\/s coding of moving video, free session, Rotterdam, Netherlands."},{"key":"18_CR53","first-page":"653","volume-title":"IEEE international conference on acoustics, speech and signal processing","author":"O. D. Escoda","year":"2007","unstructured":"Escoda, O. D., Yin, P., Dai, C., & Li, X. (2007). Geometry-adaptive block partitioning for video coding. In IEEE international conference on acoustics, speech and signal processing (pp.\u00a0653\u2013657). Los Alamitos: IEEE."},{"issue":"2","key":"18_CR54","doi-asserted-by":"publisher","first-page":"338","DOI":"10.1109\/TCSVT.2012.2203743","volume":"23","author":"Q. Wang","year":"2012","unstructured":"Wang, Q., Ji, X., Sun, M.-T., Sullivan, G. J., Li, J., & Dai, Q. (2012). Complexity reduction and performance improvement for geometry partitioning in video coding. IEEE Transactions on Circuits and Systems for Video Technology, 23(2), 338\u2013352.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR55","unstructured":"Guo, L., Yin, P., & Francois, E. (2010). Te:3 simplified geometry block partitioning. Technical report JCTVC-B085. Geneva, Switzerland, 2nd meeting: joint collaborative team on video coding (JCT-VC) of ITU-T VCEG and ISO\/IEC MPEG."},{"issue":"3","key":"18_CR56","doi-asserted-by":"publisher","first-page":"336","DOI":"10.1109\/PROC.1967.5490","volume":"55","author":"D. N. Graham","year":"1967","unstructured":"Graham, D. N. (1967). Image transmission by two-dimensional contour coding. Proceedings of the IEEE, 55(3), 336\u2013346.","journal-title":"Proceedings of the IEEE"},{"key":"18_CR57","first-page":"121","volume-title":"IEEE proceedings-f (communications, radar and signal processing)","author":"M. J. Biggar","year":"1988","unstructured":"Biggar, M. J., Morris, O. J., & Constantinides, A. G. (1988). Segmented-image coding: performance comparison with the discrete cosine transform. In IEEE proceedings-f (communications, radar and signal processing) (Vol.\u00a0135, pp.\u00a0121\u2013132). Los Alamitos: IEEE."},{"issue":"2","key":"18_CR58","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1016\/0923-5965(89)90007-6","volume":"1","author":"M. Gilge","year":"1989","unstructured":"Gilge, M., Engelhardt, T., & Mehlan, R. (1989). Coding of arbitrarily shaped image segments based on a generalized orthogonal transform. Signal Processing. Image Communication, 1(2), 153\u2013180.","journal-title":"Signal Processing. Image Communication"},{"issue":"6","key":"18_CR59","doi-asserted-by":"publisher","first-page":"1284","DOI":"10.1109\/TIP.2009.2017339","volume":"18","author":"A. A. Kassim","year":"2009","unstructured":"Kassim, A. A., Lee, W. S., & Zonoobi, D. (2009). Hierarchical segmentation-based image coding using hybrid quad-binary trees. IEEE Transactions on Image Processing, 18(6), 1284\u20131291.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR60","first-page":"1649","volume-title":"IEEE international conference on image processing","author":"S. Milani","year":"2011","unstructured":"Milani, S., & Calvagno, G. (2011). Segmentation-based motion compensation for enhanced video coding. In IEEE international conference on image processing (pp.\u00a01649\u20131652). Los Alamitos: IEEE."},{"key":"18_CR61","doi-asserted-by":"publisher","first-page":"1622","DOI":"10.1109\/ICASSP.2019.8682641","volume-title":"ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP)","author":"D. Chen","year":"2019","unstructured":"Chen, D., Chen, Q., & Zhu, F. (2019). Pixel-level texture segmentation based AV1 video compression. In ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp.\u00a01622\u20131626). Los Alamitos: IEEE."},{"issue":"10","key":"18_CR62","doi-asserted-by":"publisher","first-page":"5091","DOI":"10.1109\/TIP.2019.2910382","volume":"28","author":"Z. Wang","year":"2019","unstructured":"Wang, Z., Wang, S., Zhang, X., Wang, S., & Ma, S. (2019). Three-zone segmentation-based motion compensation for video compression. IEEE Transactions on Image Processing, 28(10), 5091\u20135104.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR63","doi-asserted-by":"publisher","first-page":"554","DOI":"10.1117\/12.185998","volume-title":"SPIE visual communications and image processing","author":"F. Marques","year":"1994","unstructured":"Marques, F., Vera, V., & Gasull, A. (1994). Hierarchical image sequence model for segmentation: application to region-based sequence coding. In SPIE visual communications and image processing (Vol.\u00a02308, pp.\u00a0554\u2013563). Bellingham, Washington, SPIE."},{"key":"18_CR64","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1109\/ICIP.1994.413604","volume-title":"IEEE international conference on image processing","author":"C. Gu","year":"1994","unstructured":"Gu, C., & Kunt, M. (1994). Very low bit-rate video coding using multi-criterion segmentation. In IEEE international conference on image processing (Vol.\u00a02, pp.\u00a0418\u2013422). Los Alamitos: IEEE."},{"issue":"2","key":"18_CR65","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1016\/0923-5965(89)90005-2","volume":"1","author":"H. G. Musmann","year":"1989","unstructured":"Musmann, H. G., H\u00f6tter, M., & Ostermann, J. (1989). Object-oriented analysis-synthesis coding of moving images. Signal Processing. Image Communication, 1(2), 117\u2013138.","journal-title":"Signal Processing. Image Communication"},{"issue":"1","key":"18_CR66","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1109\/76.554418","volume":"7","author":"P. Salembier","year":"1997","unstructured":"Salembier, P., Marqu\u00e9s, F., Pardas, M., Morros, J.\u00a0R., Corset, I., Jeannin, S., et al. (1997). Segmentation-based video coding system allowing the manipulation of objects. IEEE Transactions on Circuits and Systems for Video Technology, 7(1), 60\u201374.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR67","first-page":"65","volume-title":"Erlangen symposium, advances in digital image communication","author":"L. Torres","year":"1997","unstructured":"Torres, L., Garc\u00eda, D., & Mates, A. (1997). On the use of layers for video coding and object manipulation. In Erlangen symposium, advances in digital image communication (pp.\u00a065\u201373). Erlangen: University of Erlangen-Nuremberg."},{"key":"18_CR68","first-page":"II","volume-title":"IEEE international conference on multimedia and expo","author":"A. Vetro","year":"2003","unstructured":"Vetro, A., Haga, T., Sumi, K., & Sun, H. (2003). Object-based coding for long-term archive of surveillance video. In IEEE international conference on multimedia and expo (Vol.\u00a02, pp.\u00a0II\u2013417). Los Alamitos: IEEE."},{"issue":"3","key":"18_CR69","doi-asserted-by":"publisher","first-page":"1699","DOI":"10.1109\/TCE.2009.5278045","volume":"55","author":"D. V. S. X. De Silva","year":"2009","unstructured":"De Silva, D. V. S. X., Fernando, W. A. C., & Yasakethu, S. L. P. (2009). Object based coding of the depth maps for 3d video coding. IEEE Transactions on Consumer Electronics, 55(3), 1699\u20131706.","journal-title":"IEEE Transactions on Consumer Electronics"},{"key":"18_CR70","first-page":"1221","volume-title":"IEEE international conference on image processing","author":"M. Asikuzzaman","year":"2020","unstructured":"Asikuzzaman, M., Ahmmed, A., Pickering, M. R., & Sikora, T. (2020). Edge oriented hierarchical motion estimation for video coding. In IEEE international conference on image processing (pp.\u00a01221\u20131225). Los Alamitos: IEEE."},{"issue":"9","key":"18_CR71","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1109\/MCG.1982.1674492","volume":"2","author":"F. I. Parke","year":"1982","unstructured":"Parke, F. I. (1982). Parameterized models for facial animation. IEEE Computer Graphics and Applications, 2(9), 61\u201368.","journal-title":"IEEE Computer Graphics and Applications"},{"issue":"2","key":"18_CR72","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1016\/0923-5965(89)90006-4","volume":"1","author":"K. Aizawa","year":"1989","unstructured":"Aizawa, K., Harashima, H., & Saito, T. (1989). Model-based analysis synthesis image coding (mbasic) system for a person\u2019s face. Signal Processing. Image Communication, 1(2), 139\u2013152.","journal-title":"Signal Processing. Image Communication"},{"issue":"4","key":"18_CR73","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1049\/ip-vis:19982153","volume":"145","author":"P. M. Antoszczyszyn","year":"1998","unstructured":"Antoszczyszyn, P. M., Hannah, J.\u00a0M., & Grant, P. M. (1998). Reliable tracking of facial features in semantic-based video coding. IET Vision, Image and Signal Processing, 145(4), 257\u2013263.","journal-title":"IET Vision, Image and Signal Processing"},{"issue":"5","key":"18_CR74","doi-asserted-by":"publisher","first-page":"421","DOI":"10.1016\/j.image.2004.02.003","volume":"19","author":"M. Hu","year":"2004","unstructured":"Hu, M., Worrall, S., Sadka, A. H., & Kondoz, A. M. (2004). Automatic scalable face model design for 2D model-based video coding. Signal Processing. Image Communication, 19(5), 421\u2013436.","journal-title":"Signal Processing. Image Communication"},{"issue":"9","key":"18_CR75","doi-asserted-by":"publisher","first-page":"1541","DOI":"10.1109\/TCSVT.2014.2313890","volume":"24","author":"J. Hou","year":"2014","unstructured":"Hou, J., Chau, L.-P., Zhang, M., Magnenat-Thalmann, N., & He, Y. (2014). A highly efficient compression framework for time-varying 3D facial expressions. IEEE Transactions on Circuits and Systems for Video Technology, 24(9), 1541\u20131553.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"19","key":"18_CR76","doi-asserted-by":"publisher","first-page":"12021","DOI":"10.1007\/s11042-016-3368-4","volume":"75","author":"J. Yu","year":"2016","unstructured":"Yu, J., Luo, C., Yu, L., Li, L., & Wang, Z. (2016). Facial video coding\/decoding at ultra-low bit-rate: a 2D\/3D model-based approach. Multimedia Tools and Applications, 75(19), 12021\u201312041.","journal-title":"Multimedia Tools and Applications"},{"key":"18_CR77","doi-asserted-by":"publisher","first-page":"1945","DOI":"10.1109\/ICASSP.1989.266837","volume-title":"IEEE international conference on acoustics, speech, and signal processing","author":"R. J. Safranek","year":"1989","unstructured":"Safranek, R. J., & Johnston, J.\u00a0D. (1989). A perceptually tuned sub-band image coder with image dependent quantization and post-quantization data compression. In IEEE international conference on acoustics, speech, and signal processing (pp.\u00a01945\u20131948). Los Alamitos: IEEE."},{"issue":"10","key":"18_CR78","doi-asserted-by":"publisher","first-page":"1385","DOI":"10.1109\/5.241504","volume":"81","author":"N. Jayant","year":"1993","unstructured":"Jayant, N., Johnston, J., & Safranek, R. (1993). Signal compression based on models of human perception. Proceedings of the IEEE, 81(10), 1385\u20131422.","journal-title":"Proceedings of the IEEE"},{"key":"18_CR79","first-page":"582","volume-title":"Pacific-RIM conference on multimedia","author":"Y. Zhang","year":"2006","unstructured":"Zhang, Y., Ji, X., Zhao, D., & Gao, W. (2006). Video coding by texture analysis and synthesis using graph cut. In Pacific-RIM conference on multimedia (pp.\u00a0582\u2013589). Berlin: Springer."},{"issue":"3","key":"18_CR80","first-page":"162","volume":"26","author":"L. Ma","year":"2011","unstructured":"Ma, L., Ngan, K. N., Zhang, F., & Li, S. (2011). Adaptive block-size transform based just-noticeable difference model for images\/videos. Signal Processing: Image Communication, 26(3), 162\u2013174.","journal-title":"Signal Processing: Image Communication"},{"key":"18_CR81","first-page":"1657","volume-title":"IEEE international conference on image processing","author":"S. Wang","year":"2011","unstructured":"Wang, S., Rehman, A., Wang, Z., Ma, S., & Gao, W. (2011). SSIM-inspired divisive normalization for perceptual video coding. In IEEE international conference on image processing (pp.\u00a01657\u20131660). Los Alamitos: IEEE."},{"issue":"4","key":"18_CR82","doi-asserted-by":"publisher","first-page":"1418","DOI":"10.1109\/TIP.2012.2231090","volume":"22","author":"S. Wang","year":"2012","unstructured":"Wang, S., Rehman, A., Wang, Z., Ma, S., & Gao, W. (2012). Perceptual video coding based on SSIM-inspired divisive normalization. IEEE Transactions on Image Processing, 22(4), 1418\u20131429.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR83","first-page":"1","volume-title":"IEEE signal and information processing association annual summit and conference","author":"S. Wang","year":"2014","unstructured":"Wang, S., Ma, S., Zhao, D., & Gao, W. (2014). Lagrange multiplier based perceptual optimization for high efficiency video coding. In IEEE signal and information processing association annual summit and conference (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"key":"18_CR84","first-page":"2687","volume-title":"IEEE international symposium on circuits and systems","author":"F. Luo","year":"2016","unstructured":"Luo, F., Wang, S., Zhang, N., Ma, S., & Gao, W. (2016). GPU based sample adaptive offset parameter decision and perceptual optimization for HEVC. In IEEE international symposium on circuits and systems (pp.\u00a02687\u20132690). Los Alamitos: IEEE."},{"issue":"1","key":"18_CR85","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1109\/LSP.2016.2641456","volume":"24","author":"X. Zhang","year":"2016","unstructured":"Zhang, X., Wang, S., Gu, K., Lin, W., Ma, S., & Gao, W. (2016). Just-noticeable difference-based perceptual optimization for jpeg compression. IEEE Signal Processing Letters, 24(1), 96\u2013100.","journal-title":"IEEE Signal Processing Letters"},{"key":"18_CR86","first-page":"565","volume-title":"IEEE data compression conference","author":"J. Cui","year":"2019","unstructured":"Cui, J., Xiong, R., Zhang, X., Wang, S., & Ma, S. (2019). Perceptual video coding based on visual saliency modulated just noticeable distortion. In IEEE data compression conference (pp.\u00a0565\u2013565). Los Alamitos: IEEE."},{"key":"18_CR87","doi-asserted-by":"publisher","first-page":"182","DOI":"10.1117\/12.956684","volume-title":"Digital image processing II","author":"C. F. Hall","year":"1978","unstructured":"Hall, C. F., & Andrews, H. C. (1978). Digital color image compression in a perceptual space. In Digital image processing II (Vol.\u00a0149, pp.\u00a0182\u2013188). Bellingham: SPIE."},{"key":"18_CR88","first-page":"845","volume-title":"IEEE international conference on image processing","author":"P. Ndjiki-Nya","year":"2003","unstructured":"Ndjiki-Nya, P., Makai, B., Blattermann, G., Smolic, A., Schwarz, H., & Wiegand, T. (2003). Improved H.264\/AVC coding using texture analysis and synthesis. In IEEE international conference on image processing (pp.\u00a0845\u2013849). Los Alamitos: IEEE."},{"key":"18_CR89","first-page":"1608","volume-title":"IEEE international conference on image processing","author":"A. Stojanovic","year":"2008","unstructured":"Stojanovic, A., Wien, M., & Ohm, J.-R. (2008). Dynamic texture synthesis for H.264\/AVC inter coding. In IEEE international conference on image processing (pp.\u00a01608\u20131611). Los Alamitos: IEEE."},{"key":"18_CR90","first-page":"2045","volume-title":"IEEE international conference on image processing","author":"A. Stojanovic","year":"2010","unstructured":"Stojanovic, A., & Kosse, P. (2010). Extended dynamic texture prediction for H.264\/AVC inter coding. In IEEE international conference on image processing (pp.\u00a02045\u20132048). Los Alamitos: IEEE."},{"key":"18_CR91","first-page":"112","volume-title":"IEEE international conference on multimedia and expo","author":"C. Zhu","year":"2007","unstructured":"Zhu, C., Sun, X., Wu, F., & Li, H. (2007). Video coding with spatio-temporal texture synthesis. In IEEE international conference on multimedia and expo (pp.\u00a0112\u2013115). Los Alamitos: IEEE."},{"key":"18_CR92","first-page":"2593","volume-title":"IEEE international conference on image processing","author":"F. Zhang","year":"2010","unstructured":"Zhang, F., Bull, D. R., & Canagarajah, N. (2010). Region-based texture modelling for next generation video codecs. In IEEE international conference on image processing (pp.\u00a02593\u20132596). Los Alamitos: IEEE."},{"key":"18_CR93","unstructured":"ISO\/IEC (1995). IS 144496-2: information technology - coding of audio-visual objects - part 2: visual (MPEG-4 video)."},{"issue":"6","key":"18_CR94","doi-asserted-by":"publisher","first-page":"688","DOI":"10.1109\/76.927421","volume":"11","author":"S.-F. Chang","year":"2001","unstructured":"Chang, S.-F., Sikora, T., & Purl, A. (2001). Overview of the MPEG-7 standard. IEEE Transactions on Circuits and Systems for Video Technology, 11(6), 688\u2013695.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR95","unstructured":"Lan, C., Xu, J., Wu, F., & Sullivan, G. J. (2010). Screen content coding. Document JCTVC-B084, Geneva, Switzerland. Retrieved March 30, 2023, from https:\/\/www.itu.int\/wftp3\/av-arch\/jctvc-site\/2010_07_B_Geneva\/JCTVC-B200_r0.doc."},{"key":"18_CR96","unstructured":"Dong, S., Zhao, L., Xing, P., & Zhang, X. (2013). Surveillance video coding platform for AVS2. [Paper presentation]. The 47th AVS meeting, Shenzhen, China."},{"key":"18_CR97","first-page":"729","volume-title":"Visual communications and image processing 2010","author":"X. Zhang","year":"2010","unstructured":"Zhang, X., Liang, L., Huang, Q., Liu, Y., Huang, T., & Gao, W. (2010). An efficient coding scheme for surveillance videos captured by stationary cameras. In Visual communications and image processing 2010 (Vol.\u00a07744, pp.\u00a0729\u2013738). Bellingham: SPIE."},{"key":"18_CR98","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1007\/978-1-4899-7641-3_9","volume-title":"Machine learning models and algorithms for big data classification: thinking with examples for effective learning","author":"S. Suthaharan","year":"2016","unstructured":"Suthaharan, S. (2016). Support vector machine. In Machine learning models and algorithms for big data classification: thinking with examples for effective learning (pp.\u00a0207\u2013235). Berlin: Springer."},{"issue":"1","key":"18_CR99","doi-asserted-by":"publisher","first-page":"164","DOI":"10.1016\/j.neuron.2019.09.037","volume":"104","author":"W. J. Ma","year":"2019","unstructured":"Ma, W. J. (2019). Bayesian decision models: a primer. Neuron, 104(1), 164\u2013175.","journal-title":"Neuron"},{"key":"18_CR100","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L. Breiman","year":"2001","unstructured":"Breiman, L. (2001). Random forests. Machine Learning, 45, 5\u201332.","journal-title":"Machine Learning"},{"key":"18_CR101","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1007\/BF00116251","volume":"1","author":"J.\u00a0R. Quinlan","year":"1986","unstructured":"Quinlan, J.\u00a0R. (1986). Induction of decision trees. Machine Learning, 1, 81\u2013106.","journal-title":"Machine Learning"},{"issue":"3","key":"18_CR102","doi-asserted-by":"publisher","first-page":"349","DOI":"10.4310\/SII.2009.v2.n3.a8","volume":"2","author":"T. Hastie","year":"2009","unstructured":"Hastie, T., Rosset, S., Zhu, J., & Zou, H. (2009). Multi-class adaboost. Statistics and Its Interface, 2(3), 349\u2013360.","journal-title":"Statistics and Its Interface"},{"issue":"7","key":"18_CR103","doi-asserted-by":"publisher","first-page":"2225","DOI":"10.1109\/TIP.2015.2417498","volume":"24","author":"Y. Zhang","year":"2015","unstructured":"Zhang, Y., Kwong, S., Wang, X., Yuan, H., Pan, Z., & Xu, L. (2015). Machine learning-based coding unit depth decisions for flexible complexity allocation in high efficiency video coding. IEEE Transactions on Image Processing, 24(7), 2225\u20132238.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR104","first-page":"1","volume-title":"2019 picture coding symposium (PCS)","author":"M. Wang","year":"2019","unstructured":"Wang, M., Li, J., Zhang, L., Zhang, K., Liu, H., Wang, S., & Ma, S. (2019). Fast coding unit splitting decisions for the emergent AVS3 standard. In 2019 picture coding symposium (PCS) (pp.\u00a01\u20135). Los Alamitos: IEEE."},{"key":"18_CR105","doi-asserted-by":"publisher","first-page":"1313","DOI":"10.1109\/TIP.2019.2938670","volume":"29","author":"T. Amestoy","year":"2019","unstructured":"Amestoy, T., Mercat, A., Hamidouche, W., Menard, D., & Bergeron, C. (2019). Tunable VVC frame partitioning based on lightweight machine learning. IEEE Transactions on Image Processing, 29, 1313\u20131328.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"6","key":"18_CR106","doi-asserted-by":"publisher","first-page":"1668","DOI":"10.1109\/TCSVT.2019.2904198","volume":"30","author":"H. Yang","year":"2019","unstructured":"Yang, H., Shen, L., Dong, X., Ding, Q., An, P., & Jiang, G. (2019). Low-complexity CTU partition structure decision and fast intra mode decision for versatile video coding. IEEE Transactions on Circuits and Systems for Video Technology, 30(6), 1668\u20131682.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"10","key":"18_CR107","doi-asserted-by":"publisher","first-page":"7092","DOI":"10.1109\/TCSVT.2022.3174214","volume":"32","author":"J. Zhang","year":"2022","unstructured":"Zhang, J., Wang, M., Jia, C., Wang, S., Ma, S., & Gao, W. (2022). Scalable intra coding optimization for video coding. IEEE Transactions on Circuits and Systems for Video Technology, 32(10), 7092\u20137106.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"4","key":"18_CR108","doi-asserted-by":"publisher","first-page":"270","DOI":"10.1016\/j.jvcir.2008.03.001","volume":"19","author":"B. Ori","year":"2008","unstructured":"Ori, B., & Elad, M. (2008). Compression of facial images using the K-SVD algorithm. Journal of Visual Communication and Image Representation, 19(4), 270\u2013282.","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"18_CR109","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1109\/ICIP.2008.4711713","volume-title":"2008 15th IEEE international conference on image processing","author":"O. G. Sezer","year":"2008","unstructured":"Sezer, O. G., Harmanci, O., & Guleryuz, O. G. (2008). Sparse orthonormal transforms for image compression. In 2008 15th IEEE international conference on image processing (pp.\u00a0149\u2013152). Los Alamitos: IEEE."},{"issue":"5","key":"18_CR110","doi-asserted-by":"publisher","first-page":"1061","DOI":"10.1109\/JSTSP.2011.2135332","volume":"5","author":"J. Zepeda","year":"2011","unstructured":"Zepeda, J., Guillemot, C., & Kijak, E. (2011). Image compression using sparse representations and the iteration-tuned and aligned dictionary. IEEE Journal of Selected Topics in Signal Processing, 5(5), 1061\u20131073.","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"key":"18_CR111","unstructured":"Salvatierra, J. Z. (2010). New sparse representation methods; application to image compression and indexing. PhD thesis, Universit\u00e9 de Rennes 1."},{"issue":"6","key":"18_CR112","doi-asserted-by":"publisher","first-page":"1683","DOI":"10.1109\/TCSVT.2019.2910119","volume":"30","author":"S. Ma","year":"2020","unstructured":"Ma, S., Zhang, X., Jia, C., Zhao, Z., Wang, S., & Wang, S. (2020). Image and video compression with neural networks: a review. IEEE Transactions on Circuits and Systems for Video Technology, 30(6), 1683\u20131698.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"10","key":"18_CR113","doi-asserted-by":"publisher","first-page":"3095","DOI":"10.1109\/TCSVT.2018.2873102","volume":"29","author":"S. Ma","year":"2018","unstructured":"Ma, S., Zhang, X., Wang, S., Zhang, X., Jia, C., & Wang, S. (2018). Joint feature and texture coding: toward smart video representation via front-end intelligence. IEEE Transactions on Circuits and Systems for Video Technology, 29(10), 3095\u20133105.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR114","first-page":"1866","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Huang","year":"2021","unstructured":"Huang, Z., Lin, K., Jia, C., Wang, S., & Ma, S. (2021). Beyond VVC: towards perceptual quality optimized video compression using multi-scale hybrid approaches. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a01866\u20131869). Los Alamitos: IEEE."},{"issue":"3","key":"18_CR115","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1002\/cta.4490160304","volume":"16","author":"L. O. Chua","year":"1988","unstructured":"Chua, L. O., & Lin, T. (1988). A neural network approach to transform image coding. International Journal of Circuit Theory and Applications, 16(3), 317\u2013324.","journal-title":"International Journal of Circuit Theory and Applications"},{"key":"18_CR116","volume-title":"Models of cognition: a review of cognitive science","author":"G.W. Cottrell","year":"1989","unstructured":"Cottrell, G.W., Munro, P., & Zipser, D. (1989). Image compression by back propagation: an example of extensional programming. In N.E. Sharkey (Ed.), Models of cognition: a review of cognitive science. Hillsdale: Ablex."},{"key":"18_CR117","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1109\/IJCNN.1989.118675","volume-title":"International joint conference on neural networks","author":"N. Sonehara","year":"1989","unstructured":"Sonehara, N., Kawato, M., Miyake, S., & Nakane, K., (1989). Image data compression using a neural network model. In International joint conference on neural networks (Vol. 2, pp.\u00a035\u201341). Los Alamitos: IEEE."},{"key":"18_CR118","first-page":"5306","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Toderici","year":"2017","unstructured":"Toderici, G., Vincent, D., Johnston, N., Hwang, S. J., Minnen, D., Shor, J., & Covell, M. (2017). Full resolution image compression with recurrent neural networks. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a05306\u20135314). Los Alamitos: IEEE."},{"key":"18_CR119","first-page":"2796","volume-title":"IEEE international conference on image processing","author":"D. Minnen","year":"2017","unstructured":"Minnen, D., Toderici, G., Covell, M., Chinen, T., Johnston, N., Shor, J., et al. (2017). Spatially adaptive image compression using a tiled deep network. In IEEE international conference on image processing (pp.\u00a02796\u20132800). Los Alamitos: IEEE."},{"key":"18_CR120","unstructured":"Ball\u00e9, J., Laparra, V., & Simoncelli, E. P. (2016). End-to-end optimized image compression. arXiv preprint. arXiv:1611.01704."},{"key":"18_CR121","first-page":"1","volume-title":"IEEE picture coding symposium","author":"J. Ball\u00e9","year":"2016","unstructured":"Ball\u00e9, J., Laparra, V., & Simoncelli, E. P. (2016). End-to-end optimization of nonlinear transform codes for perceptual quality. In IEEE picture coding symposium (pp.\u00a01\u20135). Los Alamitos: IEEE."},{"key":"18_CR122","unstructured":"Ball\u00e9, J., Minnen, D., Singh, S., Hwang, S. J., & Johnston, N. (2018). Variational image compression with a scale hyperprior. arXiv preprint. arXiv:1802.01436."},{"key":"18_CR123","first-page":"10771","volume-title":"Advances in neural information processing systems","author":"D. Minnen","year":"2018","unstructured":"Minnen, D., Ball\u00e9, J., & Toderici, G. D. (2018). Joint autoregressive and hierarchical priors for learned image compression. In S. Bengio, H. M. Wallach, H. Larochelle, et al. (Eds.), Advances in neural information processing systems (Vol. 31, pp.\u00a010771\u201310780). Red Hook: Curran Associates."},{"key":"18_CR124","first-page":"2617","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition workshops","author":"L. Zhou","year":"2018","unstructured":"Zhou, L., Cai, C., Gao, Y., Su, S., & Wu, J. (2018). Variational autoencoder for low bit-rate image compression. In IEEE\/CVF conference on computer vision and pattern recognition workshops (pp.\u00a02617\u20132620). Los Alamitos: IEEE."},{"key":"18_CR125","first-page":"7939","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Cheng","year":"2020","unstructured":"Cheng, Z., Sun, H., Takeuchi, M., & Katto, J. (2020). Learned image compression with discretized Gaussian mixture likelihoods and attention modules. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a07939\u20137948). Los Alamitos: IEEE."},{"key":"18_CR126","first-page":"2922","volume-title":"International conference on machine learning","author":"O. Rippel","year":"2017","unstructured":"Rippel, O., & Bourdev, L. (2017). Real-time adaptive image compression. In International conference on machine learning (pp.\u00a02922\u20132930). PMLR."},{"issue":"1","key":"18_CR127","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1109\/JETCAS.2018.2886642","volume":"9","author":"C. Jia","year":"2018","unstructured":"Jia, C., Zhang, X., Wang, S., Wang, S., & Ma, S. (2018). Light field image compression using generative adversarial network-based view synthesis. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 9(1), 177\u2013189.","journal-title":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems"},{"key":"18_CR128","first-page":"3549","volume-title":"Advances in neural information processing systems","author":"K. Gregor","year":"2016","unstructured":"Gregor, K., Besse, F., Rezende, D. J., Danihelka, I., & Wierstra, D. (2016). Towards conceptual compression. In D. D. Lee, M. Sugiyama, U. von Luxburg, et al. (Eds.), Advances in neural information processing systems (Vol. 29, pp.\u00a03549\u20133557). Red Hook: Curran Associates."},{"key":"18_CR129","first-page":"2587","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition workshops","author":"E. Agustsson","year":"2018","unstructured":"Agustsson, E., Tschannen, M., Mentzer, F., Timofte, R., & van Gool, L. (2018). Extreme learned image compression with GANs. In IEEE\/CVF conference on computer vision and pattern recognition workshops (pp.\u00a02587\u20132590). Los Alamitos: IEEE."},{"key":"18_CR130","first-page":"1","volume-title":"IEEE visual communications and image processing","author":"T. Chen","year":"2017","unstructured":"Chen, T., Liu, H., Shen, Q., Yue, T., Cao, X., & Ma, Z. (2017). Deepcoder: a deep neural network based video compression. In IEEE visual communications and image processing (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"issue":"2","key":"18_CR131","doi-asserted-by":"publisher","first-page":"566","DOI":"10.1109\/TCSVT.2019.2892608","volume":"30","author":"Z. Chen","year":"2019","unstructured":"Chen, Z., He, T., Jin, X., & Wu, F. (2019). Learning for video compression. IEEE Transactions on Circuits and Systems for Video Technology, 30(2), 566\u2013576.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR132","first-page":"843","volume-title":"International conference on machine learning","author":"N. Srivastava","year":"2015","unstructured":"Srivastava, N., Mansimov, E., & Salakhudinov, R. (2015). Unsupervised learning of video representations using lstms. In International conference on machine learning (pp.\u00a0843\u2013852). PMLR."},{"key":"18_CR133","first-page":"7033","volume-title":"IEEE\/CVF international conference on computer vision","author":"A. Habibian","year":"2019","unstructured":"Habibian, A., van Rozendaal, T., Tomczak, J.\u00a0M., & Cohen, T. S. (2019). Video compression with rate-distortion autoencoders. In IEEE\/CVF international conference on computer vision (pp.\u00a07033\u20137042). Los Alamitos: IEEE."},{"key":"18_CR134","volume-title":"Advances in neural information processing systems","author":"S. Lombardo","year":"2019","unstructured":"Lombardo, S., Han, J., Schroers, C., & Mandt, S. (2019). Deep generative video compression. In Advances in neural information processing systems (Vol.\u00a032)."},{"key":"18_CR135","first-page":"701","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"B. Liu","year":"2021","unstructured":"Liu, B., Chen, Y., Liu, S., & Kim, H.-S. (2021). Deep learning in latent space for video prediction and compression. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a0701\u2013710). Los Alamitos: IEEE."},{"key":"18_CR136","doi-asserted-by":"crossref","unstructured":"Liu, H., Lu, M., Chen, Z., Cao, X., Ma, Z., & Wang, Y. (2022). End-to-end neural video coding using a compound spatiotemporal representation. IEEE Transactions on Circuits and Systems for Video Technology.","DOI":"10.1109\/TCSVT.2022.3150014"},{"key":"18_CR137","first-page":"168","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition workshops","author":"V. Veerabadran","year":"2020","unstructured":"Veerabadran, V., Pourreza, R., Habibian, A., & Cohen, T. S. (2020). Adversarial distortion for learned video compression. In IEEE\/CVF conference on computer vision and pattern recognition workshops (pp.\u00a0168\u2013169). Los Alamitos: IEEE."},{"key":"18_CR138","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"T.-C. Wang","year":"2021","unstructured":"Wang, T.-C., Mallya, A., & Liu, M.-Y. (2021). One-shot free-view neural talking-head synthesis for video conferencing. In IEEE\/CVF conference on computer vision and pattern recognition. Los Alamitos: IEEE."},{"key":"18_CR139","first-page":"1","volume-title":"IEEE international conference on multimedia and expo","author":"R. Wang","year":"2022","unstructured":"Wang, R., Mao, Q., Wang, S., Jia, C., Wang, R., & Ma, S. (2022). Disentangled visual representations for extreme human body video compression. In IEEE international conference on multimedia and expo (pp.\u00a01\u20136). Los Alamitos: IEEE."},{"issue":"14\u201315","key":"18_CR140","doi-asserted-by":"publisher","first-page":"2627","DOI":"10.1016\/S1352-2310(97)00447-0","volume":"32","author":"M. W. Gardner","year":"1998","unstructured":"Gardner, M. W., & Dorling, S. R. (1998). Artificial neural networks (the multilayer perceptron)\u2014a review of applications in the atmospheric sciences. Atmospheric Environment, 32(14\u201315), 2627\u20132636.","journal-title":"Atmospheric Environment"},{"key":"18_CR141","first-page":"2793","volume-title":"IEEE international conference on acoustics, speech, and signal processing","author":"S. A. Dianat","year":"1991","unstructured":"Dianat, S. A., Nasrabadi, N. M., & Venkataraman, S. (1991). A non-linear predictor for differential pulse-code encoder (dpcm) using artificial neural networks. In IEEE international conference on acoustics, speech, and signal processing (pp.\u00a02793\u20132794). Los Alamitos: IEEE."},{"issue":"1","key":"18_CR142","doi-asserted-by":"publisher","first-page":"326","DOI":"10.1109\/7.481272","volume":"32","author":"A. Namphol","year":"1996","unstructured":"Namphol, A., Chin, S. H., & Arozullah, M. (1996). Image compression with a hierarchical neural network. IEEE Transactions on Aerospace and Electronic Systems, 32(1), 326\u2013338.","journal-title":"IEEE Transactions on Aerospace and Electronic Systems"},{"issue":"4","key":"18_CR143","doi-asserted-by":"publisher","first-page":"502","DOI":"10.1162\/neco.1989.1.4.502","volume":"1","author":"E. Gelenbe","year":"1989","unstructured":"Gelenbe, E. (1989). Random neural networks with negative and positive signals and product form solution. Neural Computation, 1(4), 502\u2013510.","journal-title":"Neural Computation"},{"key":"18_CR144","first-page":"3996","volume-title":"IEEE international conference on neural networks","author":"E. Gelenbe","year":"1994","unstructured":"Gelenbe, E., & Sungur, M. (1994). Random network learning and image compression. In IEEE international conference on neural networks (Vol.\u00a06, pp.\u00a03996\u20133999). Los Alamitos: IEEE."},{"key":"18_CR145","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1117\/12.420926","volume-title":"SPIE applications of artificial neural networks in image processing VI","author":"F. Hai","year":"2001","unstructured":"Hai, F., Hussain, K. F., Gelenbe, E., & Guha, R. K. (2001). Video compression with wavelets and random neural network approximations. In SPIE applications of artificial neural networks in image processing VI (Vol.\u00a04305, pp.\u00a057\u201364)."},{"issue":"7553","key":"18_CR146","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y. LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436\u2013444.","journal-title":"Nature"},{"key":"18_CR147","unstructured":"Kingma, D. P., & Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint. arXiv:1312.6114."},{"key":"18_CR148","first-page":"1462","volume-title":"International conference on machine learning","author":"K. Gregor","year":"2015","unstructured":"Gregor, K., Danihelka, I., Graves, A., Rezende, D., & Draw, D. W. (2015). A recurrent neural network for image generation. In International conference on machine learning (pp.\u00a01462\u20131471). PMLR."},{"key":"18_CR149","first-page":"221","volume-title":"IEEE\/CVF international conference on computer vision","author":"E. Agustsson","year":"2019","unstructured":"Agustsson, E., Tschannen, M., Mentzer, F., Timofte, R., & Van Gool, L. (2019). Generative adversarial networks for extreme learned image compression. In IEEE\/CVF international conference on computer vision (pp.\u00a0221\u2013231)."},{"key":"18_CR150","first-page":"586","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Zhang","year":"2018","unstructured":"Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a0586\u2013595). Los Alamitos: IEEE."},{"key":"18_CR151","first-page":"416","volume-title":"European conference on computer vision","author":"C.-Y. Wu","year":"2018","unstructured":"Wu, C.-Y., Singhal, N., & Krahenbuhl, P. (2018). Video compression through image interpolation. In European conference on computer vision (pp.\u00a0416\u2013431)."},{"key":"18_CR152","unstructured":"Ranzato, M., Szlam, A., Bruna, J., Mathieu, M., Collobert, R., & Chopra, S. (2014). Video (language) modeling: a baseline for generative models of natural videos. arXiv preprint. arXiv:1412.6604."},{"key":"18_CR153","first-page":"8503","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"E. Agustsson","year":"2020","unstructured":"Agustsson, E., Minnen, D., Johnston, N., Balle, J., Hwang, S. J., & Toderici, G. (2020). Scale-space flow for end-to-end optimized video compression. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a08503\u20138512). Los Alamitos: IEEE."},{"key":"18_CR154","first-page":"436","volume-title":"IEEE data compression conference","author":"W. Cui","year":"2017","unstructured":"Cui, W., Zhang, T., Zhang, S., Jiang, F., Zuo, W., Wan, Z., & Zhao, D. (2017). Convolutional neural networks based intra prediction for hevc. In IEEE data compression conference (pp.\u00a0436\u2013436). Los Alamitos: IEEE."},{"issue":"7","key":"18_CR155","doi-asserted-by":"publisher","first-page":"3236","DOI":"10.1109\/TIP.2018.2817044","volume":"27","author":"J. Li","year":"2018","unstructured":"Li, J., Li, B., Xu, J., Xiong, R., & Gao, W. (2018). Fully connected network-based intra prediction for image coding. IEEE Transactions on Image Processing, 27(7), 3236\u20133247.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"9","key":"18_CR156","doi-asserted-by":"publisher","first-page":"2316","DOI":"10.1109\/TCSVT.2017.2727682","volume":"28","author":"Y. Li","year":"2017","unstructured":"Li, Y., Liu, D., Li, H., Li, L., Wu, F., Zhang, H., & Yang, H. (2017). Convolutional neural network-based block up-sampling for intra frame coding. IEEE Transactions on Circuits and Systems for Video Technology, 28(9), 2316\u20132330.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR157","first-page":"600","volume-title":"Pacific rim conference on multimedia","author":"L. Feng","year":"2018","unstructured":"Feng, L., Zhang, X., Zhang, X., Wang, S., Wang, R., & Ma, S. (2018). A dual-network based super-resolution for compressed high definition video. In Pacific rim conference on multimedia (pp.\u00a0600\u2013610). Berlin: Springer."},{"key":"18_CR158","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1007\/978-3-319-51811-4_3","volume-title":"International conference on multimedia modeling","author":"Y. Dai","year":"2017","unstructured":"Dai, Y., Liu, D., & Wu, F. (2017). A convolutional neural network approach for post-processing in hevc intra coding. In International conference on multimedia modeling (pp.\u00a028\u201339). Berlin: Springer."},{"key":"18_CR159","first-page":"1","volume-title":"IEEE international symposium on circuits and systems","author":"S. Huo","year":"2018","unstructured":"Huo, S., Liu, D., Wu, F., & Li, H. (2018). Convolutional neural network-based motion compensation refinement for video coding. In IEEE international symposium on circuits and systems (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"issue":"3","key":"18_CR160","doi-asserted-by":"publisher","first-page":"840","DOI":"10.1109\/TCSVT.2018.2816932","volume":"29","author":"N. Yan","year":"2018","unstructured":"Yan, N., Liu, D., Li, H., Li, B., Li, L., & Wu, F. (2018). Convolutional neural network-based fractional-pixel motion compensation. IEEE Transactions on Circuits and Systems for Video Technology, 29(3), 840\u2013853.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR161","first-page":"206","volume-title":"IEEE international conference on image processing","author":"L. Zhao","year":"2018","unstructured":"Zhao, L., Wang, S., Zhang, X., Wang, S., Ma, S., & Gao, W. (2018). Enhanced ctu-level inter prediction with deep frame rate up-conversion for high efficiency video coding. In IEEE international conference on image processing (pp.\u00a0206\u2013210). Los Alamitos: IEEE."},{"issue":"10","key":"18_CR162","doi-asserted-by":"publisher","first-page":"4832","DOI":"10.1109\/TIP.2019.2913545","volume":"28","author":"L. Zhao","year":"2019","unstructured":"Zhao, L., Wang, S., Zhang, X., Wang, S., Ma, S., & Gao, W. (2019). Enhanced motion-compensated video coding with deep virtual reference frame generation. IEEE Transactions on Image Processing, 28(10), 4832\u20134844.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR163","first-page":"261","volume-title":"IEEE\/CVF international conference on computer vision","author":"S. Niklaus","year":"2017","unstructured":"Niklaus, S., Mai, L., & Liu, F. (2017). Video frame interpolation via adaptive separable convolution. In IEEE\/CVF international conference on computer vision (pp.\u00a0261\u2013270). Los Alamitos: IEEE."},{"key":"18_CR164","first-page":"1","volume-title":"IEEE international symposium on circuits and systems","author":"Z. Zhao","year":"2018","unstructured":"Zhao, Z., Wang, S., Wang, S., Zhang, X., Ma, S., & Yang, J. (2018). Cnn-based bi-directional motion compensation for high efficiency video coding. In IEEE international symposium on circuits and systems (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"key":"18_CR165","first-page":"6628","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Yang","year":"2020","unstructured":"Yang, R., Mentzer, F., van Gool, L., & Timofte, R. (2020). Learning for video compression with hierarchical quality and recurrent enhancement. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a06628\u20136637). Los Alamitos: IEEE."},{"key":"18_CR166","first-page":"6421","volume-title":"IEEE\/CVF international conference on computer vision","author":"A. Djelouah","year":"2019","unstructured":"Djelouah, A., Campos, J., Schaub-Meyer, S., & Schroers, C. (2019). Neural inter-frame compression for video coding. In IEEE\/CVF international conference on computer vision (pp.\u00a06421\u20136429). Los Alamitos: IEEE."},{"issue":"3","key":"18_CR167","doi-asserted-by":"publisher","first-page":"1592","DOI":"10.1109\/TCSVT.2021.3073114","volume":"32","author":"L. Zhao","year":"2022","unstructured":"Zhao, L., Wang, S., Wang, S., Ye, Y., Ma, S., & Gao, W. (2022). Enhanced surveillance video compression with dual reference frames generation. IEEE Transactions on Circuits and Systems for Video Technology, 32(3), 1592\u20131606.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"18_CR168","first-page":"395","volume-title":"Applications of digital image processing XXXVIII","author":"M. M. Alam","year":"2015","unstructured":"Alam, M. M., Nguyen, T. D., Hagan, M. T., & Chandler, D. M. (2015). A perceptual quantization strategy for HEVC based on a convolutional neural network trained on natural images. In Applications of digital image processing XXXVIII (Vol.\u00a09599, pp.\u00a0395\u2013408). Bellingham: SPIE."},{"key":"18_CR169","first-page":"1","volume-title":"IEEE visual communications and image processing","author":"R. Song","year":"2017","unstructured":"Song, R., Liu, D., Li, H., & Wu, F. (2017). Neural network-based arithmetic coding of intra prediction modes in HEVC. In IEEE visual communications and image processing (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"issue":"8","key":"18_CR170","doi-asserted-by":"publisher","first-page":"3827","DOI":"10.1109\/TIP.2018.2815841","volume":"27","author":"Y. Zhang","year":"2018","unstructured":"Zhang, Y., Shen, T., Ji, X., Zhang, Y., Xiong, R., & Dai, Q. (2018). Residual highway convolutional neural networks for in-loop filtering in HEVC. IEEE Transactions on Image Processing, 27(8), 3827\u20133841.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR171","first-page":"1","volume-title":"IEEE visual communications and image processing","author":"C. Jia","year":"2017","unstructured":"Jia, C., Wang, S., Zhang, X., Wang, S., & Ma, S. (2017). Spatial-temporal residue network based in-loop filter for video coding. In IEEE visual communications and image processing (pp.\u00a01\u20134). Los Alamitos: IEEE."},{"issue":"7","key":"18_CR172","doi-asserted-by":"publisher","first-page":"3343","DOI":"10.1109\/TIP.2019.2896489","volume":"28","author":"C. Jia","year":"2019","unstructured":"Jia, C., Wang, S., Zhang, X., Wang, S., Liu, J., Pu, S., & Ma, S. (2019). Content-aware convolutional neural network for in-loop filtering in high efficiency video coding. IEEE Transactions on Image Processing, 28(7), 3343\u20133356.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR173","first-page":"3205","volume-title":"IEEE international symposium on circuits and systems","author":"Y. Zhao","year":"2022","unstructured":"Zhao, Y., Lin, K., Wang, S., & Ma, S. (2022). Joint luma and chroma multi-scale CNN in-loop filter for versatile video coding. In IEEE international symposium on circuits and systems (pp.\u00a03205\u20133209). Los Alamitos: IEEE."},{"issue":"4","key":"18_CR174","first-page":"1","volume":"18","author":"K. Lin","year":"2022","unstructured":"Lin, K., Jia, C., Zhang, X., Wang, S., Ma, S., & Gao, W. (2022). NR-CNN: nested-residual guided CNN in-loop filtering for video coding. ACM Transactions on Multimedia Computing Communications and Applications, 18(4), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"key":"18_CR175","first-page":"576","volume-title":"IEEE\/CVF international conference on computer vision","author":"C. Dong","year":"2015","unstructured":"Dong, C., Deng, Y., Loy, C. C., & Tang, X. (2015). Compression artifacts reduction by a deep convolutional network. In IEEE\/CVF international conference on computer vision (pp.\u00a0576\u2013584). Los Alamitos: IEEE."},{"issue":"11","key":"18_CR176","doi-asserted-by":"publisher","first-page":"5365","DOI":"10.1109\/TIP.2018.2858022","volume":"27","author":"L. Zhu","year":"2018","unstructured":"Zhu, L., Zhang, Y., Wang, S., Yuan, H., Kwong, S., & Ip, H. H.-S. (2018). Convolutional neural network-based synthesized view quality enhancement for 3D video coding. IEEE Transactions on Image Processing, 27(11), 5365\u20135377.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"8","key":"18_CR177","doi-asserted-by":"publisher","first-page":"1847","DOI":"10.1109\/TPAMI.2012.272","volume":"35","author":"N. Kruger","year":"2012","unstructured":"Kruger, N., Janssen, P., Kalkan, S., Lappe, M., Leonardis, A., Piater, J., et al. (2012). Deep hierarchies in the primate visual cortex: what can we learn for computer vision? IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8), 1847\u20131871.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"1","key":"18_CR178","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41467-019-13993-7","volume":"11","author":"Y. Zhang","year":"2020","unstructured":"Zhang, Y., Han, K., Worth, R., & Liu, Z. (2020). Connecting concepts in the brain by mapping cortical representations of semantic relations. Nature Communications, 11(1), 1\u201313.","journal-title":"Nature Communications"},{"key":"18_CR179","first-page":"694","volume-title":"IEEE international conference on image processing","author":"J. Chang","year":"2019","unstructured":"Chang, J., Mao, Q., Zhao, Z., Wang, S., Wang, S., Zhu, H., & Ma, S. (2019). Layered conceptual image compression via deep semantic synthesis. In IEEE international conference on image processing (pp.\u00a0694\u2013698). Los Alamitos: IEEE."},{"key":"18_CR180","doi-asserted-by":"publisher","first-page":"2809","DOI":"10.1109\/TIP.2022.3159477","volume":"31","author":"J. Chang","year":"2022","unstructured":"Chang, J., Zhao, Z., Jia, C., Wang, S., Yang, L., Mao, Q., et al. (2022). Conceptual compression via deep structure and texture synthesis. IEEE Transactions on Image Processing, 31, 2809\u20132823.","journal-title":"IEEE Transactions on Image Processing"},{"key":"18_CR181","first-page":"1","volume-title":"IEEE international conference on multimedia and expo (ICME)","author":"J. Chang","year":"2021","unstructured":"Chang, J., Zhao, Z., Yang, L., Jia, C., Zhang, J., & Ma, S. (2021). Thousand to one: semantic prior modeling for conceptual coding. In IEEE international conference on multimedia and expo (ICME) (pp.\u00a01\u20136). Los Alamitos: IEEE."},{"key":"18_CR182","first-page":"1","volume-title":"IEEE international conference on multimedia and expo","author":"Y. Hu","year":"2020","unstructured":"Hu, Y., Yang, S., Yang, W., Duan, L.-Y., & Liu, J. (2020). Towards coding for human and machine vision: a scalable image coding approach. In IEEE international conference on multimedia and expo (pp.\u00a01\u20136). Los Alamitos: IEEE."},{"key":"18_CR183","volume-title":"Vision: a\u00a0computational investigation into the human representation and processing of visual information","author":"D. Marr","year":"1982","unstructured":"Marr, D. (1982). Vision: a\u00a0computational investigation into the human representation and processing of visual information. San Francisco: W. H. Freeman and Company."},{"issue":"1","key":"18_CR184","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1016\/j.cviu.2005.09.004","volume":"106","author":"C. Guo","year":"2007","unstructured":"Guo, C., Zhu, S.-C., & Wu, Y. N. (2007). Primal sketch: integrating structure and texture. Computer Vision and Image Understanding, 106(1), 5\u201319.","journal-title":"Computer Vision and Image Understanding"},{"key":"18_CR185","first-page":"1","volume-title":"ACM international conference on multimedia","author":"J. Chang","year":"2022","unstructured":"Chang, J., Zhang, J., Xu, Y., Li, J., Ma, S., & Gao, W. (2022). Consistency-contrast learning for conceptual coding. In ACM international conference on multimedia (pp.\u00a01\u20136). New York: ACM."},{"key":"18_CR186","first-page":"4297","volume-title":"ACM international conference on multimedia","author":"Y. Li","year":"2021","unstructured":"Li, Y., Wang, S., Zhang, X., Wang, S., Ma, S., & Wang, Y. (2021). Quality assessment of end-to-end learned image compression: the benchmark and objective measure. In ACM international conference on multimedia (pp.\u00a04297\u20134305). New York: ACM."},{"key":"18_CR187","first-page":"4401","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Karras","year":"2019","unstructured":"Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a04401\u20134410). Los Alamitos: IEEE."},{"key":"18_CR188","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"B. Zhou","year":"2017","unstructured":"Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., & Torralba, A. (2017). Scene parsing through ADE20K dataset. In IEEE\/CVF conference on computer vision and pattern recognition. Los Alamitos: IEEE."},{"key":"18_CR189","first-page":"7135","volume-title":"Advances in neural information processing systems","author":"A. Siarohin","year":"2019","unstructured":"Siarohin, A., Lathuili\u00e8re, S., Tulyakov, S., Ricci, E., & Sebe, N. (2019). First order motion model for image animation. In H. M. Wallach, H. Larochelle, A. Beygelzimer, et al. (Eds.), Advances in neural information processing systems (Vol.\u00a032, pp.\u00a07135\u20137145). Red Hook: Curran Associates."},{"key":"18_CR190","first-page":"4210","volume-title":"IEEE international conference on acoustics, speech and signal processing","author":"G. Konuko","year":"2021","unstructured":"Konuko, G., Valenzise, G., & Lathuili\u00e8re, S. (2021). Ultra-low bitrate video conferencing using deep image animation. In IEEE international conference on acoustics, speech and signal processing (pp.\u00a04210\u20134214). Los Alamitos: IEEE."},{"key":"18_CR191","volume-title":"H. 264 and MPEG-4 video compression: video coding for next-generation multimedia","author":"I. E. Richardson","year":"2004","unstructured":"Richardson, I. E. (2004). H. 264 and MPEG-4 video compression: video coding for next-generation multimedia. New York: Wiley."},{"key":"18_CR192","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"K. He","year":"2020","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In IEEE\/CVF conference on computer vision and pattern recognition. Los Alamitos: IEEE."},{"key":"18_CR193","doi-asserted-by":"crossref","unstructured":"Cao, Z., Hidalgo, G., Simon, T., Wei, S.-E., & Sheikh, Y. (2021). Openpose: realtime multi-person 2d pose estimation using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(1), 172\u2013186.","DOI":"10.1109\/TPAMI.2019.2929257"},{"key":"18_CR194","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Men","year":"2020","unstructured":"Men, Y., Mao, Y., Jiang, Y., Ma, W.-Y., & Lian, Z. (2020). Controllable person image synthesis with attribute-decomposed gan. In IEEE\/CVF conference on computer vision and pattern recognition. Los Alamitos: IEEE."},{"key":"18_CR195","unstructured":"van den Oord, A., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint. arXiv:1807.03748."},{"key":"18_CR196","unstructured":"Zablotskaia, P., Siarohin, A., Zhao, B., & Dwnet, L. S. (2019). Dense warp-based network for pose-guided human video generation. arXiv preprint. arXiv:1910.09139."},{"key":"18_CR197","first-page":"4230","volume-title":"ACM international conference on multimedia","author":"J. Li","year":"2021","unstructured":"Li, J., Jia, C., Zhang, X., Ma, S., & Gao, W. (2021). Cross modal compression: towards human-comprehensible semantic compression. In ACM international conference on multimedia (pp.\u00a04230\u20134238). New York: ACM."},{"key":"18_CR198","first-page":"3156","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"O. Vinyals","year":"2015","unstructured":"Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: a neural image caption generator. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a03156\u20133164). Los Alamitos: IEEE."},{"key":"18_CR199","first-page":"1316","volume-title":"IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Xu","year":"2018","unstructured":"Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., & He, X. (2018). AttnGAN: fine-grained text to image generation with attentional generative adversarial networks. In IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a01316\u20131324). Los Alamitos: IEEE."},{"key":"18_CR200","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1117\/12.19537","volume-title":"Image processing algorithms and techniques","author":"G. K. Wallace","year":"1990","unstructured":"Wallace, G. K. (1990). Overview of the JPEG (ISO\/CCITT) still image compression standard. In Image processing algorithms and techniques (Vol.\u00a01244, pp.\u00a0220\u2013233). Bellingham: SPIE."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-023-00018-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-023-00018-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-023-00018-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,2]],"date-time":"2023-08-02T09:25:30Z","timestamp":1690968330000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-023-00018-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,2]]},"references-count":200,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["18"],"URL":"https:\/\/doi.org\/10.1007\/s44267-023-00018-7","relation":{},"ISSN":["2731-9008"],"issn-type":[{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,2]]},"assertion":[{"value":"19 December 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 March 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 June 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 August 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"15"}}