{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:05:52Z","timestamp":1781103952865,"version":"3.54.1"},"reference-count":53,"publisher":"IGI Global Scientific Publishing","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,10,1]]},"abstract":"<p>Deep Neural Networks (DNNs) are best known for being the state-of-the-art in artificial intelligence (AI) applications including natural language processing (NLP), speech processing, computer vision, etc. In spite of all recent achievements of deep learning, it has yet to achieve semantic learning required to reason about the data. This lack of reasoning is partially imputed to the boorish memorization of patterns and curves from millions of training samples and ignoring the spatiotemporal relationships. The proposed framework puts forward a novel approach based on variational autoencoders (VAEs) by using the potential outcomes model and developing the counterfactual autoencoders. The proposed framework transforms any sort of multimedia input distributions to a meaningful latent space while giving more control over how the latent space is created. This allows us to model data that is better suited to answer inference-based queries, which is very valuable in reasoning-based AI applications.<\/p>","DOI":"10.4018\/ijmdem.2018100101","type":"journal-article","created":{"date-parts":[[2019,3,27]],"date-time":"2019-03-27T14:51:04Z","timestamp":1553698264000},"page":"1-20","source":"Crossref","is-referenced-by-count":2,"title":["Counterfactual Autoencoder for Unsupervised Semantic Learning"],"prefix":"10.4018","volume":"9","author":[{"given":"Saad","family":"Sadiq","sequence":"first","affiliation":[{"name":"University of Miami, Coral Gables, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0902-0844","authenticated-orcid":true,"given":"Mei-Ling","family":"Shyu","sequence":"additional","affiliation":[{"name":"University of Miami, Coral Gables, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel J.","family":"Feaster","sequence":"additional","affiliation":[{"name":"University of Miami - Miller School of Medicine, Miami, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"IJMDEM.2018100101-0","first-page":"39","article-title":"Neural module networks.","author":"J.Andreas","year":"2016","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"IJMDEM.2018100101-1","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1145\/1866029.1866080","article-title":"VizWiz: Nearly real-time answers to visual questions.","author":"J. P.Bigham","year":"2010","journal-title":"Proceedings of the 23rd Annual ACM Symposium on User Interface Software and Technology"},{"key":"IJMDEM.2018100101-2","first-page":"993","article-title":"Latent Dirichlet allocation.","volume":"3","author":"D. M.Blei","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"IJMDEM.2018100101-3","unstructured":"Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., & Lerchner, A. (2018). Understanding disentangling in \u03b2-VAE. arXiv:1804.03599"},{"key":"IJMDEM.2018100101-4","doi-asserted-by":"publisher","DOI":"10.1142\/S1793351X12400053"},{"key":"IJMDEM.2018100101-5","doi-asserted-by":"publisher","DOI":"10.1145\/2508037.2508042"},{"key":"IJMDEM.2018100101-6","doi-asserted-by":"publisher","DOI":"10.4018\/jmdem.2010111201"},{"key":"IJMDEM.2018100101-7","first-page":"441","article-title":"Temporal and spatial semantic models for multimedia presentations.","author":"S.-C.Chen","year":"1997","journal-title":"1997 International Symposium on Multimedia Information Processing"},{"issue":"6","key":"IJMDEM.2018100101-8","first-page":"772","article-title":"A dynamic user concept pattern learning framework for content-based image retrieval. IEEE Transactions on Systems, Man, and Cybernetics","volume":"36","author":"S.-C.Chen","year":"2006","journal-title":"Part C"},{"issue":"1","key":"IJMDEM.2018100101-9","first-page":"9","article-title":"Augmented transition network as a semantic model for video data.","volume":"3","author":"S.-C.Chen","year":"2000","journal-title":"International Journal of Networking and Information Systems"},{"key":"IJMDEM.2018100101-10","doi-asserted-by":"publisher","DOI":"10.1142\/S0218213001000738"},{"key":"IJMDEM.2018100101-11","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.1999.809783"},{"key":"IJMDEM.2018100101-12","first-page":"97","article-title":"How good are simple heuristics?","author":"J.Czerlinski","year":"1999","journal-title":"Simple Heuristics That Make Us Smart"},{"key":"IJMDEM.2018100101-13","unstructured":"Dilokthanakul, N., Mediano, P. A., Garnelo, M., Lee, M. C., Salimbeni, H., Arulkumaran, K., & Shanahan, M. (2016). Deep unsupervised clustering with gaussian mixture variational autoencoders. arXiv:1611.02648"},{"key":"IJMDEM.2018100101-14","first-page":"85","article-title":"Semantic technologies in IBM Watson.","author":"A.Gliozzo","year":"2013","journal-title":"Proceedings of the Fourth Workshop on Teaching NLP and CL"},{"key":"IJMDEM.2018100101-15","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., . . . Bengio, Y. (2014). Generative adversarial nets. In Advances in Neural Information Processing Systems (pp. 2672-2680)."},{"key":"IJMDEM.2018100101-16","unstructured":"Graves, A., Wayne, G., & Danihelka, I. (2014). Neural turing machines. Retrieved from http:\/\/arxiv.org\/abs\/1410.5401"},{"key":"IJMDEM.2018100101-17","doi-asserted-by":"publisher","DOI":"10.1162\/neco.2006.18.7.1527"},{"issue":"1","key":"IJMDEM.2018100101-18","first-page":"1303","article-title":"Stochastic variational inference.","volume":"14","author":"M. D.Hoffman","year":"2013","journal-title":"Journal of Machine Learning Research"},{"key":"IJMDEM.2018100101-19","unstructured":"Hosseini, H., & Poovendran, R. (2017). Deep neural networks do not recognize negative images. Retrieved from http:\/\/arxiv.org\/abs\/1703.06857"},{"key":"IJMDEM.2018100101-20","unstructured":"Kingma, D. P., Mohamed, S., Rezende, D. J., & Welling, M. (2014). Semi-supervised learning with deep generative models. In Advances in Neural Information Processing Systems (pp. 3581-3589)."},{"key":"IJMDEM.2018100101-21","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"IJMDEM.2018100101-22","doi-asserted-by":"publisher","DOI":"10.1214\/14-AOS1274"},{"key":"IJMDEM.2018100101-23","first-page":"1","article-title":"Image retrieval by color, texture, and spatial information.","author":"X.Li","year":"2002","journal-title":"Proceedings of the 8th International Conference on Distributed Multimedia Systems"},{"key":"IJMDEM.2018100101-24","doi-asserted-by":"publisher","DOI":"10.1504\/IJIDS.2012.047073"},{"key":"IJMDEM.2018100101-25","doi-asserted-by":"publisher","DOI":"10.4018\/jmdem.2013010103"},{"key":"IJMDEM.2018100101-26","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btp421"},{"key":"IJMDEM.2018100101-27","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2017.01.012"},{"key":"IJMDEM.2018100101-28","unstructured":"Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., & Frey, B. (2015). Adversarial autoencoders. arXiv:1511.05644"},{"issue":"2","key":"IJMDEM.2018100101-29","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1080\/10618600.2000.10474879","article-title":"Markov chain sampling methods for Dirichlet process mixture models.","volume":"9","author":"R. M.Neal","year":"2000","journal-title":"Journal of Computational and Graphical Statistics"},{"key":"IJMDEM.2018100101-30","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298640"},{"key":"IJMDEM.2018100101-31","unstructured":"Radford, A., Metz, L., & Chintala, S. (2015). Unsupervised representation learning with deep convolutional generative adversarial networks. Retrieved from http:\/\/arxiv.org\/abs\/1511.06434"},{"key":"IJMDEM.2018100101-32","first-page":"1009","article-title":"Factorial models and refiltering for speech separation and denoising.","author":"S. T.Roweis","year":"2003","journal-title":"Proceedings of the Eighth European Conference on Speech Communication and Technology"},{"key":"IJMDEM.2018100101-33","doi-asserted-by":"publisher","DOI":"10.1109\/BigMM.2017.56"},{"key":"IJMDEM.2018100101-34","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2016.87"},{"key":"IJMDEM.2018100101-35","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2017.25"},{"key":"IJMDEM.2018100101-36","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2018.00070"},{"key":"IJMDEM.2018100101-37","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijar.2008.11.006"},{"key":"IJMDEM.2018100101-38","unstructured":"Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., & Lillicrap, T. (2017). A simple neural network module for relational reasoning. In Advances in Neural Information Processing Systems (pp. 4967-4976). Retrieved from http:\/\/arxiv.org\/abs\/1706.01427"},{"key":"IJMDEM.2018100101-39","unstructured":"S\u00f8nderby, C. K., Raiko, T., Maal\u00f8e, L., S\u00f8nderby, S. K., & Winther, O. (2016). Ladder variational autoencoders. In Advances in Neural Information Processing Systems (pp. 3738-3746)."},{"key":"IJMDEM.2018100101-40","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177012031"},{"key":"IJMDEM.2018100101-41","unstructured":"Sukhbaatar, S., Weston, J., & Fergus, R. (2015). End-to-end memory networks. In Advances in Neural Information Processing Systems (pp. 2440-2448)."},{"key":"IJMDEM.2018100101-42","doi-asserted-by":"publisher","DOI":"10.1017\/S0140525X00003447"},{"key":"IJMDEM.2018100101-43","article-title":"Refinement of approximate domain theories by knowledge-based neural networks.","author":"G. G.Towell","year":"1990","journal-title":"Proceedings of the Eighth National Conference on Artificial Intelligence"},{"key":"IJMDEM.2018100101-44","doi-asserted-by":"publisher","DOI":"10.1126\/science.185.4157.1124"},{"key":"IJMDEM.2018100101-45","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_51"},{"key":"IJMDEM.2018100101-46","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2606428"},{"key":"IJMDEM.2018100101-47","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.96.022140"},{"key":"IJMDEM.2018100101-48","doi-asserted-by":"publisher","DOI":"10.4018\/IJMDEM.2017010101"},{"key":"IJMDEM.2018100101-49","doi-asserted-by":"publisher","DOI":"10.1109\/ISM.2015.126"},{"key":"IJMDEM.2018100101-50","unstructured":"Zhao, S., Song, J., & Ermon, S. (2017). Infovae: Information maximizing variational autoencoders. arXiv:1706.02262"},{"key":"IJMDEM.2018100101-51","doi-asserted-by":"publisher","DOI":"10.1109\/IRI.2011.6009579"},{"key":"IJMDEM.2018100101-52","doi-asserted-by":"publisher","DOI":"10.1109\/TETC.2014.2384992"}],"container-title":["International Journal of Multimedia Data Engineering and Management"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=226226","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T01:58:05Z","timestamp":1721095085000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/IJMDEM.2018100101"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2018,10,1]]},"references-count":53,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2018,10]]}},"URL":"https:\/\/doi.org\/10.4018\/ijmdem.2018100101","relation":{},"ISSN":["1947-8534","1947-8542"],"issn-type":[{"value":"1947-8534","type":"print"},{"value":"1947-8542","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,10,1]]}}}