{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:09:45Z","timestamp":1750219785543,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"9","license":[{"start":{"date-parts":[[2023,6,15]],"date-time":"2023-06-15T00:00:00Z","timestamp":1686787200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>The lack of scalability of data annotation translates to the need to decrease dependency on labels. Self-supervision offers a solution with data training themselves. However, it has received relatively less attention on tabular data, data that drive a large proportion of business and application domains. This work, which we name the Statistical Self-Supervisor (SSS), proposes a method for self-supervision on tabular data by defining a continuous perturbation as pretext. It enables a neural network to learn representations by learning to predict the level of additive isotropic Gaussian noise added to inputs. The choice of the pretext transformation is motivated by intrinsic characteristics of a neural network fundamentally performing linear fits under the widely adopted assumption of Gaussianity in its fitting error and the preservation of locality of a data example on the data manifold in the presence of small random perturbations. The transform condenses information in the generated representations, making them better employable for further task-specific prediction as evidenced by performance improvement of the downstream classifier. To evaluate the persistence of performance under low-annotation settings, SSS is evaluated against different levels of label availability to the downstream classifier (1% to 100%) and benchmarked against self- and semi-supervised methods. At the most label-constrained, 1% setting, we report a maximum increase of at least 2.5% against the next-best semi-supervised competing method. We report an increase of more than 1.5% against self-supervised state of the art. Ablation studies also reveal that increasing label availability from 0% to 1% results in a maximum increase of up to 50% on either of the five performance metrics and up to 15% thereafter, indicating diminishing returns in additional annotation.<\/jats:p>","DOI":"10.1145\/3594720","type":"journal-article","created":{"date-parts":[[2023,5,1]],"date-time":"2023-05-01T12:07:56Z","timestamp":1682942876000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Self-supervision for Tabular Data by Learning to Predict Additive Homoskedastic Gaussian Noise as Pretext"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0638-9689","authenticated-orcid":false,"given":"Tahir","family":"Syed","sequence":"first","affiliation":[{"name":"Institute of Business Administration Karachi, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0899-1526","authenticated-orcid":false,"given":"Behroz","family":"Mirza","sequence":"additional","affiliation":[{"name":"National University of Computer and Emerging Sciences, Karachi, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,15]]},"reference":[{"unstructured":"Muhammad Ahmad Behroz Mirza Behraj Khan and Tahir Syed. 2022. Task memorization for incremental learning with a common neural network. Retrieved from https:\/\/www.researchgate.net\/profile\/Tahir-Syed\/publication\/339165067_Task_memorization_for_incremental_learning_with_a_common_neural_network_architecture\/links\/5e9f300292851c2f52ba40ef\/Task-memorization-for-incremental-learning-with-a-common-neural-network-architecture.pdf.","key":"e_1_3_1_2_2"},{"unstructured":"E. Alpaydin. 1996. Pen based Recognition Dataset. http:\/\/archive.ics.uci.edu\/ml\/datasets\/pen-based+recognition+of+handwritten+digits. [Online; accessed -2019].","key":"e_1_3_1_3_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_4_2","DOI":"10.1609\/aaai.v35i8.16826"},{"key":"e_1_3_1_5_2","article-title":"Mixmatch: A holistic approach to semi-supervised learning","author":"Berthelot David","year":"2019","unstructured":"David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A. Raffel. 2019. Mixmatch: A holistic approach to semi-supervised learning. InAdvances in Neural Information Processing Systems (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_6_2","first-page":"92","article-title":"Combining labeled and unlabeled data with co-training","author":"Blum Avrim","year":"1998","unstructured":"Avrim Blum and Tom Mitchell. 1998. Combining labeled and unlabeled data with co-training. In Proceedings of the 11th Annual Conference on Computational Learning Theory (1998), 92\u2013100.","journal-title":"Proceedings of the 11th Annual Conference on Computational Learning Theory"},{"unstructured":"Chen Ting Simon Kornblith Mohammad Norouzi and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning . PMLR 1597\u20131607.","key":"e_1_3_1_7_2"},{"unstructured":"Noam Chomsky et\u00a0al. 2006. On cognitive structures and their development: A reply to Piaget. In Philosophy of Mind: Classical Problems\/Contemporary Issues . Routledge and Kegan Paul 751\u2013755.","key":"e_1_3_1_8_2"},{"key":"e_1_3_1_9_2","first-page":"9268","article-title":"Class-balanced loss based on effective number of samples","author":"Cui Yin","year":"2019","unstructured":"Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie. 2019. Class-balanced loss based on effective number of samples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9268\u20139277.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_10_2","first-page":"1422","article-title":"Unsupervised visual representation learning by context prediction","author":"Doersch Carl","year":"2015","unstructured":"Carl Doersch, Abhinav Gupta, and Alexei Efros. 2015. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE International Conference on Computer Vision (2015), 1422\u20131430.","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"e_1_3_1_11_2","article-title":"DREAM Architecture: A developmental approach to open-ended learning in robotics","author":"Doncieux Stephane","year":"2020","unstructured":"Stephane Doncieux, Nicolas Bredeche, L\u00e9ni Le Goff, Beno\u00eet Girard, Alexandre Coninx, Olivier Sigaud, Mehdi Khamassi, Natalia D\u00edaz-Rodr\u00edguez, David Filliat, Timothy Hospedales, et\u00a0al. 2020. DREAM Architecture: A developmental approach to open-ended learning in robotics. arXiv:2005.06223. Retrieved from https:\/\/arxiv.org\/abs\/2005.06223.","journal-title":"arXiv:2005.06223"},{"issue":"4","key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"817","DOI":"10.1016\/j.patcog.2012.09.023","article-title":"Tree ensembles for predicting structured outputs","volume":"46","author":"Dragi Kocev","year":"2013","unstructured":"Kocev Dragi, Celine Vens, Jan Struyf, and Sa\u0161o D\u017eeroski. 2013. Tree ensembles for predicting structured outputs. Pattern Recogn. 46, 4 (2013), 817\u2013833.","journal-title":"Pattern Recogn."},{"unstructured":"Nikos Komodakis and Spyros Gidaris. 2018. Unsupervised representation learning by predicting image rotations. In International Conference on Learning Representations (ICLR\u201918) .","key":"e_1_3_1_13_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_14_2","DOI":"10.1016\/j.neucom.2012.12.056"},{"unstructured":"Chuan Guo Jared Frank and Kilian Weinberger. 2020. Low frequency adversarial perturbation In Uncertainty in Artificial Intelligence . 1127\u20131137.","key":"e_1_3_1_15_2"},{"issue":"2","key":"e_1_3_1_16_2","first-page":"95","article-title":"Co-training by committee: A generalized framework for semi-supervised learning with committees","volume":"2","author":"Hady Mohamed","year":"2008","unstructured":"Mohamed Hady, Abdel Farouk, and Friedhelm Schwenker. 2008. Co-training by committee: A generalized framework for semi-supervised learning with committees. Int. J. Softw. Inf. 2, 2 (2008), 95\u2013124.","journal-title":"Int. J. Softw. Inf."},{"unstructured":"Kaiming He Haoqi Fan Yuxin Wu Saining Xie and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 9729\u20139738.","key":"e_1_3_1_17_2"},{"unstructured":"Kaggle. 2011. Give Me Some Credit. https:\/\/www.kaggle.com\/datasets\/brycecf\/give-me-some-credit-dataset. [Online; accessed - 2019].","key":"e_1_3_1_18_2"},{"unstructured":"Machine Learning Group. 2013. Credit Card Fraud Detection. https:\/\/www.kaggle.com\/datasets\/mlg-ulb\/creditcardfraud. [Online; accessed -2019].","key":"e_1_3_1_19_2"},{"unstructured":"Jeff Schlimmer. 1987. Mushroom Dataset. https:\/\/archive.ics.uci.edu\/ml\/datasets\/Mushroom. [Online; accessed -2019].","key":"e_1_3_1_20_2"},{"unstructured":"Ronny Kohavi. 1996. Adult Income Dataset. https:\/\/archive.ics.uci.edu\/ml\/datasets\/Adult\/. [Online; accessed -2019].","key":"e_1_3_1_21_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_22_2","DOI":"10.3233\/JIFS-169571"},{"unstructured":"Yann LeCun. 2019. Retrieved from https:\/\/twitter.com\/ylecun\/status\/1140445577408327683.","key":"e_1_3_1_23_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_24_2","DOI":"10.1007\/s10844-017-0457-4"},{"doi-asserted-by":"publisher","key":"e_1_3_1_25_2","DOI":"10.1109\/TPAMI.2015.2452921"},{"key":"e_1_3_1_26_2","article-title":"From softmax to sparsemax: A sparse model of attention and multi-label classification","author":"Martins Andre","year":"2016","unstructured":"Andre Martins and Ramon Astudillo. 2016. From softmax to sparsemax: A sparse model of attention and multi-label classification. In Proceedings of the International Conference on Machine Learning.","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"e_1_3_1_27_2","article-title":"In proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Misra Ishan","year":"2020","unstructured":"Ishan Misra and Laurens van der Maaten. 2020. In proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. In Advances in Neural Information Processing Systems.","journal-title":"Advances in Neural Information Processing Systems"},{"doi-asserted-by":"publisher","key":"e_1_3_1_28_2","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"e_1_3_1_29_2","first-page":"2536","article-title":"Context encoders: Feature learning by in painting","author":"Pathak Deepak","year":"2016","unstructured":"Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. 2016. Context encoders: Feature learning by in painting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2536\u20132544.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_30_2","article-title":"Context encoders: Feature learning by inpainting","author":"Pathak Deepak","year":"2016","unstructured":"Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei Efros. 2016. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_31_2","article-title":"The Synthetic Data Vault: Generative Modeling for Relational Databases","author":"Patki Neha","year":"2016","unstructured":"Neha Patki. 2016. The Synthetic Data Vault: Generative Modeling for Relational Databases. Ph.D. Dissertation.","journal-title":"Ph.D. Dissertation"},{"unstructured":"Oliver Roesler. 2013. EEG Eye State Dataset. https:\/\/archive.ics.uci.edu\/ml\/datasets\/EEG+Eye+State. [Online; accessed -2019].","key":"e_1_3_1_32_2"},{"key":"e_1_3_1_33_2","first-page":"1163","article-title":"Regularization with stochastic transformations and perturbations for deep semi-supervised learning","author":"Sajjadi Mehdi","year":"2016","unstructured":"Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. 2016. Regularization with stochastic transformations and perturbations for deep semi-supervised learning. In Advances in Neural Information Processing Systems. 1163\u20131171.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"e_1_3_1_34_2","first-page":"335","article-title":"Semi-supervised self-training for decision tree classifiers","volume":"8","author":"Tanha Jafar","year":"2018","unstructured":"Jafar Tanha, Maarten van Someren, and Hamideh Afsarmanesh. 2018. Semi-supervised self-training for decision tree classifiers. Int. J. Mach. Learn. Cybernet. 8, 1 (2018), 335\u2013370.","journal-title":"Int. J. Mach. Learn. Cybernet."},{"key":"e_1_3_1_35_2","first-page":"1195","article-title":"Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results","author":"Tarvainen Antti","year":"2017","unstructured":"Antti Tarvainen and Harri Valpola. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in Neural Information Processing Systems. 1195\u20131204.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_36_2","article-title":"SubTab: Subsetting features of tabular data for self-supervised representation learning","author":"Ucar Talip","year":"2021","unstructured":"Talip Ucar, Ehsan Hajiramezanali, and Lindsay Edwards. 2021. SubTab: Subsetting features of tabular data for self-supervised representation learning. In Proceedings of the 35th Conference on Neural Information Processing Systems.","journal-title":"Proceedings of the 35th Conference on Neural Information Processing Systems"},{"unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. (unpublished).","key":"e_1_3_1_37_2"},{"key":"e_1_3_1_38_2","article-title":"Extracting and composing robust features with denoising autoencoders","author":"Vincent Pascal","year":"2008","unstructured":"Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th International Conference on Machine Learning. ACM.","journal-title":"Proceedings of the 25th International Conference on Machine Learning"},{"key":"e_1_3_1_39_2","first-page":"8814","article-title":"PRNet: Self-supervised learning for partial-to-partial registration","author":"Wang Yue","year":"2019","unstructured":"Yue Wang and Justin M. Solomon. 2019. PRNet: Self-supervised learning for partial-to-partial registration. In Advances in Neural Information Processing Systems. 8814\u20138826.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_40_2","first-page":"3733","article-title":"Unsupervised feature learning via non-parametric instance discrimination","author":"Wu Zhirong","year":"2018","unstructured":"Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. 2018. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3733\u201313742.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_41_2","article-title":"S4l: Self-supervised semi-supervised learning","author":"Xiaohua Zhai","year":"2019","unstructured":"Zhai Xiaohua, Avital Oliver, Alexander Kolesnikov, and Lucas Beyer. 2019. S4l: Self-supervised semi-supervised learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision.","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"e_1_3_1_42_2","article-title":"VIME: Extending the success of self-and semi-supervised learning to tabular domain","author":"Yoon Jinsung","year":"2020","unstructured":"Jinsung Yoon, Yao Zhang, James Jordon, and Mihaela van der Schaar. 2020. VIME: Extending the success of self-and semi-supervised learning to tabular domain. In Advances in Neural Information Processing Systems.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","article-title":"mixup: Beyond empirical risk minimization","author":"Zhang Hongyi","year":"2017","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann Dauphin, and David Lopez-Paz. 2017. mixup: Beyond empirical risk minimization. arXiv:1710.09412. Retrieved from https:\/\/arxiv.org\/abs\/1710.09412.","journal-title":"arXiv:1710.09412"},{"key":"e_1_3_1_44_2","first-page":"649","article-title":"Colorful image colorization","author":"Zhang Richard","year":"2016","unstructured":"Richard Zhang, Phillip Isola, and Alexei A. Efros. 2016. Colorful image colorization. InProceedings of the European Conference on Computer Vision. 649\u2013666.","journal-title":"Proceedings of the European Conference on Computer Vision"},{"key":"e_1_3_1_45_2","first-page":"1058","article-title":"Split-brain autoencoders: Unsupervised learning by cross-channel prediction","author":"Zhang Richard","year":"2017","unstructured":"Richard Zhang, Phillip Isola, and Alexei A. Efros. 2017. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1058\u20131067.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_1_46_2","volume-title":"Semi-supervised Learning Literature Survey","author":"Zhu Xiaojin Jerry","year":"2005","unstructured":"Xiaojin Jerry Zhu. 2005. Semi-supervised Learning Literature Survey. Technical Report. Department of Computer Sciences, University of Wisconsin\u2014Madison."}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594720","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3594720","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:50Z","timestamp":1750178270000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594720"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,15]]},"references-count":45,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3594720"],"URL":"https:\/\/doi.org\/10.1145\/3594720","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2023,6,15]]},"assertion":[{"value":"2021-12-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}