{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T13:07:36Z","timestamp":1781096856590,"version":"3.54.1"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T00:00:00Z","timestamp":1781049600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/100020595","name":"National Science and Technology Council of Taiwan","doi-asserted-by":"crossref","award":["113-2221-E-008-100-MY3"],"award-info":[{"award-number":["113-2221-E-008-100-MY3"]}],"id":[{"id":"10.13039\/100020595","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Manage. Inf. Syst."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Adopting large-scale AI models in enterprise information systems is often hindered by high training costs and long development cycles, posing a significant managerial challenge. The standard end-to-end backpropagation (BP) algorithm is a primary driver of modern AI, but it is also the source of inefficiency in training deep networks. This article introduces a new training methodology, Supervised Contrastive Parallel Learning (SCPL), that addresses this issue by decoupling BP and transforming a long gradient flow into multiple short ones. This design enables the simultaneous computation of parameter gradients in different layers, achieving superior model parallelism and enhancing training throughput. Detailed experiments are presented to demonstrate the efficiency and effectiveness of our model compared to BP, Early Exit, GPipe, and Associated Learning (AL), a state-of-the-art method for decoupling backpropagation. By mitigating a fundamental performance bottleneck, SCPL provides a practical pathway for organizations to develop and deploy advanced information systems more cost-effectively and with greater agility. The experimental code is released for reproducibility.\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                  <\/jats:p>","DOI":"10.1145\/3793534","type":"journal-article","created":{"date-parts":[[2026,1,26]],"date-time":"2026-01-26T12:00:56Z","timestamp":1769428856000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["SCPL: Enhancing Neural Network Training Throughput with Decoupled Local Losses and Model Parallelism"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-9918-5559","authenticated-orcid":false,"given":"Ming-Yao","family":"Ho","sequence":"first","affiliation":[{"name":"Computer Science and Information Engineering, National Central University","place":["Taoyuan City, Taiwan"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6669-600X","authenticated-orcid":false,"given":"Cheng-Kai","family":"Wang","sequence":"additional","affiliation":[{"name":"Computer Science and Information Engineering, National Central University","place":["Taoyuan City, Taiwan"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3528-905X","authenticated-orcid":false,"given":"You-Teng","family":"Lin","sequence":"additional","affiliation":[{"name":"Computer Science and Information Engineering, National Central University","place":["Taoyuan, Taiwan"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5137-4449","authenticated-orcid":false,"given":"Hung-Hsuan","family":"Chen","sequence":"additional","affiliation":[{"name":"Computer Science and Information Engineering, National Central University","place":["Taoyuan, Taiwan"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,10]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"244","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Arora Sanjeev","year":"2018","unstructured":"Sanjeev Arora, Nadav Cohen, and Elad Hazan. 2018. On the optimization of deep networks: Implicit acceleration by overparameterization. In Proceedings of the International Conference on Machine Learning. PMLR, 244\u2013253."},{"key":"e_1_3_2_3_2","article-title":"Learning representations by maximizing mutual information across views","author":"Bachman Philip","year":"2019","unstructured":"Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019. Learning representations by maximizing mutual information across views. Advances in Neural Information Processing Systems. Curran Associates Inc., Vancouver, BC, Canada. 15509\u201315519.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_4_2","unstructured":"Yoshua Bengio. 2014. How auto-encoders could provide credit assignment in deep networks via target propagation. Retrieved from https:\/\/arxiv.org\/abs\/1407.7906"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3285954"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"89","DOI":"10.5220\/0009885600890097","volume-title":"Proceedings of the International Conference on Deep Learning Theory and Applications","author":"Chen Pu","year":"2020","unstructured":"Pu Chen and Hung-Hsuan Chen. 2020. Accelerating matrix factorization by overparameterization. In Proceedings of the International Conference on Deep Learning Theory and Applications. 89\u201397."},{"key":"e_1_3_2_7_2","first-page":"1597","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_2_8_2","unstructured":"Ben Cottier Tamay Besiroglu David Owen Lennart Heim and Jaime Sevilla. 2024. How Much Does It Cost to Train Frontier AI Models? Epoch AI. Retrieved January 10 2026 from https:\/\/epoch.ai\/blog\/how-much-does-it-cost-to-train-frontier-ai-models"},{"key":"e_1_3_2_9_2","first-page":"4182","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Henaff Olivier","year":"2020","unstructured":"Olivier Henaff. 2020. Data-efficient image recognition with contrastive predictive coding. In Proceedings of the International Conference on Machine Learning. PMLR, 4182\u20134192."},{"key":"e_1_3_2_10_2","article-title":"Gpipe: Efficient training of giant neural networks using pipeline parallelism","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, et\u00a0al. 2019. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA. 103\u2013112.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_11_2","article-title":"DeInfoReg: A decoupled learning framework for better training throughput","volume":"651","author":"Huang Zih-Hao","year":"2025","unstructured":"Zih-Hao Huang, You-Teng Lin, and Hung-Hsuan Chen. 2025. DeInfoReg: A decoupled learning framework for better training throughput. Neurocomputing 651, C(2025), 130813.","journal-title":"Neurocomputing"},{"key":"e_1_3_2_12_2","volume-title":"The CEO\u2019s Guide to Generative AI: Cost of Compute","author":"Value IBM Institute for Business","year":"2024","unstructured":"IBM Institute for Business Value. 2024. The CEO\u2019s Guide to Generative AI: Cost of Compute. Technical Report. IBM Corporation. Retrieved January 10, 2026 from https:\/\/www.ibm.com\/think\/insights\/ai-economics-compute-cost"},{"key":"e_1_3_2_13_2","first-page":"1627","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jaderberg Max","year":"2017","unstructured":"Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu. 2017. Decoupled neural interfaces using synthetic gradients. In Proceedings of the International Conference on Machine Learning. PMLR, 1627\u20131635."},{"issue":"1","key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"174","DOI":"10.1162\/neco_a_01335","article-title":"Associated learning: Decomposing end-to-end backpropagation based on autoencoders and target propagation","volume":"33","author":"Kao Yu-Wei","year":"2021","unstructured":"Yu-Wei Kao and Hung-Hsuan Chen. 2021. Associated learning: Decomposing end-to-end backpropagation based on autoencoders and target propagation. Neural Computation 33, 1 (2021), 174\u2013193.","journal-title":"Neural Computation"},{"key":"e_1_3_2_15_2","article-title":"Supervised contrastive learning","author":"Khosla Prannay","year":"2020","unstructured":"Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. Advances in Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA. 33 (2020), 18661\u201318673.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_16_2","first-page":"498","volume-title":"Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases","author":"Lee Dong-Hyun","year":"2015","unstructured":"Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio. 2015. Difference target propagation. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 498\u2013515."},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Shen Li Yanli Zhao Rohan Varma Omkar Salpekar Pieter Noordhuis Teng Li Adam Paszke Jeff Smith Brian Vaughan Pritam Damania and Soumith Chintala. 2020. Pytorch distributed: Experiences on accelerating data parallel training. In Proceedings of the VLDB Endow 13 12 (Aug. 2020) 3005\u20133018.","DOI":"10.14778\/3415478.3415530"},{"key":"e_1_3_2_18_2","article-title":"Putting an end to end-to-end: Gradient-isolated learning of representations","author":"L\u00f6we Sindy","year":"2019","unstructured":"Sindy L\u00f6we, Peter O\u2019Connor, and Bastiaan Veeling. 2019. Putting an end to end-to-end: Gradient-isolated learning of representations. Advances in Neural Information Processing Systems. Curran Associates Inc., Red Hook, NY, USA. 3033\u20133045.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"7","key":"e_1_3_2_19_2","first-page":"1","article-title":"Target propagation in recurrent neural networks.","volume":"21","author":"Manchev Nikolay","year":"2020","unstructured":"Nikolay Manchev and Michael W. Spratling. 2020. Target propagation in recurrent neural networks. Journal of Machine Learning Research 21, 7 (2020), 1\u201333.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_20_2","first-page":"20024","article-title":"A theoretical framework for target propagation","volume":"33","author":"Meulemans Alexander","year":"2020","unstructured":"Alexander Meulemans, Francesco Carzaniga, Johan Suykens, Jo\u00e3o Sacramento, and Benjamin F. Grewe. 2020. A theoretical framework for target propagation. Advances in Neural Information Processing Systems 33 (2020), 20024\u201320036.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_21_2","series-title":"Proceedings of Machine Learning Research","first-page":"70","volume-title":"\u201cI Can\u2019t Believe It\u2019s Not Better!\u201d at NeurIPS Workshops, Virtual","volume":"137","author":"Mitrovic Jovana","year":"2020","unstructured":"Jovana Mitrovic, Brian McWilliams, and M\u00e9lanie Rey. 2020. Less can be more in contrastive learning. In \u201cI Can\u2019t Believe It\u2019s Not Better!\u201d at NeurIPS Workshops, Virtual(Proceedings of Machine Learning Research, Vol. 137). PMLR, 70\u201375."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_2_23_2","first-page":"1532","volume-title":"Proceedings of the Empirical Methods in Natural Language Processing","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the Empirical Methods in Natural Language Processing. 1532\u20131543."},{"key":"e_1_3_2_24_2","first-page":"1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Rajbhandari Samyam","year":"2020","unstructured":"Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1\u201316."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406703"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1038\/323533a0"},{"key":"e_1_3_2_27_2","unstructured":"Christopher J. Shallue Jaehoon Lee Joseph Antognini Jascha Sohl-Dickstein Roy Frostig and George E. Dahl. 2019. Measuring the effects of data parallelism on neural network training. Journal of Machine Learning Research 20 112 (2019) 1\u201349."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_29_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Wu Dennis YH","year":"2022","unstructured":"Dennis YH Wu, Dinan Lin, Vincent Chen, and Hung-Hsuan Chen. 2022. Associated learning: An alternative to end-to-end backpropagation that works on CNN, RNN, and transformer. In Proceedings of the International Conference on Learning Representations. 18 pages."},{"key":"e_1_3_2_30_2","first-page":"11142","article-title":"Loco: Local contrastive representation learning","volume":"33","author":"Xiong Yuwen","year":"2020","unstructured":"Yuwen Xiong, Mengye Ren, and Raquel Urtasun. 2020. Loco: Local contrastive representation learning. Advances in Neural Information Processing Systems 33 (2020), 11142\u201311153.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2023.04.033"}],"container-title":["ACM Transactions on Management Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793534","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T12:36:31Z","timestamp":1781094991000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793534"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,10]]},"references-count":30,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3793534"],"URL":"https:\/\/doi.org\/10.1145\/3793534","relation":{},"ISSN":["2158-656X","2158-6578"],"issn-type":[{"value":"2158-656X","type":"print"},{"value":"2158-6578","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,10]]},"assertion":[{"value":"2024-11-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}