{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T21:45:03Z","timestamp":1778535903348,"version":"3.51.4"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,2]]},"abstract":"<jats:p>The process of training deep learning models produces a huge amount of meta-data, including but not limited to losses, hidden feature embeddings, and gradients. Model diagnosis tools have been developed to analyze losses and feature embeddings with the aim to improve the performance of these models. However, gradients, despite carrying rich information that is potentially relevant for model interpretation and data debugging, have yet to be fully explored due to their size and complexity. Each single gradient has a size as large as the number of parameters of the neural net - often measured in the tens of millions. This makes it extremely challenging to efficiently collect, store, and analyze large numbers of gradients in these models. In this work, we develop MetaStore to fill this gap. MetaStore leverages our observation that storing certain compact intermediate results produced in the back propagation process, namely, the prefix and suffix gradients, is sufficient for the exact restoration of the original gradient. These prefix and suffix gradients are much more compact than the original gradients, thus allowing us to address the gradient collection and storage challenges. Furthermore, MetaStore features a rich set of analytics operators that allow the users to analyze the gradients for data debugging or model interpretation. Rather than first having to restore the original gradients and then run analytics on top of this decompressed view, MetaStore directly executes these operators on the compact prefix and suffix structures, making gradient-based analytics efficient and scalable. Our experiments on popular deep learning models such as VGG, BERT, and ResNet and benchmark image and text datasets demonstrate that MetaStore outperforms strong baseline methods from 4 to 678x in storage costs and from 2 to 1000x in running time.<\/jats:p>","DOI":"10.14778\/3648160.3648182","type":"journal-article","created":{"date-parts":[[2024,5,3]],"date-time":"2024-05-03T21:52:53Z","timestamp":1714773173000},"page":"1446-1459","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["MetaStore: Analyzing Deep Learning Meta-Data at Scale"],"prefix":"10.14778","volume":"17","author":[{"given":"Huayi","family":"Zhang","sequence":"first","affiliation":[{"name":"WPI, Data Science, Worcester, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Binwei","family":"Yan","sequence":"additional","affiliation":[{"name":"MIT, Cambridge, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Cao","sequence":"additional","affiliation":[{"name":"U of Arizona, CS; MIT, CSAIL, Cambridge, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Samuel","family":"Madden","sequence":"additional","affiliation":[{"name":"MIT, CSAIL, Cambridge, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elke","family":"Rundensteiner","sequence":"additional","affiliation":[{"name":"WPI, Computer Science, Worcester, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2976749.2978318"},{"key":"e_1_2_1_2_1","volume-title":"Sparse communication for distributed gradient descent. arXiv preprint arXiv:1704.05021","author":"Aji A. F.","year":"2017","unstructured":"A. F. Aji and K. Heafield. Sparse communication for distributed gradient descent. arXiv preprint arXiv:1704.05021, 2017."},{"key":"e_1_2_1_3_1","volume-title":"Qsgd: Communication-efficient sgd via gradient quantization and encoding. Advances in neural information processing systems, 30","author":"Alistarh D.","year":"2017","unstructured":"D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic. Qsgd: Communication-efficient sgd via gradient quantization and encoding. Advances in neural information processing systems, 30, 2017."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702509"},{"key":"e_1_2_1_5_1","first-page":"32","article-title":"Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations","author":"Basu D.","year":"2019","unstructured":"D. Basu, D. Data, C. Karakus, and S. Diggavi. Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations. Advances in Neural Information Processing Systems, 32, 2019.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_6_1","first-page":"560","volume-title":"International Conference on Machine Learning","author":"Bernstein J.","year":"2018","unstructured":"J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar. signsgd: Compressed optimisation for non-convex problems. In International Conference on Machine Learning, pages 560--569. PMLR, 2018."},{"key":"e_1_2_1_7_1","first-page":"22234","article-title":"Evograd: Efficient gradient-based meta-learning and hyperparameter optimization","volume":"34","author":"Bohdal O.","year":"2021","unstructured":"O. Bohdal, Y. Yang, and T. Hospedales. Evograd: Efficient gradient-based meta-learning and hyperparameter optimization. Advances in Neural Information Processing Systems, 34:22234--22246, 2021.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_8_1","unstructured":"L. Cao Y. Yan Y. Wang S. Madden and E. A. Rundensteiner. Autood: Automatic outlier detection. In SIGMOD."},{"key":"e_1_2_1_9_1","volume-title":"Distribution density, tails, and outliers in machine learning: Metrics and applications. arXiv preprint arXiv:1910.13427","author":"Carlini N.","year":"2019","unstructured":"N. Carlini, U. Erlingsson, and N. Papernot. Distribution density, tails, and outliers in machine learning: Metrics and applications. arXiv preprint arXiv:1910.13427, 2019."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11728"},{"key":"e_1_2_1_11_1","volume-title":"On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259","author":"Cho K.","year":"2014","unstructured":"K. Cho, B. Van Merri\u00ebnboer, D. Bahdanau, and Y. Bengio. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259, 2014."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019","volume":"1","author":"Devlin J.","year":"2019","unstructured":"J. Devlin, M. Chang, K. Lee, and K. Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171--4186, 2019."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1162\/089976600300015015"},{"key":"e_1_2_1_14_1","first-page":"2242","volume-title":"International conference on machine learning","author":"Ghorbani A.","year":"2019","unstructured":"A. Ghorbani and J. Zou. Data shapley: Equitable valuation of data for machine learning. In International conference on machine learning, pages 2242--2251. PMLR, 2019."},{"key":"e_1_2_1_15_1","volume-title":"Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572","author":"Goodfellow I. J.","year":"2014","unstructured":"I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485460"},{"key":"e_1_2_1_17_1","volume-title":"Deep residual learning for image recognition. CoRR, abs\/1512.03385","author":"He K.","year":"2015","unstructured":"K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. CoRR, abs\/1512.03385, 2015."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/3554821.3554880"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","first-page":"5250","DOI":"10.1109\/BigData59044.2023.10386694","volume-title":"2023 IEEE International Conference on Big Data (BigData)","author":"Hu R.","year":"2023","unstructured":"R. Hu, D. Zhang, D. Tao, H. Zhang, H. Feng, and E. Rundensteiner. Uce-fid: Using large unlabeled, medium crowdsourced-labeled, and small expert-labeled tweets for foodborne illness detection. In 2023 IEEE International Conference on Big Data (BigData), pages 5250--5259. IEEE, 2023."},{"key":"e_1_2_1_21_1","first-page":"448","volume-title":"International conference on machine learning","author":"Ioffe S.","year":"2015","unstructured":"S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448--456. PMLR, 2015."},{"key":"e_1_2_1_22_1","series-title":"Proceedings of Machine Learning Research","first-page":"1167","volume-title":"The 22nd International Conference on Artificial Intelligence and Statistics, AISTATS","author":"Jia R.","year":"2019","unstructured":"R. Jia, D. Dao, B. Wang, F. A. Hubis, N. Hynes, N. M. G\u00fcrel, B. Li, C. Zhang, D. Song, and C. J. Spanos. Towards efficient data valuation based on the shapley value. In K. Chaudhuri and M. Sugiyama, editors, The 22nd International Conference on Artificial Intelligence and Statistics, AISTATS 2019, 16-18 April 2019, Naha, Okinawa, Japan, volume 89 of Proceedings of Machine Learning Research, pages 1167--1176. PMLR, 2019."},{"key":"e_1_2_1_23_1","volume-title":"Autolrs: Automatic learning-rate schedule by bayesian optimization on the fly. arXiv preprint arXiv:2105.10762","author":"Jin Y.","year":"2021","unstructured":"Y. Jin, T. Zhou, L. Zhao, Y. Zhu, C. Guo, M. Canini, and A. Krishnamurthy. Autolrs: Automatic learning-rate schedule by bayesian optimization on the fly. arXiv preprint arXiv:2105.10762, 2021."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939502.2939503"},{"key":"e_1_2_1_25_1","volume-title":"Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492","author":"Kone\u010dn\u1ef3 J.","year":"2016","unstructured":"J. Kone\u010dn\u1ef3, H. B. McMahan, F. X. Yu, P. Richt\u00e1rik, A. T. Suresh, and D. Bacon. Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492, 2016."},{"key":"e_1_2_1_26_1","volume-title":"Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887","author":"Lin Y.","year":"2017","unstructured":"Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally. Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887, 2017."},{"key":"e_1_2_1_27_1","article-title":"Activated gradients for deep neural networks","author":"Liu M.","year":"2021","unstructured":"M. Liu, L. Chen, X. Du, L. Jin, and M. Shang. Activated gradients for deep neural networks. IEEE Transactions on Neural Networks and Learning Systems, 2021.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_2_1_28_1","volume-title":"ECCV","author":"Matthew Zeiler D.","year":"2014","unstructured":"D. Matthew Zeiler and F. Rob. Visualizing and understanding convolutional neural networks. ECCV, 2014."},{"key":"e_1_2_1_29_1","first-page":"17044","article-title":"Identifying mislabeled data using the area under the margin ranking","volume":"33","author":"Pleiss G.","year":"2020","unstructured":"G. Pleiss, T. Zhang, E. Elenberg, and K. Q. Weinberger. Identifying mislabeled data using the area under the margin ranking. Advances in Neural Information Processing Systems, 33:17044--17056, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_30_1","first-page":"19920","article-title":"Estimating training data influence by tracing gradient descent","volume":"33","author":"Pruthi G.","year":"2020","unstructured":"G. Pruthi, F. Liu, S. Kale, and M. Sundararajan. Estimating training data influence by tracing gradient descent. Advances in Neural Information Processing Systems, 33:19920--19930, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_31_1","first-page":"4334","volume-title":"International conference on machine learning","author":"Ren M.","year":"2018","unstructured":"M. Ren, W. Zeng, B. Yang, and R. Urtasun. Learning to reweight examples for robust deep learning. In International conference on machine learning, pages 4334--4343. PMLR, 2018."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00445"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_2_1_34_1","first-page":"5739","volume-title":"International Conference on Machine Learning","author":"Shen Y.","year":"2019","unstructured":"Y. Shen and S. Sanghavi. Learning with bad training data via iterative trimmed loss minimization. In International Conference on Machine Learning, pages 5739--5748. PMLR, 2019."},{"key":"e_1_2_1_35_1","volume-title":"Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034","author":"Simonyan K.","year":"2013","unstructured":"K. Simonyan, A. Vedaldi, and A. Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034, 2013."},{"key":"e_1_2_1_36_1","volume-title":"3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015."},{"key":"e_1_2_1_37_1","volume-title":"Dataset cartography: Mapping and diagnosing datasets with training dynamics. arXiv preprint arXiv:2009.10795","author":"Swayamdipta S.","year":"2020","unstructured":"S. Swayamdipta, R. Schwartz, N. Lourie, Y. Wang, H. Hajishirzi, N. A. Smith, and Y. Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. arXiv preprint arXiv:2009.10795, 2020."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626729"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","first-page":"1285","DOI":"10.1145\/3183713.3196934","volume-title":"Proceedings of the 2018 International Conference on Management of Data","author":"Vartak M.","year":"2018","unstructured":"M. Vartak, J. M. F. da Trindade, S. Madden, and M. Zaharia. Mistique: A system to store and query model intermediates for model diagnosis. In Proceedings of the 2018 International Conference on Management of Data, pages 1285--1300, 2018."},{"key":"e_1_2_1_40_1","first-page":"1","volume-title":"Proceedings of the Workshop on Human-In-the-Loop Data Analytics","author":"Vartak M.","year":"2016","unstructured":"M. Vartak, H. Subramanyam, W.-E. Lee, S. Viswanathan, S. Husnoo, S. Madden, and M. Zaharia. Modeldb: a system for machine learning model management. In Proceedings of the Workshop on Human-In-the-Loop Data Analytics, pages 1--3, 2016."},{"key":"e_1_2_1_41_1","first-page":"32","article-title":"Powersgd: Practical low-rank gradient compression for distributed optimization","author":"Vogels T.","year":"2019","unstructured":"T. Vogels, S. P. Karimireddy, and M. Jaggi. Powersgd: Practical low-rank gradient compression for distributed optimization. Advances in Neural Information Processing Systems, 32, 2019.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_42_1","first-page":"31","article-title":"Atomo: Communication-efficient learning via atomic sparsification","author":"Wang H.","year":"2018","unstructured":"H. Wang, S. Sievert, S. Liu, Z. Charles, D. Papailiopoulos, and S. Wright. Atomo: Communication-efficient learning via atomic sparsification. Advances in Neural Information Processing Systems, 31, 2018.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_43_1","volume-title":"Terngrad: Ternary gradients to reduce communication in distributed deep learning. Advances in neural information processing systems, 30","author":"Wen W.","year":"2017","unstructured":"W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li. Terngrad: Ternary gradients to reduce communication in distributed deep learning. Advances in neural information processing systems, 30, 2017."},{"key":"e_1_2_1_44_1","first-page":"1317","volume-title":"SIGMOD","author":"Wu W.","year":"2020","unstructured":"W. Wu, L. Flokas, E. Wu, and J. Wang. Complaint-driven training data debugging for query 2.0. In SIGMOD, pages 1317--1334, 2020."},{"key":"e_1_2_1_45_1","volume-title":"Less: Selecting influential data for targeted instruction tuning. arXiv preprint arXiv:2402.04333","author":"Xia M.","year":"2024","unstructured":"M. Xia, S. Malladi, S. Gururangan, S. Arora, and D. Chen. Less: Selecting influential data for targeted instruction tuning. arXiv preprint arXiv:2402.04333, 2024."},{"key":"e_1_2_1_46_1","first-page":"73","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Xiao T.","year":"2021","unstructured":"T. Xiao, X.-Y. Zhang, H. Jia, M.-M. Cheng, and M.-H. Yang. Semi-supervised learning with meta-gradient. In International Conference on Artificial Intelligence and Statistics, pages 73--81. PMLR, 2021."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00021"},{"key":"e_1_2_1_48_1","first-page":"635","volume-title":"European Conference on Computer Vision","author":"Yong H.","year":"2020","unstructured":"H. Yong, J. Huang, X. Hua, and L. Zhang. Gradient centralization: A new optimization technique for deep neural networks. In European Conference on Computer Vision, pages 635--652. Springer, 2020."},{"key":"e_1_2_1_49_1","first-page":"10842","volume-title":"International Conference on Machine Learning","author":"Yoon J.","year":"2020","unstructured":"J. Yoon, S. Arik, and T. Pfister. Data valuation using reinforcement learning. In International Conference on Machine Learning, pages 10842--10851. PMLR, 2020."},{"key":"e_1_2_1_50_1","volume-title":"Understanding neural networks through deep visualization. arXiv preprint arXiv:1506.06579","author":"Yosinski J.","year":"2015","unstructured":"J. Yosinski, J. Clune, A. Nguyen, T. Fuchs, and H. Lipson. Understanding neural networks through deep visualization. arXiv preprint arXiv:1506.06579, 2015."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00765"},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the VLDB Endowment, 14(11)","author":"Zhang H.","year":"2021","unstructured":"H. Zhang, L. Cao, S. Madden, and E. Rundensteiner. Lancet: labeling complex data at scale. Proceedings of the VLDB Endowment, 14(11), 2021."},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","first-page":"2174","DOI":"10.1145\/3447548.3467320","volume-title":"Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining","author":"Zhang H.","year":"2021","unstructured":"H. Zhang, L. Cao, P. VanNostrand, S. Madden, and E. A. Rundensteiner. Elite: Robust deep anomaly detection with meta gradient. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 2174--2182, 2021."},{"key":"e_1_2_1_54_1","volume-title":"NIPS","author":"Zhang X.","year":"2015","unstructured":"X. Zhang, J. J. Zhao, and Y. LeCun. Character-level convolutional networks for text classification. In NIPS, 2015."},{"key":"e_1_2_1_55_1","first-page":"725","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Z.","year":"2021","unstructured":"Z. Zhang and T. Pfister. Learning fast sample re-weighting without reward data. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, pages 725--734, 2021."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3648160.3648182","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,3]],"date-time":"2024-05-03T21:59:42Z","timestamp":1714773582000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3648160.3648182"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2]]},"references-count":56,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,2]]}},"alternative-id":["10.14778\/3648160.3648182"],"URL":"https:\/\/doi.org\/10.14778\/3648160.3648182","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,2]]},"assertion":[{"value":"2024-05-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}