{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T08:45:16Z","timestamp":1777106716962,"version":"3.51.4"},"reference-count":140,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>\n            Deep learning accelerators efficiently train over vast and growing amounts of data, placing a newfound burden on commodity networks and storage devices. A common approach to conserve bandwidth involves resizing or compressing data prior to training. We introduce\n            <jats:italic>Progressive Compressed Records<\/jats:italic>\n            (PCRs), a data format that uses compression to reduce the overhead of fetching and transporting data, effectively reducing the training time required to achieve a target accuracy. PCRs deviate from previous storage formats by combining progressive compression with an efficient storage layout to view a single dataset at multiple fidelities---all without adding to the total dataset size. We implement PCRs and evaluate them on a range of datasets, training tasks, and hardware architectures. Our work shows that: (i) the amount of compression a dataset can tolerate exceeds 50% of the original encoding for many DL training tasks; (ii) it is possible to automatically and efficiently select appropriate compression levels for a given task; and (iii) PCRs enable tasks to readily access compressed data at runtime---\n            <jats:italic>utilizing as little as half the training bandwidth<\/jats:italic>\n            and thus potentially doubling training speed.\n          <\/jats:p>","DOI":"10.14778\/3476249.3476308","type":"journal-article","created":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T16:46:23Z","timestamp":1635353183000},"page":"2627-2641","source":"Crossref","is-referenced-by-count":9,"title":["Progressive compressed records"],"prefix":"10.14778","volume":"14","author":[{"given":"Michael","family":"Kuchnik","sequence":"first","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"George","family":"Amvrosiadis","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Virginia","family":"Smith","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,27]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1142473.1142548"},{"key":"e_1_2_1_2_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems."},{"key":"e_1_2_1_3_1","volume-title":"YouTube-8M: A Large-Scale Video Classification Benchmark. arXiv preprint arXiv:1609.08675","author":"Abu-El-Haija Sami","year":"2016"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/2789770.2789796"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData47090.2019.9005703"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294934"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1465482.1465560"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178146"},{"key":"e_1_2_1_9_1","unstructured":"ApacheBeam [n.d.]. ApacheBeam. https:\/\/www.tensorflow.org\/datasets\/beam_datasets. Accessed: 07-26-2021.  ApacheBeam [n.d.]. ApacheBeam. https:\/\/www.tensorflow.org\/datasets\/beam_datasets. Accessed: 07-26-2021."},{"key":"e_1_2_1_10_1","unstructured":"Apex [n.d.]. NVIDIA Apex. https:\/\/github.com\/NVIDIA\/apex. Accessed: 07-26-2021.  Apex [n.d.]. NVIDIA Apex. https:\/\/github.com\/NVIDIA\/apex. Accessed: 07-26-2021."},{"key":"e_1_2_1_11_1","volume-title":"Practical coreset constructions for machine learning. arXiv preprint arXiv:1703.06476","author":"Bachem Olivier","year":"2017"},{"key":"e_1_2_1_12_1","volume-title":"Efficient Nonlinear Transforms for Lossy Image Compression. In Picture Coding Symposium. 248--252","author":"Ball\u00e9 Johannes","year":"2018"},{"key":"e_1_2_1_13_1","volume-title":"International Conference on Learning Representations.","author":"Ball\u00e9 Johannes","year":"2018"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454715"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098021"},{"key":"e_1_2_1_16_1","volume-title":"Neural Information Processing Systems Workshop on Machine Learning Systems.","author":"Chen Tianqi","year":"2015"},{"key":"e_1_2_1_17_1","volume-title":"A survey of model compression and acceleration for deep neural networks. arXiv preprint arXiv:1710.09282","author":"Cheng Yu","year":"2017"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW-DISCS.2018.00011"},{"key":"e_1_2_1_19_1","volume-title":"Machine Learning and Systems.","author":"Chin Ting-Wu"},{"key":"e_1_2_1_20_1","volume-title":"Faster neural network training with data echoing. arXiv preprint arXiv:1907.05550","author":"Choi Dami","year":"2019"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3337821.3337902"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/2643634.2643639"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901323"},{"key":"e_1_2_1_24_1","volume-title":"Short and Deep: Sketching and Neural Networks. arXiv preprint arXiv:1710.07850","author":"Daniely Amit","year":"2017"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219910"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999271"},{"key":"e_1_2_1_27_1","volume-title":"ImageNet: A Large-Scale Hierarchical Image Database. In IEEE Conference on Computer Vision and Pattern Recognition. 248--255","author":"Deng Jia","year":"2009"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/2968826.2968968"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/QoMEX.2016.7498955"},{"key":"e_1_2_1_30_1","volume-title":"Benchmarking Adversarial Robustness on Image Classification. In IEEE Conference on Computer Vision and Pattern Recognition. 318--328","author":"Dong Yinpeng","year":"2020"},{"key":"e_1_2_1_31_1","volume-title":"Band-limited Training and Inference for Convolutional Neural Networks. In International Conference on Machine Learning","volume":"97","author":"Dziedzic Adam","year":"2019"},{"key":"e_1_2_1_32_1","volume-title":"Roy","author":"Dziugaite Gintare Karolina","year":"2016"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627817.2627920"},{"key":"e_1_2_1_35_1","unstructured":"Flatbuffers [n.d.]. Flatbuffers. https:\/\/google.github.io\/flatbuffers\/flatbuffers_benchmarks.html. Accessed: 07-26-2021.  Flatbuffers [n.d.]. Flatbuffers. https:\/\/google.github.io\/flatbuffers\/flatbuffers_benchmarks.html. Accessed: 07-26-2021."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3307441.3307464"},{"key":"e_1_2_1_38_1","unstructured":"GCPDiskBandwidth 2021. GCPDiskBandwidth. https:\/\/cloud.google.com\/compute\/docs\/disks\/performance. Accessed: 07-26-2021.  GCPDiskBandwidth 2021. GCPDiskBandwidth. https:\/\/cloud.google.com\/compute\/docs\/disks\/performance. Accessed: 07-26-2021."},{"key":"e_1_2_1_39_1","unstructured":"GCPNetworkBandwidth 2021. GCPNetworkBandwidth. https:\/\/cloud.google.com\/compute\/docs\/network-bandwidth. Accessed: 07-26-2021.  GCPNetworkBandwidth 2021. GCPNetworkBandwidth. https:\/\/cloud.google.com\/compute\/docs\/network-bandwidth. Accessed: 07-26-2021."},{"key":"e_1_2_1_40_1","volume-title":"International Conference on Learning Representations.","author":"Geirhos Robert","year":"2019"},{"key":"e_1_2_1_41_1","unstructured":"Google. 2021. Google Cloud. https:\/\/cloud.google.com\/. Accessed: 07-26-2021.  Google. 2021. Google Cloud. https:\/\/cloud.google.com\/. Accessed: 07-26-2021."},{"key":"e_1_2_1_42_1","volume-title":"large minibatch SGD: Training ImageNet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327144.3327308"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_2_1_45_1","volume-title":"International Conference on Learning Representations","author":"Han Song","year":"2016"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969366"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.5555\/2462638"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/JRPROC.1952.273898"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/SiPS.2014.6986082"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454299"},{"key":"e_1_2_1_52_1","volume-title":"Neural Information Processing Systems Workshop on Systems for ML.","author":"Jia Xianyan","year":"2018"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3360307"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080246"},{"key":"e_1_2_1_55_1","unstructured":"JPEGTran 2015. JPEGTran libjpeg.txt. https:\/\/github.com\/cloudflare\/jpegtran\/blob\/master\/libjpeg.txt. Accessed: 07-26-2021.  JPEGTran 2015. JPEGTran libjpeg.txt. https:\/\/github.com\/cloudflare\/jpegtran\/blob\/master\/libjpeg.txt. Accessed: 07-26-2021."},{"key":"e_1_2_1_56_1","unstructured":"JPEGTranManPage [n.d.]. jpegtran(1) - Linux max page. https:\/\/linux.die.net\/man\/1\/jpegtran. Accessed: 7-26-2021.  JPEGTranManPage [n.d.]. jpegtran(1) - Linux max page. https:\/\/linux.die.net\/man\/1\/jpegtran. Accessed: 7-26-2021."},{"key":"e_1_2_1_57_1","volume-title":"Conference on Learning Theory","volume":"99","author":"Karnin Zohar","year":"2019"},{"key":"e_1_2_1_58_1","volume-title":"International Conference on Learning Representations","author":"Karras Tero","year":"2018"},{"key":"e_1_2_1_59_1","unstructured":"Mahesh Khadatare Zoheb Khan and Harun Bayraktar. 2020. Leveraging the Hardware JPEG Decoder and NVIDIA nvJPEG Library on NVIDIA A100 GPUs. https:\/\/developer.nvidia.com\/blog\/leveraging-hardware-jpeg-decoder-and-nvjpeg-on-a100. Accessed: 07-26-2021.  Mahesh Khadatare Zoheb Khan and Harun Bayraktar. 2020. Leveraging the Hardware JPEG Decoder and NVIDIA nvJPEG Library on NVIDIA A100 GPUs. https:\/\/developer.nvidia.com\/blog\/leveraging-hardware-jpeg-decoder-and-nvjpeg-on-a100. Accessed: 07-26-2021."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/46.5.487"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.77"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.5555\/3386691.3386718"},{"key":"e_1_2_1_64_1","unstructured":"Sameer Kumar Victor Bitorff Dehao Chen Chiachen Chou Blake Hechtman HyoukJoong Lee Naveen Kumar Peter Mattson Shibo Wang Tao Wang etal 2019. Scale MLPerf-0.6 models on Google TPU-v3 Pods. arXiv preprint arXiv:1909.09756 (2019).  Sameer Kumar Victor Bitorff Dehao Chen Chiachen Chou Blake Hechtman HyoukJoong Lee Naveen Kumar Peter Mattson Shibo Wang Tao Wang et al. 2019. Scale MLPerf-0.6 models on Google TPU-v3 Pods. arXiv preprint arXiv:1909.09756 (2019)."},{"key":"e_1_2_1_65_1","unstructured":"Sameer Kumar James Bradbury Cliff Young Yu Emma Wang Anselm Levskaya Blake Hechtman Dehao Chen HyoukJoong Lee Mehmet Deveci Naveen Kumar Pankaj Kanwar Shibo Wang Skye Wanderman-Milne Steve Lacy Tao Wang Tayo Oguntebi Yazhou Zu Yuanzhong Xu and Andy Swing. 2021. Exploring the limits of Concurrency in ML Training on Google TPUs.  Sameer Kumar James Bradbury Cliff Young Yu Emma Wang Anselm Levskaya Blake Hechtman Dehao Chen HyoukJoong Lee Mehmet Deveci Naveen Kumar Pankaj Kanwar Shibo Wang Skye Wanderman-Milne Steve Lacy Tao Wang Tayo Oguntebi Yazhou Zu Yuanzhong Xu and Andy Swing. 2021. Exploring the limits of Concurrency in ML Training on Google TPUs."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00054"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2018.00021"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350874"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007328.3007331"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2487623"},{"key":"e_1_2_1_71_1","volume-title":"Conference on Systems and Machine Learning.","author":"Lim Hyeontaek","year":"2019"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_2_1_73_1","volume-title":"European conference on computer vision.","author":"Lin Tsung-Yi"},{"key":"e_1_2_1_74_1","volume-title":"International Conference on Learning Representations.","author":"Lin Yujun"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1287\/opre.9.3.383"},{"key":"e_1_2_1_76_1","volume-title":"Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"Liu Zihao","year":"2019"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196022"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339559"},{"key":"e_1_2_1_80_1","unstructured":"Jim McDonnell. 2012. Large Directory Causes LS to Hang. http:\/\/unixetc.co.uk\/2012\/05\/20\/large-directory-causes-ls-to-hang\/. Accessed: 07-26-2021.  Jim McDonnell. 2012. Large Directory Causes LS to Hang. http:\/\/unixetc.co.uk\/2012\/05\/20\/large-directory-causes-ls-to-hang\/. Accessed: 07-26-2021."},{"key":"e_1_2_1_81_1","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Meng Qi","year":"2017"},{"key":"e_1_2_1_82_1","volume-title":"High-Fidelity Generative Image Compression. Advances in Neural Information Processing Systems 33","author":"Mentzer Fabian","year":"2020"},{"key":"e_1_2_1_83_1","volume-title":"International Conference on Learning Representations","author":"Micikevicius Paulius","year":"2018"},{"key":"e_1_2_1_84_1","unstructured":"MLPerfHPCv0.7 2020. MLPerfHPCv0.7 Results. https:\/\/mlcommons.org\/en\/news\/mlperf-hpc-v07. Accessed: 07-26-2021.  MLPerfHPCv0.7 2020. MLPerfHPCv0.7 Results. https:\/\/mlcommons.org\/en\/news\/mlperf-hpc-v07. Accessed: 07-26-2021."},{"key":"e_1_2_1_85_1","unstructured":"MLPerfv0.7 2020. MLPerfTraining v0.7 Results. https:\/\/mlcommons.org\/en\/news\/mlperf-training-v07\/. Accessed: 07-26-2021.  MLPerfv0.7 2020. MLPerfTraining v0.7 Results. https:\/\/mlcommons.org\/en\/news\/mlperf-training-v07\/. Accessed: 07-26-2021."},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.5555\/364682.364691"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.14778\/3446095.3446100"},{"key":"e_1_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476374"},{"key":"e_1_2_1_89_1","unstructured":"NVIDIA. 2018. DALI. https:\/\/github.com\/NVIDIA\/DALI. Accessed: 07-26-2021.  NVIDIA. 2018. DALI. https:\/\/github.com\/NVIDIA\/DALI. Accessed: 07-26-2021."},{"key":"e_1_2_1_90_1","unstructured":"nvJPEG [n.d.]. nvJPEG. https:\/\/developer.nvidia.com\/nvjpeg. Accessed: 07-26-2021.  nvJPEG [n.d.]. nvJPEG. https:\/\/developer.nvidia.com\/nvjpeg. Accessed: 07-26-2021."},{"key":"e_1_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455008"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.5555\/3277355.3277385"},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370870"},{"key":"e_1_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2016.7533047"},{"key":"e_1_2_1_96_1","unstructured":"Protobuf [n.d.]. Protocol Buffers. https:\/\/developers.google.com\/protocol-buffers\/. Accessed: 07-26-2021.  Protobuf [n.d.]. Protocol Buffers. https:\/\/developers.google.com\/protocol-buffers\/. Accessed: 07-26-2021."},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/3331526"},{"key":"e_1_2_1_98_1","unstructured":"PytorchWebDataset 2020. [RFC] Add tar-based IterableDataset implementation to PyTorch. https:\/\/github.com\/pytorch\/pytorch\/issues\/38419. Accessed: 07-26-2021.  PytorchWebDataset 2020. [RFC] Add tar-based IterableDataset implementation to PyTorch. https:\/\/github.com\/pytorch\/pytorch\/issues\/38419. Accessed: 07-26-2021."},{"key":"e_1_2_1_99_1","unstructured":"RecordIODataset [n.d.]. Create a Dataset Using RecordIO. https:\/\/mxnet.apache.org\/api\/faq\/recordio and https:\/\/gluon-cv.mxnet.io\/build\/examples_datasets\/recordio.html. Accessed: 07-26-2021.  RecordIODataset [n.d.]. Create a Dataset Using RecordIO. https:\/\/mxnet.apache.org\/api\/faq\/recordio and https:\/\/gluon-cv.mxnet.io\/build\/examples_datasets\/recordio.html. Accessed: 07-26-2021."},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00045"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2007.905532"},{"key":"e_1_2_1_103_1","volume-title":"Riduan Khaddam-Aljameh, and Evangelos Eleftheriou.","author":"Sebastian Abu","year":"2020"},{"key":"e_1_2_1_104_1","volume-title":"A tutorial on principal component analysis. arXiv preprint arXiv:1404.1100","author":"Shlens Jonathon","year":"2014"},{"key":"e_1_2_1_105_1","volume-title":"International Conference on Learning Representations","author":"Simonyan Karen","year":"2015"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8655(01)00079-4"},{"key":"e_1_2_1_107_1","unstructured":"Stoyan Stefanov. 2021. Book of Speed.  Stoyan Stefanov. 2021. Book of Speed."},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_109_1","unstructured":"Tensorflow. 2020. imagenet_to_gcs.py. https:\/\/github.com\/tensorflow\/tpu\/tree\/master\/tools\/datasets\/imagenet_to_gcs.py. Accessed: 07-26-2021.  Tensorflow. 2020. imagenet_to_gcs.py. https:\/\/github.com\/tensorflow\/tpu\/tree\/master\/tools\/datasets\/imagenet_to_gcs.py. Accessed: 07-26-2021."},{"key":"e_1_2_1_110_1","unstructured":"tf.data 2021. tf.data: Build TensorFlow input pipelines. https:\/\/www.tensorflow.org\/guide\/data. Accessed: 07-26-2021.  tf.data 2021. tf.data: Build TensorFlow input pipelines. https:\/\/www.tensorflow.org\/guide\/data. Accessed: 07-26-2021."},{"key":"e_1_2_1_111_1","unstructured":"TFOp 2021. Create an op. https:\/\/www.tensorflow.org\/guide\/create_op. Accessed: 07-26-2021.  TFOp 2021. Create an op. https:\/\/www.tensorflow.org\/guide\/create_op. Accessed: 07-26-2021."},{"key":"e_1_2_1_112_1","unstructured":"TFRecords 2021. TFRecord and tf.train.Example. https:\/\/www.tensorflow.org\/tutorials\/load_data\/tfrecord. Accessed: 07-26-2021.  TFRecords 2021. TFRecord and tf.train.Example. https:\/\/www.tensorflow.org\/tutorials\/load_data\/tfrecord. Accessed: 07-26-2021."},{"key":"e_1_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.577"},{"key":"e_1_2_1_114_1","volume-title":"Towards image understanding from deep compression without decoding. arXiv preprint arXiv:1803.06131","author":"Torfason Robert","year":"2018"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455028"},{"key":"e_1_2_1_116_1","volume-title":"The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data","author":"Tschandl Philipp","year":"2018"},{"key":"e_1_2_1_117_1","volume-title":"Irish Machine Vision and Image Processing Conference.","author":"Ulicny Matej","year":"2017"},{"key":"e_1_2_1_118_1","volume-title":"Examining the impact of blur on recognition by convolutional networks. arXiv preprint arXiv:1611.05760","author":"Vasiljevic Igor","year":"2016"},{"key":"e_1_2_1_119_1","doi-asserted-by":"publisher","DOI":"10.1109\/30.125072"},{"key":"e_1_2_1_120_1","volume-title":"High-Frequency Component Helps Explain the Generalization of Convolutional Neural Networks. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"Wang Haohan"},{"key":"e_1_2_1_121_1","unstructured":"Yu Emma Wang Gu-Yeon Wei and David Brooks. 2020. A Systematic Methodology for Analysis of Deep Learning Hardware and Software Platforms. In Machine Learning and Systems.  Yu Emma Wang Gu-Yeon Wei and David Brooks. 2020. A Systematic Methodology for Analysis of Deep Learning Hardware and Software Platforms. In Machine Learning and Systems."},{"key":"e_1_2_1_122_1","doi-asserted-by":"publisher","DOI":"10.14778\/3317315.3317322"},{"key":"e_1_2_1_123_1","volume-title":"Conference on Signals, Systems & Computers","volume":"2","author":"Wang Zhou","year":"2003"},{"key":"e_1_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3327063"},{"key":"e_1_2_1_125_1","doi-asserted-by":"publisher","DOI":"10.5555\/1298455.1298485"},{"key":"e_1_2_1_126_1","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225076"},{"key":"e_1_2_1_127_1","doi-asserted-by":"publisher","DOI":"10.5555\/1364813.1364815"},{"key":"e_1_2_1_128_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294915"},{"key":"e_1_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1561\/0400000060"},{"key":"e_1_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"},{"key":"e_1_2_1_131_1","volume-title":"AAAI Conference on Artificial Intelligence","volume":"32","author":"Xu Yuhui","year":"2018"},{"key":"e_1_2_1_132_1","volume-title":"International Conference on Neural Information Processing.","author":"John Xu Zhi-Qin","year":"2019"},{"key":"e_1_2_1_133_1","volume-title":"Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds. arXiv preprint arXiv:1903.12650","author":"Yamazaki Masafumi","year":"2019"},{"key":"e_1_2_1_134_1","doi-asserted-by":"publisher","DOI":"10.5555\/3154601.3154627"},{"key":"e_1_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455476"},{"key":"e_1_2_1_136_1","volume-title":"Neural Information Processing Systems Workshop on Systems for ML.","author":"Ying Chris","year":"2018"},{"key":"e_1_2_1_137_1","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225069"},{"key":"e_1_2_1_138_1","first-page":"10282","article-title":"Stochastic Optimization with Laggard Data Pipelines","volume":"33","author":"Zhang Cyril","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.5555\/3154690.3154708"},{"key":"e_1_2_1_140_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.485"},{"key":"e_1_2_1_141_1","volume-title":"TBD: Benchmarking and analyzing deep neural network training. arXiv preprint arXiv:1803.06905","author":"Zhu Hongyu","year":"2018"},{"key":"e_1_2_1_142_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.150"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476249.3476308","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:11:59Z","timestamp":1672222319000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476249.3476308"}},"subtitle":["taking a byte out of deep learning data"],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":140,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476249.3476308"],"URL":"https:\/\/doi.org\/10.14778\/3476249.3476308","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}