{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:06:34Z","timestamp":1750309594605,"version":"3.41.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T00:00:00Z","timestamp":1746662400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Queue"],"published-print":{"date-parts":[[2025,5,8]]},"abstract":"<jats:p>As the scaling of pretraining is reaching a plateau of diminishing returns, model inference is quickly becoming an important driver for model performance. Today, test-time compute scaling offers a new, exciting avenue to increase model performance beyond what can be achieved with training, and test-time compute techniques cover a fertile area for many more breakthroughs in AI. Innovations using ensemble methods, iterative refinement, repeated sampling, retrieval augmentation, chain-of-thought reasoning, search, and agentic ensembles are already yielding improvements in model quality performance and offer additional opportunities for future growth.<\/jats:p>","DOI":"10.1145\/3733701","type":"journal-article","created":{"date-parts":[[2025,5,24]],"date-time":"2025-05-24T22:30:16Z","timestamp":1748125816000},"page":"40-78","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["AI: It's All About Inference Now"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-4963-4915","authenticated-orcid":false,"given":"Michael","family":"Gschwind","sequence":"first","affiliation":[{"name":"Nvidia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,5,24]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Ainslie J. Lee-Thorp J. de Jong M. Zemlyanskiy Y. Lebr\u00f3n F. Sanghai S. 2023. GQA: training generalized multi-query transformer models from multi-head checkpoints. arXiv:2305.13245; https:\/\/arxiv.org\/abs\/2305.13245.","DOI":"10.18653\/v1\/2023.emnlp-main.298"},{"key":"e_1_2_1_2_1","unstructured":"Anderson M. et al. 2021. First-generation inference accelerator deployment at Facebook. arXiv:2107.04140; https:\/\/arxiv.org\/abs\/2107.04140."},{"key":"e_1_2_1_3_1","unstructured":"Beeching E. Tunstall L. Rush S. 2024. Scaling test time compute with open models: tutorial and experiments to outperform Llama 3.1 70B on MATH-500 with a 3B Model. Hugging Face blog; https:\/\/huggingface.co\/spaces\/HuggingFaceH4\/blogpost-scaling-test-time-compute."},{"key":"e_1_2_1_4_1","unstructured":"Belkada Y. Marty F. Benayoun M. Han E. Shojanazeri H. Puhrsch C. Guessous D. Gschwind M. Chauhan G. 2022. BetterTransformer out of the box performance for Hugging Face transformers; https:\/\/medium.com\/pytorch\/bettertransformer-out-of-the-box-performance-for-huggingface-transformers-3fbe27d50ab2."},{"key":"e_1_2_1_5_1","unstructured":"Bengio Y. L\u00e9onard N. Courville A. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv:1308.3432; https:\/\/arxiv.org\/abs\/1308.3432."},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Bondarenko Y. Nagel M. Blankevoort T. 2021. Understanding and overcoming the challenges of efficient transformer quantization. arXiv:2109.12948; https:\/\/arxiv.org\/abs\/2109.12948.","DOI":"10.18653\/v1\/2021.emnlp-main.627"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150464"},{"key":"e_1_2_1_8_1","volume-title":"PIP: Perturbation-based Iterative Pruning for large language models. arXiv:2501.15278","author":"Cao Y.","year":"2025","unstructured":"Cao, Y., Xu, W.-J., Shen, Y., Shi, W., Chan, C.-M., Xu, J. 2025. PIP: Perturbation-based Iterative Pruning for large language models. arXiv:2501.15278; https:\/\/arxiv.org\/abs\/2501.15278."},{"key":"e_1_2_1_9_1","unstructured":"Chang C.-C. Lin W.-C. Lin C.-Y. Chen C.-Y. Hu Y.-F. Wang P.-S. Huang N.-C. Ceze L. Abdelfattah M. S. Wu K.-C. 2024. Palu: compressing KV-Cache with low-rank projection. arXiv:2407.21118; https:\/\/arxiv.org\/abs\/2407.21118."},{"key":"e_1_2_1_10_1","unstructured":"Chen C. et al. 2023. Accelerating large language model inference with speculative sampling. arXiv:2302.01318; https:\/\/arxiv.org\/abs\/2302.01318."},{"key":"e_1_2_1_11_1","unstructured":"Chen M. Tworek J. et al. 2021. Evaluating large language models trained on code. arXiv:12107.03374; https:\/\/arxiv.org\/abs\/2107.03374."},{"key":"e_1_2_1_12_1","volume-title":"DeepSeek-V3 technical report. arXiv:2412.19437","author":"DeepSeek AI.","year":"1943","unstructured":"DeepSeek-AI. 2024. DeepSeek-V3 technical report. arXiv:2412.19437; https:\/\/arxiv.org\/abs\/2412.19437."},{"key":"e_1_2_1_13_1","unstructured":"DeepSeek-AI. 2025. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv:2501.12948; https:\/\/arxiv.org\/abs\/2501.12948."},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Elhoushi M. Shrivastava A. Liskovich D. Hosmer B. Wasti B. Lai L. Mahmoud A. Acun B. et al. 2024. LayerSkip: enabling early exit inference and self-speculative decoding. arXiv:2404.16710; https:\/\/arxiv.org\/abs\/2404.16710.","DOI":"10.18653\/v1\/2024.acl-long.681"},{"key":"e_1_2_1_15_1","volume-title":"International Conference on Learning Representations. arXiv:1803","author":"Frankle J.","year":"2018","unstructured":"Frankle, J., Carbin, M. 2018. The lottery ticket hypothesis: finding sparse, trainable neural networks. International Conference on Learning Representations. arXiv:1803.03635; https:\/\/arxiv.org\/abs\/1803.03635."},{"key":"e_1_2_1_16_1","unstructured":"Frantar E. Ashkboos S. Hoefler T. Alistarh D. 2022. GPTQ: accurate post-training quantization for generative pre-trained Transformers. arXiv:2210.17323; https:\/\/arxiv.org\/abs\/2210.17323."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 33rd International Conference on Machine Learning, 1050?1059; https:\/\/dl.acm.org\/doi\/10","author":"Gal Y.","year":"2016","unstructured":"Gal, Y., Ghahramani, Z. 2016. Dropout as a Bayesian approximation: representing model uncertainty in deep learning. Proceedings of the 33rd International Conference on Machine Learning, 1050?1059; https:\/\/dl.acm.org\/doi\/10.5555\/3045390.3045502."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning, 15706-15734; https:\/\/dl.acm.org\/doi\/10","author":"Gloeckle F.","year":"2024","unstructured":"Gloeckle, F., Idrissi, B. Y., Rozi\u00e8re, B., Lopez-Paz, D., Synnaeve, G. 2024. Better & faster large language models via multi-token prediction. Proceedings of the 41st International Conference on Machine Learning, 15706-15734; https:\/\/dl.acm.org\/doi\/10.5555\/3692070.3692699."},{"key":"e_1_2_1_19_1","unstructured":"Grattafiori A. et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783; https:\/\/arxiv.org\/abs\/2407.21783."},{"key":"e_1_2_1_20_1","volume-title":"MIT","author":"Gschwind K.","year":"2021","unstructured":"Gschwind, K. 2021. Model compression and AutoML for efficient click-through rate prediction. MEng. thesis, MIT; https:\/\/dspace.mit.edu\/bitstream\/handle\/1721.1\/139253\/Gschwind-gschwind-meng-eecs-2021-thesis.pdf."},{"key":"e_1_2_1_21_1","volume-title":"LLMs everywhere: acceleration from servers to mobile devices in the age of generative AI. Keynote speech at the International Conference on Supercomputing","author":"Gschwind M.","year":"2024","unstructured":"Gschwind, M. 2024. LLMs everywhere: acceleration from servers to mobile devices in the age of generative AI. Keynote speech at the International Conference on Supercomputing; https:\/\/ics2024.github.io\/keynote.html."},{"key":"e_1_2_1_22_1","article-title":"Workload acceleration with the IBM POWER vector-scalar architecture","author":"Gschwind M.","year":"2016","unstructured":"Gschwind, M. 2016. Workload acceleration with the IBM POWER vector-scalar architecture, IBM Journal of Research and Development 60(2-3); https:\/\/ieeexplore.ieee.org\/document\/7442604.","journal-title":"IBM Journal of Research and Development 60(2-3); https:\/\/ieeexplore.ieee.org\/document\/7442604."},{"key":"e_1_2_1_23_1","unstructured":"Gschwind M. Han E. Wolchok S. Zhu R. Puhrsch C. 2022. A better transformer for fast transformer inference. PyTorch blog; https:\/\/pytorch.org\/blog\/a-better-transformer-for-fast-transformer-encoder-inference\/."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.41"},{"key":"e_1_2_1_25_1","unstructured":"Gupta N. Gschwind M. Husa D. Dewan C. Khabsa M. 2023. MultiRay: optimizing efficiency for large-scale AI models. Meta AI blog. https:\/\/ai.meta.com\/blog\/multiray-large-scale-AI-models\/."},{"key":"e_1_2_1_26_1","volume-title":"International Conference on Learning Representations; https:\/\/arxiv.org\/abs\/1510","author":"Han S.","year":"2016","unstructured":"Han, S., Mao, H., Dally, W. J. 2016. Deep compression: compressing deep neural networks with pruning, trained quantization, and Huffman coding. International Conference on Learning Representations; https:\/\/arxiv.org\/abs\/1510.00149."},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 29th International Conference on Neural Information Processing Systems","volume":"1","author":"Han S.","year":"2015","unstructured":"Han, S., Pool, J., Tran, J., Dally, W. 2015. Learning both weights and connections for efficient neural networks. Proceedings of the 29th International Conference on Neural Information Processing Systems, volume 1, 1135?1143; https:\/\/dl.acm.org\/doi\/10.5555\/2969239.2969366."},{"key":"e_1_2_1_28_1","volume-title":"NIPS 2014 Deep Learning Workshop. arXiv:1503","author":"Hinton G.","year":"2015","unstructured":"Hinton, G., Vinyals, O., Dean, J. 2015. Distilling the knowledge in a neural network. NIPS 2014 Deep Learning Workshop. arXiv:1503.02531; https:\/\/arxiv.org\/abs\/1503.02531."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3602446"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.136.0675"},{"key":"e_1_2_1_31_1","volume-title":"ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 1097?1105","author":"Krizhevsky A.","year":"2012","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G. 2012. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 1097?1105; https:\/\/papers.nips.cc\/paper_files\/paper\/2012\/hash\/c399862d3b9d6b76c8436e924a68c45b-Abstract.html"},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Kwon W. Li Z. Zhuang S. Sheng Y. Zheng L. Yu C. H. Gonzalez J. E. Zhang H. Stoica I. 2023. Efficient memory management for large language model serving with paged attention. arXiv:2309.06180 https:\/\/arxiv.org\/abs\/2309.06180.","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496517"},{"key":"e_1_2_1_35_1","volume-title":"Compact language models via pruning and knowledge distillation. Advances in Neural Information Processing Systems 37","author":"Muralidharan S.","year":"2024","unstructured":"Muralidharan, S., Sreenivas, S. T., Joshi, R., Chochowski, M., Patwary, M., Shoeybi, M., Catanzaro, B., Kautz, J., Molchanov, P. 2024. Compact language models via pruning and knowledge distillation. Advances in Neural Information Processing Systems 37; https:\/\/papers.nips.cc\/paper_files\/paper\/2024\/hash\/4822991365c962105b1b95b1107d30e5-Abstract-Conference.html."},{"key":"e_1_2_1_36_1","unstructured":"NVIDIA. 2024. NVIDIA Blackwell Architecture Technical Brief: powering the new era of generative AI and accelerated computing; https:\/\/resources.nvidia.com\/en-us-blackwell-architecture."},{"key":"e_1_2_1_37_1","unstructured":"Or A. Zhang J. Smothers E. Khandelwal K. Rao S. 2023. Quantization-aware training for large language models with PyTorch. PyTorch blog; https:\/\/pytorch.org\/blog\/quantization-aware-training\/."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of Machine Learning and Systems 5; https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2023\/hash\/c4be71ab8d24cdfb45e3d06dbfca2780-Abstract-mlsys2023","author":"Pope R.","year":"2023","unstructured":"Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., Dean, J. 2023. Efficiently scaling transformer inference. Proceedings of Machine Learning and Systems 5; https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2023\/hash\/c4be71ab8d24cdfb45e3d06dbfca2780-Abstract-mlsys2023.html."},{"key":"e_1_2_1_39_1","volume-title":"Artificial Intelligence: A Modern Approach. Pearson.","author":"Russell S. J.","year":"2010","unstructured":"Russell, S. J., Norvig, P. 2010. Artificial Intelligence: A Modern Approach. Pearson."},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems, 68539-68551; https:\/\/dl.acm.org\/doi\/10","author":"Schick T.","year":"2023","unstructured":"Schick, T., Dwivedi-Yu, J., Dess\u00ec, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., Scialom, T. 2023. Toolformer: language models can teach themselves to use tools. Proceedings of the 37th International Conference on Neural Information Processing Systems, 68539-68551; https:\/\/dl.acm.org\/doi\/10.5555\/3666122.3669119."},{"key":"e_1_2_1_41_1","unstructured":"Seo S. Noh S. Lee J. Lim S. Lee W. H. Kang H. 2024. REVECA: adaptive planning and trajectory-based validation in cooperative language agents using information relevance and relative proximity. arXiv:2405.16751; https:\/\/arxiv.org\/abs\/2405.16751."},{"key":"e_1_2_1_42_1","unstructured":"Shao Z. Wang P. Zhu Q. Xu R. Song J. Bi X. Zhang H. Zhang M. Li Y. K. Wu Y. Guo D. 2024. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300; https:\/\/arxiv.org\/abs\/2402.03300."},{"key":"e_1_2_1_43_1","volume-title":"33rd International Conference on Neural Information Processing Systems. arXiv:1911","author":"Shazeer N.","year":"2019","unstructured":"Shazeer, N. 2019. Fast transformer decoding: one write-head is all you need. 33rd International Conference on Neural Information Processing Systems. arXiv:1911.02150; https:\/\/arxiv.org\/abs\/1911.02150."},{"key":"e_1_2_1_44_1","unstructured":"Snell C. Lee J. Xu K. Kumar A. 2024. Scaling LLM test-time compute optimally can be more effective than scaling LLM parameters. arXiv:2408.03314; https:\/\/arxiv.org\/abs\/2408.03314."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327546.3327673"},{"key":"e_1_2_1_46_1","unstructured":"Subramanian S. Saroufim M. Zhang J. 2022. Practical quantization in PyTorch. PyTorch blog; https:\/\/pytorch.org\/blog\/quantization-in-practice\/."},{"key":"e_1_2_1_47_1","unstructured":"Sutskever I. 2024. Sequence to sequence learning with neural networks: what a decade. Test of Time Award Talk at NeurIPS; https:\/\/www.youtube.com\/watch?v=1yvBqasHLZs."},{"key":"e_1_2_1_48_1","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems","volume":"2","author":"Sutskever I.","year":"2014","unstructured":"Sutskever, I., Vinyals, O., Le, Q. V. 2014. Sequence to sequence learning with neural networks. Proceedings of the 28th International Conference on Neural Information Processing Systems, volume 2, 3104?3112, https:\/\/dl.acm.org\/doi\/10.5555\/2969033.2969173."},{"key":"e_1_2_1_49_1","unstructured":"Szegedy C. Zaremba W. Sutskever I. Bruna J. Erhan D. Goodfellow I. Fergus R. 2013. Intriguing properties of neural networks. arXiv:1312.6199; https:\/\/arxiv.org\/abs\/1312.6199."},{"key":"e_1_2_1_50_1","unstructured":"Team PyTorch. 2023. torchtune: easily fine-tune LLMs using PyTorch. PyTorch blog; https:\/\/pytorch.org\/blog\/torchtune-fine-tune-llms\/."},{"key":"e_1_2_1_51_1","unstructured":"Team PyTorch. 2024. Introducing torchchat: accelerating local LLM inference on laptop desktop and mobile. PyTorch blog; https:\/\/pytorch.org\/blog\/torchchat-local-llm-inference\/."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2016.7900006"},{"key":"e_1_2_1_53_1","unstructured":"Touvron H. et al. 2023. Llama 2: open foundation and fine-tuned chat models. arXiv:2307.09288; https:\/\/arxiv.org\/abs\/2307.09288."},{"key":"e_1_2_1_54_1","unstructured":"Wang X. Wei J. Schuurmans D. Le Q. Chi E. Narang S. Chowdhery A. Zhou D. 2023. Self-consistency improves chain of thought reasoning in language models. arXiv:2203.11171; https:\/\/arxiv.org\/abs\/2203.11171."},{"key":"e_1_2_1_55_1","doi-asserted-by":"crossref","unstructured":"Williams S. Waterman A. Patterson D. 2009. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM 52(4) 65?76; https:\/\/dl.acm.org\/doi\/10.1145\/1498765.1498785.","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_2_1_56_1","unstructured":"Wu C.-J. Raghavendra R. Gupta U. Acun B. Ardalani N. Maeng K. Chang G. Aga F. Huang J. Bai C. Gschwind M. Gupta A. Ott M. Melnikov A. Candido S. Brooks D. et al. 2022. Sustainable AI: environmental implications challenges and opportunities. Machine Learning and Systems 4. arXiv:2109.02079; https:\/\/arxiv.org\/abs\/2111.00364."},{"key":"e_1_2_1_57_1","unstructured":"Ye Z. Chen L. Lai R. Lin W. Zhang Y. Wang S. Chen T. Kasikci B. Grover V. Krishnamurthy A. Ceze L. 2025. FlashInfer: efficient and customizable attention engine for LLM inference serving. arXiv:2501.01005 https:\/\/arxiv.org\/abs\/2501.01005."},{"key":"e_1_2_1_58_1","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation; https:\/\/www.usenix.org\/system\/files\/osdi22-yu.pdf.","author":"Yu G.","year":"2022","unstructured":"Yu, G., Jeong, J., Kim, G. Kim, S., Chun, B. 2022. Orca: a distributed serving system for transformer-based generative models, 16th USENIX Symposium on Operating Systems Design and Implementation; https:\/\/www.usenix.org\/system\/files\/osdi22-yu.pdf."},{"key":"e_1_2_1_59_1","unstructured":"Zheng L. Yin L. Xie Z. Sun C. Huang J. Yu C. H. Cao S. Kozyrakis C. Stoica I. Gonzalez J. E. Barrett C. Sheng Y. 2023. SGLang: efficient execution of structured language model programs. arXiv:2312.07104 https:\/\/arxiv.org\/abs\/2312.07104."}],"container-title":["Queue"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733701","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733701","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:56:56Z","timestamp":1750298216000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733701"}},"subtitle":["Model inference has become the critical driver for model performance."],"short-title":[],"issued":{"date-parts":[[2025,5,8]]},"references-count":59,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,5,8]]}},"alternative-id":["10.1145\/3733701"],"URL":"https:\/\/doi.org\/10.1145\/3733701","relation":{},"ISSN":["1542-7730","1542-7749"],"issn-type":[{"type":"print","value":"1542-7730"},{"type":"electronic","value":"1542-7749"}],"subject":[],"published":{"date-parts":[[2025,5,8]]},"assertion":[{"value":"2025-05-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}