{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T18:37:31Z","timestamp":1782931051188,"version":"3.54.5"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T00:00:00Z","timestamp":1746662400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Queue"],"published-print":{"date-parts":[[2025,5,8]]},"abstract":"<jats:p>Generative AI at the edge is the next phase in AI's deployment: from centralized supercomputers to ubiquitous assistants and creators operating alongside humans. The challenges are significant but so are the opportunities for personalization, privacy, and innovation. By tackling the technical hurdles and establishing new frameworks (conceptual and infrastructural), we can ensure this transition is successful and beneficial.<\/jats:p>","DOI":"10.1145\/3733702","type":"journal-article","created":{"date-parts":[[2025,5,21]],"date-time":"2025-05-21T22:21:53Z","timestamp":1747866113000},"page":"79-137","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Generative AI at the Edge: Challenges and Opportunities"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5259-7721","authenticated-orcid":false,"given":"Vijay Janapa","family":"Reddi","sequence":"first","affiliation":[{"name":"Harvard University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,5,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473","author":"Bahdanau D.","year":"2014","unstructured":"Bahdanau, D., et al. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473."},{"key":"e_1_2_1_2_1","volume-title":"NeurIPS Datasets and Benchmarks Track.","author":"Banbury C.","year":"2021","unstructured":"Banbury, C., et al. (2021). MLPerf Tiny benchmark. In NeurIPS Datasets and Benchmarks Track."},{"key":"e_1_2_1_3_1","volume-title":"On the opportunities and risks of foundation models. Stanford CRFM Report. https:\/\/crfm.stanford.edu\/report.html","author":"Bommasani R.","year":"2021","unstructured":"Bommasani, R., et al. (2021). On the opportunities and risks of foundation models. Stanford CRFM Report. https:\/\/crfm.stanford.edu\/report.html"},{"key":"e_1_2_1_4_1","volume-title":"RT-1: Robotics Transformer for real-world control at scale. arXiv preprint arXiv:2212.06817","author":"Brohan A.","year":"2022","unstructured":"Brohan, A., et al. (2022). RT-1: Robotics Transformer for real-world control at scale. arXiv preprint arXiv:2212.06817."},{"key":"e_1_2_1_5_1","volume-title":"NeurIPS","author":"Dettmers T.","year":"2022","unstructured":"Dettmers, T., et al. (2022). LLM.int8(): 8-bit matrix multiplication for transformers at scale. In NeurIPS 2022. https:\/\/dl.acm.org\/doi\/10.5555\/3600270.3602468"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 40th ICML, 8469?8488","author":"Driess D.","year":"2023","unstructured":"Driess, D., et al. (2023). PaLM-E: An embodied multimodal language model. In Proceedings of the 40th ICML, 8469?8488. https:\/\/dl.acm.org\/doi\/10.5555\/3618408.3618748"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 37th ICML. https:\/\/dl.acm.org\/doi\/10","author":"Guu K.","year":"2020","unstructured":"Guu, K., et al. (2020). REALM: Retrieval-Augmented Language Model pre-training. In Proceedings of the 37th ICML. https:\/\/dl.acm.org\/doi\/10.5555\/3524938.3525306"},{"key":"e_1_2_1_8_1","volume-title":"Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531","author":"Hinton G.","year":"2015","unstructured":"Hinton, G., et al. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531."},{"key":"e_1_2_1_9_1","volume-title":"QT-DoG: Quantization-aware training for domain generalization. arXiv preprint arXiv:2410.06020","author":"Javed S.","year":"2024","unstructured":"Javed, S., et al. (2024). QT-DoG: Quantization-aware training for domain generalization. arXiv preprint arXiv:2410.06020."},{"key":"e_1_2_1_10_1","volume-title":"LLM internal states reveal hallucination risk faced with a query. arXiv preprint arXiv:2407.03282","author":"Ji Z.","year":"2024","unstructured":"Ji, Z., et al. (2024). LLM internal states reveal hallucination risk faced with a query. arXiv preprint arXiv:2407.03282."},{"key":"e_1_2_1_11_1","volume-title":"Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492","author":"Konecny J.","year":"2016","unstructured":"Konecny, J., et al. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492."},{"key":"e_1_2_1_12_1","volume-title":"ICLR 2020","author":"Lan Z.","year":"2020","unstructured":"Lan, Z., et al. (2020). ALBERT: A lite BERT for self-supervised learning of language representations. In ICLR 2020. https:\/\/iclr.cc\/virtual_2020\/poster_H1eA7AEtvS.html"},{"key":"e_1_2_1_13_1","volume-title":"NeurIPS","author":"Lewis P.","year":"2020","unstructured":"Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In NeurIPS 2020. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/6b493230205f780e1bc26945df7481e5-Abstract.html"},{"key":"e_1_2_1_14_1","volume-title":"Holistic evaluation of language models (HELM). arXiv preprint arXiv:2211.09110","author":"Liang P.","year":"2022","unstructured":"Liang, P., et al. (2022). Holistic evaluation of language models (HELM). arXiv preprint arXiv:2211.09110."},{"key":"e_1_2_1_15_1","volume-title":"Medicine on the edge: Comparative performance of on-device LLMs for clinical reasoning. arXiv preprint arXiv:2502.08954","author":"Nissen L.","year":"2025","unstructured":"Nissen, L., et al. (2025). Medicine on the edge: Comparative performance of on-device LLMs for clinical reasoning. arXiv preprint arXiv:2502.08954."},{"key":"e_1_2_1_16_1","volume-title":"ICLR 2024","author":"Ong I.","year":"2024","unstructured":"Ong, I., et al. (2024). RouteLLM: Learning to route LLMs from preference data. In ICLR 2024. https:\/\/arxiv.org\/abs\/2406.18665"},{"key":"e_1_2_1_17_1","volume-title":"NeurIPS","author":"Ouyang L.","year":"2022","unstructured":"Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. In NeurIPS 2022, 27730?27744. https:\/\/dl.acm.org\/doi\/10.5555\/3600270.3602281"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3586183.3606763"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CODES-ISSS60120.2024.00015"},{"key":"e_1_2_1_20_1","volume-title":"Machine Learning Systems. https:\/\/mlsysbook.ai","author":"Reddi V. J.","year":"2025","unstructured":"Reddi, V. J. (2025). Machine Learning Systems. https:\/\/mlsysbook.ai"},{"key":"e_1_2_1_21_1","volume-title":"Robots that ask for help: Uncertainty alignment for large language model planners. arXiv preprint arXiv:2307.01928","author":"Ren A. Z.","year":"2023","unstructured":"Ren, A. Z., et al. (2023). Robots that ask for help: Uncertainty alignment for large language model planners. arXiv preprint arXiv:2307.01928."},{"key":"e_1_2_1_22_1","volume-title":"Green AI. Communications of the ACM, 63(12), 54?63. https:\/\/cacm.acm.org\/research\/green-ai\/","author":"Schwartz R.","year":"2020","unstructured":"Schwartz, R., et al. (2020). Green AI. Communications of the ACM, 63(12), 54?63. https:\/\/cacm.acm.org\/research\/green-ai\/"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.195"},{"key":"e_1_2_1_24_1","volume-title":"NeurIPS","author":"Sutskever I.","year":"2014","unstructured":"Sutskever, I., et al. (2014). Sequence to sequence learning with neural networks. In NeurIPS 2014, 3104?3112. https:\/\/dl.acm.org\/doi\/10.5555\/2969033.2969173"},{"key":"e_1_2_1_25_1","volume-title":"HPCA 2025","author":"Tschand A.","year":"2025","unstructured":"Tschand, A., et al. (2025). MLPerf Power: Benchmarking the energy efficiency of machine learning systems from microwatts to megawatts for sustainable AI. In HPCA 2025. https:\/\/arxiv.org\/abs\/2410.12032"},{"key":"e_1_2_1_26_1","volume-title":"NeurIPS","author":"Vaswani A.","year":"2017","unstructured":"Vaswani, A., et al. (2017). Attention is all you need. In NeurIPS 2017. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf"},{"key":"e_1_2_1_27_1","volume-title":"Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services","author":"Xu M.","year":"2024","unstructured":"Xu, M., et al. (2024). Unleashing the power of edge-cloud generative AI in mobile networks: A survey of AIGC services. IEEE Communications Surveys and Tutorials, 26(2), 1127?1170. https:\/\/dl.acm.org\/doi\/10.1109\/COMST.2024.3353265"},{"key":"e_1_2_1_28_1","volume-title":"Beyond perplexity: Multi-dimensional safety evaluation of LLM compression. arXiv preprint arXiv:2407.04965","author":"Xu Z.","year":"2024","unstructured":"Xu, Z., et al. (2024). Beyond perplexity: Multi-dimensional safety evaluation of LLM compression. arXiv preprint arXiv:2407.04965."},{"key":"e_1_2_1_29_1","volume-title":"PhoneLM: An efficient and capable small language model family through principled pre-training. arXiv preprint arXiv:2411.05046","author":"Yi R.","year":"2024","unstructured":"Yi, R., et al. (2024). PhoneLM: An efficient and capable small language model family through principled pre-training. arXiv preprint arXiv:2411.05046."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3649329.3658473"}],"container-title":["Queue"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733702","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3733702","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:56:56Z","timestamp":1750298216000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3733702"}},"subtitle":["The next phase in AI deployment"],"short-title":[],"issued":{"date-parts":[[2025,5,8]]},"references-count":30,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,5,8]]}},"alternative-id":["10.1145\/3733702"],"URL":"https:\/\/doi.org\/10.1145\/3733702","relation":{},"ISSN":["1542-7730","1542-7749"],"issn-type":[{"value":"1542-7730","type":"print"},{"value":"1542-7749","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,8]]},"assertion":[{"value":"2025-05-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}