{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T18:00:33Z","timestamp":1785952833862,"version":"3.56.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2025,9,3]]},"abstract":"<jats:p>The advent of Large Language Models (LLMs) has changed the way we process information today and has unlocked new ways of delivering intelligence to the user. One of the ways of interfacing with AI is via smart assistants and chatbots that take multimodal inputs. However, the diversity of input tasks imply the possibility of both latency-critical and complex input instructions for AI assistants. Further, LLMs cannot be deployed on the edge for low-latency outputs, as that presents challenges due to their high computational demands and memory requirements. This work explores such trade-offs and contributes a smart LLM selection policy, called SELA, that leverages a suite of LLM models with disparate characteristics to optimize overall quality of service (QoS). SELA uses a time-criticality and complexity predictor at the edge to identify the optimal LLM choice for a given input instruction. Experiments on public instruction benchmarks demonstrate that SELA provides 9% to 62% higher QoS scores compared to the state-of-the-art selection policies.<\/jats:p>","DOI":"10.1145\/3749483","type":"journal-article","created":{"date-parts":[[2025,9,3]],"date-time":"2025-09-03T17:15:45Z","timestamp":1756919745000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["SELA: Smart Edge LLM Agent to Optimize Response Trade-offs of AI Assistants"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2960-1128","authenticated-orcid":false,"given":"Shreshth","family":"Tuli","sequence":"first","affiliation":[{"name":"Happening Technology, London, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4548-7951","authenticated-orcid":false,"given":"Giuliano","family":"Casale","sequence":"additional","affiliation":[{"name":"Imperial College London, London, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7828-7687","authenticated-orcid":false,"given":"Manuel","family":"Roveri","sequence":"additional","affiliation":[{"name":"Politecnico di Milano, Milan, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,3]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"AbacusResearch\/RasGulla1-7b. Hugging Face --- huggingface.co. https:\/\/huggingface.co\/AbacusResearch\/RasGulla1-7b. [Accessed 14-07-2024]."},{"key":"e_1_2_1_2_1","unstructured":"Dahoas\/instruct-human-assistant-prompt. Datasets at Hugging Face --- huggingface.co. https:\/\/huggingface.co\/datasets\/Dahoas\/instruct-human-assistant-prompt. [Accessed 15-07-2024]."},{"key":"e_1_2_1_3_1","unstructured":"Project Astra --- deepmind.google. https:\/\/deepmind.google\/technologies\/gemini\/project-astra\/. [Accessed 14-07-2024]."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330701"},{"key":"e_1_2_1_5_1","volume-title":"Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862","author":"Bai Y.","year":"2022","unstructured":"Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324915000285"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.nlposs-1.24"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2021.3073066"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641289"},{"key":"e_1_2_1_10_1","volume-title":"et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna.lmsys.org (accessed","author":"Chiang W.-L.","year":"2023","unstructured":"W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https:\/\/vicuna.lmsys.org (accessed 14 April 2023), 2(3):6, 2023."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3638550.3641126"},{"key":"e_1_2_1_12_1","first-page":"325","article-title":"Prompt cache: Modular attention reuse for low-latency inference","volume":"6","author":"Gim I.","year":"2024","unstructured":"I. Gim, G. Chen, S.-s. Lee, N. Sarda, A. Khandelwal, and L. Zhong. Prompt cache: Modular attention reuse for low-latency inference. Proceedings of Machine Learning and Systems, 6:325--338, 2024.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_2_1_13_1","volume-title":"Jina embeddings 2: 8192-token general-purpose text embeddings for long documents. arXiv preprint arXiv:2310.19923","author":"G\u00fcnther M.","year":"2023","unstructured":"M. G\u00fcnther, J. Ong, I. Mohr, A. Abdessalem, T. Abel, M. K. Akram, S. Guzman, G. Mastrapas, S. Sturua, B. Wang, et al. Jina embeddings 2: 8192-token general-purpose text embeddings for long documents. arXiv preprint arXiv:2310.19923, 2023."},{"key":"e_1_2_1_14_1","first-page":"1861","volume-title":"International conference on machine learning","author":"Haarnoja T.","year":"2018","unstructured":"T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, pages 1861--1870. PMLR, 2018."},{"key":"e_1_2_1_15_1","volume-title":"Learning continuous control policies by stochastic value gradients. Advances in neural information processing systems, 28","author":"Heess N.","year":"2015","unstructured":"N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa. Learning continuous control policies by stochastic value gradients. Advances in neural information processing systems, 28, 2015."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.398"},{"issue":"3","key":"e_1_2_1_17_1","first-page":"3","article-title":"Phi-2: The surprising power of small language models","volume":"1","author":"Javaheripi M.","year":"2023","unstructured":"M. Javaheripi, S. Bubeck, M. Abdin, J. Aneja, S. Bubeck, C. C. T. Mendes, W. Chen, A. Del Giorno, R. Eldan, S. Gopi, et al. Phi-2: The surprising power of small language models. Microsoft Research Blog, 1(3):3, 2023.","journal-title":"Microsoft Research Blog"},{"key":"e_1_2_1_18_1","volume-title":"F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825","author":"Jiang A. Q.","year":"2023","unstructured":"A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642773"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3640543.3645200"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TST.2015.7040517"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3527155"},{"key":"e_1_2_1_23_1","volume-title":"Sentiment analysis algorithms and applications: A survey. Ain Shams engineering journal, 5(4):1093--1113","author":"Medhat W.","year":"2014","unstructured":"W. Medhat, A. Hassan, and H. Korashy. Sentiment analysis algorithms and applications: A survey. Ain Shams engineering journal, 5(4):1093--1113, 2014."},{"key":"e_1_2_1_24_1","volume-title":"Human-level control through deep reinforcement learning. nature, 518(7540):529--533","author":"Mnih V.","year":"2015","unstructured":"V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529--533, 2015."},{"key":"e_1_2_1_25_1","volume-title":"et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32","author":"Paszke A.","year":"2019","unstructured":"A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019."},{"key":"e_1_2_1_26_1","volume-title":"Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409","author":"Press O.","year":"2021","unstructured":"O. Press, N. A. Smith, and M. Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3623402"},{"key":"e_1_2_1_29_1","volume-title":"Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347","author":"Schulman J.","year":"2017","unstructured":"J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017."},{"key":"e_1_2_1_30_1","volume-title":"Pickllm: Context-aware rl-assisted large language model routing. arXiv preprint arXiv:2412.12170","author":"Sikeridis D.","year":"2024","unstructured":"D. Sikeridis, D. Ramdass, and P. Pareek. Pickllm: Context-aware rl-assisted large language model routing. arXiv preprint arXiv:2412.12170, 2024."},{"key":"e_1_2_1_31_1","first-page":"34188","article-title":"Llm-check: Investigating detection of hallucinations in large language models","volume":"37","author":"Sriramanan G.","year":"2024","unstructured":"G. Sriramanan, S. Bharti, V. S. Sadasivan, S. Saha, P. Kattakinda, and S. Feizi. Llm-check: Investigating detection of hallucinations in large language models. Advances in Neural Information Processing Systems, 37:34188--34216, 2024.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2016.7900006"},{"key":"e_1_2_1_33_1","volume-title":"Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288","author":"Touvron H.","year":"2023","unstructured":"H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3087349"},{"key":"e_1_2_1_35_1","volume-title":"Attention is all you need. Advances in neural information processing systems, 30","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017."},{"key":"e_1_2_1_36_1","unstructured":"A. Weaver. Palmyra LLMs empower secure enterprise-grade generative AI for business --- writer.com. https:\/\/writer.com\/blog\/palmyra\/. [Accessed 14-07-2024]."},{"key":"e_1_2_1_37_1","first-page":"2516","article-title":"Zero time waste: Recycling predictions in early exit neural networks","volume":"34","author":"Wo\u0142czyk M.","year":"2021","unstructured":"M. Wo\u0142czyk, B. W\u00f3jcik, K. Ba\u0142azy, I. T. Podolak, J. Tabor, M. \u015amieja, and T. Trzcinski. Zero time waste: Recycling predictions in early exit neural networks. Advances in Neural Information Processing Systems, 34:2516--2528, 2021.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_38_1","volume-title":"Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771","author":"Wolf T.","year":"2019","unstructured":"T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771, 2019."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1609\/hcomp.v4i1.13283"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.463"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i17.29922"},{"key":"e_1_2_1_42_1","first-page":"14759","article-title":"Task-oriented feature distillation","volume":"33","author":"Zhang L.","year":"2020","unstructured":"L. Zhang, Y. Shi, Z. Shi, K. Ma, and C. Bao. Task-oriented feature distillation. Advances in Neural Information Processing Systems, 33:14759--14771, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_43_1","volume-title":"Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385","author":"Zhang P.","year":"2024","unstructured":"P. Zhang, G. Zeng, T. Wang, and W. Lu. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024."},{"key":"e_1_2_1_44_1","first-page":"36","article-title":"Judging llm-as-a-judge with mt-bench and chatbot arena","author":"Zheng L.","year":"2024","unstructured":"L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36, 2024.","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3749483","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T16:28:36Z","timestamp":1758817716000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3749483"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,3]]},"references-count":44,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,3]]}},"alternative-id":["10.1145\/3749483"],"URL":"https:\/\/doi.org\/10.1145\/3749483","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,3]]},"assertion":[{"value":"2025-09-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}