PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 17, 2025The Journal of Supercomputing0 citationsOpen Access

A data-augmented model routing framework for efficient LLM deployment in edge–cloud environments

View Full Paper
MPMuhammad Syafiq Mohd PoziYSYukinori SatoMPMuhammad Syafiq Mohd Pozi

Key Points

  • Achieving efficiency with up to 16 times improvements compared to traditional cascaded approaches, while ensuring inference accuracy.
  • The proposed routing framework optimally allocates tasks between weak and strong LLMs based on prompt classification.
  • Observational analysis of computational demand highlights the benefits of the multi-LLM model approach.
  • This method may provide a pathway for practical applications in edge-cloud computing environments.

Abstract

Abstract Large language model (LLM)-based program generation tasks are hindered by high computational demands. These challenges, along with high deployment costs, often pose a barrier to practical applications. To address these, we propose a novel data-augmented multi-LLM model routing approach that classifies prompts based on whether they should be processed on a weak LLM engine or a strong LLM. Experimental results show up to 16 times better efficiency compared to the existing cascaded approaches, while preserving the inference accuracy. Thus, the proposed method optimally allocates prompts across multiple LLMs, reducing computational costs while maintaining inference accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Pozi et al. (2025) studied this question.

synapsesocial.com/papers/692509f6c0ce034ddc352c62https://doi.org/10.1007/s11227-025-08034-8
Ask AI
Helpful
Bookmark
Share
View Full Paper