PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 21, 2025Journal of Information Systems and Informatics3 citationsOpen Access

Hybrid Cloud Architecture for Efficient and Cost-Effective Large Language Model Deployment

View Full Paper
XQXin Qi

Key Points

  • Implementing a hybrid cloud architecture can reduce cloud API usage by over 60%, significantly lowering costs.
  • The system matches the accuracy of a state-of-the-art LLM while reducing average latency by approximately 40%.
  • A confidence-based routing mechanism determines when to use a cloud-hosted LLM, improving efficiency.
  • This approach enhances data privacy by processing sensitive queries on-premise rather than in the cloud.

Abstract

Large Language Models (LLMs) have achieved remarkable success across natural language tasks, but their enormous computational requirements pose challenges for practical deployment. This paper proposes a hybrid cloud–edge architecture to deploy LLMs in a cost-effective and efficient manner. The proposed system employs a lightweight on-premise LLM to handle the bulk of user requests, and dynamically offloads complex queries to a powerful cloud-hosted LLM only when necessary. We implement a confidence-based routing mechanism to decide when to invoke the cloud model. Experiments on a question-answering use case demonstrate that our hybrid approach can match the accuracy of a state-of-the-art LLM while reducing cloud API usage by over 60%, resulting in significant cost savings and a ~40% reduction in average latency. We also discuss how the hybrid strategy enhances data privacy by keeping sensitive queries on-premise. These results highlight a promising direction for organizations to leverage advanced LLM capabilities without prohibitive expense or risk, by intelligently combining local and cloud resources.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xin Qi (2025) studied this question.

synapsesocial.com/papers/68f74e597f21f73e19e5b458https://doi.org/10.51519/journalisi.v7i3.1170
Ask AI
Helpful
Bookmark
Share
View Full Paper