PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 20260 citationsOpen Access

Carbon-Aware Inference Routing for Large Language Models: A Real-Time Framework for Sustainable AI Serving

View Full Paper
PRPreethi RaghuveeranOldham Council

Key Points

  • The aim is to develop a real-time routing framework that reduces carbon emissions for large language model inference while maintaining accuracy and low latency.
  • Proposed CAIR framework routes requests based on task complexity and live grid carbon intensity.
  • Preliminary analysis conducted on a system handling 1 million prompts per day.
  • Emissions reduction targeted through intelligent request routing.
  • Achieved approximately 62% reduction in inference carbon emissions.
  • No loss in accuracy or latency observed during implementation.

Abstract

This paper proposes CAIR (Carbon-Aware Inference Router), a real-time routing framework for large language models. Requests are routed between model tiers based on task complexity and live grid carbon intensity, targeting measurable emissions reduction without accuracy or latency loss. Preliminary analysis on a 1M prompt/day system suggests ~62% reduction in inference carbon. Framework repository: https://github.com/pretzelslab/sa1-carbon-inference-router

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Preethi Raghuveeran (2026) studied this question.

synapsesocial.com/papers/69f6e5f38071d4f1bdfc6933https://doi.org/10.5281/zenodo.19934621
Ask AI
Helpful
Bookmark
Share
View Full Paper