Large language model (LLM) inference consumes substantial datacenter energy and produces both carbon emissions and criteria pollutants (PM2.5, SO2, NOx). These criteria pollutants impact public health depending on the meteorological conditions of emission regions and population exposure. We apply health impact assessment to LLM inference, a use case of growing concern due to its rapidly rising energy consumption, and characterize health impacts from datacenter energy consumption as quantifiable health cost metrics. We then present HealthServe, a systems-level framework for health-aware LLM inference serving, which jointly minimizes carbon emissions and health costs across geo-distributed GPU clusters while meeting latency SLOs. HealthServe employs a hierarchical scheduling architecture with Social Cognitive Optimization (SCO) that maintains condition-indexed decision libraries for configuration reuse under recurring operating conditions. Evaluation on a heterogeneous GPU cluster across different U.S. datacenter locations demonstrates over 50% carbon footprint and over 25% health cost reductions over the state-of-the-art energy-aware LLM inference serving solutions. Additionally, we show that HealthServe can be augmented with existing carbon-aware serving approaches to provide additive sustainability benefits.
Zhou et al. (Fri,) studied this question.