Public debate frequently portrays artificial intelligence inference as highly energy- and water- intensive, often based on opaque or inconsistent assumptions. In this paper, we provide a bottom-up estimation of the marginal electricity and water footprint of large language model inference using top level hypothesis. Using explicit computational, hardware, and infrastructure parameters, we demonstrate that a single query typically consumes around 1.3 Wh of electricity and 4.3 milliliters of water. These estimates suggest that marginal impacts may be lower than previously estimated through top-down attribution methods.
Puslecki et al. (Sun,) studied this question.