Buildings continue to represent one of the largest sectors for electricity consumption globally. A significant portion of this demand is driven by heating, ventilation, and air conditioning (HVAC) systems within these structures. Due to the complexity associated with modeling large-scale HVAC systems, traditional model-based optimal control strategies become increasingly impractical. In this work, we present a game-theoretic approach to optimal control for building HVAC systems, framing the problem as a two-player non-zero-sum cooperative game. We propose a data-driven, model-free state feedback Q-learning value iteration method that addresses the quadratic game optimization problem without requiring any prior knowledge of the zone’s dynamics. Mass flow rate and supply air temperature are treated as the twoprimary decision-making players. The building’s HVAC zone is considered as an environment in which these players interact, with its underlying dynamics remaining entirely unknown to them. The Q-learning value iteration algorithm is demonstrated to effectively learn optimal game policies for both players using input-state data under external disturbances, notably without the requirement of an initially admissible policy—a key advantage in scenarios with limitedprior information on dynamics. The convergence of the proposed value iteration algorithm to the Nash equilibrium is formally established. Numerical results validate the effectiveness of the proposed approach in maintaining temperature regulation, even in the presence of unknown zone behavior and external disturbances.Presented as a poster at the 2025 IFAC Modeling, Estimation and Control Conference (MECC), Pittsburgh, PA, USA, October 5–8, 2025.Official proceedings version available at: https://paperhost.org/proceedings/ifac/MECC25/files/0350.pdfUse conference credentials mecc2025pitt / mecc2025pitt if prompted for access.
Rizvi et al. (Sun,) studied this question.