PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 30, 2025ASME Journal of Engineering for Sustainable Buildings and Cities3 citations

From Simulation to Reality: A Study of Reinforcement Learning Control in Operational Building Environments

View Full Paper
XWXinlin WangNMNariman MahdaviVAVahid Aryai

Key Points

  • Reinforcement learning successfully maintains thermal comfort in HVAC systems during field tests with static conditions.
  • The optimal policy was trained in a data-driven simulation environment, yielding comparable energy use to conventional systems.
  • Field evaluations demonstrated reinforcement learning's effectiveness, yet sensitivity emerged under changing comfort limits.
  • A detailed analysis identified issues such as reward design vulnerabilities and limits of policy generalization in real-world applications.

Abstract

Abstract Deploying reliable and cost-effective HVAC control strategies is essential for modern buildings. While Reinforcement Learning (RL) has shown promise in simulation, its real-world effectiveness remains underexplored. This study presents one of the first end-to-end field evaluations of an RL-based HVAC controller trained entirely in a data-driven simulation environment. The optimal policy is first trained using a data-driven digital twin of an office building. This trained policy is then deployed across two air handling units (AHUs) in the building under two different scenarios: one with static and the other one with dynamic thermal comfort limits. Field results in the static case show a successful transfer from simulation to real-world, where the RL controller consistently maintains thermal comfort and achieves comparable energy use to the baseline building management system (BMS). In contrast, RL performance degrades under dynamic comfort conditions, revealing its sensitivity to nonstationary environments and real-world complexities. To unpack this, we present a detailed analysis identifying key contributing factors, such as reward design vulnerabilities, policy generalisation limits, and the sim-to-real gap, which can serve as a reference for future deployments. This work provides empirical validation and critical insights into the opportunities and current limitations of RL-based HVAC control in real-world buildings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2025) studied this question.

synapsesocial.com/papers/68dc1e308a7d58c25ebb13c4https://doi.org/10.1115/1.4070007
Ask AI
Helpful
Bookmark
Share
View Full Paper