• Multiple deep reinforcement learning agents coordinate home automation devices. • User interactions are modeled as implicit feedback to generate agent rewards • The recommender system successfully adapts to changes in users’ daily routines. • Two domain-specific action selection strategies adapted for reinforcement learning systems in environments with the possibility of routine change. • A new metric based on the Hamming Score for evaluating predictions for smart devices with a composite state. The rapid expansion of the Internet of Things (IoT) has provided the technological foundation for smart homes, in which interconnected devices enable the environment to adapt to residents’ needs. Yet, this proliferation of IoT-enabled devices also introduces a significant challenge: managing and coordinating their multiple operational states, which extend far beyond a simple “switch on/switch off”. Aiming to automate the action process and anticipate user actions, this study proposes a state recommender system for actuator devices in smart homes, developed using deep reinforcement learning algorithms integrated with implicit feedback. This approach enables composite state encoding and coordination among multiple agents, thereby anticipating user needs and adapting to dynamic routines. In addition to using a dataset captured from a real-world IoT environment, we use a smart home simulator to generate two datasets based on three different routines and perform experiments with two reinforcement learning algorithms, Deep Q-Learning and Differential Semi-gradient n-step SARSA, benchmarking their performance against an online Supervised Learning baseline based on Behavior Cloning. Additionally, we tested these algorithms using two-state approaches: simple and composite. The results confirm the system’s successful adaptation to user routines across both approaches, underscoring its potential to enhance personalization in smart home environments.
Boaventura et al. (Sun,) studied this question.