A State Feedback Bias Compensating Q-learning Value Iteration Algorithm for Model-Free Game-Theoretic HVAC Optimal Control | Synapse