Distributed resource optimisation using the Q-learning algorithm, in device-to-device communication: A reinforcement learning paradigm | Synapse