Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show t...
No takes yet. Share an insight, caveat, or question.
Mandt et al. (2017) studied this question.