An algorithm of the form Xk + 1 = Xₖ - aₖ (∇ U(Xₖ ) + ξ ₖ ) + bₖ Wₖ, where U( · ) is a smooth function on Rᵈ, \ ξ ₖ \ is a sequence of Rᵈ-valued random variables, \ Wₖ \ is a sequence of independent standard d-dimensional Gaussian random variables, aₖ = A / k and bₖ = √ B / √klog log k for k large, is considered. An algorithm of this type arises by adding slowly decreasing white Gaussian noise to a stochastic gradient algorithm. It is shown, under suitable conditions on U( · ), \ ξ ₖ \, A, and B, that Xₖ converges in probability to the set of global minima of U( · ). No prior information is assumed as to what bounded region contains a global minimum. The analysis is based on the asymptotic behavior of the related diffusion process dY(t) = - ∇ U(Y(t))dt + c(t)dW(t), where W( · ) is a standard d-dimensional Wiener process and c(t) = √ C / √log t for t large.
No takes yet. Share an insight, caveat, or question.
Gelfand et al. (1991) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: