We derive cooling schedules for the global optimization of learning in neural networks. We discuss a two-level system with one global and one local minimum. The analysis is extended to systems with many minima. The optimal cooling schedule is (asymptotically) of the form {η}(t)=η*/lnt, with {η}(t) the learning parameter at time t and η* a constant, dependent on the reference learning parameters for the various transitions. In some simple cases, η* can be calculated. Simulations confirm the theoretical results.
No takes yet. Share an insight, caveat, or question.
Heskes et al. (1993) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: