The Langevin updating rule, in which noise is added to the weights during learning, is presented and shown to improve learning on problems with initially ill-conditioned Hessians. This is particularly important for multilayer perceptrons with many hidden layers, that often have ill-conditioned Hessians. In addition, Manhattan updating is shown to have a similar effect.
No takes yet. Share an insight, caveat, or question.
Thorsteinn Rögnvaldsson (1994) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: