We show how nonlinear embedding algorithms popular for use with “shallow ” semi-supervised learning techniques such as kernel methods can be easily applied to deep multi-layer architectures, either as a regularizer at the output layer, or on each layer of the architecture. This trick provides a simple alternative to existing approaches to deep learning whilst yielding competitive error rates compared to those methods, and existing shallow semi-supervised techniques.
No takes yet. Share an insight, caveat, or question.
Weston et al. (2008) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: