Deep neural networks (DNNs) form the backbone of almost every-of-the-art technique in the fields such as computer vision, speech, and text analysis. The recent advances in computational technology made the use of DNNs more practical. Despite the overwhelming performances DNN and the advances in computational technology, it is seen that very few try to train their models from the scratch. Training of DNNs still a difficult and tedious job. The main challenges that researchers face training of DNNs are the vanishing/exploding gradient problem and the non-convex nature of the objective function which has up to million. The approaches suggested in He and Xavier solve the vanishing problem by providing a sophisticated initialization technique. These have been quite effective and have achieved good results on standard, but these same approaches do not work very well on more practical. We think the reason for this is not making use of data statistics for the network weights. Optimizing such a high dimensional loss requires careful initialization of network weights. In this work, we a data dependent initialization and analyze its performance against the initialization techniques such as He and Xavier. We performed our on some practical datasets and the results show our algorithm's classification accuracy.
No takes yet. Share an insight, caveat, or question.
Koturwar et al. (2017) studied this question.