-
Weight initialization: why your deep net trains or dies before step one
A deep network does not begin learning from a neutral state. Before the optimizer takes its first step, the initial weights have already determined:
A deep network does not begin learning from a neutral state. Before the optimizer takes its first step, the initial weights have already determined: