Deep learning has been behind incredible breakthroughs, such as voice recognition, image classification, and text generation.While these successes are undeniable, our mathematical understanding of deep learning is still developing.More rigorous insight is needed to overcome fundamental challenges, including the interpretability, robustness, and bias of deep learning, and lower its environmental and computational costs.Let us define deep learning informally as machine learning methods that use feed-forward neural networks with many (that is, more than a handful) layers.While most traditional approaches use a finite number of layers, we will focus on more recent approaches that conceptually use infinitely many layers.We will explain those approaches by defining differential equations whose dynamics are modeled by trainable neural network components and whose time roughly corresponds to the depth of the network.Using three examples from machine learning and applied mathematics, we will see how continuous-depth neural network architectures, defined by ordinary differential equations (ODEs), can provide new insights into deep learning and a foundation for more efficient algorithms.Even though many deep learning approaches used in practice today do not rely on differential equations, I find many opportunities for mathematical research in this area.As we shall see, phrasing the problem continuously in time enables one to borrow numerical techniques and analysis to gain more insight into deep learning, design new approaches, and crack open the black box of deep learning.We shall also see how neural ODEs can approximate solutions to high-dimensional, nonlinear optimal control problems.
No takes yet. Share an insight, caveat, or question.
Lars Ruthotto (2024) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: