Nonlinear random matrix theory for deep learning*This article is an updated version of: Ha J-S, Park Y-J, Chae H-J, Park S-S and Choi H-L 2017 Nonlinear random matrix theory for deep learning Advances in Neural Information Processing Systems 30 ed I Guyon, U V Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan and R Garnett (Red Hook, NY: Curran Associates, Inc.) pp 2634–46. (20th December 2019)
- Record Type:
- Journal Article
- Title:
- Nonlinear random matrix theory for deep learning*This article is an updated version of: Ha J-S, Park Y-J, Chae H-J, Park S-S and Choi H-L 2017 Nonlinear random matrix theory for deep learning Advances in Neural Information Processing Systems 30 ed I Guyon, U V Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan and R Garnett (Red Hook, NY: Curran Associates, Inc.) pp 2634–46. (20th December 2019)
- Main Title:
- Nonlinear random matrix theory for deep learning*This article is an updated version of: Ha J-S, Park Y-J, Chae H-J, Park S-S and Choi H-L 2017 Nonlinear random matrix theory for deep learning Advances in Neural Information Processing Systems 30 ed I Guyon, U V Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan and R Garnett (Red Hook, NY: Curran Associates, Inc.) pp 2634–46.
- Authors:
- Pennington, Jeffrey
Worah, Pratik - Abstract:
- Abstract: Neural network configurations with random weights play an important role in the analysis of deep learning. They define the initial loss landscape and are closely related to kernel and random feature methods. Despite the fact that these networks are built out of random matrices, the vast and powerful machinery of random matrix theory has so far found limited success in studying them. A main obstacle in this direction is that neural networks are nonlinear, which prevents the straightforward utilization of many of the existing mathematical results. In this work, we open the door for direct applications of random matrix theory to deep learning by demonstrating that the pointwise nonlinearities typically applied in neural networks can be incorporated into a standard method of proof in random matrix theory known as the moments method. The test case for our study is the Gram matrix FF T, , where W is a random weight matrix, X is a random data matrix, and is a pointwise nonlinear activation function. We derive an explicit representation for the trace of the resolvent of this matrix, which defines its limiting spectral distribution. We apply these results to the computation of the asymptotic performance of single-layer random feature networks on a memorization task and to the analysis of the eigenvalues of the data covariance matrix as it propagates through a neural network. As a byproduct of our analysis, we identify an intriguing new class of activation functions withAbstract: Neural network configurations with random weights play an important role in the analysis of deep learning. They define the initial loss landscape and are closely related to kernel and random feature methods. Despite the fact that these networks are built out of random matrices, the vast and powerful machinery of random matrix theory has so far found limited success in studying them. A main obstacle in this direction is that neural networks are nonlinear, which prevents the straightforward utilization of many of the existing mathematical results. In this work, we open the door for direct applications of random matrix theory to deep learning by demonstrating that the pointwise nonlinearities typically applied in neural networks can be incorporated into a standard method of proof in random matrix theory known as the moments method. The test case for our study is the Gram matrix FF T, , where W is a random weight matrix, X is a random data matrix, and is a pointwise nonlinear activation function. We derive an explicit representation for the trace of the resolvent of this matrix, which defines its limiting spectral distribution. We apply these results to the computation of the asymptotic performance of single-layer random feature networks on a memorization task and to the analysis of the eigenvalues of the data covariance matrix as it propagates through a neural network. As a byproduct of our analysis, we identify an intriguing new class of activation functions with favorable properties. … (more)
- Is Part Of:
- Journal of statistical mechanics. (2019:Dec.)
- Journal:
- Journal of statistical mechanics
- Issue:
- (2019:Dec.)
- Issue Display:
- Volume 1000060 (2019)
- Year:
- 2019
- Volume:
- 1000060
- Issue Sort Value:
- 2019-1000060-0000-0000
- Page Start:
- Page End:
- Publication Date:
- 2019-12-20
- Subjects:
- Statistical mechanics -- Periodicals
Mechanics -- Statistical methods -- Periodicals
530.1305 - Journal URLs:
- http://ioppublishing.org/ ↗
- DOI:
- 10.1088/1742-5468/ab3bc3 ↗
- Languages:
- English
- ISSNs:
- 1742-5468
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 14316.xml