Disordered Dynamics in High Dimensions: Connections to Random Matrices and Machine Learning

TL;DR

Using DMFT to analyze high-dimensional random matrix-driven dynamics, revealing learning and generalization mechanisms.

cond-mat.dis-nn 🔴 Advanced 2026-01-03 47 views
Blake Bordelon Cengiz Pehlevan
random matrix dynamical systems machine learning DMFT high-dimensional analysis

Key Findings

Methodology

This work employs cavity method and path integral techniques to analyze infinite-dimensional systems driven by random matrices. By constructing single-site stochastic processes characterized by correlation and response functions, the study links spectral properties of matrices to system behavior. For linear time-invariant systems, the resolvent of the random matrix is connected to the response function, revealing spectral density laws such as Wigner’s semicircle. Applications include gradient flow, stochastic gradient descent on random feature models, and deep linear networks, with bias-variance decompositions derived via subset averaging of DMFT noise variables. The study extends to non-Hermitian matrices, showing they can induce non-monotonic loss curves, contrasting with Hermitian cases. Asymptotic training and test loss dynamics for deep linear networks with high-dimensional data are also derived, with weights modeled as spiked random matrices.

Key Results

  • In GOE matrices, the response function R(τ) matches the Wigner semicircle law, with spectral density ho(\lambda) = rac{1}{2\pi}\sqrt{4 - \lambda^2} over [-2, 2], confirmed by numerical simulations with N=8000. The spectral distribution aligns with theoretical predictions, validating the spectral-response connection.
  • In random feature models, the DMFT correlation functions capture non-monotonic test loss behaviors near interpolation thresholds. Bias-variance decomposition via noise subset averaging reveals mechanisms behind non-linear generalization phenomena.
  • Spectral analysis of non-Hermitian matrices shows complex eigenvalue distributions, such as the circular law and diagonally modulated spectra, which induce non-monotonic training loss curves. Hermitian matrices with matching spectra do not produce such effects, highlighting different underlying mechanisms.

Significance

This research advances the theoretical understanding of high-dimensional stochastic systems relevant to machine learning. By linking spectral properties of random matrices to learning dynamics, it offers a unified framework to analyze generalization, training stability, and non-linear effects. These insights can inform the design of more robust algorithms, improve understanding of deep network training, and foster cross-disciplinary applications in physics, neuroscience, and ecology, addressing longstanding challenges in high-dimensional statistics and complex systems.

Technical Contribution

The paper introduces a comprehensive DMFT-based framework, combining cavity and path integral methods, to analyze the dynamics of systems driven by random matrices. It provides explicit formulas for correlation and response functions, connects spectral density laws to dynamical behavior, and extends analysis to non-Hermitian matrices and non-linear models. The approach yields rigorous asymptotic descriptions of training and test errors, revealing mechanisms behind non-monotonic loss curves and spectral influences on learning dynamics, thus enriching the theoretical toolkit for high-dimensional statistical physics and machine learning.

Novelty

This work is the first systematic application of DMFT to high-dimensional random matrix-driven systems in machine learning, especially for non-Hermitian matrices and deep linear networks. It uncovers novel mechanisms for non-monotonic loss behavior, distinct from classical eigenvalue instability, emphasizing the role of spectral structure and correlation functions. The integration of cavity and path integral techniques offers a new, rigorous perspective on the spectral-dynamical relationship, filling a significant gap in the theoretical understanding of complex learning systems.

Limitations

  • The models primarily assume linearity or near-linear regimes; real-world deep networks with nonlinear activations may exhibit additional complexities not captured here.
  • Spectral analyses are asymptotic, and finite-size effects could lead to deviations in practical settings.
  • Numerical implementations of the theoretical models are computationally intensive, limiting immediate scalability to very high dimensions or complex architectures.

Future Work

Future research will extend the framework to fully nonlinear deep networks, incorporating activation functions and multi-layer interactions. Exploring spectral evolution during training, especially in non-convex regimes, is a key direction. Additionally, integrating these insights into algorithm design could improve training stability and generalization, bridging theory and practice in high-dimensional machine learning.

AI Executive Summary

High-dimensional systems driven by randomness are fundamental in physics, neuroscience, ecology, and machine learning. Traditional tools often fall short in capturing their complex dynamics, especially in nonlinear or non-Hermitian cases. This study introduces a unified analytical framework based on dynamical mean field theory (DMFT), combining cavity and path integral methods, to systematically analyze these systems. By focusing on spectral properties of random matrices, the work reveals how eigenvalue distributions influence system responses, training, and generalization behaviors.

In the simplest case of GOE matrices, the response function aligns with Wigner’s semicircle law, confirming the spectral-density connection. Extending to models like random feature regression, the formalism captures non-monotonic test loss curves, especially near interpolation thresholds, by decomposing bias and variance through subset averaging of DMFT noise variables. For non-Hermitian matrices, the spectral analysis uncovers complex eigenvalue distributions, such as the circular law, which induce oscillatory or non-monotonic training loss behaviors, contrasting with Hermitian matrices.

Moreover, the framework provides asymptotic descriptions of training and test errors for deep linear networks trained on high-dimensional data, with weights modeled as spiked random matrices. These results deepen our understanding of how spectral structures govern learning dynamics, stability, and generalization in high-dimensional regimes. The insights gained pave the way for designing more robust algorithms, understanding the spectral evolution during training, and extending the analysis to nonlinear deep networks. Overall, this work bridges statistical physics, random matrix theory, and machine learning, offering a powerful toolset for future research in complex high-dimensional systems.

Deep Dive

Plain Language Accessible to non-experts

想象一个大型工厂,里面有许多机器(代表系统中的变量),它们相互影响、合作,生产出不同的产品(系统状态)。这些机器的影响有时是随机的,比如原料的质量不稳定,导致工厂的整体表现变得复杂难预测。科学家用一种叫DMFT的方法,就像是在每台机器上装上传感器,记录它们的工作状态和对变化的反应。通过分析这些传感器数据,可以理解工厂的整体运作规律。这个方法帮助我们知道,为什么有时候工厂的产量会突然变好或变差,甚至出现意料之外的波动。它还告诉我们,影响工厂的“噪声”有多大,以及不同机器之间的关系是怎样的。这个研究就像用数学的“放大镜”,看清了复杂系统背后的基本规律,让我们更好地设计和控制这些系统,无论是机器、神经网络还是生态环境。

ELI14 Explained like you're 14

想象你在学校,有很多学生(代表变量)一起学习。每个学生的成绩受老师布置的作业(输入)和同学们的影响(相互作用)。有时候,老师布置的题目很随机(像随机矩阵),每个学生的表现变得难以预测。科学家用一种叫DMFT的方法,就像给每个学生装上“心情传感器”,记录他们的学习状态和对变化的反应。通过分析这些“传感器”数据,可以理解整个班级的学习动态。比如,为什么有时候成绩会突然提高,有时候又会下降?这就像系统中的“噪声”在起作用。这个方法帮我们看清了复杂系统的基本规律,不管是学校、神经网络还是生态系统,都可以用它来理解和改善。它让我们知道,随机的影响其实可以用数学模型描述得很清楚,从而更好地控制和优化这些系统。就像老师用科学的方法帮助学生更好地学习一样,科学家用DMFT让复杂系统变得更可控、更理解。

Abstract

We provide an overview of high dimensional dynamical systems driven by random matrices, focusing on applications to simple models of learning and generalization in machine learning theory. Using both cavity method arguments and path integrals, we review how the behavior of a coupled infinite dimensional system can be characterized as a stochastic process for each single site of the system. We provide a pedagogical treatment of dynamical mean field theory (DMFT), a framework that can be flexibly applied to these settings. The DMFT single site stochastic process is fully characterized by a set of (two-time) correlation and response functions. For linear time-invariant systems, we illustrate connections between random matrix resolvents and the DMFT response. We demonstrate applications of these ideas to machine learning models such as gradient flow, stochastic gradient descent on random feature models and deep linear networks in the feature learning regime trained on random data. We demonstrate how bias and variance decompositions (analysis of ensembling/bagging etc) can be computed by averaging over subsets of the DMFT noise variables. From our formalism we also investigate how linear systems driven with random non-Hermitian matrices (such as random feature models) can exhibit non-monotonic loss curves with training time, while Hermitian matrices with the matching spectra do not, highlighting a different mechanism for non-monotonicity than small eigenvalues causing instability to label noise. Lastly, we provide asymptotic descriptions of the training and test loss dynamics for randomly initialized deep linear neural networks trained in the feature learning regime with high-dimensional random data. In this case, the time translation invariance structure is lost and the hidden layer weights are characterized as spiked random matrices.

cond-mat.dis-nn stat.ML