|
Cédric Gerbelot
|
Briefly
I am an assistant professor (chaire de professeur junior) in the Unité de Mathématiques Pures et Appliquées at Ecole Normale Supérieure de Lyon.
I work at the interface between machine learning theory, probability and mathematical physics. More precisely, I am interested in high-dimensional probability, statistics, optimization and mathematical methods inspired by spin glass theory.
I obtained my PhD in 2022 at Ecole Normale Supérieure de Paris under the supervision of Florent Krzakala (EPFL) and Marc Lelarge (INRIA, ENS). I was then a Courant Instructor for two years at the Courant Institute of Mathematical Sciences, NYU, where I worked with Gerard Ben Arous .
Here is a CV.
Contact
E-mail: cedric [dot] gerbelot-barrillon [at] ens-lyon [dot] fr
Physical address: UMPA, ENS Lyon, 46 Allée d'Italie - 69007 Lyon
|
Selected Publications
You can also find my publications on Google Scholar, and some code on Github.
Gerbelot, C. and Mourrat, J.C. Singular Perturbations and Hierarchical Learning in Two-Layer Neural Networks, 2026, Preprint. [arXiv] [Show Abstract]
Abstract: We study the population gradient flow of an infinitely wide two-layer neural network learning a misspecified single-index model in high dimension. The two layers are optimized
jointly, with a perturbative parameter tuning the relative training speed between the first and second layer. This setting
was considered by Berthier, Montanari and Zhou in \cite{berthier2024learning}, who conjectured a hierarchical learning scenario with explicit timescales as the second layer is trained faster than the first.
In this paper, we prove that the constant and linear components of the hidden link function are indeed recovered within the predicted timescales, at sharp explicit thresholds. We then analyze the onset of learning of the quadratic component and show that the components learned at earlier stages continue to influence the dynamics in an essential way. Our proof is based on quantitative approximation results
for singularly perturbed flows evolving near a manifold defined by integral constraints. At a phenomenological level, we also show that the empirical measure of the weights displays
singular behaviour when reaching the quadratic component of the hidden link, with a small fraction of neurons growing significantly while the remaining ones rearrange to preserve the components already learned.
Ben Arous, G., Gerbelot, C., Piccolo, V. (alphabetical order) Stochastic Gradient Descent in High Dimensions for Multi-Spiked Tensor PCA , Communications on Pure and Applied Mathematics, 2026. [arXiv] [Show Abstract]
Abstract:
Part of a series of three papers on the multi-spike tensor estimation problem in high-dimension using Langevin dynamics, gradient flow and online stochastic gradient descent.
Ben Arous, G., Gerbelot, C., Piccolo, V. (alphabetical order) Langevin Dynamics for High Dimensional Optimization: the Case of Multi-Spiked Tensor PCA , Probability Theory and Related Fields, 2026. [arXiv] [Show Abstract]
Abstract:
Part of a series of three papers on the multi-spike tensor estimation problem in high-dimension using Langevin dynamics, gradient flow and online stochastic gradient descent.
Gerbelot, C., Troiani, E., Mignacco, F., Krzakala, F., Zdeborova, L. Rigorous dynamical mean field theory for stochastic gradient descent methods, SIAM Journal on Mathematics of Data Science (SIMODS), 2024. [arXiv] [Show Abstract]
Abstract: We prove closed-form equations for the exact high-dimensional asymptotics of a family of first order gradient-based methods, learning an estimator (e.g. M-estimator, shallow neural network, ...)
from observations on Gaussian data with empirical risk minimization. This includes widely used algorithms such as stochastic gradient descent (SGD) or Nesterov acceleration. The obtained equations match those resulting from the discretization of dynamical mean-field theory (DMFT) equations from statistical physics when
applied to gradient flow. Our proof method allows us to
give an explicit description of how memory kernels build up in the effective dynamics, and to include non-separable update functions, allowing datasets with non-identity covariance matrices. Finally, we provide numerical implementations of the equations for SGD with generic extensive batch-size and with constant learning rates.
Gerbelot, C. and Berthier, R. Graph-based Approximate Message Passing Iterations, Information and Inference : A Journal of the IMA, 2023. [arXiv] [Show Abstract]
Abstract: Approximate-message passing (AMP) algorithms have become an important element of high-dimensional statistical inference, mostly due to their adaptability and concentration properties, the state evolution (SE) equations. This is demonstrated by the growing number of new iterations proposed for increasingly complex problems, ranging from multi-layer inference to low-rank matrix estimation with elaborate priors. In this paper, we address the following questions: is there a structure underlying all AMP iterations that unifies them in a common framework? Can we use such a structure to give a modular proof of state evolution equations, adaptable to new AMP iterations without reproducing each time the full argument ? We propose an answer to both questions, showing that AMP instances can be generically indexed by an oriented graph. This enables to give a unified interpretation of these iterations, independent from the problem they solve, and a way of composing them arbitrarily. We then show that all AMP iterations indexed by such a graph admit rigorous SE equations, extending the reach of previous proofs, and proving a number of recent heuristic derivations of those equations. Our proof naturally includes non-separable functions and we show how existing refinements, such as spatial coupling or matrix-valued variables, can be combined with our framework.
Gerbelot, C., Abbara, A., & Krzakala, F. Asymptotic Errors for Teacher-Student Convex Generalized Linear Models (or : How to Prove Kabashima's Replica Formula), 2020, IEEE Transactions on Information Theory. [arXiv] [Show Abstract]
Abstract: There has been a recent surge of interest in the study of asymptotic reconstruction performance in various cases of generalized linear estimation problems in the teacher-student setting, especially for the case of i.i.d standard normal matrices. Here, we go beyond these matrices, and prove an analytical formula for the reconstruction performance of convex generalized linear models with rotationally-invariant data matrices with arbitrary bounded spectrum, rigorously confirming a conjecture originally derived using the replica method from statistical physics. The formula includes many problems such as compressed sensing or sparse logistic classification. The proof is achieved by leveraging on message passing algorithms and the statistical properties of their iterates, allowing to characterize the asymptotic empirical distribution of the estimator. Our proof is crucially based on the construction of converging sequences of an oracle multi-layer vector approximate message passing algorithm, where the convergence analysis is done by checking the stability of an equivalent dynamical system. We illustrate our claim with numerical examples on mainstream learning methods such as sparse logistic regression and linear support vector classifiers, showing excellent agreement between moderate size simulation and the asymptotic prediction.
Currently Teaching
Graduate Stochastic Calculus - ENS Lyon - Fall 2025, Fall 2026
High-dimensional Gradient Dynamics in Machine Learning - ENS Lyon - Spring 2025, Spring 2026 (roughly one third of the course)
Undergraduate Introduction to Machine Learning - ENS Lyon - Spring 2025, Spring 2026 (roughly two thirds of the course)
|