Mean Activation Curvature for Scalable Second-Order Optimization in Deep Networks

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Second-order methods can accelerate deep neural network training, but their adoption is limited by the cost and instability of estimating and inverting curvature matrices. We revisit Kronecker-factored Fisher approximations via an empirical structural analysis of activa- tion and gradient statistics in modern architectures. Across a range of models, we find that activation statistics capture most of the effective curvature directions, while gradient statis- tics mainly act as a global or diagonal rescaling. Based on this observation, we propose mean activation curvature, a scalable curvature surrogate that yields two optimizers, MAC and SMAC, offering different trade-offs between expressiveness and efficiency. We further extend the construction to self-attention layers with a structured approximation that retains the role of attention scores, and provide a convergence analysis under standard assumptions. Experiments on vision and language models show that MAC/SMAC matches or improves the accuracy of existing second-order baselines while reducing training time and memory usage. Code: github.com/hseung88/mac.

키워드

Deep learning optimizationSecond-order optimizationCurvature approximationPreconditioned gradient methods
제목
Mean Activation Curvature for Scalable Second-Order Optimization in Deep Networks
저자
Seung, HyunseokLee, JaewooKo, Hyunsuk
DOI
10.1007/s10115-026-02781-7
발행일
2026-05
유형
정기학술지(Article(Perspective Article포함))
저널명
Knowledge and Information Systems
68
1
페이지
1 ~ 30