Second-order methods promise improved stability and faster convergence, yet they remain underused due to implementation overhead, tuning brittleness, and the lack of composable APIs. We introduce Somax, a composable Optax-native stack that treats curvature-aware training as a single JIT-compiled step governed by a static plan. Somax exposes first-class modules – curvature operators, estimators, linear solvers, preconditioners, and damping policies – behind a single step interface and composes with Optax by applying standard gradient transformations (e.g., momentum, weight decay, schedules) to the computed direction. This design makes typically hidden choices explicit and swappable. Somax separates planning from execution: it derives a static plan (including cadences) from module requirements, then runs the step through a specialized execution path that reuses intermediate results across modules. We report system-oriented ablations showing that (i) composition choices materially affect scaling behavior and time-to-accuracy, and (ii) planning reduces per-step overhead relative to unplanned composition with redundant recomputation.

Second-Order, First-Class: A Composable Stack for Curvature-Aware Training / Korbit, M., Zanon, M.. - 16945:(2026), pp. 90-105. [10.1007/978-3-032-37670-1_6]

Second-Order, First-Class: A Composable Stack for Curvature-Aware Training

Korbit, Mikalai
;
Zanon, Mario
2026

Abstract

Second-order methods promise improved stability and faster convergence, yet they remain underused due to implementation overhead, tuning brittleness, and the lack of composable APIs. We introduce Somax, a composable Optax-native stack that treats curvature-aware training as a single JIT-compiled step governed by a static plan. Somax exposes first-class modules – curvature operators, estimators, linear solvers, preconditioners, and damping policies – behind a single step interface and composes with Optax by applying standard gradient transformations (e.g., momentum, weight decay, schedules) to the computed direction. This design makes typically hidden choices explicit and swappable. Somax separates planning from execution: it derives a static plan (including cadences) from module requirements, then runs the step through a specialized execution path that reuses intermediate results across modules. We report system-oriented ablations showing that (i) composition choices materially affect scaling behavior and time-to-accuracy, and (ii) planning reduces per-step overhead relative to unplanned composition with redundant recomputation.
2026
9783032376695
9783032376701
Machine Learning Systems, Curvature-Aware Training, Second-Order Optimization, Gauss-Newton Methods, Matrix-Free Linear Algebra, Conjugate Gradient, Execution Planning, JAX
File in questo prodotto:
File Dimensione Formato  
2603.25976v1.pdf

Accesso aperto

Tipologia: Documento in Pre-print
Licenza: Creative commons
Dimensione 999.75 kB
Formato Adobe PDF
999.75 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11771/44298
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • OpenAlex 0
social impact