Loss landscapes and optimization

Training a language model means navigating a loss landscape in high dimensional parameter space. The landscape is non convex, but empirical evidence shows it is remarkably smooth at scale — full of wide, connected basins rather than isolated sharp minima.