MIT CSAIL researchers have proposed revisiting the classical gradient descent mechanism through the lens of dual vector spaces, bridging theory with Shampoo and µP optimization methods.