2 problems
- 0 votes0 replies0 views
Conjecture explaining the difference between fixed and dynamically learned optimizer coefficients
Let \texttt{MADA}\xspace\ be a meta-adaptive optimizer that learns interpolation coefficients, and let \texttt{MADA}\xspace\-FS denote the version using the fixed final optimiz…
- 0 votes0 replies0 views
Conjecture that AMSGrad's maximum operator impedes hyper-gradient flow
The base optimizer AMSGrad uses a maximum operator in its second-moment term, and \texttt{MADA}\xspace\ updates optimizer coefficients using hyper-gradients. The hyper-gradient o…