Repository navigation
Keep a gradient second moment when AdaBelief beta1 is 0 - #559
Open
shaneraphel wants to merge 1 commit into
Open
shaneraphel wants to merge 1 commit into
shaneraphel wants to merge 1 commit into
Conversation
With beta1 = 0 the first moment equals the current gradient, so (g - m) is 0 and the second moment stays at eps. The step is then lr * g / sqrt(eps). At eps 1e-16 a scalar quadratic reached -inf on step 10. Adam with beta1 = 0 still tracks g^2. Use that residual only when beta1 is 0. A positive beta1 still stores (1 - beta2) (g - m)^2 plus the eps that is added into the second-moment state.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
With
beta1 = 0the first moment is set to the current gradient, so the belief residualg - mis 0 and the second moment stays ateps. The step islr * g / sqrt(eps). This port's defaultepsis0.001, which hides it, but the paper'sepsis1e-16(and that is whatAdaBelief(..., eps=1e-16)accepts). A scalar quadraticx^2with learning rate1e-2and thatepsreached-infon step 10. The same overflow is juntang-zhuang/Adabelief-Optimizer#67; the fix there is the same idea, kept in that repository's latest package.Only the degenerate case changes. When
beta1 == 0the residual isg, so the second moment tracksg^2the way Adam does atbeta1 = 0. For any positivebeta1the residual is stillg - mafter the first-moment update. One step atbeta1 = 0.9still stores(1 - beta2) (g - m)^2plus theepsthis implementation adds into the second-moment state.Test plan
pytest tests/test_adabelief_beta1.pybeta1matches the residual variance, including the addedeps.x^2atbeta1 = 0,eps = 1e-16, learning rate1e-2stay finite and the parameter moves from0.5toward0.