Repository navigation
Raise Shampoo preconditioners to -1/(2k) per the paper - #548
Open
shaneraphel wants to merge 2 commits into
Open
shaneraphel wants to merge 2 commits into
shaneraphel wants to merge 2 commits into
Conversation
get_trace averaged the Hessian diagonal over the spatial dims only for 4D kernels. Conv1d (3D) and Conv3d (5D) weights left tmp_output unbound. Average over dims 2..ndim-1 instead; 4D is unchanged.
Each mode factor of an order-k tensor enters to the power -1/(2k): -1/4 per side for matrices. The code used -1/k, twice the published exponent. The 7 pre-existing test_optimizer Shampoo failures are a state-dict precision issue on this torch version and fail identically without this change.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Gupta, Koren, Singer 2018 put each mode factor of an order-
ktensor to the power-1/(2k):-1/4on each side for matrices,-1/2for vectors. The code used-1/k, twice the published exponent on every mode.On a 20×6 least-squares toy at the same learning rate, the paper exponent reaches 0.523 final MSE against 0.817 for the old exponent.
Fixes #502.
Test plan
pytest tests/test_shampoo_exponent.py:X^4 = M^{-1}for a matrix factor,X^2 = M^{-1}for a vector factor, and the toy comparison abovepytest tests/test_optimizer_with_nn.py -k "Shampoo or Adahessian": 2 passedtests/test_optimizer.py -k Shampoohas 7 failures that exist without this change (state-dict precision on this torch version)