Repository navigation
Integrate native determinism with canonical contacts and constraint allocation - #1775
Draft
huster-wgm wants to merge 22 commits into
Draft
huster-wgm wants to merge 22 commits into
huster-wgm wants to merge 22 commits into
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
…erministic arithmetic Adapt contact and constraint ordering from google-deepmind#1422 (a93c6d3); preserve current-main adhesion and discrete integration, canonicalize flex candidates before filtering, and order sparse Hessian block inputs.
Use one Hessian block per logical slot in deterministic mode so the triangular per-thread bound is independent of batch occupancy. Keep Warp RUN_TO_RUN reductions and complete new helper docstrings.
Split independent worlds without changing within-world reduction order. Add a 144-world articulated graph regression for the int32 scratch allocation limit encountered with worst-case scatter storage.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This draft combines Taylor Howell's native determinism API and Warp
RUN_TO_RUNarithmetic from #1687 with the contact ordering and count-scan-emit constraint allocation from #1422. It is rebased onto current main9a5455f5, retaining the original API commit and identifying the ordering donor commit (a93c6d3) in the integration history.The initial arithmetic implementation can produce incorrect sparse solves or lose actuator-force contributions: deferred reductions cannot substitute for writes consumed by later substitution levels, and dynamic sparse loops need explicit scatter-record bounds. The integration fixes those cases while preserving current-main adhesion, discrete integration, flex and derivative behavior.
Validation at head
e916b8e:uv run pytest -n 8 -qgives 1984 passed, 60 skipped on CPU and 2003 passed, 39 skipped on RTX PRO 6000, using locked MuJoCo3.14.1.dev990581733/ Warp1.15.0. All pre-commit/kernel-analyzer hooks pass. Thirteen heterogeneous Hessian capacity/support-boundary cases and ten sparse-force row/support cases additionally pass on Warp 1.18, including nested conditional graphs, completed/unchanged worlds, compact mappings and full-capacity/width fallback. The four companion Newton native adapter CPU/GPU tests pass against this head. The new 144-world, 43-DoF sparse graph regression verifies eager/graph equality, finite states and zero overflow. Earlier arithmetic reproductions fail on the original #1687 implementation, and the mixed-contact permutation regression fails before the ordering port.A fixed-input G1 heightfield diagnostic uses native flags only, unchanged 0.025 m terrain / 0.1 m base, no host sorting, and checks all overflow flags. Nine and 144 replicas match bitwise in qpos/qvel/qacc and applied/smooth/constraint forces at every one of 80 substeps. An independent nine-world run remains identical for 400 substeps. Saved qpos matches across independent 1/9/144-world processes on their common 80-step prefix. This is an empirical result for this workload, not a cross-batch/device API guarantee.
A separate GPU-eager closed-loop check runs 48 slope/stair clips through frozen SONIC planner/tracker components with a fixed padded tracker microbatch of 32. Splitting the same clips into 9-world batches versus one 144-world batch gives byte-identical state, reference, action and command arrays for all 144 A/B/C rollouts (44,447 controller frames), with identical termination frames/reasons and zero overflow. An independent 9-world process repeat also matches. After sparse-force dispatch, the full 144-world closed-loop rerun preserves all 48 clips byte-for-byte, including actions and termination. Its measured loop time drops from 159.50 to 141.31 seconds (278.66 to 314.54 active controller frames/s); these are single-run end-to-end measurements, separate from the repeated fixed-input benchmark. This check uses an isolated Newton 1.6 compatibility checkout, MuJoCo 3.14.1 and Warp 1.18; it does not exercise an outer captured tracker loop.
Conditional CUDA Graph execution needs Warp 1.18+ for deterministic scratch allocation inside conditional bodies. Warp 1.15 can use unrolled graphs with
graph_conditional=False, but the 144-world G1 case exhausts 95 GiB when retaining scratch for all solver iterations. The 1.18 conditional path completes without reducing solver iterations or suppressing overflow. ISLANDS remains reserved, so ALL is not a promise of full determinism with island processing.Matched 144-world G1 cost probe on Warp 1.18.0, native conditional solver graphs, unchanged capacities/solver settings: 80 steps from a reset state, one fixed PD target, five synchronized repetitions after warmup, median reported. All runs remain finite with every overflow flag clear.
Hessian support dispatch improves deterministic graph throughput by 3.32x, and sparse-force dispatch adds another 20.28%. The workload has 49 velocity DoFs, but every traced Hessian block has sparse support 15; its maximum active block count varies from 60 to 80. The additional 128-slot bucket avoids jumping directly from 64 to 256. A separate full 80-step trace saves qpos/qvel/qacc/qfrc_constraint for every world: all four complete trajectories are byte-identical before and after this optimization. Five timed repeats also retain identical final states.
A separately instrumented post-window step drops from 31.28 to 24.33 ms after sparse-force dispatch, including 13.71 to 6.76 ms in solve and 7.90 ms in collision; this is separate from the throughput average. Full 80-step qpos/qvel/qacc/qfrc_constraint trajectories remain byte-identical across this additional optimization. Native reductions along the rigid-body tree remain a significant cost. Captured branches retain scratch for all capacity/support choices: the earlier Hessian-only version had one observed total GPU usage sample of 55,871 MiB, not an isolated or matched peak-memory measurement. Peak memory was not remeasured after sparse-force dispatch. Performance and memory remain open work. This constant-target probe measures physics execution cost.
The companion Newton adapter is newton-physics/newton#4632: it enables the native option before derived model constants are computed. This PR remains draft for upstream API/design review and performance work.