Skip to content

Integrate native determinism with canonical contacts and constraint allocation - #1775

Draft
huster-wgm wants to merge 22 commits into
google-deepmind:mainfrom
huster-wgm:integrate/determinism-api-contact-ordering
Draft

huster-wgm wants to merge 22 commits into
google-deepmind:mainfrom
huster-wgm:integrate/determinism-api-contact-ordering

Conversation

@huster-wgm

@huster-wgm huster-wgm commented Oct 10, 2026 •

Copy link
Copy Markdown

This draft combines Taylor Howell's native determinism API and Warp RUN_TO_RUN arithmetic from #1687 with the contact ordering and count-scan-emit constraint allocation from #1422. It is rebased onto current main 9a5455f5, retaining the original API commit and identifying the ordering donor commit (a93c6d3) in the integration history.

The initial arithmetic implementation can produce incorrect sparse solves or lose actuator-force contributions: deferred reductions cannot substitute for writes consumed by later substitution levels, and dynamic sparse loops need explicit scatter-record bounds. The integration fixes those cases while preserving current-main adhesion, discrete integration, flex and derivative behavior.

  • Keep fixed-order writes between dependent sparse substitution levels, with native Warp reductions where valid. Bound sparse force/Hessian scatter records and partition independent worlds so captured workspaces do not exceed Warp's int32 array dimension. Select sufficient Hessian capacities on device (16/64/128/256/full capacity) to avoid sorting mostly empty buffers in conditional graphs; retain canonical source order and the full-capacity fallback. Also select a sufficient sparse-support bound (16 or full DoF width), retaining the same arithmetic and reduction order. Apply the same sufficient row/support dispatch to sparse constraint-force projection, including full-coordinate support bounds for compact solves. Allocate dispatch scratch before solver loop capture.
  • Canonicalize rigid/SDF contacts and flex candidates before filtering. Permute every allocated Contact field, including adhesion and multi-column constraint addresses. Sort external contacts before constraint scans when collision detection is disabled.
  • Allocate constraint rows and sparse entries by count/scan/emit, and order Hessian block inputs canonically. Disable the order-sensitive incremental Hessian update in deterministic mode.
  • Add numerical-reference, heterogeneous-world, repeated-step and CUDA Graph regressions; fix analyzer suppression scope and the existing zero-active-constraint Jacobian fixture. Complete new callable docstrings and rebuild the analyzer extension from the corrected sources.

Validation at head e916b8e: uv run pytest -n 8 -q gives 1984 passed, 60 skipped on CPU and 2003 passed, 39 skipped on RTX PRO 6000, using locked MuJoCo 3.14.1.dev990581733 / Warp 1.15.0. All pre-commit/kernel-analyzer hooks pass. Thirteen heterogeneous Hessian capacity/support-boundary cases and ten sparse-force row/support cases additionally pass on Warp 1.18, including nested conditional graphs, completed/unchanged worlds, compact mappings and full-capacity/width fallback. The four companion Newton native adapter CPU/GPU tests pass against this head. The new 144-world, 43-DoF sparse graph regression verifies eager/graph equality, finite states and zero overflow. Earlier arithmetic reproductions fail on the original #1687 implementation, and the mixed-contact permutation regression fails before the ordering port.

A fixed-input G1 heightfield diagnostic uses native flags only, unchanged 0.025 m terrain / 0.1 m base, no host sorting, and checks all overflow flags. Nine and 144 replicas match bitwise in qpos/qvel/qacc and applied/smooth/constraint forces at every one of 80 substeps. An independent nine-world run remains identical for 400 substeps. Saved qpos matches across independent 1/9/144-world processes on their common 80-step prefix. This is an empirical result for this workload, not a cross-batch/device API guarantee.

A separate GPU-eager closed-loop check runs 48 slope/stair clips through frozen SONIC planner/tracker components with a fixed padded tracker microbatch of 32. Splitting the same clips into 9-world batches versus one 144-world batch gives byte-identical state, reference, action and command arrays for all 144 A/B/C rollouts (44,447 controller frames), with identical termination frames/reasons and zero overflow. An independent 9-world process repeat also matches. After sparse-force dispatch, the full 144-world closed-loop rerun preserves all 48 clips byte-for-byte, including actions and termination. Its measured loop time drops from 159.50 to 141.31 seconds (278.66 to 314.54 active controller frames/s); these are single-run end-to-end measurements, separate from the repeated fixed-input benchmark. This check uses an isolated Newton 1.6 compatibility checkout, MuJoCo 3.14.1 and Warp 1.18; it does not exercise an outer captured tracker loop.

Conditional CUDA Graph execution needs Warp 1.18+ for deterministic scratch allocation inside conditional bodies. Warp 1.15 can use unrolled graphs with graph_conditional=False, but the 144-world G1 case exhausts 95 GiB when retaining scratch for all solver iterations. The 1.18 conditional path completes without reducing solver iterations or suppressing overflow. ISLANDS remains reserved, so ALL is not a promise of full determinism with island processing.

Matched 144-world G1 cost probe on Warp 1.18.0, native conditional solver graphs, unchanged capacities/solver settings: 80 steps from a reset state, one fixed PD target, five synchronized repetitions after warmup, median reported. All runs remain finite with every overflow flag clear.

CUDA Graph configuration World-steps/s Time relative to determinism off
Determinism off 15,364 1.00x
Native deterministic, before capacity dispatch 709 21.66x
Block-count dispatch 1,483 10.35x
Hessian block-count and sparse-support dispatch 4,929 3.12x
Plus sparse-force row/support dispatch 5,928 2.59x

Hessian support dispatch improves deterministic graph throughput by 3.32x, and sparse-force dispatch adds another 20.28%. The workload has 49 velocity DoFs, but every traced Hessian block has sparse support 15; its maximum active block count varies from 60 to 80. The additional 128-slot bucket avoids jumping directly from 64 to 256. A separate full 80-step trace saves qpos/qvel/qacc/qfrc_constraint for every world: all four complete trajectories are byte-identical before and after this optimization. Five timed repeats also retain identical final states.

A separately instrumented post-window step drops from 31.28 to 24.33 ms after sparse-force dispatch, including 13.71 to 6.76 ms in solve and 7.90 ms in collision; this is separate from the throughput average. Full 80-step qpos/qvel/qacc/qfrc_constraint trajectories remain byte-identical across this additional optimization. Native reductions along the rigid-body tree remain a significant cost. Captured branches retain scratch for all capacity/support choices: the earlier Hessian-only version had one observed total GPU usage sample of 55,871 MiB, not an isolated or matched peak-memory measurement. Peak memory was not remeasured after sparse-force dispatch. Performance and memory remain open work. This constant-target probe measures physics execution cost.

The companion Newton adapter is newton-physics/newton#4632: it enables the native option before derived model constants are computed. This PR remains draft for upstream API/design review and performance work.

@google-cla

google-cla Bot commented Oct 10, 2026

Copy link
Copy Markdown

Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA).

View this failed invocation of the CLA check for more information.

For the most up to date status, view the checks section at the bottom of the pull request.

…erministic arithmetic

Adapt contact and constraint ordering from google-deepmind#1422 (a93c6d3); preserve current-main adhesion and discrete integration, canonicalize flex candidates before filtering, and order sparse Hessian block inputs.
Use one Hessian block per logical slot in deterministic mode so the triangular per-thread bound is independent of batch occupancy. Keep Warp RUN_TO_RUN reductions and complete new helper docstrings.
Split independent worlds without changing within-world reduction order. Add a 144-world articulated graph regression for the int32 scratch allocation limit encountered with worst-case scatter storage.
@huster-wgm huster-wgm changed the title Rebase determinism API onto main and fix deterministic arithmetic Integrate native determinism with canonical contacts and constraint allocation Oct 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants