Skip to content

add dsl algorithms support to bench collective - #887

Draft
RJ Souza (Empyreus) wants to merge 20 commits into
mainfrom
rjsouza/multinode-ci
Draft

add dsl algorithms support to bench collective#887
RJ Souza (Empyreus) wants to merge 20 commits into
mainfrom
rjsouza/multinode-ci

Conversation

@Empyreus

@Empyreus RJ Souza (Empyreus) commented Aug 25, 2026

Copy link
Copy Markdown
Contributor
  • add dsl algorithms support to bench collective
  • replace multinode pipeline tests with bench_collective tests

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds multi-node DSL collective autotuning and CI coverage for AllReduce, AllGather, and ReduceScatter.

Changes:

  • Compiles and tunes runtime DSL variants with configurable launch geometry.
  • Adds ReduceScatter benchmarking and correctness validation.
  • Reworks multi-node CI to exercise DSL collectives.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 7 comments.

Show a summary per file
File Description
test/deploy/run_tests.sh Adds manual multi-node DSL test helpers.
python/mscclpp/default_algos/reducescatter_multi_nodes.py Adds a plan-generation CLI.
python/mscclpp_benchmark/tuner.py Supports candidate-specific tuning dimensions.
python/mscclpp_benchmark/correctness.py Validates ReduceScatter output.
python/mscclpp_benchmark/comm.py Compiles and executes DSL variants.
python/mscclpp_benchmark/bench_collective.py Adds DSL candidates and ReduceScatter benchmarking.
.azure-pipelines/multi-nodes-test.yml Adds three multi-node DSL benchmark jobs.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/mscclpp_benchmark/bench_collective.py
Comment thread python/mscclpp_benchmark/bench_collective.py
Comment thread python/mscclpp_benchmark/comm.py Outdated
Comment thread python/mscclpp_benchmark/comm.py
Comment thread python/mscclpp_benchmark/bench_collective.py
Comment thread test/deploy/run_tests.sh Outdated
Comment thread python/mscclpp_benchmark/bench_collective.py
@Empyreus RJ Souza (Empyreus) changed the title Rjsouza/multinode ci add dsl algorithms support to bench collective Aug 25, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 7 out of 7 changed files in this pull request and generated 3 comments.

Suppressed comments (2)

python/mscclpp_benchmark/comm.py:203

  • Every compiled DSL plan is hard-coded as in-place, but --enable-dsl can be combined with --buffer-mode out-of-place, and _dsl_candidate_specs does not filter by algorithm.buffer_mode. For out-of-place AllReduce/AllGather, the tuner therefore executes plans whose buffer layout assumes aliasing, causing incorrect output before falling into its default fallback. Either compile the requested mode where the builder supports it, or reject/filter DSL candidates whose mode does not match the case.
                    collective=collective_op,
                    nranks_per_node=nranks_per_node,
                    world_size=world_size,
                    in_place=True,

python/mscclpp_benchmark/bench_collective.py:405

  • The accepted buffer_mode argument is ignored for ReduceScatter: requesting out-of-place still returns an aliased input/output case. This silently benchmarks a different layout than the CLI requested. Since the registered ReduceScatter DSL algorithm is in-place-only, reject this combination explicitly.
    if collective == _REDUCESCATTER:
        # The DSL reducescatter is compiled in-place, so the per-rank output chunk always aliases the
        # matching slice of the full input buffer (mirrors python/test/executor_test.py build_bufs).
        input_buffer = _mscclpp().GpuBuffer(nelems * comm_group.nranks, dtype=dtype_spec.cupy_dtype)

Comment thread python/mscclpp/language/collectives.py
Comment on lines +408 to +411
return BenchmarkCase(
collective=collective,
message_size=output.nbytes,
total_size=input_buffer.nbytes,
Comment on lines +315 to +317
if n_nodes > 1 and not candidate.supports_multi_node:
filtered_out = True
continue
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants