StructRL is a set of patches and add-on modules for the upstream projects below, pinned to the commits used in the paper.
| Repository | Commit | Patch | Used for |
|---|---|---|---|
| RLinf | 4993639 |
rlinf_core.patch |
RL training on RoboCasa365 |
| robocasa | 8f3c96e |
robocasa_gymutils.patch |
RoboCasa365 |
| robosuite | 232ce7d (v1.5.2) |
none | RoboCasa365 |
| Isaac-GR00T | 4af2b62 |
isaacgroot_shim.patch |
GR00T-N1.5 |
| openpi (RoboCasa fork) | 5a6beda |
openpi_rc365.patch |
π0.5 SFT and evaluation on RoboCasa365 |
| RLinf, second checkout | 4993639 |
rlinf_core.patch, then libero_addon.patch |
RL training on LIBERO |
| LIBERO (RLinf fork) | 0c5e40c |
none | RL training on LIBERO |
| LIBERO | 8f1084e |
none | LIBERO evaluation client |
scripts/setup_upstream.sh clones them next to the code, applies the patches, and installs the Hydra
configs into RLinf:
git clone https://github.com/amazon-science/StructRL.git && cd StructRL
export PROJECT_ROOT=$PWD
bash scripts/setup_upstream.sh # RoboCasa365 with GR00T-N1.5
bash scripts/setup_upstream.sh --openpi # + π0.5 on RoboCasa365
bash scripts/setup_upstream.sh --libero # + LIBERO (second RLinf in $PROJECT_ROOT/libero_rl)
export LIBERO_RL_ROOT=$PROJECT_ROOT/libero_rl CKPT_ROOT=$PROJECT_ROOT/ckptsThe launchers use the virtual environments below, under $PROJECT_ROOT. envs/ has the frozen package
list of each one (Python 3.11 unless noted). Install the list, then the patched checkouts it names with
pip install --no-deps -e. RLinf itself is used from source through PYTHONPATH.
| Environment | Package list | Used by |
|---|---|---|
.venv-rlinf |
envs/rlinf-gr00t.txt |
GR00T-N1.5 training on RoboCasa365 |
.venv-groot |
envs/groot-eval.txt |
GR00T-N1.5 evaluation; policy server for LIBERO evaluation |
.venv-rlinf-openpi-train |
envs/rlinf-openpi-train.txt |
π0.5 training on RoboCasa365 |
.venv-rlinf-libero |
envs/rlinf-libero.txt |
GR00T-N1.5 training on LIBERO |
.venv-rlinf-openpi-libero |
envs/rlinf-openpi-libero.txt |
π0.5 training on LIBERO |
.venv-libero |
envs/libero-eval.txt (Python 3.10) |
LIBERO evaluation client |
openpi/.venv |
uv sync in the openpi fork |
π0.5 SFT and policy server on RoboCasa365 |
.venv-rlinf-openpi-eval |
robocasa, robosuite, openpi-client |
π0.5 evaluation client on RoboCasa365 |
| What | Source | Location |
|---|---|---|
| RoboCasa365 v1.0 datasets | RoboCasa docs | $PROJECT_ROOT/datasets (see configs/env/) |
| GR00T-N1.5, RoboCasa365 Composite-Seen SFT | robocasa/robocasa365_checkpoints | $PROJECT_ROOT/ckpts/gr00t_n1-5/foundation_model_learning/target_posttraining/composite_seen/checkpoint-60000 |
| π0.5, RoboCasa365 SFT | train with launch/sft_pi05.sh, then convert to PyTorch |
export PI05_SFT_CKPT=<dir> |
| GR00T-N1.5, LIBERO SFT | youliangtan/gr00t-n1.5-libero-long-posttrain (and -spatial, -object, -goal) |
$CKPT_ROOT/gr00t-n1.5-libero-{long,spatial,object,goal} |
| π0.5, LIBERO SFT | RLinf/RLinf-Pi05-LIBERO-130-fullshot-SFT | $CKPT_ROOT/RLinf-Pi05-LIBERO-130-fullshot-SFT |
LIBERO reads its asset paths from $LIBERO_CONFIG_PATH/config.yaml (default ~/.libero), which LIBERO
creates on first import.
The launchers are Slurm batch scripts: training uses 4 nodes with 8 GPUs each, evaluation 2 nodes with
8 GPUs each. Replace <YOUR_PARTITION>, submit from the repository root (Slurm logs go to logs/), and
set NCCL_SOCKET_IFNAME in the NCCL block of each launcher to your network interface prefix.