Blog / Docs • Email • ORCID
I am currently a Master's student in Artificial Intelligence at Shanghai Normal University. My work focuses on real-world Vision-Language-Action (VLA) systems and reinforcement-learning post-training for robotic manipulation. I build real-world pipelines from teleoperation and multi-camera data collection to policy fine-tuning, server-client deployment, and fixed-protocol evaluation. I also study how rollout, failure-recovery, and human-intervention experience can be converted into policy-update signals.
- Real-world VLA pipelines: Teleoperation, multi-camera episode collection, SFT / LoRA / full-parameter fine-tuning, and policy server-client deployment.
- VLA post-training: Reproducing and comparing noise-space steering, advantage-conditioned policy learning, online data distillation, and residual actor-critic methods.
- Evaluation & data flywheel: Tracking revisions, configurations, logs, checkpoints, and fixed evaluation protocols across LIBERO, MetaWorld, ManiSkill, and real robots.
- VLA Models & Data: LeRobot | π0.5 | GR00T N1.7 | SmolVLA | PEFT
- Robotics & Deployment: Unitree G1 | SOARM-101 | PICO / XR | LinkerHand L10 | RealSense | ROS / ROS2 | ZMQ | Action Chunking
- RL & Simulation: SAC | Residual Actor-Critic | LIBERO | MetaWorld | ManiSkill | Ray
- Deep Learning & Engineering: Python | PyTorch | Docker | Linux | Git | W&B
- 3D Vision: 3D Gaussian Splatting | Camera Localization
- 🧪 VLA Post-training Experiments
- Reproductions of DSRL, RECAP, STEAM, FlowDAgger, and RL Token with tracked configurations, checkpoints, evaluation results, and explicit protocol boundaries.
- 🦾 Real-world VLA Systems
- Built teleoperation, multi-camera data, training, and deployment pipelines for Unitree G1 + GR00T N1.7 and a dual-arm robot + π0.5; completed 20/20 in a fixed-scene G1 demo protocol and 18/20 full-task trials on the dual-arm task.
- 🤖 SOARM-101 × SmolVLA Pick-and-Place
- Completed task SFT, failure-oriented demonstration supplementation, and a real-world inference loop, reaching 17/20 full-task successes across mixed object-and-target position combinations.
- 📄 GSplatLoc: Camera Localization via 3D Gaussian Splatting — First Author
- Formulated camera pose estimation as optimization through differentiable 3DGS rendering; the accompanying GitHub repository has 90+ stars.
"From demonstrations to deployment-time policy improvement."




