Adaptive Training for Robot Learning
Tasks, rewards, and domain randomization — GACL, Reward Training Wheels, and LUCID
Overview
My research asks how a teacher can understand a robot student and adapt three parts of training: tasks, auxiliary rewards, and simulation conditions. GACL and Reward Training Wheels (IROS 2025) address the first two decisions; LUCID (preprint, 2026) extends the agenda to execution-informed domain randomization. These are complementary methods with distinct mechanisms and evaluations.
Explore the animated research story for teacher–student diagrams, classroom analogies, a scheduling demo, and linked results.
Grounded Adaptive Curriculum Learning (GACL)
GACL (Wang et al., 2025) (with Zifan Xu, Peter Stone, and Xuesu Xiao) introduces a teacher-student paradigm in which an informed teacher generates training tasks by monitoring student performance in real time, while grounding the curriculum in limited reference samples from the target task distribution so training remains relevant to deployment.
Published results: 6.8% higher success rate than state-of-the-art curriculum methods on wheeled navigation in constrained environments, and 6.1% higher on quadruped locomotion in challenging 3D confined spaces.
Reward Training Wheels (RTW)
RTW (Wang et al., 2025), co-first-authored with Tong Xu, automates auxiliary reward shaping: a teacher adaptively weights auxiliary reward components as the robot’s proficiency grows, while the primary objective stays fixed. Weights can rise or fall; they are not constrained to fade monotonically.
Published results: In simulation, RTW achieved a 2.35-point increase in navigation success and a 122.62% relative improvement in off-road mobility on vertically challenging terrain, reaching the respective performance thresholds about 35% and 3× faster. Sim-trained policies achieved 5/5 physical off-road trials versus 2/5 for expert-designed rewards, with up to 47.4% reduction in orientation angles (more stable poses).
LUCID: pacing domain randomization
LUCID is a 2026 preprint on humanoid motion tracking. A frozen temporal encoder compares issued-command and measured-execution histories. A bounded PI scheduler adjusts shared randomization intensity, with return-based backoff when training degrades.
The study reports 88.9% vs. 76.8% full-randomization simulation completion against filtered-error PI (Table III). On the physical Unitree G1 with 40 ms added delay, completion was 38/60 vs. 23/60 across four motions, three checkpoints, and five repetitions (Table VII). These are preprint study results. Only the policy and low-level controller are needed at deployment.
Platforms
- Wheeled ground robots navigating highly constrained spaces
- Quadruped locomotion in confined 3D environments
- Off-road vehicles on vertically challenging terrain
- Humanoid motion tracking in LUCID simulation and physical G1 experiments (preprint)