← Interactive presentation and experiment library

LUCID · Professor review · Revision 13 · September 22, 2026

Train with feedback.
Deploy the policy.

A 15-second montage of six varied pushes opens the 2:59 submission film. The full, uncut G1 take runs alongside it at normal speed and continues through the story: the deployment dilemma, learned training feedback, simulation stress tests, and physical evidence. Review the timed narration and visual logic below. Results are from the LUCID manuscript, a research draft for review.

The main panel opens with selected highlights; the right panel plays the full 173-second take from the start. That recording continues without cuts or speed changes and expands for the hardware chapter. It shows LUCID under 0–60 ms added randomized command delay and manual perturbations. The paper’s fixed +40 ms hardware benchmark, the +60 ms simulation aggregate, and the selected +40 ms simulation clips are separate evidence, each labeled with its own condition.

Useful review questions: Is the command–execution gap motivation clear? Do the shared encoder and next-block feedback loop communicate the contribution? Are the evidence conditions easy to distinguish? Please identify the timestamp when suggesting a change.

0:00–0:15

One policy. Repeated real-world disturbances.

Frame from the One policy. Repeated real-world disturbances. chapter

Narration

Hand pushes. Foot pushes. Downward pressure. LUCID repeats a turning motion under added randomized delay. The full, uncut take runs alongside our story.

Story and visualization

Six varied push moments play in the main panel for 15 seconds. The full 173-second physical recording starts simultaneously in the right panel and continues at normal speed without cuts. The selected montage and the continuous take are labeled distinctly.

Inspect this animation ↗
0:15–0:28

How much uncertainty should training face?

Frame from the How much uncertainty should training face? chapter

Narration

Too little randomization leaves policies unprepared. Too much makes learning harder. LUCID uses learned command–execution feedback to set training difficulty.

Story and visualization

The deployment dilemma: randomization must balance exposure with learnability. Recorded G1 states show why instantaneous joint error mixes execution mismatch with torque-generating offsets. LUCID compares temporal histories in a learned space.

Inspect this animation ↗
0:28–0:43

Compare histories. Guide the next training block.

Frame from the Compare histories. Guide the next training block. chapter

Narration

Issued commands pass through simulated delay and dynamics. Command and execution histories enter the same frozen encoder. Their normalized latent discrepancy guides the next training block.

Story and visualization

The core idea comes first: issued-command history branches before delay, measured history follows execution, and both enter one frozen encoder. Their normalized discrepancy guides the next training block.

Inspect this animation ↗
0:43–0:52

Learn the comparison. Then freeze it.

Frame from the Learn the comparison. Then freeze it. chapter

Narration

A denoising V A E learns clean motion from noisy windows. Then we freeze its temporal encoder.

Story and visualization

How the comparison is learned: denoise motion windows, then freeze the encoder before policy training. The same comparison function is used while the policy changes.

Inspect this animation ↗
0:52–1:05

One feedback loop. Six training channels.

Frame from the One feedback loop. Six training channels. chapter

Narration

Bounded P I sets randomization intensity from the latent gap. Low returns trigger a separate backoff. One intensity controls six channels for the next training block.

Story and visualization

A concise feedback loop: the upper-quantile latent gap drives bounded PI, with independent return-based backoff. A shared intensity sets six channels for the next block: actuation delay, surface contact, mass and center of mass, joint offsets, pushes, and observation noise. Recorded G1 skeletons illustrate the channels; the diagram is explanatory.

Inspect this animation ↗
1:05–1:18

Stress-test beyond the training delay range.

Frame from the Stress-test beyond the training delay range. chapter

Narration

Training uses up to forty milliseconds of added delay. Frozen policies face sixty milliseconds on full held-out motions, fifty percent above the training maximum.

Story and visualization

Stress testing extends the maximum added delay from 40 ms in training to 60 ms in the held-out test, a 50% increase in this parameter. The test uses nominal dynamics. The separate context replay shows 1,024 environments with 0–40 ms delay; the paper trains with 4,096.

Inspect this animation ↗
1:18–1:35

Unseen delay. Higher completion.

Frame from the Unseen delay. Higher completion. chapter

Narration

Under unseen sixty-millisecond delay, completion rises from fifty-two point one percent with filtered-error P I to seventy-three point eight with LUCID. That is a twenty-one point seven percentage-point gain.

Story and visualization

Completion under unseen +60 ms delay rises from 52.1% to 73.8%, a 21.7 percentage-point gain over filtered-error PI. Reported intervals and the hybrid’s highest tested means remain visible.

Inspect this animation ↗
1:35–1:59

Watch the response to delay and pushes.

Frame from the Watch the response to delay and pushes. chapter

Narration

These selected demonstrations combine forty milliseconds of added delay with scheduled pushes. LUCID is on the left; filtered-error P I is on the right. They show behavior under perturbation. Aggregate completion comes from separate benchmark trials.

Story and visualization

Full six-second walking and turning demonstrations play at 1×, with +40 ms added delay and scheduled impulses. LUCID is left, reference center, filtered-error PI right. These selected recordings differ from the +60 ms aggregate benchmark; root drift remains visible.

Inspect this animation ↗
1:59–2:10

The representation and feedback matter.

Frame from the The representation and feedback matter. chapter

Narration

Denoising raises completion over ordinary reconstruction. Live feedback also improves on one donor’s replayed schedule.

Story and visualization

Denoising and live feedback are tested separately. The replay comparison uses one donor schedule and five recipient seeds; its scope remains explicit.

Inspect this animation ↗
2:10–2:46

From simulation stress tests to physical G1.

Frame from the From simulation stress tests to physical G1. chapter

Narration

This is the same uncut recording, still at normal speed. LUCID repeats the turning reference through hand and foot perturbations, with zero-to-sixty-millisecond added randomized delay. Separately, the paper’s matched forty-millisecond test reports thirty-eight of sixty completions, versus twenty-three for filtered-error P I. The encoder and curriculum scheduler stay in training.

Story and visualization

The same continuous physical take expands without restarting. Its randomized 0–60 ms delay and manual perturbations provide qualitative deployment evidence. A separate panel reports the manuscript’s matched fixed +40 ms benchmark: 38/60 versus 23/60 full-trial completions. These counts do not come from the displayed take.

Inspect this animation ↗
2:46–2:59

Train with feedback. Deploy the policy.

Frame from the Train with feedback. Deploy the policy. chapter

Narration

LUCID uses learned execution feedback to guide domain randomization. On the robot, only the policy and low-level controller remain.

Story and visualization

The deployment path becomes the main takeaway: policy → low-level controller → G1. The encoder and curriculum scheduler remain training components. The choreographed robot lettering remains above the architecture. The continuous take finishes at 2:53, followed by a completion card.

Inspect this animation ↗