LLUCIDTrain for deployment
Review copy · v13
AN ANIMATED COMPANION TO THE PAPER

Prepare the policy for real-world delay.

Learned feedback in training.
Policy and controller onboard.

0:00 / 2:59
THE CORE IDEA

Execution becomes curriculum.

A learned motion representation turns the command–execution relationship into a signal for training difficulty.

REVIEW THE PRESENTATION

From learned feedback to physical G1.

The 2:59 submission film opens with six varied push moments. The full 173-second take starts alongside them and continues at 1× through the story: the deployment dilemma, learned feedback, simulation stress tests, and physical G1 evidence.

Review copy · 2:59 submission film · six push highlights · continuous physical take.
PHYSICAL UNITREE G1 / ONE CONTINUOUS TAKE

Looped turning under manual perturbations

LUCID policy · 0–60 ms added randomized command delay. This 173-second excerpt runs at normal speed with no internal cuts, repeats, or speed changes.

The selected demonstration is separate from the paper’s matched +40 ms benchmark. Manual perturbation forces were not measured. No matched filtered-error PI hardware recording is shown here.

Inspect the annotated hand and foot contacts

Times refer to this continuous excerpt. Arrows show approximate directions in the camera view. Sustained contacts are grouped; these labels are not measured force impulses.

LabelExcerpt timeContactView directionInspect
Editable annotation timing ↗
THE METHOD / FOLLOW THE CAUSE AND EFFECT

One feedback loop. The next training world.

Move the inputs. Watch the published controller change the ranges for the next block.

ILLUSTRATIVE INPUTS · NOT A TRAINING LOG

What should happen next?

Nominal gap reference r = 0.145. kP = 0.8; kI = 0.15; α = 0.04. Integral limited to ±0.8; controller output to ±1.

NEXT BLOCK
0.4960
PI pacing

e = 1 − y/r
I ← clip(I + e, −0.8, 0.8)
u = clip(0.8e + 0.15I, −1, 1)
λ_next = clip(λ + 0.04u, 0, 1)
λ stays fixed within the block. Environment parameters still follow their separate sampling clocks.
ONE INTENSITY → SIX CHANNELS

Representative ranges from the manuscript. Continuous ranges interpolate from nominal settings; discrete delay bounds advance in 5 ms ticks.

Explore the supplied 90-second process-film study

A cinematic method illustration: recorded evaluation poses, MuJoCo geometry, and staged signal diagrams. The close-up ghost is a representative reference; histories, latent vectors, controller inputs, and formation lettering are illustrative. This optional study has its own ambient soundtrack and is separate from the narrated 179-second presentation.

Design choices and evidence boundaries ↗
THE DATA / RECORDED STATES → G1 FORWARD KINEMATICS

The skeleton comes from the robot model.

Walk, turn, and side-step through all 300 captured samples. The curves and skeletons share one clock.

3 motions · 50 Hz · full 6 s captures
0.02 / 6.00 s
WHAT YOU ARE SEEING

Forward kinematics from the original G1 joint states and source references, using the same MuJoCo model as the videos. Skeletons align the pelvis in XY for pose comparison; height and orientation stay intact. Root error uses the original world positions. Dashed lines mark the 2 s and 4 s impulse schedule.

KEEP THE EVIDENCE CONNECTED

The supplied bundle and author identify LUCID on the left and filtered-error PI on the right, with the shared reference in the center. The curves use those corresponding tracks; aggregate performance estimates come from the manuscript. The actual issued-command histories and matching paper encoder outputs are not in these captures, so the method equations are shown symbolically.

THE EXPERIMENTS / PARALLEL EVALUATION

One policy. A thousand different experiences.

From one robot’s actuator delays to a field of 1,024 independent environments.

24 s replay · 0.25× simulation time
Isaac Lab physics / state replayFrozen policy · no resets
FOLLOW THE EXPERIMENT

See what each robot experiences.

Blue bars → actuator delay

The visible delay range is 0–40 ms. Each robot’s bar shows its maximum group delay.

Orange arrows → applied pushes

Direction and magnitude are visible when a velocity impulse occurs.

Green / red → failure status

Green means no failure yet. Red persists after the first detected failure.

MEASURED IN THE PAPER

Scale in the replay.
Robustness in the results.

88.9%

Full-range simulation completion

Reported LUCID mean
73.8%

Unseen 60 ms delay completion

Reported LUCID mean
+25.0 pp

G1 gain with +40 ms added delay

Paired difference vs filtered-error PI

The footage visualizes evaluation with a frozen policy; it does not show training updates. The replay’s live status counts are separate from the paper’s completion estimates above. Original source ↗

THE EXPERIMENTS / MOTION LIBRARY

One motion. Three synchronized views.

Explore all five original excerpts. Walking and turning also include their complete six-second recordings.

Original MP4 ↗24-second presentation cut ↗
HOW TO READ THE COLORS

Red: pelvis below 60% of reference height for 0.10 s. Amber: displayed tracking error above 0.50 m. Cyan: kinematic reference, without perturbations.

WHAT THESE CLIPS SHOW

Post-hoc balance excerpts, selected from eight motion pairs. Some LUCID runs fall after the excerpt ends. These clips do not establish overall success rates. Isaac Lab clips show PhysX poses rendered in MuJoCo.

Selection record and policy provenance

Full six-second MuJoCo pose captures for walking, turning, and side stepping are available in the data explorer. The MP4 library retains the selected excerpts. Timings below come from the supplied selection audit; “none” means no height-threshold fall within those six seconds.

DynamicsMotionExcerptLUCID fallFiltered-error PI fall
Clip conditions, source hashes, and timing ↗
ADDITIONAL PAPER EVIDENCE

Explore the details at your own pace.

The film emphasizes the main effect and two design tests. These figures retain the policy-response curves, return-backoff ablation and matched allocations.

The complete training architectureCommand and measured histories enter a shared frozen encoder; latent discrepancy feeds PI pacing and return backoff sets the next training block
Policy response across injected delaysPaper Figure 2b policy-dependent latent discrepancy with Table III completion results
Return backoff and matched channel allocationsPaper component ablations and matched channel allocation results
From the paperTraining rolloutExecution feedbackFeedback controlNext training blockImage & figure sources ↗