Humanoid motion tracking · simulation study · ICRA 2027 submission

When failure is not difficulty

A humanoid that keeps failing on a motion may be facing a hard problem, or it may be chasing a reference that no admissible contact could ever support. Those two cases look identical to a training curriculum, and they call for opposite responses.

2,442 / 10,705
references flagged by the screen in one retargeting pipeline
39 / 40
decisions agreeing across two independent implementations
Inconclusive
registered allocation benefit, at three paired seeds

Simulation only · . Every number on this page comes from simulation. This project has not run a physical robot experiment of any kind. What that would take is set out below.

01 · The problem

Persistent failure is an ambiguous instruction.

Large motion banks are retargeted from human capture to a robot, and trainers spend extra updates wherever the policy currently fails. That works only if more practice can change the outcome.

A retargeted trajectory can demand a base wrench that no admissible contact can supply, or joint effort beyond the modeled actuator range. We call this reference–physics misalignment. Under it, failure conflates three different things: an inadmissible reference, a motion the bank barely covers, and a motion that is genuinely hard.

Those call for opposite responses — repair or exclude the reference, improve coverage, or practice more. CLIMB separates reference admission from practice allocation so each can be measured on its own terms.

02 · The paper

Screening misalignment, and a matched test of allocation.

Submission draft · complete simulation result

When Failure Is Not Difficulty: Screening Reference–Physics Misalignment and Testing Adaptive Allocation on Exact Support for Humanoid Motion Tracking

Persistent tracking failure need not identify useful practice: a retargeted reference may demand support unavailable under the declared robot and scene. We study this reference–physics misalignment through model-relative screening and controlled allocation experiments. In a three-seed Unitree G1 campaign, top-1 exposure peaks at 87–89%, an unsupported kneel-and-crawl reference repeatedly attracts practice, and uniform sampling yields higher held-out survival than that sampler in each seed. A final-trajectory contact-capacity screen flags 2,442 of 10,705 AMASS-derived references in one retargeting pipeline and 7 of 4,950 in a separate production pairing; two implementations agree on 39 of 40 enriched-panel decisions. Feasibility features improve cross-policy difficulty transfer from Spearman correlation 0.567 to 0.609 over 100 held-out clips. We then bind training to 1,184 exact temporal units and compare four allocators at equal support and 49.2 million transitions per policy. Relative-progress allocation changes exposure but does not establish the registered tracking benefit: feasible-hard R−U TrackingScore is −0.0151, with paired seed-level 95% interval [−0.0944, +0.0642]. The broad non-regression criterion also fails, and exploratory learning curves favor uniform. These results separate reference admission, sampling intervention, and policy utility; they establish neither allocation equivalence nor a general advantage of gating.

The paper makes two contributions. First, a misalignment diagnosis and a model-relative contact-capacity screen, evaluated at bank scale and across two implementations, with feasibility features that carry difficulty information between policies sharing a learner. Second, an exact temporal-support training interface together with a matched four-arm, three-seed allocation study that returns a complete, inconclusive policy result and explicit exposure measurements.

A third contribution was planned and deliberately removed rather than left pending. The admission on/off experiment that would have tested it has no result yet; why is described below. Nothing in the paper depends on it.

Progress-driven curricula are not new here. ALP-GMM establishes learning-progress precedent, LIMMT studies physics-aware motion curation, and GMT and EGM include adaptive motion sampling. What is tested here is the final-trajectory admissibility screen coupled to exact temporal support, and whether allocation inside that support buys anything.

03 · Screening

The screen asks a physical question, not a kinematic one.

For each frame, contact-free inverse dynamics gives the wrench the environment must supply. A linear program then asks whether admissible contacts can supply it within the modeled actuator limits. What is left over is unsupported force.

01
Measured · bank scale

Misalignment is common in one pipeline and rare in another.

Under a fixed rule, 2,442 of 10,705 AMASS-derived references are flagged in one retargeting pipeline (22.8%), against 7 of 4,950 in a separately filtered production pairing (0.14%). Source rates inside the first corpus range from 0.1% to 100%.

How we use this: prevalence is a property of a corpus-and-pipeline pairing, never of retargeted motion in general. The two counts are never pooled, and the contrast is not a causal comparison of retargeters: corpus, filtering, robot file and friction all differ.

02
Measured · cross-implementation

Two independent implementations agree.

On a deterministic, flag-enriched 20+20 panel passed through both codebases, strict decisions agree for 39 of 40 clips (Cohen's κ = 0.948), with rank correlation 0.984 on unsupported fraction and 0.997 on airborne fraction.

How we use this: this closes an implementation-agreement question. It is not a prevalence sample, and it does not remove the corpus and filtering confounds between the two full-bank rates.

03
Sealed · cross-policy transfer

Feasibility carries difficulty information between policies.

An eleven-feature intrinsic atlas fit to one policy's per-clip difficulty ranks another policy's difficulty at Spearman 0.567. Adding three screen features raises it to 0.609, above 198 of 200 control designs in which i.i.d. standard-normal columns replace them (one-sided p = 0.010).

How we use this: the control establishes that the gain is not from added design width alone. It is not a comparison against three other real reference features, and the transfer is across policies sharing an architecture, not across architectures or robots.

04
Measured · the motivating case

An adaptive sampler spent its budget on a reference nothing could support.

A failure-driven sampler concentrated on the same kneel-and-crawl reference in all three seeds. Campaign peak top-1 exposure reached 0.884, 0.870 and 0.893, though in two of the seeds that maximum belonged to a different clip. During the attractor’s descent the nearest collision geometry sits 7–10 cm above the floor while a median 329 N of support demand — against a 327 N model weight — has no modeled source. The reference passes ordinary kinematic checks.

How we use this: the same campaign's third arm, a sampler holding exactly 10% of mass on the uniform prior, peaked at 0.568/0.649/0.696 instead and beat uniform in all three paired seeds. The collapse therefore tracks the sampler's additive-floor construction, not outcome-driven allocation as such.

05
Measured · actuator-model sensitivity

The screen barely moves when the actuator model tightens.

The screen compiles the same ±139 N·m knee and hip-roll ranges as the training model, so we re-screened the 900-clip bank behind our exact support at 120 and 90 N·m, changing nothing else. The baseline reproduced the published screen exactly on all 900 clips. Tightening left 888 of 900 clips unchanged and moved the flagged count from 99 to 98; no clip became newly flagged.

How we use this: on this bank the screen’s decisions are dominated by the unsupported-wrench residual rather than the actuator channel. The run also refuted the prediction we registered before running it. We argued a tighter limit could only raise the flagged fraction; it fell, because the reported quantity is the translational component of a residual whose total the program minimizes, so tightening redistributes slack between components. Four clips fall, by at most 0.024, and the lower-bound claim that argument supported is withdrawn.

04 · The controlled experiment

Changed practice did not produce a better tracker.

Four allocation rules, one learner, identical references and support, 49,152,000 simulator transitions per policy, three paired training seeds. Final feasible-hard R−U TrackingScore is −0.01507, with two-sided seed-level 95% interval [−0.09436, +0.06422]. The registered decision is inconclusive.

Final checkpoint 3999 · three independently trained seed pairs
R−U panelSeed 21Seed 22Seed 23Mean95% t interval
Feasible-hard · primary−0.03400−0.03299+0.02178−0.01507[−0.09436, +0.06422]
All-panel · guard−0.04892−0.01402+0.03500−0.00932[−0.11405, +0.09542]

Two seeds favor uniform and one favors relative progress. The +0.02 point target and the positive lower bound do not pass, and the all-panel lower bound fails its −0.01 non-regression guard. These intervals establish neither benefit nor harm nor equivalence, and they do not rule out a target-sized benefit.

Final hard and all-panel paired R minus U results have negative means and wide confidence intervals spanning zero. Uniform sampling has a higher mean feasible-hard score at all four observed checkpoints.
Points are paired seeds; diamonds and bars show means and two-sided 95% t intervals (df=2). Dashed lines mark the hard point target and all-panel lower-bound margin. Thin learning curves are seeds; thick curves are means. Training transitions are the cost axis. Curve data · PDF.

The design was underpowered for its own target. The observed seed standard deviation of 0.03192 gives a 95% half-width of 0.07929 at three pairs, far larger than the +0.02 the study set out to detect. Thousands of evaluation conditions do not create more independent policies. This is an observed precision limit, not a retrospective power guarantee and not evidence of equivalence.

No efficiency claim either. Exploratory normalized learning-curve area is lower for relative progress in all three seeds (mean −0.02898). Two of its seeds never reach their paired uniform final score on the observed grid.

Exposure moved; that is not the same as tracking learning. Relative progress and conditional failure both changed allocation (mean total variation 0.0825–0.0859), while the absolute-progress arm stayed near uniform at 0.0293–0.0298. But a learning-free reference for the same functional form yields total variation 0.0605 at 1,184 units and exceeds the 0.05 separation gate in every one of 2,000 replicates. Clearing that gate records that allocation moved, not that it followed learning progress.

05 · What else we learned

Each result narrows what the next experiment must prove.

06
Measured · simulator interface

Admissible support can be enforced throughout a trial.

The interface exposes 1,184 admitted motion units and 368,951 legal starts from 800 training motions, each bound by SHA-256. A 50-step trial cannot wrap or leave its admitted interval. All twelve confirmation runs replay with zero invalid-start, invalid-reference or censored-reset events.

How we use this: hold support fixed while comparing samplers. Model admissibility under a declared screen is not a guarantee of physical feasibility on a robot.

07
Measured · exploratory, one fixed policy

Repairing a reference did not make it better training data.

Contact-manifold repair qualified 22 of 26 flagged candidates under residual, displacement, joint-limit and contact-IK gates, leaving four byte-identical controls untouched. But a fixed-policy comparison gave repaired-minus-raw TrackingScore −0.0010, interval [−0.0086, +0.0080]. On the 22 repaired clips alone it is −0.0015; the four unchanged controls return +0.0020, which bounds evaluator reproducibility at twice the headline magnitude.

How we use this: screen qualification and policy utility are different questions. Repair changes the tracking target, so it needs its own paired evaluation before being treated as an improvement.

08
Pending · no result

The admission test trained completely and then stopped before measuring anything.

A five-paired-seed experiment comparing the same allocator with admission on and off finished all ten full-budget training runs, each passing its gate. It then failed on its first evaluation cell: the evaluation path omitted an adapter that reconciles two provenance keys in the stored condition manifest. Every scientific parameter, all 2,800 conditions and all 100 motion records matched exactly.

How we use this: the failure occurred before any checkpoint was loaded or any result written, so no outcome has been observed. The seal makes a stopped campaign terminal by design, so finishing it requires a contract revision rather than a restart. Until then the paper carries no admission claim.

06 · From simulation to a real robot

What this is, and what it is not.

CLIMB is a data-and-training-interface contribution with a bank-scale measurement result. It is not a controller. Everything above is simulation, and the honest distance to a physical Unitree G1 is two or three programmes, not a port.

No hardware experiment has been run in this project. No physical G1, no tethered bench test, no sim-to-real transfer. The stages below are the route, not a record of progress along it.

Five gaps between this model and that robot

AxisWhat the simulator is configured to doWhat is unmeasured
ActuationHip-roll and knee effort clamped at ±139 N·m with force limiting on; joint damping, friction loss and backlash are all zero, and the PD gains are synthesized from a reflected-inertia calculation rather than measured.Our configuration audit records vendor-published knee maxima of 90 N·m and 120 N·m for two G1 revisions, which would make the compiled limit 1.54× and 1.16× those figures. We archive no snapshot of that source and cannot confirm the numbers describe the same point in the drivetrain, so we treat this as a recorded discrepancy rather than a measurement. There is no torque–speed curve and no power or thermal limit. The screen compiles the same ranges as training, so we re-screened the 900-clip bank behind our exact support at 120 and 90 N·m. It barely moved: 888 clips unchanged, flagged count 99 → 98 of 900, no clip newly flagged. On this bank the screen's decisions are dominated by the unsupported-wrench residual, not by the actuator channel.
LatencyZero command delay on all six actuator groups; 5 ms physics, 20 ms control.Command delay, motor response lag and observation latency are three separate perturbations. None is modeled.
ContactRigid capsule feet with sliding friction 0.6, sampled over 0.3–1.2 at startup. Every non-foot collision geometry — shins, thighs, torso, hands — is frictionless.Configured friction is not measured contact fidelity, and a robot that kneels or crawls makes exactly the non-foot contacts the model treats as frictionless.
RandomizationExactly four events: torso centre-of-mass offset, encoder bias, foot friction, and a velocity push every one to three seconds. Three are retained at evaluation; the push is removed.Nothing randomizes mass, inertia, actuator strength, control gains or latency. The robustness a deployment would rely on is largely untrained.
TerrainA flat, level, perfectly rigid infinite plane.Uneven terrain is untested, and terrain-bearing references would need their terrain present in the scene before the screen could judge them.
ObservationThe policy reads an exact pelvis linear velocity from a simulated velocimeter, and the reference anchor pose expressed in the robot’s own torso frame.A physical G1 measures neither. Base velocity must be inferred by an estimator fusing inertial and kinematic data, and the anchor pose presupposes the robot knows where it is relative to the reference. On hardware this adds state estimation, reference synchronization and re-synchronization after disturbance, on-robot compute at the control rate, and a safety envelope. None of it is built here.

The screen itself declares electrical, thermal, compliance, latency and structural limits out of scope. Its verdict is admissible under this declared model and scene, which is a useful and checkable statement, and a strictly weaker one than physical feasibility.

Execution is short and segmented

A training trial is capped at one second of simulated time — 50 control steps at 50 Hz — though failure terminations end many trials sooner, and a held-out evaluation window is three seconds. The longest genuinely continuous execution ever demonstrated is 8.58 seconds, on a development checkpoint that is not one of the sealed comparison seeds, over two fully admitted training references with no active resets. All long-horizon evidence in the project totals eight attempts on those two references. That is a pipeline check, not a comparison and not a generalization result.

A deployment needs more than longer windows. An attempt currently begins by writing the reference state directly into the simulator, so there is no entry controller, no approach and no settling phase; a world is retired at the first termination, so there is no exit controller and no recovery; and references are never joined across a rejected transition, so there is no transition policy. Each of those is a component that does not exist yet.

The ladder

Ordered by what can invalidate what, not by what is exciting. The first two rungs cost no GPU time at all, and either could change what the paper claims.

  1. RUNG 1
    Desk work

    Audit the observation interface, before anything else.

    The upstream hardware-oriented configuration of this same G1 tracking task deliberately removes two observation terms — base linear velocity, and the reference anchor pose in the robot's torso frame — because a physical robot cannot measure either directly. Our configuration never sets that flag, so our policies consume both. Every trained policy in this project therefore depends on information a real G1 would have to infer from a state estimator. This rung enumerates every observation term, decides for each whether it is reconstructible on-robot, and then specifies the estimator, the reference synchronization, the control-rate compute and the safety envelope. It can invalidate every rung below it, and it costs desk work rather than compute.

  2. RUNG 2 · DONE
    Desk work

    Re-screen the bank under tighter actuator limits.

    Run on 8 September; the result is in the screening section above. Tightening the knee and hip-roll limit from ±139 to 120 and 90 N·m leaves 888 of 900 clips completely unchanged and moves the flagged count from 99 to 98. The screen is far more stable on this axis than we expected. It also refuted the prediction we registered before running: we argued the flagged fraction could only rise, and it fell, because the reported quantity is the translational component of a residual whose total the program minimizes, so tightening redistributes slack between components. The lower-bound claim that argument supported has been withdrawn from the paper and from this page. What remains on this rung is establishing provenance for the vendor figures, which our audit records without an archived source, and re-running the sweep on the full 10,705-clip corpus, which is not held locally.

  3. RUNG 3
    Simulation

    Finish the admission test.

    Does excluding model-inadmissible practice protect learning under a common allocator? Ten training runs are already complete and gate-passing; only the evaluation is missing, and finishing it needs a revised contract rather than a restart. Roughly one to six GPU-hours. This sits on the training-value line rather than the transfer path, so the rungs below it do not wait on the answer.

  4. RUNG 4
    Simulation

    Measure sensitivity to the assumptions we know are wrong.

    Re-evaluate frozen policies under declared perturbations: command delays of 5, 10 and 20 ms, knee caps at the two published figures, and foot friction at 0.3 and 1.2. The design exists and passes its CPU checks, and a corrected GPU queue did complete all nine development cells — but it then failed an aggregate exact output-parity check against the unchanged condition, and that cause is unresolved. The binding constraint is that failure, not compute. Until parity is restored the study reports nothing rather than something approximate.

  5. RUNG 5
    Simulation

    Score whole references instead of windows.

    Nothing on a robot is scored in one-second training trials or three-second windows. This rung runs whole admitted references with one initialization, natural failure and no re-initialization, which is the shape a hardware session actually has. Sequences must be chosen from reference kinematics and screen support alone, never by which policy succeeds, and admitted intervals must never be joined across a rejected transition. It is currently blocked on missing frame-level screen provenance for the held-out panel, not on compute.

  6. RUNG 6
    Hardware

    Tethered bench, then system identification.

    With the robot held or gantried, first confirm that the exported policy produces the same actions from the same observations as the simulator. Then measure what the model currently assumes: actual command delay, the torque–speed envelope, saturation and thermal behaviour. Those measurements are what let rung 2 be re-run against measured limits instead of recorded ones. Any uncommanded actuator behaviour stops the rung.

  7. RUNG 7
    Hardware

    Test the screen, not the curriculum.

    The defensible physical claim this project can reach is about its instrument: on a real G1, under one policy and one protocol, do references the screen flags fail earlier or more often than matched unflagged references? That tests the measurement line, which is the project's positive result. It does not presuppose that admission or allocation helps training, and it does not require the allocation comparison to have succeeded.

The allocation claim is not reachable on hardware, and it is worth saying why. The independent unit for any allocation result is the training seed pair, and robot time does not manufacture training seeds. At the seed spread we observed, resolving the +0.02 effect the study set out to detect would take roughly a dozen paired seeds in simulation before a hardware version of the question would even be meaningful. A hardware demonstration of a policy whose simulation comparison is inconclusive would produce a video, not a result.

Several of these rungs should be expected to return nulls, and this project's record says so plainly: four consecutive allocation arms failed their manipulation or implementation gates before two passed, a sealed predictive gate failed on valid data and its follow-on study was cancelled outright, and the headline comparison came back inconclusive. Each was preserved rather than rescued. That is the standard these rungs are held to.

07 · Inspect the evidence

Measured results, preserved decisions and open questions.

Research ledger

Every paper-bound number and its artifact path, the standing red-team audit, and the dated record of what stopped and why.

Results ledger · Red-team audit · Status

Evidence labels: measured = observed in the stated experiment; exploratory = limited or secondary inference; sealed = preserved protocol or decision; pending = evidence not yet available. Public exports contain aggregate measurements and source identities, not trained weights or licensed motion payloads. Earlier manuscript drafts document the project's history, are not synchronized with this page, and use absolute wording such as “impossible” that the current manuscript has replaced with model-relative language. Each carries a banner saying so: flagship draft · companion draft · 5 September archive.