CLIMB is a method for generalist humanoid tracking that screens the final robot-space motion, routes admissible, repairable, missing-context, and quarantined intervals, and applies learning progress only inside exact feasible support. Its motivating finding is reference–physics misalignment: persistent policy error can describe either useful control difficulty or a reference the declared robot/scene model cannot realize.
The new sampler passes its allocation gate at mean TV 0.0832 and 0.0827. Held-out tracking benefit remains pending. The fixed failure baseline has passed its first calibration seed; the second is running at this dated snapshot.
Why this direction: E4’s seed-1 mean TV fell to 0.029659, below its frozen 0.05 gate; policy endpoints stayed closed and seeds 2–3 stopped. Normalizing the same progress signal by its weighted mean sustains the intended change in practice. The next test compares uniform, absolute progress, relative progress, and conditional failure on fresh paired seeds.
Read the latest results, method, fixed experiment plan and evidence →
refeasContact-free inverse dynamics plus a torque-limited contact LP tests whether the declared robot and scene can supply each reference's demanded wrench.
bank-scale · cross-implementationA contact-manifold projection restores eligible support geometry, then residual, distortion, IK, integrity, and exact-start gates re-qualify the complete trajectory.
22/26 stratified candidates · 4/4 no-op controlsA binary feasibility mask assigns rejected intervals zero mass; capped absolute-learning-progress allocation acts only over hash-bound, non-wrapping legal starts.
1,184 units · 368,951 starts
refeas screens, DFRP repairs and re-qualifies eligible contacts, and exact-support ALP allocates only over legal starts. The 22/26 result is a frozen stratified repair-panel qualification rate; the 1,184 units and 368,951 starts are support properties, not policy-performance claims.Generalist trackers retarget large motion banks, adapt exposure from policy outcomes, and summarize performance over sampled starts. CLIMB makes four contracts explicit so reference admissibility is decided before a curriculum interprets policy error:
Joint limits, velocities, and smoothness do not reveal a missing support source. refeas evaluates the final robot-space wrench against modeled contacts and actuator limits.
Whole-clip pruning discards legal maneuvers embedded beside flawed frames. Exact segmentation preserves every full-horizon-safe start; DFRP can recover eligible contact geometry.
support · bank relativeFailure, completion, and learning progress rank admitted trials; they never override the admission decision. Hard zero mass and unit/clip caps make that separation executable.
intrinsic demand + policy progressPaired starts and liveness-weighted scores prevent late offsets or early termination from making an unsupported prefix look partially learnable.
2,800 paired held-out conditionsThree seeds of failure-adaptive training under-perform plain uniform (held-out survival 0.780 vs 0.810; per-seed Δ = +0.030/+0.028/+0.030) while pouring 87–89 % of exposure onto a single kneel-and-crawl clip. Deriving the sampler's math exposes the non-floor; both upstream repos notified.
sealed-confirmatory · Exp-1 · reports/campaign_summary_3arm.jsonNormalise-then-mix (a true 10 % floor on the simplex) keeps the same failure signal but bounds what it can spend: entropy 0.61 vs 0.39, and it beats the broken sampler by +0.045 endpoint (3/3 seeds) while matching uniform on the sealed primary.
sealed-confirmatory · Exp-2 "Branch B" · plan/BRANCH_DECISION.mdrefeas localizes reference--physics misalignmentContact-free inverse dynamics + a torque-limited contact LP: for a full second of the retargeted descent, no collision geometry is within 6 cm of the floor while the pelvis falls 0.35 m — ~329 N (the robot weighs 327 N) unsupported. The kneel itself is feasible; the bank just never contains it (3.2 % of training duration). The human sat back on their heels; the robot's legs can't fold that far; the retargeter lifted the legs instead of lowering the root.
measurement · N1 · reports/N1_clip44_knee_id.jsonScreening all 10,705 clips (~1 CPU-second each): 22.8 % exceed 10 % infeasible frames — 39 % of ground-contact motions, 59 % of dynamic ones — and the rate spans 0.1 % to 100 % across source datasets under one retargeter. A second production corpus/pipeline measures 0.14 %, so 22.8 % is not a generic retargeted-bank rate. 29 of our own 100 held-out evaluation clips are contaminated.
measurement · reports/feasibility_all/prevalence_report.txtOn a frozen, source-diverse CPU panel, DFRP qualifies 22/26 flagged candidates and leaves 4/4 feasible controls byte-identical. Two residual-dynamics and two contact-IK failures remain quarantined.
measured implementation · plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.mdtools/n1_knee_id.py; every policy dies exactly here, at every start offset,
with ~0 actuator saturation.Each factor has its own measurement and intervention. The product is a staged decision rule, not an independence assumption: infeasible references are routed first; support and intrinsic demand rank only admitted trials.
Measured transfer boundary: adding feasibility features raises cross-policy difficulty-transfer Spearman correlation from 0.567 to 0.609 on 100 held-out clips (permutation p = 0.010). All 6/6 directed policy pairs move positively and 4/6 clear the random-feature baseline; the policies share an architecture, so this does not establish cross-architecture or cross-robot transfer.
Adaptive < uniform 3/3 seeds; same attractor 6/6 adaptive+grounded runs; ε/(Σq+ε) derivation filed upstream.
campaign_summary_3arm.json · A5 · A7 · #1153 · #73Rescues adaptivity from collapse (AULC +0.055); matches uniform at 100 clips (Branch B, honestly adjudicated on the sealed primary); its edge lives entirely on feasible clips (+0.025 vs −0.009).
BRANCH_DECISION.md · N_atlas_v21.jsonThe pre-registered explanation failed its own criteria — and we kept it. The negative is what forced the feasibility screen.
PREREGISTRATION_G1_clip44.md 41e4b20c · G1_RESULT.md~329 N unsupported, 86 % of descent frames; kneel phase fully supportable even with frictionless knees; control clip supported at every frame.
N1_clip44_knee_id.json · N1_CMU76_knee_id.json22.8 % of the AMASS→G1 corpus/pipeline flagged versus 0.14 % in a second production stack; eval set 29/100 contaminated → sealed stratified-endpoint policy.
prevalence_report.txt · GLOBAL_EVAL_ADDENDUM a93a87a0Feasibility features lift cross-policy difficulty prediction beyond a random-feature baseline on 4/6 pairs (p = 0.010–0.045); support alone did not (sealed null, kept).
N_atlas_v21.json · N2_atlas_support.jsonPaired |Δφ| is chaos-dominated (identical physics differs 2.5–8 mm in seconds). Signed replicate-means over a published identical-physics floor resolve 6–14 mm effects; replication across IC seeds r = 0.92, 6/6 sign agreement above 5 mm.
g1_v2_summary.json · N5_RESULT.mdP-SIGN failed: family 7/12, clean controls 4/12, airborne localisation 2/7. The clip-44 reversal remains exploratory, but the rollout-only infeasibility detector claim is rejected.
P_SIGN_RESULT.md · p_sign_summary.jsonA validated dual-stack, same-engine instrument (MuJoCo Warp 3.11.0 via mjlab v1.6.0 and via Newton 7bb6d02d; classic MuJoCo 3.11.0 C as third referee): per-substep |Δq̇| ≤ 3×10⁻⁵ after eliminating four silent integration errors — with the elimination protocol (paired substeps, third referee, shadow solver) as reusable method.
S1_RESULT.md · S1_*_absorb.json| bank category | clips | >10 % infeasible frames | >25 % | duration share flagged |
|---|---|---|---|---|
| dynamic | 804 | 58.6 % | 32.6 % | 54.5 % |
| ground-contact | 175 | 39.4 % | 18.3 % | 44.1 % |
| locomotion | 5,591 | 24.5 % | 17.4 % | 31.2 % |
| quiet | 4,135 | 12.9 % | 7.5 % | 19.8 % |
| all (AMASS → G1) | 10,705 | 22.8 % | 14.8 % | 27.4 % |
infeasible_frac > 0.10 rule; they are not a causal retargeter comparison. The 40-clip, flag-enriched panel instead checks the two implementations on identical inputs (39/40 decisions agree, ρ = 0.984, κ = 0.948) and is not a prevalence sample.By source dataset: GRAB 0.1 % · TCD 1.6 % · KIT 17 % · BMLmovi 22 % · Eyes-Japan 23 % · BMLhandball 27 % · CMU 40 % · ACCAD 42 % · HUMAN4D 55 % · Transitions 90 % · CNRS 100 %. A three-orders-of-magnitude spread under one pipeline and one robot is a source-corpus × pipeline property, not a difficulty gradient.
The hash-bound record retains every negative control, but the method claim rests on the interfaces they discriminate. Whole-clip filtering tests whether coarse removal preserves diversity; soft weighting tests whether a penalty actually removes rejected support; Newton tests an alternative predictive instrument.
| seal | hash | prediction | outcome |
|---|---|---|---|
| A2 · grounded arm | 37daa8a9 | grounded ≥ uniform, ≫ adaptive (AULC primary) | ✅ ≫ adaptive · ≈ uniform (Branch B); co-primary reported uninformative (registration defect documented) |
| G1 · fragility gate | 41e4b20c | contact/CoM fragility ≥ 2× controls, ≥ 5× floor | ❌ failed as sealed — kept, published |
| S1 · conformance verdict | — | "contact-event fork = fragility finding" | ⬅ withdrawn: four integration errors; harness fixed to |Δq̇| ≤ 3e-5 |
| N2 · support features | pre-stated | residuals→low-support ✚ transfer lift | ✅ residuals (ρ +0.60) · ❌ transfer (inside noise) — split kept |
| Atlas v2.1 · feasibility | 9b1a2c78 | within-bank lift · transfer lift · residual anatomy | ❌ F1 · ✅ F2 (0.567→0.609, p=.01) · half F3 |
| P-TAX · reward tax | 7960057a | tax→difficulty beyond feasibility flag | ❌ null (0/3 arms) — hygiene finding only |
| D1 · eval policy | a93a87a0 | feasible-only primary endpoints, threshold provenance cited | — policy, sealed before any E3 number |
| N3 · composition causality | af1b7c9f | attractor's feasible phase 0.00 → ≥ 0.25 (2/2 seeds); random-16 control < 0.10 | ⚠ mixed — targeted endpoints pass (0.750/0.750; random 0), adaptive regression triggers stop; descent prediction misses |
| E-HYG · clip pruning | a5494b7c | feasible heldout Δ ≥ +0.015; no excessive coverage cost | ❌ null — feasible heldout Δ −0.0101; zero-shot ground −0.0354 inside bracket |
| E3 v2 · support moderation | 2c38845b | named clips predicted to get worse at 800 (all 22 dynamic held-out lose support) | 🕐 bidirectional & risky, post-N3 |
| P-SIGN · signature generality | c7916e8c | ≥8/12 family clips ≥ +5 mm airborne; controls < 2 mm; 3× localisation | ❌ failed — 7/12 family, 4/12 controls, 2/7 localised; runtime-detector claim rejected |
| FGAS · soft eligibility | 3521c80e | feasible-hard20 Δ ≥ +0.05 with lower CI > 0; heldout no-regression | ❌ Δ −0.0196; no-regression passes, but rejected-start mass 0.199 fails the implementation gate |
| N7 · contact repair | 90da8a08 | R/repaired − K/raw ≥ +0.05 with lower CI > 0; heldout and coverage guards | ❌ +0.0397, CI [+0.0153,+0.0658], below SESOI; coverage fails |
Contact-free inverse dynamics + torque-limited contact LP. ~1 CPU-second per clip; MuJoCo + SciPy only; Apache-2.0; G1 worked example (a synthetic hover the screen flags at 45 % infeasible). Screen before training; ship per-clip flags with datasets.
github.com/linjiw/refeas · v0.1.0A root-translation plus contact-IK operator projects eligible candidates and then reruns residual, displacement, joint-limit, IK-residual, provenance, and exact-support gates. The stratified exact panel qualifies 22/26 repairs and keeps 4/4 feasible controls byte-identical; legacy bank-wide projections remain qualification-incomplete.
tools/repair_contact_projection.py · plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.mdThe sealed fixed-offset harness exposed the "0.31 survival" artifact. A post-outcome audit found unpaired startup randomization, clipped duplicate offsets, and reset-contaminated terminal quality metrics; a paired per-episode v2 is now required.
tools/eval_stratified.py · plan/SEGMENT_NATIVE_FOLLOWUP_2026-08-20.mdPer-substep paired stepping, an independent third engine as referee, and a shadow solver stepping the other engine's own trajectory. Any second physics implementation coupled to an RL harness deserves this — end-of-episode metrics matched while the harness was broken.
tools/s1_newton_conformance.pyThe ICRA-sized manuscript carries the three-interface framework, bank-scale screen, DFRP qualification, and exact-support allocator; two upstream drafts document retargeting and dataset implications without making them the paper's main contribution.
paper/companion/ · reports/upstream_drafts/The screen, routing contract, and exact-support gate are implemented. Whether ALP adds policy value inside the gate remains the current controlled question.
The three-seed soft arm gives Δ −0.0196 on feasible-hard20 and fails its rejected-start-mass gate (0.199). Diagnostics confirm that failure weighting overwhelms a clip-mean soft multiplier; a segment-native follow-up requires a new seal.
R/repaired improves over K/raw by +0.0397, but misses the +0.05 SESOI and the coverage rule. R/raw is −0.0036: the gain is reference easing plus co-adaptation, concentrated in over-budget edits, not better raw-reference policy skill.
800-clip bank with bidirectional named predictions — including 22 clips we claim will get worse. Predicted harms are the sharpest test the support story can face.
Newton v1.5 conformance passed, but the valid 40-unit no-training gate did not: adaptive partial ρ is +0.141 (p = 0.158) and LOCO lift is −0.006. Newton remains an instrument; G3 is killed.
One robot, one reward configuration, survival-centric endpoints, simulation only, 4k-iteration horizon. The solver-ensemble program was explicitly descoped: the second engine earned its keep as a referee and measurement instrument, not an oracle.
The GPU lifecycle, bounded curriculum, and paired evaluator pass. One seed shows a motion-quality
signal but no established survival benefit. All 900 licensed payload identities pass, calibration and independent validation are complete,
but the completed sealed seed-1 confirmation failed its allocation gate. Its decision is not_tested; policy endpoints remain closed and seeds 2–3 stopped. The newer relative-progress studies are summarized above.
An 8-env lifecycle trace completed 24/24 exact trials with no wrap. Two 512-env arms then ran 200 iterations each (2.46 M steps/arm). A diagnosed early-failure incentive was corrected symmetrically with a failure-only event cost; invalid starts, invalid frames, and censored resets stayed at zero.
reports/segment_v2_smoke/timeline_trace.json · autoresearch/segment-native-260820-2259/research_log.mdOn 504 identical worlds per policy, adaptive minus exact-uniform success is +0.0079 (unit-bootstrap 95 % CI −0.0536 to +0.0714); survival is +0.0115 s (−0.0100 to +0.0346). Seeds, starts, initial state, and startup randomization hashes match.
reports/segment_v2_pilot/result.jsonAcross 22,321 paired common-survivor frames, adaptive training reduces body-position error by 4.20 mm (95 % CI 1.90–6.63 mm) and anchor-orientation error by 0.0280 rad (0.0163–0.0397). Anchor position, joint error, and work remain unresolved; this is one-seed exploratory evidence.
tools/analyze_segment_pilot.py · reports/segment_v2_pilot/The adaptive distribution ended only 1.40 % total-variation from its control (correlation 0.998), despite healthy entropy. The subsequent 12-candidate ALP calibration selected ρ = 0.40 and λ = 0.05. Independent validation measured mean TV 0.1056, at least 700.1 effective units, maximum unit mass 0.0134, and zero invalid/censored events over its 50-iteration run. E4 compared Exact ALP and Exact Uniform over 1,184 units / 368,951 legal starts. Its short calibration did not predict sustained long-run contrast: the completed seed-1 allocation gate failed.
reports/g_segment/calibration/result.json · measured calibrationExact ALP and Exact Uniform share the same feasible support, 512 environments, and 4,000 iterations. The three-seed protocol checks sampler manipulation before measuring TrackingScore, survival, and common-survivor fidelity. Both seed-1 arms completed training. Mean post-warm-up adaptive TV was 0.029659, below the required 0.05. The frozen gate stopped continuation to seeds 2–3; no E4 policy-null claim follows.
plan/G_SEGMENT_FREEZE.sha256 · plan/E4_CONTINUATION_2026-09-05.mdUpdated disposition, 5 September 2026: the E4 allocation decision is not_tested. The separate fixed-policy DFRP comparison is complete: 26 clips, 656 paired conditions per reference arm, repaired minus raw TrackingScore −0.001003 (clip-bootstrap 95% CI −0.008588 to +0.008020). No aggregate tracking gain was demonstrated. This does not change the separate 22/26 physics-screen qualification result. See the latest research update for the full interpretation, downloadable evidence and next experiment.
Full implementation note: see the segment-native training and next-direction page for the lifecycle, training telemetry, paired evaluation, limits, and next causal design.
DFRP v1 connects strict repair entry, two-tier routing, source-motion-bound exact support, and segment-native starts in one fail-closed contract.
The strict >10% rule flags 2,442 clips. Legacy root projection puts 644 (26.4%) inside the new 8 cm tier and another 962 in the exploratory 8–15 cm tier.
reports/dfrp_v0/census/summary.jsonA source-diverse CPU panel admits 22/26 flagged repairs and all four byte-identical controls. Four failures remain excluded: two residual-dynamics and two IK-qualification cases.
84.6% panel result · not a census estimateExact sidecars must match the selected motion hash and partition every frame. Missing or mismatched evidence fails closed; legacy root-only repairs are still not promoted.
dfrp_bank_manifest/1 · exact source bindingThe curated 26-clip view yields 36 units and 10,561 legal 50-step starts. Runtime verifies motion, sidecar, DFRP payload, and unit-table identities before sampling.
plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.mdThe final Exact Uniform seed-1 policy is preselected for raw-versus-repaired evaluation. The paired design includes all 22 qualified repairs and four controls; failures, short windows, and two training-overlap clips remain in the accounting.
Design and execution record · no policy result yetClaim boundary: 22/26 is a stratified exact-panel implementation gate, not a bank-wide recovery rate. It establishes a trustworthy handoff artifact; policy and hardware benefits remain untested.
Newton v1.5 passed same-state conformance, then the tested three-axis fragility vector failed its sealed no-training predictive gate on valid data. This bounds the alternative instrument's predictive use; E4 independently compares Exact ALP with Exact Uniform on identical feasible support.
On one easy and one contact-rich hash-bound unit, placement, first observation, first action, resynchronized state, contact timing, and deterministic repeats agree after seven live-model import residuals are mirrored.
plan/NEWTON15_RECERT_RESULT.mdAcross 40 valid units, the three-axis Newton vector misses both predeclared gates: partial ρ ≥ 0.25 and held-out-clip LOCO lift ≥ 0.05. The grounded replication also fails.
plan/NEWTON_PRED_RESULT.md · reports/newton15_pred/result.jsonThe registered decision was explicit: valid-data failure leaves Newton as an analysis instrument and forbids Newton-fragility-weighted training. A raw per-axis correlation does not override the joint gate.
adaptive: p = 0.158 · LOCO lift −0.006