Historical research record. This page preserves an earlier study or working draft; its status statements are dated. It also uses absolute wording such as “impossible” and “dynamically infeasible” that the current manuscript has replaced with model-relative language: the screen’s verdict is that no admissible contact supplies the demanded wrench under the declared robot and scene, which is checkable and strictly weaker than a claim about physics. Read the latest findings and research plan · Current confirmation progress.
CLIMB · Robotixx · project page · updated 2026-09-05

CLIMB: feasibility-gated motion tracking for generalist humanoid controllers.

CLIMB is a method for generalist humanoid tracking that screens the final robot-space motion, routes admissible, repairable, missing-context, and quarantined intervals, and applies learning progress only inside exact feasible support. Its motivating finding is reference–physics misalignment: persistent policy error can describe either useful control difficulty or a reference the declared robot/scene model cannot realize.

87–89 %
campaign peak top-1 allocation; the same unsupported attractor becomes dominant in 3/3 seeds, but is not the peak holder in every seed
~329 N
≈ full body weight unsupported for 1 s in that clip's retargeted descent: no contact within 6 cm
22.8 %
of one 10,705-clip AMASS→G1 corpus/pipeline is dynamically infeasible; a second production stack measures 0.14 %
22 / 26
flagged candidates pass every exact repair gate in a stratified CPU panel; 4/4 feasible controls remain byte-identical
0.567 → 0.609
cross-policy difficulty transfer once feasibility features join the atlas (perm p = 0.01)
Latest research · 5 September 2026

Relative progress sustains allocation in two full runs measured · exploratory

The new sampler passes its allocation gate at mean TV 0.0832 and 0.0827. Held-out tracking benefit remains pending. The fixed failure baseline has passed its first calibration seed; the second is running at this dated snapshot.

Why this direction: E4’s seed-1 mean TV fell to 0.029659, below its frozen 0.05 gate; policy endpoints stayed closed and seeds 2–3 stopped. Normalizing the same progress signal by its weighted mean sustains the intended change in practice. The next test compares uniform, absolute progress, relative progress, and conditional failure on fresh paired seeds.

Read the latest results, method, fixed experiment plan and evidence →

00 · The CLIMB method

Screen, repair, and allocate only on verified support

I · Screen with refeas

Contact-free inverse dynamics plus a torque-limited contact LP tests whether the declared robot and scene can supply each reference's demanded wrench.

bank-scale · cross-implementation

II · Repair with DFRP

A contact-manifold projection restores eligible support geometry, then residual, distortion, IK, integrity, and exact-start gates re-qualify the complete trajectory.

22/26 stratified candidates · 4/4 no-op controls

III · Gate exact-support ALP

A binary feasibility mask assigns rejected intervals zero mass; capped absolute-learning-progress allocation acts only over hash-bound, non-wrapping legal starts.

1,184 units · 368,951 starts
CLIMB pipeline from uncurated robot-space references through refeas screening, DFRP repair, and exact-support allocation
The CLIMB data-to-policy framework. Difficulty is separated into model-relative feasibility, bank-relative support, and intrinsic demand before policy outcomes affect allocation. refeas screens, DFRP repairs and re-qualifies eligible contacts, and exact-support ALP allocates only over legal starts. The 22/26 result is a frozen stratified repair-panel qualification rate; the 1,184 units and 368,951 starts are support properties, not policy-performance claims.
01 · The interfaces

One failure signal currently carries four different meanings

Generalist trackers retarget large motion banks, adapt exposure from policy outcomes, and summarize performance over sampled starts. CLIMB makes four contracts explicit so reference admissibility is decided before a curriculum interprets policy error:

Reference admission

Joint limits, velocities, and smoothness do not reveal a missing support source. refeas evaluates the final robot-space wrench against modeled contacts and actuator limits.

feasibility · robot + scene relative

Support construction

Whole-clip pruning discards legal maneuvers embedded beside flawed frames. Exact segmentation preserves every full-horizon-safe start; DFRP can recover eligible contact geometry.

support · bank relative

Compute allocation

Failure, completion, and learning progress rank admitted trials; they never override the admission decision. Hard zero mass and unit/clip caps make that separation executable.

intrinsic demand + policy progress

Evaluation conditioning

Paired starts and liveness-weighted scores prevent late offsets or early termination from making an unsupported prefix look partially learnable.

2,800 paired held-out conditions
02 · Motivating evidence

Measurements that define the three CLIMB interfaces

Persistent errors can become allocation attractors

Three seeds of failure-adaptive training under-perform plain uniform (held-out survival 0.780 vs 0.810; per-seed Δ = +0.030/+0.028/+0.030) while pouring 87–89 % of exposure onto a single kneel-and-crawl clip. Deriving the sampler's math exposes the non-floor; both upstream repos notified.

sealed-confirmatory · Exp-1 · reports/campaign_summary_3arm.json

A true mixture bounds the legacy failure sampler

Normalise-then-mix (a true 10 % floor on the simplex) keeps the same failure signal but bounds what it can spend: entropy 0.61 vs 0.39, and it beats the broken sampler by +0.045 endpoint (3/3 seeds) while matching uniform on the sealed primary.

sealed-confirmatory · Exp-2 "Branch B" · plan/BRANCH_DECISION.md

refeas localizes reference--physics misalignment

Contact-free inverse dynamics + a torque-limited contact LP: for a full second of the retargeted descent, no collision geometry is within 6 cm of the floor while the pelvis falls 0.35 m — ~329 N (the robot weighs 327 N) unsupported. The kneel itself is feasible; the bank just never contains it (3.2 % of training duration). The human sat back on their heels; the robot's legs can't fold that far; the retargeter lifted the legs instead of lowering the root.

measurement · N1 · reports/N1_clip44_knee_id.json

Feasibility prevalence is corpus-and-pipeline specific

Screening all 10,705 clips (~1 CPU-second each): 22.8 % exceed 10 % infeasible frames — 39 % of ground-contact motions, 59 % of dynamic ones — and the rate spans 0.1 % to 100 % across source datasets under one retargeter. A second production corpus/pipeline measures 0.14 %, so 22.8 % is not a generic retargeted-bank rate. 29 of our own 100 held-out evaluation clips are contaminated.

measurement · reports/feasibility_all/prevalence_report.txt

DFRP restores exact support without silently admitting failures

On a frozen, source-diverse CPU panel, DFRP qualifies 22/26 flagged candidates and leaves 4/4 feasible controls byte-identical. Two residual-dynamics and two contact-IK failures remain quarantined.

measured implementation · plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.md
Stick-figure frames of the kneel descent floating above the ground plane, with the unsupported-force trace pegged at body weight during the airborne window
The impossible descent. Retargeted reference frames (top; red panels = no contact available) and the torque-limited unsupported force (bottom): the descent (0.75–1.75 s) and the rise (8.0–8.5 s) demand ≈ body weight with nothing to push on. Generated by tools/n1_knee_id.py; every policy dies exactly here, at every start offset, with ~0 actuator saturation.
03 · The claim

Difficulty = feasibility × support × intrinsic

Each factor has its own measurement and intervention. The product is a staged decision rule, not an independence assumption: infeasible references are routed first; support and intrinsic demand rank only admitted trials.

Where a tracking policy fails on a clip: difficulty = feasibility × support × intrinsic route on feasibility first; then measure bank support and intrinsic demand inside the admitted set Feasibility can any controller do this on this robot? measure contact-free inverse dynamics + torque-limited contact LP → unsupported wrench, airborne fraction (≈1 s/clip) fix repair the transition (project onto contact) or exclude; upstream note to the retargeting pipeline #44: descent 0.75–1.75 s airborne, 329 N unsupported Support has the training bank anything like it? measure kNN distance and duration-weighted density in atlas space, relative to the bank; category mass fix composition: add feasible neighbours (N3), then coverage-grounded sampling then preserve deployment-prior coverage #44: kneel/crawl = 3.2 % of the bank; 8th-lowest support Intrinsic how hard is it once it is feasible and seen? measure reference atlas: kinematics, contact switching, required μ, GRF, support margin; calibrated fragility (N5) fix curriculum on a capped simplex; robustness training on the resolved mechanisms (delay, motor strength) #44: no ±δ changes survival; ρ atlas→difficulty 0.57–0.83 admission, corpus support, and learnable demand become separate decisions with separate evidence
The decomposition. Feasibility is measured by the contact LP and fixed by repair/exclusion (and an upstream retargeting fix); support by bank-relative kNN density and fixed by composition; intrinsic difficulty by the reference atlas and addressed by curricula and robustness training.

Measured transfer boundary: adding feasibility features raises cross-policy difficulty-transfer Spearman correlation from 0.567 to 0.609 on 100 held-out clips (permutation p = 0.010). All 6/6 directed policy pairs move positively and 4/6 clear the random-feature baseline; the policies share an architecture, so this does not establish cross-architecture or cross-robot transfer.

04 · The evidence

Core findings and their status

Historical allocation concentration sealed ✓

Adaptive < uniform 3/3 seeds; same attractor 6/6 adaptive+grounded runs; ε/(Σq+ε) derivation filed upstream.

campaign_summary_3arm.json · A5 · A7 · #1153 · #73

Grounded repair sealed ✓

Rescues adaptivity from collapse (AULC +0.055); matches uniform at 100 clips (Branch B, honestly adjudicated on the sealed primary); its edge lives entirely on feasible clips (+0.025 vs −0.009).

BRANCH_DECISION.md · N_atlas_v21.json

Physics-fragility gate sealed ✗

The pre-registered explanation failed its own criteria — and we kept it. The negative is what forced the feasibility screen.

PREREGISTRATION_G1_clip44.md 41e4b20c · G1_RESULT.md

Airborne-reference verdict measured ✓

~329 N unsupported, 86 % of descent frames; kneel phase fully supportable even with frictionless knees; control clip supported at every frame.

N1_clip44_knee_id.json · N1_CMU76_knee_id.json

Prevalence at scale measured ✓

22.8 % of the AMASS→G1 corpus/pipeline flagged versus 0.14 % in a second production stack; eval set 29/100 contaminated → sealed stratified-endpoint policy.

prevalence_report.txt · GLOBAL_EVAL_ADDENDUM a93a87a0

Transfer lift sealed ✓

Feasibility features lift cross-policy difficulty prediction beyond a random-feature baseline on 4/6 pairs (p = 0.010–0.045); support alone did not (sealed null, kept).

N_atlas_v21.json · N2_atlas_support.json

Calibrated instrument method

Paired |Δφ| is chaos-dominated (identical physics differs 2.5–8 mm in seconds). Signed replicate-means over a published identical-physics floor resolve 6–14 mm effects; replication across IC seeds r = 0.92, 6/6 sign agreement above 5 mm.

g1_v2_summary.json · N5_RESULT.md

Sign-reversal signature sealed ✗

P-SIGN failed: family 7/12, clean controls 4/12, airborne localisation 2/7. The clip-44 reversal remains exploratory, but the rollout-only infeasibility detector claim is rejected.

P_SIGN_RESULT.md · p_sign_summary.json

Conformance harness closed ✓

A validated dual-stack, same-engine instrument (MuJoCo Warp 3.11.0 via mjlab v1.6.0 and via Newton 7bb6d02d; classic MuJoCo 3.11.0 C as third referee): per-substep |Δq̇| ≤ 3×10⁻⁵ after eliminating four silent integration errors — with the elimination protocol (paired substeps, third referee, shadow solver) as reusable method.

S1_RESULT.md · S1_*_absorb.json
Per-axis fragility over time for the impossible clip versus a matched easy clip, with the same-solver noise floor
Why physics didn't explain it. Paired ±δ divergence per intervention axis on the attractor (left) vs its matched-easy control (right), against the same-solver floor (dotted). Every axis sits at the floor until the airborne descent; only the motor axis rises before the fall — the sign-reversal lead now under sealed test.
bank categoryclips>10 % infeasible frames>25 %duration share flagged
dynamic80458.6 %32.6 %54.5 %
ground-contact17539.4 %18.3 %44.1 %
locomotion5,59124.5 %17.4 %31.2 %
quiet4,13512.9 %7.5 %19.8 %
all (AMASS → G1)10,70522.8 %14.8 %27.4 %
Separate bank-scale feasibility rates and a same-clip cross-implementation agreement plot
Two questions, two valid denominators. The full-bank bars report separate corpus/pipeline measurements under the strict infeasible_frac > 0.10 rule; they are not a causal retargeter comparison. The 40-clip, flag-enriched panel instead checks the two implementations on identical inputs (39/40 decisions agree, ρ = 0.984, κ = 0.948) and is not a prevalence sample.

By source dataset: GRAB 0.1 % · TCD 1.6 % · KIT 17 % · BMLmovi 22 % · Eyes-Japan 23 % · BMLhandball 27 % · CMU 40 % · ACCAD 42 % · HUMAN4D 55 % · Transitions 90 % · CNRS 100 %. A three-orders-of-magnitude spread under one pipeline and one robot is a source-corpus × pipeline property, not a difficulty gradient.

05 · Ablation and analysis

Alternative routes reveal which contract each module must enforce

The hash-bound record retains every negative control, but the method claim rests on the interfaces they discriminate. Whole-clip filtering tests whether coarse removal preserves diversity; soft weighting tests whether a penalty actually removes rejected support; Newton tests an alternative predictive instrument.

sealhashpredictionoutcome
A2 · grounded arm37daa8a9grounded ≥ uniform, ≫ adaptive (AULC primary)✅ ≫ adaptive · ≈ uniform (Branch B); co-primary reported uninformative (registration defect documented)
G1 · fragility gate41e4b20ccontact/CoM fragility ≥ 2× controls, ≥ 5× floor❌ failed as sealed — kept, published
S1 · conformance verdict"contact-event fork = fragility finding"withdrawn: four integration errors; harness fixed to |Δq̇| ≤ 3e-5
N2 · support featurespre-statedresiduals→low-support ✚ transfer lift✅ residuals (ρ +0.60) · ❌ transfer (inside noise) — split kept
Atlas v2.1 · feasibility9b1a2c78within-bank lift · transfer lift · residual anatomy❌ F1 · ✅ F2 (0.567→0.609, p=.01) · half F3
P-TAX · reward tax7960057atax→difficulty beyond feasibility flag❌ null (0/3 arms) — hygiene finding only
D1 · eval policya93a87a0feasible-only primary endpoints, threshold provenance cited— policy, sealed before any E3 number
N3 · composition causalityaf1b7c9fattractor's feasible phase 0.00 → ≥ 0.25 (2/2 seeds); random-16 control < 0.10⚠ mixed — targeted endpoints pass (0.750/0.750; random 0), adaptive regression triggers stop; descent prediction misses
E-HYG · clip pruninga5494b7cfeasible heldout Δ ≥ +0.015; no excessive coverage cost❌ null — feasible heldout Δ −0.0101; zero-shot ground −0.0354 inside bracket
E3 v2 · support moderation2c38845bnamed clips predicted to get worse at 800 (all 22 dynamic held-out lose support)🕐 bidirectional & risky, post-N3
P-SIGN · signature generalityc7916e8c≥8/12 family clips ≥ +5 mm airborne; controls < 2 mm; 3× localisation❌ failed — 7/12 family, 4/12 controls, 2/7 localised; runtime-detector claim rejected
FGAS · soft eligibility3521c80efeasible-hard20 Δ ≥ +0.05 with lower CI > 0; heldout no-regression❌ Δ −0.0196; no-regression passes, but rejected-start mass 0.199 fails the implementation gate
N7 · contact repair90da8a08R/repaired − K/raw ≥ +0.05 with lower CI > 0; heldout and coverage guards❌ +0.0397, CI [+0.0153,+0.0658], below SESOI; coverage fails
06 · The tools

Released instruments, not just claims

refeas — the feasibility screen

Contact-free inverse dynamics + torque-limited contact LP. ~1 CPU-second per clip; MuJoCo + SciPy only; Apache-2.0; G1 worked example (a synthetic hover the screen flags at 45 % infeasible). Screen before training; ship per-clip flags with datasets.

github.com/linjiw/refeas · v0.1.0

Contact-projection repair

A root-translation plus contact-IK operator projects eligible candidates and then reruns residual, displacement, joint-limit, IK-residual, provenance, and exact-support gates. The stratified exact panel qualifies 22/26 repairs and keeps 4/4 feasible controls byte-identical; legacy bank-wide projections remain qualification-incomplete.

tools/repair_contact_projection.py · plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.md

Stratified-start evaluation

The sealed fixed-offset harness exposed the "0.31 survival" artifact. A post-outcome audit found unpaired startup randomization, clipped duplicate offsets, and reset-contaminated terminal quality metrics; a paired per-episode v2 is now required.

tools/eval_stratified.py · plan/SEGMENT_NATIVE_FOLLOWUP_2026-08-20.md

Dual-engine conformance protocol

Per-substep paired stepping, an independent third engine as referee, and a shadow solver stepping the other engine's own trajectory. Any second physics implementation coupled to an RL harness deserves this — end-of-episode metrics matched while the harness was broken.

tools/s1_newton_conformance.py

ICRA manuscript + upstream notes

The ICRA-sized manuscript carries the three-interface framework, bank-scale screen, DFRP qualification, and exact-support allocator; two upstream drafts document retargeting and dataset implications without making them the paper's main contribution.

paper/companion/ · reports/upstream_drafts/
07 · Design boundaries

What the current evidence establishes—and what it does not

The screen, routing contract, and exact-support gate are implemented. Whether ALP adds policy value inside the gate remains the current controlled question.

FGAS — soft eligibility sealed ✗

The three-seed soft arm gives Δ −0.0196 on feasible-hard20 and fails its rejected-start-mass gate (0.199). Diagnostics confirm that failure weighting overwhelms a clip-mean soft multiplier; a segment-native follow-up requires a new seal.

N7 — repair the impossible sealed ✗

R/repaired improves over K/raw by +0.0397, but misses the +0.05 SESOI and the coverage rule. R/raw is −0.0036: the gain is reference easing plus co-adaptation, concentrated in over-budget edits, not better raw-reference policy skill.

E3 — support at scale post-Sept 15

800-clip bank with bidirectional named predictions — including 22 clips we claim will get worse. Predicted harms are the sharpest test the support story can face.

Newton predictive gate sealed ✗

Newton v1.5 conformance passed, but the valid 40-unit no-training gate did not: adaptive partial ρ is +0.141 (p = 0.158) and LOCO lift is −0.006. Newton remains an instrument; G3 is killed.

Honest scope

One robot, one reward configuration, survival-centric endpoints, simulation only, 4k-iteration horizon. The solver-ensemble program was explicitly descoped: the second engine earned its keep as a referee and measurement instrument, not an oracle.

08 · E4 exact-support tracking

Exact-support curriculum: calibrated allocation and paired tracking measured + exploratory

The GPU lifecycle, bounded curriculum, and paired evaluator pass. One seed shows a motion-quality signal but no established survival benefit. All 900 licensed payload identities pass, calibration and independent validation are complete, but the completed sealed seed-1 confirmation failed its allocation gate. Its decision is not_tested; policy endpoints remain closed and seeds 2–3 stopped. The newer relative-progress studies are summarized above.

Exact-support lifecycle validation

An 8-env lifecycle trace completed 24/24 exact trials with no wrap. Two 512-env arms then ran 200 iterations each (2.46 M steps/arm). A diagnosed early-failure incentive was corrected symmetrically with a failure-only event cost; invalid starts, invalid frames, and censored resets stayed at zero.

reports/segment_v2_smoke/timeline_trace.json · autoresearch/segment-native-260820-2259/research_log.md

Paired survival is inconclusive

On 504 identical worlds per policy, adaptive minus exact-uniform success is +0.0079 (unit-bootstrap 95 % CI −0.0536 to +0.0714); survival is +0.0115 s (−0.0100 to +0.0346). Seeds, starts, initial state, and startup randomization hashes match.

reports/segment_v2_pilot/result.json

Motion-quality signal

Across 22,321 paired common-survivor frames, adaptive training reduces body-position error by 4.20 mm (95 % CI 1.90–6.63 mm) and anchor-orientation error by 0.0280 rad (0.0163–0.0397). Anchor position, joint error, and work remain unresolved; this is one-seed exploratory evidence.

tools/analyze_segment_pilot.py · reports/segment_v2_pilot/

E4: calibrated allocation

The adaptive distribution ended only 1.40 % total-variation from its control (correlation 0.998), despite healthy entropy. The subsequent 12-candidate ALP calibration selected ρ = 0.40 and λ = 0.05. Independent validation measured mean TV 0.1056, at least 700.1 effective units, maximum unit mass 0.0134, and zero invalid/censored events over its 50-iteration run. E4 compared Exact ALP and Exact Uniform over 1,184 units / 368,951 legal starts. Its short calibration did not predict sustained long-run contrast: the completed seed-1 allocation gate failed.

reports/g_segment/calibration/result.json · measured calibration

E4 confirmation sealed · not_tested

Exact ALP and Exact Uniform share the same feasible support, 512 environments, and 4,000 iterations. The three-seed protocol checks sampler manipulation before measuring TrackingScore, survival, and common-survivor fidelity. Both seed-1 arms completed training. Mean post-warm-up adaptive TV was 0.029659, below the required 0.05. The frozen gate stopped continuation to seeds 2–3; no E4 policy-null claim follows.

plan/G_SEGMENT_FREEZE.sha256 · plan/E4_CONTINUATION_2026-09-05.md

Updated disposition, 5 September 2026: the E4 allocation decision is not_tested. The separate fixed-policy DFRP comparison is complete: 26 clips, 656 paired conditions per reference arm, repaired minus raw TrackingScore −0.001003 (clip-bootstrap 95% CI −0.008588 to +0.008020). No aggregate tracking gain was demonstrated. This does not change the separate 22/26 physics-screen qualification result. See the latest research update for the full interpretation, downloadable evidence and next experiment.

Full implementation note: see the segment-native training and next-direction page for the lifecycle, training telemetry, paired evaluation, limits, and next causal design.

09 · DFRP v1 exact CPU gate

Repair candidates are not training data until provenance and exact support agree measured implementation

DFRP v1 connects strict repair entry, two-tier routing, source-motion-bound exact support, and segment-native starts in one fail-closed contract.

Strict routing audit

The strict >10% rule flags 2,442 clips. Legacy root projection puts 644 (26.4%) inside the new 8 cm tier and another 962 in the exploratory 8–15 cm tier.

reports/dfrp_v0/census/summary.json

Frozen exact panel

A source-diverse CPU panel admits 22/26 flagged repairs and all four byte-identical controls. Four failures remain excluded: two residual-dynamics and two IK-qualification cases.

84.6% panel result · not a census estimate

Nothing leaks into training

Exact sidecars must match the selected motion hash and partition every frame. Missing or mismatched evidence fails closed; legacy root-only repairs are still not promoted.

dfrp_bank_manifest/1 · exact source binding

Exact MJLab handoff

The curated 26-clip view yields 36 units and 10,561 legal 50-step starts. Runtime verifies motion, sidecar, DFRP payload, and unit-table identities before sampling.

plan/DFRP_V1_EXACT_PANEL_RESULT_2026-08-21.md

Fixed-policy tracking validation queued

The final Exact Uniform seed-1 policy is preselected for raw-versus-repaired evaluation. The paired design includes all 22 qualified repairs and four controls; failures, short windows, and two training-overlap clips remain in the accounting.

Design and execution record · no policy result yet

Claim boundary: 22/26 is a stratified exact-panel implementation gate, not a bank-wide recovery rate. It establishes a trustworthy handoff artifact; policy and hardware benefits remain untested.

10 · Alternative instrument sanity check

Newton remains a validated referee, not a curriculum signal sealed ✗

Newton v1.5 passed same-state conformance, then the tested three-axis fragility vector failed its sealed no-training predictive gate on valid data. This bounds the alternative instrument's predictive use; E4 independently compares Exact ALP with Exact Uniform on identical feasible support.

Conformance passed measured ✓

On one easy and one contact-rich hash-bound unit, placement, first observation, first action, resynchronized state, contact timing, and deterministic repeats agree after seven live-model import residuals are mirrored.

plan/NEWTON15_RECERT_RESULT.md

Prediction failed sealed ✗

Across 40 valid units, the three-axis Newton vector misses both predeclared gates: partial ρ ≥ 0.25 and held-out-clip LOCO lift ≥ 0.05. The grounded replication also fails.

plan/NEWTON_PRED_RESULT.md · reports/newton15_pred/result.json

No Newton-weighted allocation

The registered decision was explicit: valid-data failure leaves Newton as an analysis instrument and forbids Newton-fragility-weighted training. A raw per-axis correlation does not override the joint gate.

adaptive: p = 0.158 · LOCO lift −0.006