Historical research record. This page preserves an earlier study or working draft; its status statements are dated. It also uses absolute wording such as “impossible” and “dynamically infeasible” that the current manuscript has replaced with model-relative language: the screen’s verdict is that no admissible contact supplies the demanded wrench under the declared robot and scene, which is checkable and strictly weaker than a claim about physics. Read the latest findings and research plan · Current confirmation progress.
Segment-native v2 · exploratory result + next gate

Train on the hard motion—not on an impossible transition.

We turned offline physics screening into an exact motion-tracking lifecycle: feasible start segments, fixed-horizon trials, conditional outcome attribution, bounded sampling, and paired terminal-safe evaluation. The pipeline now works on GPU. The first policy result is promising on two quality channels and inconclusive on survival. Newton's later predictive gate failed, leaving exact-support learning progress as the next causal test.

4.92 M
total PPO transitions across two matched 512-environment arms
0
invalid starts, escaped reference frames, or censored resets
+0.79 pp
adaptive success delta; interval crosses zero
−4.20 mm
common-survivor body-position error delta; exploratory

Historical experiment record. For the completed E4 gate decision, repair comparison, two relative-progress replications and the next four-arm experiment, see the research update of 5 September 2026. The pilot measurements below retain their original scope.

01 · Claim boundary

What changed—and what has not been established

The engineering result and the policy result are different claims. Keeping them separate is the point of the audit.

Keep: a valid training experiment now exists

The command never samples a known-invalid start, never teleports through a continuing PPO transition, attributes each outcome to a stable segment unit, caps concentration, and records terminal metrics before reset.

Hold: no confirmed survival benefit yet

One seed improved success by 0.79 percentage points, but the unit-clustered interval spans −5.36 to +7.14. The treatment itself was only 1.40% total-variation from control.

02 · Method in the training loop

From a physics screen to a clean PPO transition

The method acts before and during training. Offline physics defines admissible support; online learning allocates exposure only inside that support.

Screen reference forces

Contact-free inverse dynamics and contact capacity identify unsupported frames. This is hygiene, not a policy score.

Build exact units

Contiguous feasible frame intervals become stable units. The table retains 42 units and 4,679 legal 50-step starts.

Sample a safe horizon

Every proposal satisfies the complete future window. Known-invalid probability is exactly zero.

Track and terminate

The reference advances from s+1 through s+50, then emits an explicit timeout before any wrap or teleport.

Credit and rebalance

Completed outcomes update stable unit statistics on a fixed clock. Exploration and 5%/25% unit/clip caps remain enforced.

Why exact segments matter

A 90%-feasible one-second bin still sampled its known-bad 10% in the old design. Frame-exact support recovers the good boundary material while making invalid exposure zero.

support contract

Why fixed horizons matter

Changing start distributions changed time-to-clip-end. The old command could teleport at a wrap without a terminal transition. Every new 1.0-second trial ends explicitly.

MDP contract

Why conditional outcomes matter

Raw failure arrivals mix difficulty with sampling frequency and episode duration. Attempts and failures now estimate conditional difficulty for the unit actually visited.

attribution contract
03 · GPU training quality

The implementation trained; the first reward did not

A bounded pilot caught a second failure mode before scale-up: the original net reward made early termination cheaper than tracking the full horizon.

Debug finding. The first policy drove mean episode length from roughly 21 to 5.5 of 50 steps while its return improved. That was reward hacking, not curriculum learning. A symmetric, failure-only −10 event cost restored segment completions and was used in both final arms.

Training telemetryExact-uniformAdaptiveInterpretation
PPO transitions2,457,6002,457,600matched compute
Environments × iterations512 × 200512 × 200same seed/config
Final-batch episode length / 5036.5234.68training telemetry, not evaluation
Cumulative failure rate0.91410.9162still a hard panel
Invalid / censored counters0 / 00 / 0mechanism gate passed
Final top unit / clip mass0.05 / 0.250.05 / 0.25caps held exactly
Adaptive TV from control00.0140treatment too weak

The checkpoint stores sampler statistics and RNG state, but not a bit-exact full simulator continuation. Resume is correctly labeled sampler-equivalent only.

04 · Paired evaluation

Identical worlds expose a quality signal, not a survival conclusion

Each policy saw the same 42 units, three phases, four replicates, seed, initial state, and startup domain randomization: 504 worlds per policy. Quality is compared only on 22,321 frames where both policies were still active.

EndpointUniformAdaptiveAdaptive − uniformUnit-bootstrap 95% interval
Success0.67460.6825+0.0079[−0.0536,+0.0714]
Mean survival0.9011 s0.9126 s+0.0115 s[−0.0100,+0.0346]
Body-position errorpaired common-survivor frames−0.00420 m[−0.00663,−0.00190]
Anchor-orientation errorpaired common-survivor frames−0.02795 rad[−0.03966,−0.01628]
Anchor-position errorpaired common-survivor frames−0.00517 m[−0.01247,+0.00233]
Mechanical work / actuatorpaired common-survivor frames−0.00246 J[−0.00545,+0.00046]

Best current interpretation. Exact segment training is operational, and adaptive allocation may improve local pose/orientation quality. It has not yet shown a reliable survival advantage. More seeds would be premature until the adaptive distribution materially differs from control.

05 · DFRP v1 exact CPU gate

Make repair provenance and exact support one training contract

The next step is not another seed. DFRP now routes every clip under explicit 10%, 5%, and 8/15 cm thresholds, binds exact support to the selected motion hash, and promotes only repairs with full-horizon-safe starts.

Frozen exact panel

A source-diverse 30-clip CPU panel admits 22/26 flagged repairs and all four controls. The 84.6% result is panel-specific, not a bank-wide recovery estimate.

gate passed

Two-tier repair

644/2,442 strict-flagged legacy clips fit an 8 cm displacement budget; another 962 fit 8–15 cm. All remain qualification-incomplete until newly repaired and exactly rescreened.

26.4% primary budget

Four failures stay out

Two clips remain above 5% infeasibility; two more clear that screen but exceed the 10 mm IK bound. Residual dynamics and geometric qualification remain separate gates.

fail closed

Exact MJLab handoff

The curated 26-clip view yields 36 verified units and 10,561 legal 50-step starts. Runtime rechecks motion, sidecar, DFRP, and unit-table identities.

hash-bound support

Claim boundary. The 8 cm bank-wide count reclassifies legacy root-only files, while 22/26 measures a stratified exact panel. Neither is a bank-wide root+IK recovery rate, and no policy or hardware benefit has been tested.

06 · Newton result

Conformance passed; predictive value did not

Newton earned a role as a reproducible analysis instrument, then failed the predeclared gate required for a curriculum feature. These are separate findings.

Canonical-state conformance passed

One easy and one contact-rich hash-bound unit pass placement, observation, action, resynchronized state, contact-timing, and deterministic-repeat checks after seven import residuals are mirrored.

measured ✓

Predictive panel was valid

The final run retained 40 units after the sealed alive exclusion. Determinism, pairing, reference bounds, contact checks, and manipulation checks support reading the outcome.

valid measurement

Joint predictive gate failed

Adaptive partial Spearman is +0.141 (p = 0.158) and held-out-clip LOCO lift is −0.006, below the registered +0.25 and +0.05 thresholds. Grounded replication also fails.

sealed ✗

G3 is not eligible

The kill rule was fixed before outcomes: valid-data failure keeps Newton as an instrument and forbids Newton-fragility-weighted training. Individual axis correlations do not reopen it.

G3 killed

Claim boundary. This failure rejects the tested three-axis Newton vector as an incremental curriculum signal on this panel. It does not undo the conformance result or claim that Newton is broadly uninformative.

07 · Next causal experiment

One controlled contrast isolates allocation

The pre-seal comparator audit removed both confounded G0 and ineligible G3. Phase G now changes one thing: how training exposure is allocated inside identical exact support.

G1

Exact-feasible control

Deployment-uniform over 368,951 legal starts; 1,184 variable-length units provide attribution.

G2

Absolute learning progress

The same support and prior, reweighted by ALP with endpoint-blind calibrated ρ = 0.40 and λ = 0.05. G2−G1 is allocation value.

Why G0 was removed. G0 samples clips uniformly and then frames within a clip; G1 is uniform over legal starts and therefore weights clips by admissible duration. Their difference would mix feasibility hygiene with duration exposure, so it cannot identify either effect.

Manipulation calibration passed

Two of 12 predeclared settings passed. Sampler ledgers selected ρ = 0.40, λ = 0.05 (mean TV 0.1079); its one independent seed also passed (mean TV 0.1056). No policy endpoint was read.

Score precision and liveness

The primary is a horizon-weighted pose/orientation TrackingScore on 25 outcome-blind reference-hard clips. Survival and mechanical work decompose the result; a passed-gate two-endpoint null recommends G1 only in this setting.

08 · Research fit

Recent work sharpens the claim instead of replacing the experiment

Physics-aware curation and adaptive sampling are active research areas. CLIMB's remaining question is narrower: does this allocation rule help after analytic support and the evaluation contract are fixed?

Feasibility is already a data axis

LIMMT combines heuristic physical scores with diversity and complexity selection. CLIMB does not claim that data quality first became important here; it contributes a policy-independent contact-capacity test on final robot-space trajectories.

LIMMT · 2026

Exposure may not fix capability

Athena-WBC reports residual feasible clips that targeted training still cannot solve. Therefore a Phase-G null would reject this ALP treatment, not prove that the remaining motions are intrinsically unlearnable.

Athena-WBC · 2026

One-factor comparisons matter

YAHMP changes individual design choices on one fixed Unitree G1 motion set. Phase G follows the same causal discipline by holding robot, support, PPO, reward, caps, compute, and evaluator fixed.

YAHMP · 2026

Contact deserves its own evidence

HumanTracker shows that kinematic averages can miss support and contact defects. CLIMB keeps contact timing exploratory until its own blinded held-out manual-label gate passes.

HumanTracker · 2026
09 · Current gate

The treatment is calibrated; confirmation awaits its seal

A licensed local reconstruction passes all 800 training and 100 disjoint evaluation SHA-256 identities. The endpoint-blind selector and independent validation both pass, the 512-environment footprint is measured, and contact timing is frozen exploratory-only because independent human labels are absent.

Measured manipulation only

Independent validation retains at least 700.1 effective units, at most 0.0134 top-1 mass, zero invalid/censored events, and 0.2365 saturation. These establish treatment separation—not tracking benefit.

Fail closed on outcomes

A different retarget would confound allocation with preprocessing. The advisor contract forbids a unilateral confirmation; no reward, survival, or tracking endpoint is opened before review, seal, and explicit approval.

Next executable action. Review and seal the hash-complete Phase-G contract, then launch the matched G1/G2 seed-1 manipulation gate only with explicit approval. G2 policy benefit remains pending; the calibration is not promoted into an endpoint result.

10 · Evidence and implementation

Trace every number