Historical experiment record. For the completed E4 gate decision, repair comparison, two relative-progress replications and the next four-arm experiment, see the research update of 5 September 2026. The pilot measurements below retain their original scope.
What changed—and what has not been established
The engineering result and the policy result are different claims. Keeping them separate is the point of the audit.
Keep: a valid training experiment now exists
The command never samples a known-invalid start, never teleports through a continuing PPO transition, attributes each outcome to a stable segment unit, caps concentration, and records terminal metrics before reset.
Hold: no confirmed survival benefit yet
One seed improved success by 0.79 percentage points, but the unit-clustered interval spans −5.36 to +7.14. The treatment itself was only 1.40% total-variation from control.
From a physics screen to a clean PPO transition
The method acts before and during training. Offline physics defines admissible support; online learning allocates exposure only inside that support.
Screen reference forces
Contact-free inverse dynamics and contact capacity identify unsupported frames. This is hygiene, not a policy score.
Build exact units
Contiguous feasible frame intervals become stable units. The table retains 42 units and 4,679 legal 50-step starts.
Sample a safe horizon
Every proposal satisfies the complete future window. Known-invalid probability is exactly zero.
Track and terminate
The reference advances from s+1 through s+50, then emits an explicit timeout before any wrap or teleport.
Credit and rebalance
Completed outcomes update stable unit statistics on a fixed clock. Exploration and 5%/25% unit/clip caps remain enforced.
Why exact segments matter
A 90%-feasible one-second bin still sampled its known-bad 10% in the old design. Frame-exact support recovers the good boundary material while making invalid exposure zero.
support contractWhy fixed horizons matter
Changing start distributions changed time-to-clip-end. The old command could teleport at a wrap without a terminal transition. Every new 1.0-second trial ends explicitly.
MDP contractWhy conditional outcomes matter
Raw failure arrivals mix difficulty with sampling frequency and episode duration. Attempts and failures now estimate conditional difficulty for the unit actually visited.
attribution contractThe implementation trained; the first reward did not
A bounded pilot caught a second failure mode before scale-up: the original net reward made early termination cheaper than tracking the full horizon.
Debug finding. The first policy drove mean episode length from roughly 21 to 5.5 of 50 steps while its return improved. That was reward hacking, not curriculum learning. A symmetric, failure-only −10 event cost restored segment completions and was used in both final arms.
| Training telemetry | Exact-uniform | Adaptive | Interpretation |
|---|---|---|---|
| PPO transitions | 2,457,600 | 2,457,600 | matched compute |
| Environments × iterations | 512 × 200 | 512 × 200 | same seed/config |
| Final-batch episode length / 50 | 36.52 | 34.68 | training telemetry, not evaluation |
| Cumulative failure rate | 0.9141 | 0.9162 | still a hard panel |
| Invalid / censored counters | 0 / 0 | 0 / 0 | mechanism gate passed |
| Final top unit / clip mass | 0.05 / 0.25 | 0.05 / 0.25 | caps held exactly |
| Adaptive TV from control | 0 | 0.0140 | treatment too weak |
The checkpoint stores sampler statistics and RNG state, but not a bit-exact full simulator continuation. Resume is correctly labeled sampler-equivalent only.
Identical worlds expose a quality signal, not a survival conclusion
Each policy saw the same 42 units, three phases, four replicates, seed, initial state, and startup domain randomization: 504 worlds per policy. Quality is compared only on 22,321 frames where both policies were still active.
| Endpoint | Uniform | Adaptive | Adaptive − uniform | Unit-bootstrap 95% interval |
|---|---|---|---|---|
| Success | 0.6746 | 0.6825 | +0.0079 | [−0.0536,+0.0714] |
| Mean survival | 0.9011 s | 0.9126 s | +0.0115 s | [−0.0100,+0.0346] |
| Body-position error | paired common-survivor frames | −0.00420 m | [−0.00663,−0.00190] | |
| Anchor-orientation error | paired common-survivor frames | −0.02795 rad | [−0.03966,−0.01628] | |
| Anchor-position error | paired common-survivor frames | −0.00517 m | [−0.01247,+0.00233] | |
| Mechanical work / actuator | paired common-survivor frames | −0.00246 J | [−0.00545,+0.00046] | |
Best current interpretation. Exact segment training is operational, and adaptive allocation may improve local pose/orientation quality. It has not yet shown a reliable survival advantage. More seeds would be premature until the adaptive distribution materially differs from control.
Make repair provenance and exact support one training contract
The next step is not another seed. DFRP now routes every clip under explicit 10%, 5%, and 8/15 cm thresholds, binds exact support to the selected motion hash, and promotes only repairs with full-horizon-safe starts.
Frozen exact panel
A source-diverse 30-clip CPU panel admits 22/26 flagged repairs and all four controls. The 84.6% result is panel-specific, not a bank-wide recovery estimate.
gate passedTwo-tier repair
644/2,442 strict-flagged legacy clips fit an 8 cm displacement budget; another 962 fit 8–15 cm. All remain qualification-incomplete until newly repaired and exactly rescreened.
26.4% primary budgetFour failures stay out
Two clips remain above 5% infeasibility; two more clear that screen but exceed the 10 mm IK bound. Residual dynamics and geometric qualification remain separate gates.
fail closedExact MJLab handoff
The curated 26-clip view yields 36 verified units and 10,561 legal 50-step starts. Runtime rechecks motion, sidecar, DFRP, and unit-table identities.
hash-bound supportClaim boundary. The 8 cm bank-wide count reclassifies legacy root-only files, while 22/26 measures a stratified exact panel. Neither is a bank-wide root+IK recovery rate, and no policy or hardware benefit has been tested.
Conformance passed; predictive value did not
Newton earned a role as a reproducible analysis instrument, then failed the predeclared gate required for a curriculum feature. These are separate findings.
Canonical-state conformance passed
One easy and one contact-rich hash-bound unit pass placement, observation, action, resynchronized state, contact-timing, and deterministic-repeat checks after seven import residuals are mirrored.
measured ✓Predictive panel was valid
The final run retained 40 units after the sealed alive exclusion. Determinism, pairing, reference bounds, contact checks, and manipulation checks support reading the outcome.
valid measurementJoint predictive gate failed
Adaptive partial Spearman is +0.141 (p = 0.158) and held-out-clip LOCO lift is −0.006, below the registered +0.25 and +0.05 thresholds. Grounded replication also fails.
sealed ✗G3 is not eligible
The kill rule was fixed before outcomes: valid-data failure keeps Newton as an instrument and forbids Newton-fragility-weighted training. Individual axis correlations do not reopen it.
G3 killedClaim boundary. This failure rejects the tested three-axis Newton vector as an incremental curriculum signal on this panel. It does not undo the conformance result or claim that Newton is broadly uninformative.
One controlled contrast isolates allocation
The pre-seal comparator audit removed both confounded G0 and ineligible G3. Phase G now changes one thing: how training exposure is allocated inside identical exact support.
Exact-feasible control
Deployment-uniform over 368,951 legal starts; 1,184 variable-length units provide attribution.
Absolute learning progress
The same support and prior, reweighted by ALP with endpoint-blind calibrated ρ = 0.40 and λ = 0.05. G2−G1 is allocation value.
Why G0 was removed. G0 samples clips uniformly and then frames within a clip; G1 is uniform over legal starts and therefore weights clips by admissible duration. Their difference would mix feasibility hygiene with duration exposure, so it cannot identify either effect.
Manipulation calibration passed
Two of 12 predeclared settings passed. Sampler ledgers selected ρ = 0.40, λ = 0.05 (mean TV 0.1079); its one independent seed also passed (mean TV 0.1056). No policy endpoint was read.
Score precision and liveness
The primary is a horizon-weighted pose/orientation TrackingScore on 25 outcome-blind reference-hard clips. Survival and mechanical work decompose the result; a passed-gate two-endpoint null recommends G1 only in this setting.
Recent work sharpens the claim instead of replacing the experiment
Physics-aware curation and adaptive sampling are active research areas. CLIMB's remaining question is narrower: does this allocation rule help after analytic support and the evaluation contract are fixed?
Feasibility is already a data axis
LIMMT combines heuristic physical scores with diversity and complexity selection. CLIMB does not claim that data quality first became important here; it contributes a policy-independent contact-capacity test on final robot-space trajectories.
LIMMT · 2026Exposure may not fix capability
Athena-WBC reports residual feasible clips that targeted training still cannot solve. Therefore a Phase-G null would reject this ALP treatment, not prove that the remaining motions are intrinsically unlearnable.
Athena-WBC · 2026One-factor comparisons matter
YAHMP changes individual design choices on one fixed Unitree G1 motion set. Phase G follows the same causal discipline by holding robot, support, PPO, reward, caps, compute, and evaluator fixed.
YAHMP · 2026Contact deserves its own evidence
HumanTracker shows that kinematic averages can miss support and contact defects. CLIMB keeps contact timing exploratory until its own blinded held-out manual-label gate passes.
HumanTracker · 2026The treatment is calibrated; confirmation awaits its seal
A licensed local reconstruction passes all 800 training and 100 disjoint evaluation SHA-256 identities. The endpoint-blind selector and independent validation both pass, the 512-environment footprint is measured, and contact timing is frozen exploratory-only because independent human labels are absent.
Measured manipulation only
Independent validation retains at least 700.1 effective units, at most 0.0134 top-1 mass, zero invalid/censored events, and 0.2365 saturation. These establish treatment separation—not tracking benefit.
Fail closed on outcomes
A different retarget would confound allocation with preprocessing. The advisor contract forbids a unilateral confirmation; no reward, survival, or tracking endpoint is opened before review, seal, and explicit approval.
Next executable action. Review and seal the hash-complete Phase-G contract, then launch the matched G1/G2 seed-1 manipulation gate only with explicit approval. G2 policy benefit remains pending; the calibration is not promoted into an endpoint result.
Trace every number
- Exploratory pilot result and decision
- Machine-readable paired result
- Vectorized GPU lifecycle trace
- Newton-grounded direction addendum
- Sealed Newton predictive-gate result
- Current direction and machine-readiness audit
- Exact 900-file payload recovery and verification record
- Endpoint-blind G2 manipulation calibration result
- Draft Phase-G preregistration
- Feasibility-first ICRA execution plan
- Paper-completion checkpoint and stop rules
- ICRA-sized claim and eight-page argument map (withheld during review)
- ICRA-sized methods manuscript companion (withheld during review)
- Current primary-literature positioning audit
- Five silent harness traps
- DFRP v0 CPU result and claim boundaries
- DFRP v1 frozen exact-panel result
- Newton v1.5.0 release, solver feature matrix, and collision pipeline