Superseded snapshot. This page documents the protocol and a dated 6 September progress snapshot taken while training was still running. The campaign has since completed: all twelve training runs and all forty-eight held-out cells finished, and the registered result is inconclusive. Sentences below that describe runs as queued or in progress refer to that snapshot, not to the current state. Read the completed result.
New complete result · 6 September, 23:47 EDT: all 12 training runs and 48 held-out cells finished. The registered policy result is inconclusive; the all-panel guard does not pass. Read paired policy results and the revised research plan. The allocation progress snapshots below retain their earlier timestamps.

Confirmation update · 6 September 2026

Changed exposure. Tracking benefit still to test.

All three relative-progress confirmation seeds pass the allocation gate. The fixed four-arm campaign will test whether that exposure improves held-out tracking.

12 / 12
confirmation training runs complete and independently verified
492
saved checkpoint states replayed across all completed runs
Inconclusive
held-out policy benefit in the four-arm confirmation
Current checkpoint

The intervention is measurable; the policy comparison remains unopened.

All U/A/R/D runs for seeds 21 and 22, plus R23 and D23, have completed 4,000 iterations and passed their declared training gates. Every saved sampler state replays exactly, with zero recorded invalid-start, invalid-reference or censored-reset events. A is retained under its frozen comparator rule even when allocation contrast is weak.

Static snapshot · 6 September 2026, 22:17 EDT. U23 is running, with checkpoint 1100 saved; A23 is queued. The scheduler and follow-on workers remain active. No held-out evaluation cell has started. Both D calibration seeds and all four seed-51 confirmation-entrypoint smokes passed before the immutable freeze. The original overnight attempt timed out waiting for GPU capacity before any scientific job; its record is preserved. The active operational re-entry uses the same scientific contract and schedule.

Read how this experiment fits the wider research plan →

01 · Current confirmation measurements

Relative progress and conditional failure sustain allocation contrast.

Each completed run below has 512 environments, 4,000 PPO iterations and 41 verified saved states. Mean TV averages the 37 saved states at iterations 400–3999. Incomplete U23/A23 runs are excluded from the measured table and figure; no value is imputed.

Measured training allocation · no policy-utility inference
Arm and seedMean TVCompleted trialsDecision
U · seed 210.0000000001,086,702Training gate pass
U · seed 220.0000000001,086,071Training gate pass
A · seed 210.0297591701,084,430Training gate pass
A · seed 220.0293385941,082,658Training gate pass
R · seed 210.0834849491,083,411Training gate pass
R · seed 220.0824919971,087,418Training gate pass
R · seed 230.0823328441,089,129Training gate pass
D · seed 210.0839106891,116,479Training gate pass
D · seed 220.0858650931,116,999Training gate pass
D · seed 230.0839838531,117,687Training gate pass
Allocation TV histories by training seed. Relative-progress R and conditional-failure D sustain higher exposure contrast than absolute-progress A; U remains at the prior. Seed 23 currently includes only completed R and D runs.
All 410 verified confirmation snapshots. Lines show within-run histories, not independent replications or tracking learning curves. The dashed 0.05 line is a mean-TV gate for R/D, not a per-checkpoint requirement and not A's gate. Similar mean TV does not make R and D identical treatments. CSV data · PDF figure · Verified summary and source hashes.

Interpretation before policy outcomes

R passes in all three confirmation seeds (mean TV 0.083485, 0.082492, 0.082333); D also passes in all three. A's completed runs remain near 0.029–0.030. D begins near the prior and increases its allocation contrast later; R changes exposure earlier and follows a different trajectory. This reproduces the exposure distinction that motivated relative normalization. It does not show that R selects learnable motions better than D or that either improves tracking.

Development evidence · before confirmation

The failed gate gave us a specific next hypothesis.

E4: sealed decision not_tested

Both seed-1 training arms finished, but absolute-progress sampling reached mean post-warm-up TV of only 0.029659, below the frozen 0.05 minimum. Policy endpoints stayed closed and seeds 2–3 stopped. This is an inadequate intervention, not a measured policy null.

Repair: measured exploratory result

The fixed-policy comparison finished on 26 clips and 656 paired conditions per reference arm. Repaired minus raw TrackingScore was −0.001003, with clip-bootstrap 95% CI [−0.008588, +0.008020]. It demonstrated no aggregate tracking gain. The separate 22/26 repair qualification rate describes the physics screen.

Relative progress: replicated allocation

One fixed candidate completed two fresh runs of 512 environments and 4,000 PPO iterations. Both passed full sampler replay and the allocation gate. All recorded invalid-start, invalid-reference, and censored-reset counters stayed at zero.

Replay of E4’s saved sampler states showed that the fixed additive floor increasingly dominated the focus distribution as absolute progress shrank. We therefore normalized the existing progress signal by its deployment-weighted mean, leaving the estimator and admitted starts fixed.

Measured exploratory allocation · 4,000 iterations per run
RunMean TVMinimum effective unitsCompleted trialsDecision
Relative R · seed 110.083211616.251,087,814Manipulation pass
Relative R · seed 120.082728618.611,086,192Manipulation pass
Failure D · seed 310.084955625.061,113,773Calibration pass
Failure D · seed 320.085939624.291,118,388Calibration pass

TV is total variation: half the summed absolute probability differences from the deployment prior. Effective units summarize distribution breadth using entropy. Each run has 41 saved snapshots, including 37 from iteration 400 onward. Those correlated snapshots are not independent training seeds. E4 used a different saved cadence, so its mean is a diagnostic reference rather than a matched estimate.

Three development histories show sustained relative-progress allocation contrast; the failure baseline starts with lower contrast and grows later. Effective support remains broad and final saturation is below its limit.
Complete verified R11, R12 and D31 histories. Similar average TV conceals different exposure schedules: D starts near the prior and changes more late in training. The dashed TV line is a minimum for the post-warm-up mean, not a per-checkpoint requirement. Replication is unequal and development seeds differ; this figure is not a paired method comparison. Download all 123 snapshots (CSV) · PDF figure.
02 · Method design

Keep the signal’s scale from erasing its ranking.

Let b be the deployment prior over admitted motion units and g the existing nonnegative absolute change in smoothed conditional success. Set its deployment-weighted mean to μ = Σu bugu. With complete history and positive μ, use:

wᵤ = gᵤ / μ + 2
qᵤ = bᵤ wᵤ / 3
pᵤ = 0.40 bᵤ + 0.60 qᵤ
Apply the existing joint unit and clip caps.

Before active caps, this is exactly 80% deployment prior + 20% normalized progress distribution. A common positive rescaling of all progress values leaves it unchanged. With incomplete history or zero mean progress, the existing prior/cap fallback applies. This is a different normalization of the same information source.

The controlled setting stays fixed: Unitree G1 in pinned mjlab 1.6.0, a flat scene, 800 training motions, 1,184 admitted units, 368,951 exact legal starts, and 50-step trials that cannot wrap or leave an admitted interval. PPO, observations, actions, rewards, the progress estimator, and the 0.05 unit / 0.25 clip probability caps remain matched.

Main unresolved mechanism: absolute progress includes both improvement and decline. About 46.6% and 46.8% of the positive excess allocation mass in the two R histories went to declining success estimates, averaged over post-warm-up snapshots. That might reflect useful recovery from forgetting or estimation noise. These percentages describe excess probability mass, not the fraction of all training samples. More TV alone cannot resolve this.

Learning-progress curricula and failure-based motion sampling already have precedents. The proposed contribution is an auditable composition with exact support and a controlled policy test; the normalization alone is a small engineering change. No novelty or best-performance claim is established here.

03 · Fixed research plan

Four allocation rules. The same training budget and held-out conditions.

ArmAllocation ruleQuestion
UDeployment-uniform priorDoes adapting exposure help at all?
AExisting absolute-progress floorDoes the scale correction matter?
RFixed relative-progress rule aboveDoes sustained progress allocation improve tracking?
DConditional failure, fixed ρ = 0.8, power 1, floor 0Is progress more useful than practicing failures?

Run all four arms on fresh paired training seeds 21, 22 and 23, with 512 environments and 4,000 iterations per run. Evaluate checkpoints 1000, 2000, 3000 and 3999 on 100 held-out clips, including 25 feasible-hard clips, using 2,800 paired conditions per evaluation cell. This gives 12 training jobs and 48 evaluation cells. All 12 training manipulation checks precede evaluation.

The primary endpoint is the final-checkpoint R−U difference in feasible-hard clip-mean, liveness-weighted TrackingScore. A benefit decision requires a mean gain of at least +0.02, a positive lower two-sided seed-level 95% t-confidence bound (df = 2), and an all-panel lower bound above −0.01. With valid preceding gates, a hard-panel upper bound below +0.02 rules out the target-sized benefit; other unresolved cases remain inconclusive under the fixed decision rules. Gate failures retain their declared not_tested or invalid status.

R−A, R−D, learning curves, survivor-conditioned pose error with coverage, all-condition mechanical work with exposure, and elapsed cost are secondary descriptions. R−D is not a pure ranking-only ablation: startup behavior and floor/cap composition differ. A paired hierarchical bootstrap is supplementary to the primary seed-level interval.

Complete: calibrate and freeze

D31/D32 and all four seed-51 confirmation-entrypoint smokes passed. Source bindings, replay logs and all 12 configuration hashes were verified before sealing the current contract. Scientific settings remain fixed.

Running: finish all training

Ten training gates pass at this snapshot. U23 is running, followed by A23. The scheduler opens policy evaluation only after all twelve gates pass; a failure retains its declared disposition.

Next: measure policy utility

Run the 48 evaluation cells and validate all training, pairing and provenance records before aggregation. The queued efficiency postprocessor requires exact reproduction of the original complete analysis.

Precision is a real limitation. With three paired seeds, the primary 95% t-interval half-width is 2.484 times the observed seed-difference standard deviation. At an observed mean of +0.02, a positive lower bound needs that SD below about 0.00805. A prospective precision audit explores hypothetical normal-distribution scenarios; it is not measured power or a forecast. An inconclusive result is plausible and will be retained.

The implementation includes checkpoint-linked sampler telemetry, complete history replay, paired-evaluation provenance checks and a strict campaign analyzer. All four actual seed-51 smokes passed before the contract freeze. Sample-efficiency analysis, the frozen-policy pilot and CUDA lifecycle checks are queued with separate prerequisites; these queued stages are not completed evidence.

04 · Inspect the record

Evidence, design and reproducible boundaries

The downloads below export already measured results and identify their original artifacts by repository-relative path and SHA-256. They contain no trained weights or licensed motion payloads. Incomplete training runs are excluded from the completed-run comparison. No held-out policy outcomes were read by the public exporter.

Labels: sealed means the protocol or decision is preserved; measured means observed in the stated experiment; exploratory limits generalization; pending means the evidence does not yet exist. Simulator qualification does not establish hardware feasibility. Repair uncertainty here resamples clips for one fixed policy, not training seeds; numerical repeatability and repair-specific effects remain separate questions.