# Shallow-soil practice and transfer

The previous study had no full reference completions on soil. Greater leg correction capacity and firmer constitutive presets did not improve the original task. This follow-up tests whether limiting sink with a shallower physical substrate creates useful practice at the original motion speed.

The pilot compares 4 cm and original 14 cm beds, keeping the top surface at z=0 and all sand/mud constitutive parameters, motion timing, rewards, failure thresholds, torque limits, reference placement and policy weights unchanged. The rigid bed top and particle bounds change together in both native USD and the MPM collider. Ground remains unchanged. This deliberately adds support through a shallower substrate; it is not a claim of improved performance on deep soil.

Use the retained 400-rollout checkpoint, three existing training clips, reset seed 23, 500 steps and four terrain families, at 2 cm MPM cells. A practice depth is admitted only with at least 50% fewer pooled soil task failures than the fresh 14 cm control, zero ground failures, at least one full reference completion per soil family with at least one second of actual material contact, at least 7.5 seconds of aggregate contact per soil family, and no soil-family root/planar-velocity RMSE regression. Each family has 30 seconds of evaluation. Reference completions may include rigid-apron traversal; contact-qualified completion does not mean every frame was on soil.

If 4 cm fails, a predeclared second 2 cm practice pilot may be tested. The deepest passing practice depth is preferred. Reject nonfinite runs and geometry mismatches. Foot load-and-lift probes at 2 and 1 cm grids separately screen dependence on resolution and rigid support; they do not calibrate real material mechanics.

If practice is admitted, compare depth-scheduled training against equal-budget original-depth continuation, both starting from the common retained policy with the same architecture, rewards and eligible 84-motion pool. Specify the exact schedule and budget before training outcomes are observed. The 20 reserved motions remain unused during selection. Policy acceptance is evaluated on the original 14 cm task: at least 20% fewer pooled soil failures than both start and the matched control, zero ground failures, no per-soil-family root/velocity RMSE regression against start and at least 95% retained soil contact. Report full reference completions separately. Freeze a passing checkpoint before additional validation; no automatic promotion follows from easier practice.

This is a numerical training study with uncalibrated, explicitly coupled material physics. Thin layers and contact resolution are consequential limitations. One training seed and reused tuning clips do not establish generalization or convergence. Preserve rejected attempts and avoid selecting by training loss or reset-biased RMSE alone.
