Can a frozen generative motion model become a reliable humanoid local navigation controller through adaptation informed by physical execution?
What this is. KimoNav compiles a typed navigation program into a metric SE(2) path, generates whole-body motion for it with a frozen ARDY diffusion model, repairs that reference where it is about to violate the task, and hands it to a frozen SONIC tracker driving a simulated Unitree G1 on flat ground. Those stages are studied and scored separately; no single script yet runs the whole chain end to end. Every robot execution here is physics simulation, and several studies on this page are entirely offline — geometry, forecasts and model fits that add no simulator captures at all, each labelled as such. No hardware has been run.
We ask the robot to walk, turn, and stop on a metric path and on time, then score the original request — not the reference. KimoNav studies the gap between a motion that looks right and a robot that completes that request.
Ongoing research by Linji Wang. Not a published paper. The working manuscripts stay local pending author review; this page is their public summary. Upstream Kimodo, ARDY and SONIC are NVIDIA systems used frozen; integrating them is not our contribution.
LATEST / 9 SEPTEMBER 2026
The inspected governor gain now reproduces live: 10/24 complete tasks versus 8/24, two gains and no losses. Those 24 are the same development programs the rule was designed from, under one simulator seed, so this is implementation reproducibility rather than a validated controller — and 10/24 is the arithmetic ceiling any two-arm rule could reach on this panel. A new 18-proposal geometry study advances stopping holds but yields no complete-task nominal admission.
The same robot, the same request, the same simulator seed. On the left the reference the system would have kept; on the right the reference the governor put in its place. Both panels are replays of recorded states — no physics was re-run to make this video, and no hardware is involved.
Read the denominator first. These are 2 of 24 inspected development programs — specifically the only two whose outcome changed. The governor kept the current reference in the other 22. Complete-task success goes from 8/24 to 10/24 with no regressions; repairing every program unconditionally would have stayed at 8/24, gaining the same two and losing two others. 10/24 is also the arithmetic ceiling here: ten of the 24 programs have at least one successful execution across the two references, so no rule choosing between them could have done better.
What you are looking at. The robots look almost identical, and that is the point: neither falls, both stay on the requested path, both reach the stop marker. The only failed gate is the requested final STOP — one second on the left program, 0.96 s on the right — where residual speed must stay strictly below 0.10 m/s. Left to right, peak STOP speed goes 0.217 → 0.072 m/s and 0.118 → 0.081 m/s. In fairness to the comparison, the repaired runs end slightly farther from the requested stop point than the runs they replace.
What "live" does and does not mean. It means the choice was computed during the run, from that run's own measured history — not from hindsight. It does not mean real time: the simulation was paused for 18.1 s and 11.3 s of wall clock while each decision was prepared, and the run records make no real-time claim. The rule itself was designed after these same programs were inspected, and the reserved evaluation set has not been opened, so nothing here is out-of-sample validation.
Replayed from recorded Isaac Lab root and joint states through MuJoCo forward kinematics; the two panels are bit-identical up to the decision tick and time-aligned on the original clock throughout. Source captures, hashes and per-frame claims →
What changed inside the reference
The repair moves the whole-body reference within a five-centimetre position tube and a three-degree heading tube of the compiled request, clipped to the robot's own joint limits, and brings the terminal hold forward so the stop is reachable. The request itself is never relaxed: position, heading, speed and every STOP interval stay exactly as compiled.
Why it is only two programs
The governor is deliberately minimal — it leaves the reference alone unless the current one is already forecast to fail. On this panel that condition fired twice. That is the intended behaviour, and it is also why the result is small: two changed programs cannot distinguish a working method from a lucky one. The reserved 36-program set exists to answer that, and remains sealed.
CURRENT / REFERENCE REPAIR IN SIMULATION
A live gain. A remaining tradeoff.
The original request and strict STOP threshold stay fixed. We separate complete task success, STOP-speed compliance and motion quality, and distinguish new simulator execution from selection of already measured outcomes.
8 → 10 / 24live governor, 2 gains 0 losses — on the same inspected panel the rule was designed from; repairing everything stays 8/24
10 / 24ceiling any two-arm rule could reach here
0 / 16offline repair forecasts admitting the task — no execution
36reserved programs still sealed and unevaluated
Same 24 inspected development programs, one simulator seed. The four columns are not the same tier of evidence. Current and repair-always are archived paired captures; governor replay is a selection over already-measured outcomes with no new physics; live binary is 24 qualified captures — 22 new, two reused, one declared recovery. Note that repair-always ties the governor at 7/24 on task-and-quality, so that row is not a governor-specific gain.
Metric
Current archived capture
Repair always archived capture
Governor replay saved-outcome selection
Live binary 22 new + 2 reused
Complete task
8/24
8/24
10/24
10/24
STOP speed with coverage
12/24
16/24
14/24
14/24
Complete STOP phase
9/24
9/24
11/24
11/24
Runtime quality
14/24
11/24
14/24
14/24
Task and runtime quality
5/24
7/24
7/24
7/24
Unconditional repair rescues two tasks and loses two successes. The governor selects two repairs and retains current in 22 programs. Its rule was designed after inspecting those outcomes. The two task rescues use the sequential projection branch; the coupled branch does not produce a complete success in this panel. One early termination remains a task failure with unobserved STOP.
READ 10/24 CORRECTLY
10/24 is a ceiling, not a frontier. Across the two references executed per program, ten of the 24 have at least one successful execution: six succeed with both references, two only with the current one, two only with the repair. No rule that picks between these two arms can exceed 10/24 here, however good its forecast. The governor reaches that ceiling; it does not move it. And the rule was designed after these outcomes were inspected, so this panel cannot also test it.
HOLDING STILL IS ALSO HARD
A perfectly static reference is not automatically enough. Given six perfectly static three-second hold references, the frozen tracker still fails the full STOP check 0/6, with peak speeds of 0.1338, 0.1475, 0.1091, 0.1516, 0.1634 and 0.1791 m/s against a strictly-below-0.10 m/s requirement. But every one of those violations is a startup transient: each peak occurs between 0.02 s and 0.38 s, only four to nine samples per capture exceed the limit, and all six then hold below 0.10 m/s continuously from 0.16–0.44 s onward. These are six initial-pose conditions on one standing task, not six navigation designs — and on the 24 navigation programs the same frozen tracker does pass the strict STOP check on 12 of 24 current and 16 of 24 repaired arms. Read it as a warning about the transient into a hold, not as a floor the tracker cannot go below.
The binary rule reproduces its inspected replay
All 24 live choices and task outcomes match replay. The two gains are right pivot–walk 10 s and right strafe–walk 3 s. All fresh repair and forecast arrays match their archived counterparts, and the selected first-episode execution matches all 33 recorded arrays per program. This establishes implementation reproducibility under the same simulator seed, with no new-program validation.
The original interrupted attempt remains unavailable. The completed panel has 24 captures and 25 total starts; its original-first-attempt view retains 10 successes, 23 available outcomes and one unavailable outcome among 24 allocated programs. One actual early termination remains a task failure with unobserved STOP. Preparation takes 11.30–18.13 s with physics paused.
Stopping on time is not one problem. This four-arm comparison holds the program fixed and varies only the intervention, so the two failure modes separate: braking alone brings the stop under the speed limit but loses the requested position, and stride compensation alone preserves position but does not stop. Only the combination satisfies both — on this program. The identical four-arm construction was also run on a second program (fresh-bank strafe–walk right, candidate seed 2703) and the combination failed there too, at a 0.1341 m/s STOP peak; across the eight fixed first candidates it has been tried on, it produced one complete success.
One program, four replayed arms. Complete success requires the STOP peak below 0.10 m/s and maximum position error at most 0.20 m — not either one alone.
Arm
STOP peak (m/s)
Complete task
What it loses
Original continuation
0.2139
Fail
stop speed
Braking only
0.0605
Fail
position (0.2480 m against a 0.20 m limit)
Stride compensation only
0.2142
Fail
stop speed
Stride + braking
0.0756
Pass
—
Wide figure — scroll it sideways to read the labels.
A post-hoc mechanism illustration, not a result. All four arms are the same program (m3k_strafe_walk_stop_left_5s) at simulator seed 1701, replayed from already-saved trajectories: no new physics, zero independent programs, and no generalization claim. It shows how the two failure modes trade against each other; it does not show a success rate — and it does not expand what the reference bank could already do. On this same program a different original candidate (draw 1704) already succeeds with no repair at all, at a 0.0937 m/s STOP peak; the panel audit's own words are that the rescue "does not establish a new program beyond the original finite-bank support ceiling".
Earlier stopping can lose timed progress
This study is entirely offline: it adds zero robot-simulation captures. Its "0/16" is a forecast result, not an execution outcome. The registered 18-proposal grid varies the reference path prior and settling time on two inspected programs. Sixteen repairs pass geometry and rate checks, but none of their forecasts passes the complete task. The executed position tolerance stays 0.20 m; STOP stays strictly below 0.10 m/s.
All 18 geometry proposals retained. Hold times follow requested settling 0.12 / 0.24 / 0.40 s; starred times belong to rejected references.
Side / prior
Geometry admitted
Proposed hold starts (s)
Full-task nominal admission
Left / 5 cm
3/3
5.96 / 5.84 / 5.84
0/3
Left / 10 cm
3/3
5.96 / 5.84 / 5.72
0/3
Left / 15 cm
3/3
5.96 / 5.84 / 5.68
0/3
Right / 5 cm
3/3
5.96 / 5.84 / 5.84
0/3
Right / 10 cm
3/3
5.96 / 5.84 / 5.72
0/3
Right / 15 cm
1/3
5.96 / 5.84* / 5.68*
0/1
At 0.40 s requested settling, the 5 cm prior delays the hold to 5.84 s; 10 cm advances it to 5.72 s; 15 cm reaches 5.68 s. Four left candidates predict compliant STOP speed while failing position. The lowest STOP forecast is 0.043712 m/s with 0.326540 m maximum position error. The closest right forecast is 0.100476 m/s, still failing strict STOP.
The 18 completed forecasts cover 16 admitted candidates and two current comparators. A zero-state import-path failure is preserved; a registered forecast-only recovery reused the geometry without another solve. This grid adds zero robot-simulation captures. Its low-speed candidates are not promoted as successful task repairs.
Wide figure — scroll it sideways to read the labels.
Offline terminal-reference mechanism, complete unsmoothed original-clock traces. The case order was fixed in the registration before any new outcome was accessible; no simulation was run to produce this figure. Three paired instances are plotted — two distinct programs (left and right strafe–walk–stop 5 s), the left one from two different draws; the registration's fourth case is a disabled null control and is not a plotted row. The registration records no independent generalization claim. This mechanism is offline — it is not online, live, or reactive control.Vector PDF for zooming →
The separate pelvis pilot: a wrong-sign selection
Three fresh captures on two inspected strafe–walk programs: two governor arms (0/2 successes) and one right zero-offset control (0/1). Each uses its own measured history and a freshly constructed reference bank. STOP must remain strictly below 0.10 m/s.
Live simulation pilot. Maximum speed over the original STOP interval; all three captures fail STOP and the complete task.
Arm
Predicted (m/s)
Executed (m/s)
Task
Left governor · retain current
0.264323
0.213887
Fail
Right governor · +20 mm X
0.096005
0.158513
Fail
Right control · zero offset
0.118264
0.150212
Fail
Prediction error changes the choice. On the right, the offset predicts a 0.022260 m/s benefit over zero, but measured STOP peak rises by 0.008301 m/s. Both repairs improve on current's 0.221494 m/s peak, yet both fail. The paired banks, forecasts and pre-update histories match exactly.
Latency remains unresolved. Total preparation takes 23.28–27.54 s with physics paused. These are fresh simulator interventions, not real-time execution. No falls, early terminations or missing pilot captures occurred; three trials are only two program designs.
Existing source-bound waveforms from the live pilot. No new physics or rendering-based outcome is added by this page. Download figure PDF.
Nominal forecasts capture some response
The original model evaluates 48 forecasts over 24 paired programs. At 1.5 s it achieves 0.02050 m position RMSE on 23 completely observed program pairs; at 2 s, 0.02256 m on 13. Predicted STOP-change direction agrees on 17/23, including an exact identity fallback. These average errors do not certify STOP compliance.
Existing paired development forecast analysis, including censoring and the original STOP limit. Download figure PDF.
Better action fit did not improve the forecast
The causal joint-default comparison keeps all 48 slots: 24 new forecasts and 24 exact nominal reuses. At 1 s, position RMSE worsens from 0.016395 to 0.016822 m on 23 complete program pairs; at 2 s, from 0.022559 to 0.022585 m on 13. Speed error worsens at all five horizons. Five predicted-task-pass/actual-fail slots and six predicted-STOP-pass/actual-fail slots remain. The offsets are not adopted.
The task prediction denominator is 48 arms over 24 programs; observed STOP covers 46 arms over 23 programs. Missing intervals are not filled. The binary live collection is now complete under its declared recovery accounting. Its two selected-task false admissions and two observed-STOP false admissions remain explicit; nominal predictions are not guarantees.
The title is the draft's, and it is not yet earned: preparation runs with physics paused, so today this is offline repair. See the method section.
The active method draft now includes the completed live binary check and the reference-envelope timing–position tradeoff, alongside the terminal comparison, forecast validation, live pelvis failure and rejected offset model. Its abstract and conclusion retain these limits. It is not submission-ready.
What is implemented
Whole-body terminal projection, a nominal tracker-conditioned execution forecast, original-clock task-error accounting and fresh own-history reference replacement. The method must still show that these pieces yield reliable complete-task improvements under actual computation delay.
Two distinct paper scopes
The earlier diagnostic manuscript remains an account of requested, generated and executed motion and selection limits. The current method draft develops repair from those consumed results. The full manuscript stays local pending author review; this page provides its public summary.
All execution is simulated G1 motion. No hardware reliability, real-time operation, calibrated stopping guarantee or broad generalization is established. OpenAI Codex assisted with analysis, code and substantive drafting; human author review remains required.
01 / THE MODELS WE BUILD ON
Three components. Different responsibilities.
Kimodo and ARDY generate kinematic motion. SONIC turns a reference into physical behavior. These are upstream systems; KimoNav is the research around their navigation interface, adaptation, and evaluation.
OFFLINE MOTION GENERATION
Kimodo
A text- and constraint-conditioned diffusion model. Its two-stage denoiser separates root motion from body motion, supporting paths, waypoints, and joint constraints.
Role here: the broader controllable-motion foundation and an offline motion/data source. It is a separate model, not a stage that runs inside our current ARDY→SONIC execution loop.
Autoregressive diffusion with explicit root features and a compact body representation. It generates continuations from history while accepting text and kinematic constraints.
Role here: the frozen G1 generator used in our experiments. Our setup uses 48 history frames and up to 52 new frames per chunk at 25 Hz; later chunks can use generated history.
A learned whole-body tracking system for humanoids. Reference motion and robot observations determine low-level actions in physics.
Role here: the frozen tracking policy for a simulated Unitree G1. Our pinned setup uses a 50 Hz control clock and ten reference samples spanning 0–0.9 seconds.
ARDY already demonstrates integration with SONIC on G1. Combining these components is not our novelty claim. Upstream capabilities do not, by themselves, establish our navigation performance.
NEW / FAILURE ANATOMY
What can the next intervention repair?
A passing kinematic reference does not prove physical feasibility. This audit separates reference defects from execution failures across the same 24 development contexts.
192 candidate executions across 24 programs, two generation banks and four candidates per bank. Candidates share program dependencies; these are not 192 independent programs. STOP checks fail in 73 references and 109 executions; failure types overlap.
RECORDED EXECUTION / WHY THE CLOCK MATTERS
A better continuation cannot erase the past.
At 3.54 seconds, measured-history replanning puts this robot closer to the requested path. Both complete programs still fail: their shared prefix first violates the path at 1.52 seconds, before the update at 2.08 seconds.
A preselected development failure, replayed from recorded Isaac Lab root and joint states. Rendering adds no physics and is not hardware footage. The dots show the original requested path. Frame times and source provenance →
02 / METHOD DESIGN
Repair the reference. Only when it fails.
Both generative models stay frozen. Our contribution sits in the interface between them: compile the request, bound how far a reference may be moved, forecast what the fixed tracker will do with it, and replace it only when its own forecast already violates the task.
Wide figure — scroll it sideways to read the labels.
The evaluation pipeline and its three clocks. A timed request compiles to 25 Hz generated motion under frozen ARDY, passes through the G1 reference interface — named joint projection and a time-preserving 25→50 Hz conversion — and is tracked by frozen SONIC at 50 Hz over 200 Hz physics. The repair, forecast and governor stages described below sit inside that reference interface. Both the reference and the execution are scored against the request, whose phase deadlines never move. Offline replay does not establish live generation deadlines.Figure SVG.
01 · OURS
Typed program
Five primitives — turn in place, walk, arc turn, strafe, STOP — compiled to a dense root path and heading with trapezoidal profiles at 1.5 m/s² and 180 °/s².
02 · FROZEN
ARDY generation
Whole-body continuations from causal history under the compiled root constraints. 52-frame chunks at 25 Hz, 10 denoising steps, CFG (2, 2). Parameters unchanged and hash-checked. In the terminal-repair studies these continuations are generated ahead of the run and replayed, not sampled inside the decision.
RESEARCH INTERVENTION
Repair · forecast · governor
A bounded whole-body terminal repair, a nominal MuJoCo forward model of the fixed tracker, and a rule that replaces the reference only when the current one is already forecast to fail.
03 · FROZEN
SONIC + simulated G1
A 50 Hz tracking policy over 200 Hz physics with a 0–0.9 s reference preview, one environment, flat ground. Executed states, contacts, clocks and terminations are recorded.
↶
Two different events, one clock. The causal history bridge returns 48 measured frames at 25 Hz and regenerates the continuation at a fixed 2.08 s boundary. The governor's terminal decision is a separate, later event — one second before the requested stop, which across the 24 programs falls between 3.08 s and 10.12 s. Each program gets one of each; repeated live replanning remains an unrun test.
The governor rule, in full
If the current reference has no eligible forecast, retain current.
If the current reference's own forecast is already admissible, retain current.
Otherwise take the admissible repair with the smallest predicted STOP peak speed, ties broken by index.
If no repair is admissible, retain current.
On the 24 inspected programs it retained the current reference 22 times and replaced it twice. Both replacements turned a failed task into a passed one; nothing regressed. Note the scale honestly: each decision compared the current reference against exactly one candidate repair, so rule 3 never had to break a tie.
What the method must still demonstrate
A gain on programs the rule was not designed from. The 36 reserved programs remain sealed and unevaluated.
Control computed while the simulation keeps running. Today physics is paused for 11.30–18.13 s per decision.
A forecast that predicts the sign of its own intervention. In the pelvis pilot it did not.
Motion quality alongside task gains: STOP behaviour, sliding, and termination.
A task-success gain that survives strong simple controls. Repairing every program unconditionally reaches the same 7/24 on task-and-quality.
A benefit that is consistent across repeats. Seed sensitivity is a live issue here: the historical selection arms move +1, +1, −2 across three simulator seeds.
Where the reference may move, and how far
The working repair is a sequential projection with hard bounds: a 5 cm position tube and a 3° heading tube around the compiled request, root height within ±4 cm, per-sample slew limits, and every pose clipped to the G1's named joint limits. The same joint-limit projection is applied one stage earlier to frozen ARDY's own output: 13 of the 24 original references needed it, by at most 0.1804 rad, and 29 of the 96 candidates by at most 0.0869 rad. Stance segments are geometric hypotheses selected by height, speed and angular speed; no measured contact force enters the projection. The executed entry point wraps a coupled solver around that projection, but on this panel the coupled branch changed only 2 of 24 references and rescued zero tasks — every task change came through the sequential projection.
Why the measurement clock changes the model
At the continuation boundary the predictor initializes from a measured state at 2.04 s and the new reference becomes active at 2.08 s. The intervening 40 ms must run on the original reference — a candidate cannot retroactively change earlier preview. The same 40 ms rule governs the later terminal decision. The reconstruction of the 25→50 Hz conversion matches 6,912 saved transaction digests with zero difference from the recorded references, and future-state poisoning leaves every input unchanged.
This interface reads privileged simulator root state and contacts. Hardware deployment would require a state estimator, frame alignment, and latency and noise qualification. None of that has been done.
Wide figure — scroll it sideways to read the labels.
Two inspected development traces at simulator seed 1701 illustrate the selection and timing effects. The light-blue interval is the 5.68–6.08 s reference transition; grey is the unchanged 6.12–7.08 s STOP interval. The dotted horizontal line is the strict 0.10 m/s threshold. Left: the governor retains the original reference, whereas unconditional repair creates a STOP violation. Right: immediate repair creates a complete success, but the measured 40 ms activation delay loses the narrow STOP margin. Dashed state-triggered curves coincide with original on the left and unconditional repair on the right. These two illustrations are part of the complete four-program table; they are not an independent test set or evidence of general stopping robustness.Reproduced from the figure's registered caption, with spacing and spelling normalised.
MEASURED, AND UNRESOLVED
The 40 ms budget is not met. All 44 full repair transactions exceed it, and preparation currently runs with physics paused. In the one study where activation age was varied directly, a 0 ms age gives a STOP peak of 0.098216 m/s and the task passes; a 40 ms injected age gives 0.101329 m/s and the same task fails. Until control is computed while the simulation continues, this work claims offline repair with paused preparation — not deadline-aware control.
Success is scored on the original request, not on tracking error
A program succeeds only if all of the following hold on the original clock: maximum position error ≤ 0.20 m; maximum heading error ≤ 10°; speed MAE ≤ 0.15 m/s; ordered phase boundaries with per-phase position and heading checks; complete duration and coverage; STOP dwell at least the requested length, with speed strictly below 0.10 m/s throughout including its boundaries, and overshoot ≤ 0.20 m from STOP onset; and no operational fall, defined as root height below 0.25 m or roll/pitch beyond 1.4 rad. A later sample cannot erase an earlier violation, and an actual early termination remains a failure with any unobserved STOP interval explicitly recorded as unavailable.
NEW / COMMON-INTERFACE PLANNER COMPARISON
A practical alternative, under the same evaluator.
Native planner pose replay succeeds on 0/24 development programs; archived ARDY fixed references succeed on 6/24. This measures two command-to-reference pipelines with the frozen G1 tracker. The decisive caveat: all 24 planner references already violate the request before any physics runs, so 0/24 measures the reference, not the planner — it does not establish that ARDY beats native SONIC deployment.
24 development programs, one planner generation per program, simulation seed 1701. All 24 new physics captures were independently rescored. This is not a new confirmation set.
Original-clock execution · equal program weighting · lower errors and sliding are better
Method
Success
Path (m)
Heading (°)
STOP overshoot (m)
Sliding ratio
ARDY fixed
6/24
0.116
1.736
0.175
0.123
Native planner · pose replay
0/24
0.393
5.869
0.888
0.123
What is held fixed. The G1 pose-replay interface, frozen tracking checkpoint, original request clock, evaluator and simulation seed. Planner commands use requested speed, movement direction and facing; idle explicitly implements STOP. No tracker update or command tuning was performed.
What this does not isolate. The planner receives sampled commands; ARDY receives root constraints. Initial states depend on their references. The common interface derives velocities from poses, whereas native C++ deployment blends velocity channels separately. This is not full native deployment parity or a tracker-only comparison.
Planner outcomes include 0 early endings and 0 operational falls. 0/24 generated references pass the request checks; passing kinematics does not certify dynamics. Missing or partial task channels retain their availability counts in the complete measured report →
HISTORICAL / COMPLETED SELECTION DIAGNOSTIC
The missing repeats change the conclusion.
Completing the first-candidate arms separates selection from regeneration. Minimum jerk adds one success in two simulator seeds but loses two in the third. These are the same twelve designs, already consumed by analysis.
Twelve designs per seed. The first-candidate repeat extension adds 18 captures and reuses six exact earlier captures; it was registered after inspecting the original results.
Seed
ARDY fixed
First candidate
Minimum jerk
Change vs first
Slide regressions vs first
1701
2/12
3/12
4/12
+1
0
1702
2/12
1/12
2/12
+1
2
1703
1/12
5/12
3/12
-2
1
All gains over fixed reference are left reverse-pair designs. At seed 1703, selection loses the right reverse-pair 3 s and 5 s first-candidate successes. The extension has no new falls or early endings. Primary first candidate retains its two early terminations and only ten observed STOP intervals. The later repair study's 36 reserved programs are a separate set.
Eight candidates per context across two four-seed draws. Candidates and frames are repeated observations, not independent tasks. Both banks are development data; all rows of a context stay in the same family split.
PHYSICS EXECUTION · M3O
Can we select a better reference?
A twelve-feature risk ranker and a 32-feature phase/state ranker choose among four continuations. We execute candidates from matched prefixes and compare selected outcomes. The oracle uses hindsight and quality constraints; it is a ceiling, not a deployable method.
M3o selected physics outcomes · task success and quality are evaluated together
Selection
Original success
Fresh success
Original early
Fresh early
Original slide
Fresh slide
K=1 reference
8/24
9/24
1
0
0
0
Risk ranker · 12 features
8/24
10/24
1
0
5
5
Phase ranker · 32 features
7/24
10/24
2
0
6
4
Qualified oracle · hindsight
14/24
13/24
0
0
0
0
Fresh draws: both rankers achieve 10/24 successes. The phase model lowers mean task loss by 2.62% relative to the risk model, but has four foot-slide regressions. Its original-bank result also fails.
Coverage matters: none of the 24 fresh pivot–walk–stop candidate executions succeeds. A better ranker cannot select a successful reference that the bank does not contain.
The table summarizes selected outcomes from the candidate banks. M3o adds 96 fresh candidate attempts plus four fixed/null controls, for 100 new physical attempts. No physical fall was recorded; early termination and task failures still count. Sliding regressions are relative to the matched K=1 control.
M3p's learned rate-change models look accurate with a measured-state reset at each step, but accumulate error when feeding back their own predictions. M3q replaces unrestricted state feedback with a bounded convex blend of predicted and reference rates.
Bounded responsenext rate = ρ × predicted rate + (1 − ρ) × reference rate0 ≤ ρ < 1 · known active preview · no later measured-state resets
Primary window: 24 contexts, 192 candidate records; 28,820/28,992 target states observed. The 172 missing suffix targets belong to two early endings.
Common first 3 seconds · free prediction · lower is better
Predictor
Composite
Displacement (m)
Velocity (m/s)
Heading (°)
Reference following
0.5845
0.0983
0.1810
2.3499
M3p · learned, no preview
3.7234
0.3739
0.4128
3.1909
M3p · learned preview
5.8309
0.5614
0.4563
5.9862
M3q · bounded, no preview
0.5561
0.0989
0.1754
2.3688
M3q · bounded preview
0.4710
0.0962
0.1592
2.3688
Lower is better. Each row predicts previously measured execution; these are not controller-success scores. The composite scales displacement, velocity, and heading MSE by .20 m, .15 m/s, and 10° squared, with equal context weighting. Endpoint horizons have different requested denominators. Download all windows, families, and both banks ↗
PROMOTION GATE: CLOSED
19.41% lower composite, with a retained heading regression. Bounded preview improves displacement and velocity prediction, and every rate bound passes. Heading RMSE rises from 2.350° to 2.369°, violating the frozen no-channel-regression rule. We keep that rule and report the tradeoff.
M3p: 128 regression fits, 768 audited forecasts. M3q: 3,840 grid forecasts and 32 finite parameter selections; no new regression training or physics. Every outer fold selects .10 s decay/.10 s preview for horizontal velocity. Yaw uses zero preview, with identity for the pivot holdout and .04 s decay for the other three families.
ONE RETAINED EXAMPLE
A reference path and a physical path can disagree.
Recorded G1 execution: right pivot–walk–stop, 5 s continuation, fresh seed 2703. Chosen after evaluation to illustrate a failure; retained in all quantitative denominators. This context first violates its timed path at 1.52 s, before the 2.08 s continuation decision. Lines show geometric paths; no time warping is used in task scoring. Trace and provenance ↗
04 / HOW THE RESEARCH EVOLVED
Each stage changes the next question.
The research began with a small reconstruction adapter. Physical evaluation changed the objective: learn from executed outcomes, then test whether the intervention remains useful when the system runs forward.
M2 / RECONSTRUCTION
Does the model use the path?
Token-local features reduced joint MSE by 7.21%, but mean and shuffled-path controls explained much of the gain. The useful-path gate failed. Explore the archived study →
M3A–M3H / EXECUTION INTERFACE
Make the intervention measurable.
Move into autoregressive generation and simulated tracking. Qualify measured history, physical clocks, contacts, and persistent reference replacement. Preserve engineering failures and negative transfer.
M3I–M3O / EXECUTED OUTCOMES
Measure headroom before learning.
Build matched candidate banks and outcome ledgers. Ranking gains fail quality/transfer gates, and some families have too few successful references to select.
M3P / RESPONSE IDENTIFICATION
Free prediction is the harder test.
Qualify the exact preview. Linear models achieve low one-step error yet drift in free rollout, rejecting them as a correction teacher.
M3Q / BOUNDED DYNAMICS
Constrain the response structure.
A bounded conventional model improves predictive composite by 19.41%. The small heading regression remains visible; compensation is not promoted.
M3R / REPAIRABILITY & OBSERVATIONS
Separate missing references from tracking drift.
Audit all 192 outcomes and the actor information map. Split 512 local retention motions into 304 training, 103 development and 105 final examples by source. Tracker training remains unstarted while interface and runtime qualification are incomplete.
R-A / R-B · HISTORICAL DIAGNOSTIC
Separate feedback from selection.
The first-candidate repeat extension is complete. Minimum jerk changes first-candidate successes by +1, +1, −2 across three seeds; the compiler braking lead was not promoted.
SEPTEMBER 9 / REFERENCE REPAIR
Test the predicted intervention.
The completed live binary governor reproduces 10/24 versus 8/24 current on inspected programs. The next 18-proposal grid advances holds but offers no full-task nominal admission. Prediction transfer, actual delay handling and final36 evaluation remain unfinished.
Watch the original motion-bank illustrationSaved native ARDY motion at 25 fps, ±90° turns. A kinematic dataset illustration, not a physical success demonstration or output of a promoted adapter. It is a skeleton line drawing of frozen ARDY output — no simulator and no physics replay, unlike the recorded execution clips above. Original render provenance ↗
05 / REPRODUCIBILITY & OPEN QUESTIONS
Make the claim as strong as the evidence.
KimoNav is ongoing research by Linji Wang, not a published paper. The current evidence covers terminal repair, forecast errors and the limits of selection; the earlier benchmark remains a diagnostic contribution. The controller claim remains to be earned.
Verified for this update
This continuation adds 22 simulator captures with 44 binary forecasts, then 18 geometry attempts and 18 completed grid forecasts. Separate saved-output audits recompute task/model scores and inspect paired execution arrays, raw projections, limits, rates and clocks. Both audits pass. Original interruptions, the zero-state forecast failure and rejected references remain recorded. Audit verification itself reruns no producer.
Still to establish
Reliable aggregate task improvement from repair and selection, prediction of new repair effects, actual delayed repeated feedback, broad motion retention, sensing noise, and hardware transfer. Torque and ZMP have not been measured here. A bounded predictor is not a physical stability guarantee.
This public notebook contains curated results, protocols, figures, and source identities. The full training workspace and checkpoints are not distributed through the website. Upstream credit: Kimodo, ARDY, and GEAR-SONIC, developed by NVIDIA and their collaborators.
NEXT / REMAINING WORK
What would change the conclusion?
Finish the current evidence chain before making a controller claim. Each step below names the missing measurement and which contribution it serves; hardware, onboard estimation and environmental transfer remain later scopes.
COMPLETE / LIVE CHECK
The 24-program allocation is closed.
Two reuses plus 22 new captures reproduce the replay gain. The interrupted original attempt stays unavailable; a single declared recovery completes coverage. Both final study audits pass.
2 / PREDICTION
Validate repair-effect prediction.
Use separate development variation to measure STOP residuals, wrong-sign effects and false admissions. Keep the rejected offset calibration as a negative result; separate margin calibration from evaluation.
3 / REFERENCE GEOMETRY
Repair stopping and progress jointly.
The 18-proposal envelope ablation is complete: earlier holds can lose position, and no candidate predicts complete-task admission. Next, preserve timed progress while varying braking or terminal motion; validate any full-task proposal in a registered matched live comparison.
4 / LATENCY
Let physics continue during compute.
Measure readiness, activation age, stale-update rejections and missed deadlines through repeated updates. Preserve the original task and consumed error budget.
5 / FINAL EVALUATION
Freeze before opening final36.
Fix the method, baselines, compute budget and analysis first. Report program-level paired outcomes and uncertainty, including every failure and quality regression. This is parameter variation, not unseen-family transfer.
6 / MANUSCRIPT
Complete author review.
Align the final claims with prospective results, credit prior work, complete the factual AI-assistance inventory and check the actual submission PDF. The historical PDF-test receipt does not cover the new method draft.
What a method contribution still needs
One frozen, executable composition. There is currently no single script that runs the complete method end to end; its own claim table records "no current evidence" for the composed method.
A registered evaluation on the 36 reserved programs, at two nested simulator seeds reported separately, with every regression printed. Roughly 150 captures, about two hours of simulator wall time.
Control computed under a measured reference-age distribution rather than a paused clock.
A forecast whose predicted intervention effect has the right sign at the decision boundary.
What a system or benchmark contribution still needs
The language benchmark is fully specified and fully unrun: 1,440 designed trials with no results, and 660 authoring slots with no collected instructions. Its blocker is host memory, not GPU memory.
The 1,000-clip corpus collapses to 129 source groups — one holding 827 clips — leaving 14 independent validation clips. The split does not survive its own dependence audit.
An installable release: the evidence tree is 6.8 GB and excluded from version control, and 55 scripts hard-code absolute local paths.
A resolved licence chain for generated clips derived from proprietary motion capture.
Judgement. The finished, defensible contribution today is the diagnostic account — the benchmark with three clocks, the error anatomy, the selection headroom with its failed learned comparators, the held-out interventions, and the native-planner baseline. The repair method is real ongoing work and is not yet a controller claim. A combined system-and-dataset paper is not reachable from the current evidence, because the benchmark reports no results yet.