# KimoNav research review — 9 September 2026

**Current conclusion:** whole-body reference repair can change simulated task outcomes, but a reliable execution-aware controller has not been demonstrated. Unconditional terminal repair has no net task-success gain. A favorable saved-outcome selection result does not establish live control, and the latest live pelvis-offset experiment exposes a wrong-sign prediction of the intervention effect.

The active local paper is a method draft; its public summary is [paper-summary.md](paper-summary.md). The earlier diagnostic manuscript remains a separate historical argument. The full manuscript is not distributed here. All execution evidence below is **flat-ground G1 simulation**, not hardware. This review runs no new generation, fitting, optimization, or physics.

## Completed evidence

| Study | Unit and allocation | Result | What it establishes |
| --- | --- | --- | --- |
| Original-24 terminal repair | 24 matched development programs; one simulator seed; two additional nulls excluded | Current 8/24 → repair 8/24 complete tasks; two rescues and two losses. STOP-speed compliance 12/24 → 16/24, full STOP phase 9/24 → 9/24. Runtime quality 14/24 → 11/24. | Repair affects execution; improved stopping alone does not imply improved navigation. |
| Minimal-intervention governor replay | Same 24 inspected programs, two saved reference outcomes per program; zero new captures | 10/24 selected successes versus 8/24 current and unconditional repair. Selects two repairs, retains current 22 times; avoids both observed losses. | A post-hoc development selection result. The rule was designed after inspecting the outcomes. |
| Nominal execution forecast | 48 forecasts over 24 paired development programs | At 1.5 s, XY RMSE 0.02050 m on 23 fully observed pairs; at 2 s, 0.02256 m on 13. STOP-effect direction agrees on 17/23, including one identity fallback. | Useful average response prediction, with consequential errors near the stopping threshold. |
| Live pelvis-reference pilot | Three new captures on two inspected strafe–walk 5 s programs; two governor arms and one right zero-offset control | Governor 0/2 successes; zero-offset control 0/1. All three fail STOP speed. | Fresh own-history construction and selection execute; no successful support is added by this pilot. |
| Causal joint-default offset model | 24 programs, 48 slots: 24 new forecasts for 12 accepted fits, 24 exact nominal reuses | 1 s XY RMSE worsens 0.016395 → 0.016822 m; 2 s 0.022559 → 0.022585 m. Speed MAE worsens at all five horizons. | Better past-action fit did not improve multi-step prediction overall; offsets are not adopted. |
| Earlier selection repeat completion | Same 12 historical designs at simulator seeds 1701/1702/1703; extension adds 18 captures and reuses six | Fixed / first / minimum jerk: 2/3/4, 2/1/2, 1/5/3, each out of 12. | Minimum jerk changes first-candidate successes by +1, +1, −2. These consumed designs are not a fresh method holdout. |

Original-24 terminal repair retains one actual early termination with unobserved STOP in its all-program denominator. Two coupled-optimization returns improve STOP-speed compliance without yielding complete task successes; the two task rescues use the older sequential branch. One rescue, right pivot–walk 10 s, succeeds where both previously executed four-candidate banks had no success. This is evidence of added executable support on **one inspected program**, not broad support expansion or evidence that the coupled branch caused the rescue.

Sources: terminal result, governor replay, nominal forecast, live pilot, offset comparison, and repeat extension.

## The most consequential new result

| Right 5 s reference | Predicted STOP peak (m/s) | Executed STOP peak (m/s) | Complete task |
| --- | ---: | ---: | --- |
| Current, qualified null | — | 0.221494 | Fail |
| Zero-offset repair | 0.118264 | 0.150212 | Fail |
| Selected +20 mm world-X repair | 0.096005 | 0.158513 | Fail |

The strict threshold remains **below 0.10 m/s throughout the original STOP interval**. The offset's predicted advantage over zero is −0.022260 m/s; its measured effect is +0.008301 m/s. The selected peak is underpredicted by 0.062508 m/s. The paired right arms have identical fresh candidate banks, forecast arrays, and pre-update state histories, so this comparison isolates a wrong-sign selection decision within the tested bank. Both repairs improve stopping relative to current, but neither completes the task.

The left governor retains current; all 33 recorded state/action arrays exactly reproduce its null. There are no falls, early terminations, or missing captures in this three-trial pilot. Worker time is 21.46–24.91 s and total runtime preparation is 23.28–27.54 s, with physics paused. Immediate activation after a pause is not a measured real-time capability. [Original-clock waveform](terminal_pelvis_live.png).

## Analysis and claim boundaries

1. **Geometry and execution remain different problems.** Foot/joint/rate acceptance does not certify stopping or full-task success. Sequential and coupled branches must be reported separately when attributing gains.
2. **Prediction error at the decision boundary matters more than average trajectory error.** The default-offset model leaves five predicted-task-pass/actual-fail slots across three programs, and six predicted-STOP-pass/actual-fail slots across four programs. The STOP denominator is 46 observed arms; the full task denominator is 48, including the terminated program. These are observed errors, not calibrated probabilities.
3. **The 10/24 replay is development evidence.** It uses outcomes already inspected while choosing the rule. The live pelvis experiment changes the action bank, so it tests transfer of nominal admission to new repairs; it does not directly re-execute the original two-choice 24-program rule.
4. **Late repair cannot recover a failed past sample.** Seven original-24 programs already violate the task before the terminal event. More aggressive terminal geometry alone cannot rescue their full original-clock result.
5. **The prospective contribution remains narrow.** Contact-aware reference filtering, stopping prediction and reference governors have precedents. The research opportunity is demonstrated repair of an immutable timed task under retained error budgets and actual execution delays. None of those precedents validates this implementation's guarantees. See the primary-source comparison.

## Incomplete work at the review snapshot

The original-24 live binary-governor allocation is **incomplete**: three start records and two qualified execution records are present; slot 2 has no terminal ledger record and 21 slots have no start record. There is no completion receipt and no matching active experiment process was found during inspection. The reason for the interrupted collection is not established. Do not report an all-24 success count or count a start record as a completed capture. The [snapshot JSON](progress-2026-09-09.json) binds the exact files inspected.

The reference-envelope proposal specifies 18 geometry attempts varying a reference prior and settling time. It is a design, not a result. It must retain the executed task's 0.20 m position tolerance and original STOP clock. No new generation or access to the final 36 programs occurs in this review; existing records report that set sealed. Those 36 parameter-variation programs are distinct from the earlier consumed 12-design confirmation set.

## Remaining work, in order

| Priority | Work | Evidence needed to close it |
| --- | --- | --- |
| 1 | Reconcile the partial binary live allocation and its unfinished slot under the frozen protocol. Preserve first attempts and failures; version any necessary repair. | A complete 24-slot accounting, causal-input and delivered-reference checks, and an aggregate audit. Missing slots remain missing until actually executed. This resolves whether the saved two-choice rule reproduces live. |
| 2 | Diagnose prediction of intervention differences on separate development variation. Keep the rejected offset model as a negative result. | Paired current/zero/offset outcomes, STOP-peak residuals, wrong-sign and false-admission rates, censoring and full-task scores. Any empirical margin must be chosen separately from its evaluation set. This tests whether selection errors can be reduced without erasing useful repairs. |
| 3 | Test one registered geometry/timing ablation after the current collection is reconciled. | All declared proposals and rejection reasons, then matched executions of qualified variants; branch-specific effects and whole-task success. A looser reference prior must not change task tolerances. |
| 4 | Qualify computation during continuing simulation and repeated updates. | Measured ready/activation times, stale-update rejection, deadline misses, persistent state and original budgets. Report success with actual delay; compare compute-matched baselines. Paused 23–28 s preparation is insufficient. |
| 5 | Freeze the method, comparator budgets, metrics and analysis before opening final36. | Prospective program-level paired evaluation with all failures, quality regressions and intervals retained. Mirrored pairs and repeats remain dependent units. No final-panel tuning or unseen-family claim from parameter variation. |
| 6 | Finish the scientific manuscript and author review. | Reconcile all claims with the frozen evaluation, credit upstream systems and related work, complete the factual AI-assistance inventory, and perform a new PDF/venue check on the actual submission candidate. The historical PDF-test receipt does not cover the method draft. |

Hardware, onboard estimation, obstacles, payloads and broad transfer remain later research scopes. They are not demonstrated by the present simulation experiments.

## Verification for this update

The review exporter recounts the 24 task pairs after excluding two nulls, the two gains/two losses, all 24 saved governor selections, the three live outcomes, and the signed offset effect. It extracts the full-horizon model metrics, preserves varying denominators, and records source SHA-256 hashes. It does not rerun the prior independent auditors or raw simulator trajectories. Reproduce with `python3 knav/scripts/build_progress_review_20260909.py`; rerunning produces a newly timestamped snapshot, not a replacement experimental registration. Original ledgers and failed attempts remain unchanged.
