Literature that changes the traversal plan
Checked 18 September 2026. This is a targeted primary-source reassessment, not an exhaustive survey or a guarantee of novelty. Full-text access below means the relevant method, evaluation and limitation sections were inspected; it does not mean an external method was reproduced. Paper claims are attributed to their authors. Our design decisions are inferences.
The closest work now
| Primary source and inspected scope | What is established in that source | Decision for KimoNav |
|---|---|---|
| TANGO, 2609.09158v1, §§3–5 and deployment/metrics appendices; full text | RGB, language and proprioception condition whole-body reference chunks executed by SONIC. Its Plan–Edit–Track pipeline filters references through physical execution; action labels remain the edited references, not the tracked states. It also compares another tracker and identifies RGB-only perception and tracker capability as limitations. | Direct overlap with our broad destination. Compare actual reference inputs and physical task definitions. “Whole-body VLA,” frozen-tracker composition, and swapping a tracker are insufficient novelty. Its outputs are references for a tracker, not direct motor torques. |
| PASSAGE, 2609.18732v1, §§III–V and map/runtime appendices; full text | Destination and multi-layer geometric observations condition a flow planner and perceptive tracker. Scene-aligned data, chunk continuity and planner-side post-training support onboard traversal; goal completion and contact-free completion are distinguished. | Essential new comparator. Scene-conditioned planning, layered overhead geometry, frozen-tracker refinement and reusable interfaces are already occupied. Evaluate our actual remaining question: task preservation and reliable adaptation under controlled command/observation changes. |
| CAT / HumanoidPF, 2601.16035v1, §§III–IV; full text | Per-body geometric fields guide traversal learning; hybrid scenes cover ground, lateral and overhead obstacles. Deployment uses LiDAR–inertial mapping and target selection. | Strong geometry/control comparator. Match sensing and motor capabilities before attributing a result to our interface. Per-body features and 3D clearance are established tools. Official repository was inspected for availability, not executed. |
| DWMP, 2609.12347v1, method and experimental sections; full text | Separate dynamics and depth world-model representations feed a traversal policy, with simulation and G1 demonstrations. | A predictive visual representation is a plausible later comparison, not a necessary first component. Compare with a simple geometry/history encoder before adopting two world models. |
These sources make a timing-only differentiator unconvincing unless a real application requires it. We should study an operationally meaningful failure, with matched baselines, rather than create a narrower score merely because existing systems were not evaluated on it. The proposed task/observation intervention study remains a hypothesis; this review does not establish that no prior work studies the same issue.
Motion, scene data, and contact
| Source and scope | Relevant finding | Design implication |
|---|---|---|
| Moving Through Clutter, 2603.05993v1, §§III–V; full text | VR collection couples geometry with whole-body traversal. The paper reports 348 trajectories across 145 scenes and discusses scene-agnostic retargeting and omitted contact-assisted progression. | Connect our program to this existing scene-aligned resource. Preserve human motion, retargeted reference and executed robot data as separate levels. Reuse accessible assets before building another scene generator. |
| SceneBot, 2606.27581v1, §§3–6; full text | Motion-conditioned scene reconstruction and contact labels train a scene-interaction tracker. Its conclusion leaves autonomous perceptive goal-driven control to future work. | Strong contact-aware motor and data precedent. A task-level planner remains a different obligation. New support contacts require motor training or qualification, not only collision-free animation. |
| Perceptive BFM, 2606.08059v1, synthesis, distillation and limitations; full text | Terrain-conformal reference synthesis and perceptual distillation adapt supplied motion. Its stated foot-centric assumptions can leave upper-body collisions. | Use as a terrain-adaptation comparison. Preserving upper-body style is unsuitable when shoulder/head clearance requires changing that style. |
| Learning from Hallucination family, author project synthesis | Open-space navigation experience is paired with constructed obstacle configurations; later variants change hallucination and deployment requirements. | Motion-first scene construction is established. Our data branch must test whole-body feasibility, alternative actions and downstream learning against scene-first or existing scene-aligned data. |
| SUMMON, INFERACT, SceMoS, MotionBricks | Previously reviewed in the September 15 hindsight/tokenization notes; not independently re-read in this revision. | Keep as secondary prior art for scene reconstruction, motion vocabularies and root/body factorization. Verify the relevant full-text passages before a specific novelty claim. |
What the behavior-model papers do and do not imply
| Source and inspected scope | Relevant finding | Design implication |
|---|---|---|
| Behavior Foundation Model, 2509.13780v1, §§III–V; full text | A tracking proxy, masked geometric control, CVAE and online student-state supervision provide reusable control. | These are methodological precedents. Our offline context masking and limited stored queries are not a full reproduction. Judge the public prior independently of privileged reconstruction. |
| Scaling BFM, 2607.15163v1, architecture, scaling and deployment sections; full text plus existing source audit | Time-offset motion targets and observation/action histories support global whole-body control; motion diversity and rollout quantity are distinct. | Time tokens and cross-attention are established. Repeating the same clips cannot replace missing body/scene behavior. Relative features still need an explicit localization contract. |
| BFM-Zero, 2511.04131v1, formulation and inference; full text plus September 15 code audit | Forward–backward learning supports objective-conditioned behavior and inference through a learned representation. | A separate pretraining route. Its latent is not SONIC code64, and importing its name does not give our student zero-shot properties. Reconsider only against a specific task and compute budget. |
| SONIC, 2511.07820v4, control/planning and deployment sections; full text | A motion tracking system and kinematic planning interface provide reusable humanoid motion execution. | Keep both original and repaired checkpoints as explicit baselines. Reproduce the actual reference context, initialization and controller implementation; adapted pose replay does not establish full native deployment parity. |
| ADAPT, 2609.00677v1, §§3–4; full text | Closed-loop text-driven control, robustness adaptation and goal steering include a reaching-and-stationary objective. | Neither text control nor stopping is new. Compare command semantics, allowed observations and execution latency. Our strict historical timing results are not directly comparable with its goal criterion. |
| MaskedMimic, HOVER, CLAW, Humanoid-LLA, GenTrack, RTC | Covered in existing September 8/15 primary-source reviews; those local reviews were inspected here. No new full-text verification in this revision. | Retain as comparator families for sparse control, parameterized planning, language and execution-aware generation. Check implementation and availability before selecting an experimental baseline. |
A focused next reading pass
- Task and runtime contracts: inspect released TANGO/PASSAGE/CAT code, if available, for goal acceptance, body/reference representation, collision accounting, recovery, sensor delay and initialization. Do not treat a project page as proof that a runnable checkpoint exists. This revision does not certify the release completeness of TANGO or PASSAGE.
- Task preservation under partial observation: seek body-aware planning and constrained/uncertainty-aware navigation methods with explicit false-admission, abstention and intervention outcomes. Include conventional search/MPC and geometric checks, not only learned policies.
- Supervision that distinguishes decisions: compare scene-aligned demonstration, scene-first generation, and motion-first augmentation at matched real behavior coverage. Ask whether changing an obstacle changes the desired action under a fixed destination.
- Compression only after capability: revisit online masked distillation and multimodal prediction after a planner–tracker provides useful actions. Test whether compression improves latency/data cost while retaining task and contact outcomes.
Stop a reading branch when it resolves the design decision. Continue the nearest-work search before a paper claim, especially because the September 2026 traversal literature has changed the comparison set substantially.