1Australian Institute for Machine Learning, Adelaide University · 2Zhejiang Wanli University · 3Fudan University · 4A*STAR
†Joint first authors *Corresponding author · Sihao Lin and Qi Wu are also with the Responsible AI Research Centre, AIML.
Single-turn VLN agents run open-loop: one instruction, no way to verify progress or recover once sensor noise, odometry drift, or perceptual aliasing push them into a lost state. Talk2Escape (T2E) is a model-agnostic dialogue intervention module. A lightweight kinematic monitor watches for localized looping and trajectory divergence; when a trigger fires, a vision-language translator turns raw egocentric views into a concise grounded query, and a corrective hint from an oracle or human guide is injected back into the agent's prompt. No fine-tuning, no learned dialogue policy — pure zero-shot MLLM reasoning, on simulated benchmarks and on a physical Unitree Go2.
When people get lost in an unfamiliar building, they don't keep wandering — they describe what they see and ask a guide. This instinctive error-recovery loop is exactly what standard VLN frameworks lack. Prior dialogue-driven VLN work treats conversation as supplementary planning context inside simulators that assume perfect communication and zero physical execution error. T2E instead treats dialogue as an operational safety net for real deployment.
On R2R-CE Val-Unseen (100-episode protocol of Open-Nav), T2E + GTA reaches 66.0% SR — past the strongest supervised method (Efficient-VLN, 64.2%) without a single in-domain training trajectory. Even a vanilla NavGPT base jumps to 60.0% SR.
| Method | Type | NE ↓ | OSR ↑ | SR ↑ | SPL ↑ |
|---|---|---|---|---|---|
| Efficient-VLN | Supervised | 4.18 | 73.7 | 64.2 | 55.9 |
| NavFoM | Supervised | 4.61 | 72.1 | 61.7 | 55.3 |
| GTA | Zero-shot | 4.95 | 56.2 | 48.8 | 41.8 |
| VLN-Zero | Zero-shot | 5.97 | 51.6 | 42.4 | 26.3 |
| Talk2Escape + NavGPT | Zero-shot | 4.84 | 60.0 | 64.0 | 46.4 |
| Talk2Escape + GTA | Zero-shot | 4.80 | 66.0 | 72.0 | 49.4 |
R2R-CE Val-Unseen, 100 sampled episodes following the Open-Nav protocol. Full RxR-CE and complete baselines in the paper.
On the high-fidelity VLNVerse benchmark, passive Chain-of-Thought lifts NavGPT by only 9.4 points; active dialogue lifts it by 65.3.
| Agent | NE ↓ | SR ↑ | OSR ↑ | SPL ↑ | nDTW ↑ |
|---|---|---|---|---|---|
| NavGPT | 6.10 | 19.30 | 61.50 | 8.13 | 52.60 |
| NavGPT + CoT | 5.25 | 28.70 | 50.50 | 9.91 | 44.10 |
| NavGPT + T2E | 2.92 | 84.60 | 88.30 | 35.20 | 60.50 |
| MapGPT | 5.62 | 25.53 | 55.32 | 7.57 | 43.38 |
| MapGPT + CoT | 4.51 | 42.19 | 72.40 | 11.32 | 45.20 |
| MapGPT + T2E | 4.14 | 67.02 | 74.47 | 30.46 | 53.62 |
Fine-grained instruction task on VLNVerse.
The deviation trigger is the best autonomous strategy (SR 66.0, SPL 49.4). The hybrid trigger fires too often (23.1% of steps) and drops to SR 54.0 — excessive querying disrupts the agent's contextual memory and causes "intervention fatigue".
Weaker agents degrade linearly as guidance gets sparser; a strong agent like GTA shows an inverted U-shape — too-frequent polling interferes with its internal planning, too-sparse polling deprives it of rescue.
MapGPT's dense absolute 3D coordinates conflict with the oracle's egocentric hints; a clean, action-centric history aligns far better with natural-language corrections than over-engineered geometric prompts.
Human guides receiving only static frames struggle to reconstruct the robot's true orientation — the "asymmetric context" bottleneck. Future work: spatio-temporal video buffers and multi-view memory in the dialogue state.