IROS 2026

Talk2Escape: Conversational Grounding for Vision-and-Language Navigation

Zerui Li1†, Sihao Lin1†, Yanyan Shao2, Jiwen Zhang3, Xiangyu Shi1, Shijie Li4, Qi Wu1*

1Australian Institute for Machine Learning, Adelaide University  ·  2Zhejiang Wanli University  ·  3Fudan University  ·  4A*STAR

†Joint first authors   *Corresponding author   ·  Sihao Lin and Qi Wu are also with the Responsible AI Research Centre, AIML.

Talk2Escape on a Unitree Go2 — the robot detects its own failure, asks for help, and escapes.
TL;DR

Navigation agents shouldn't wander until they fail. They should ask for help.

Single-turn VLN agents run open-loop: one instruction, no way to verify progress or recover once sensor noise, odometry drift, or perceptual aliasing push them into a lost state. Talk2Escape (T2E) is a model-agnostic dialogue intervention module. A lightweight kinematic monitor watches for localized looping and trajectory divergence; when a trigger fires, a vision-language translator turns raw egocentric views into a concise grounded query, and a corrective hint from an oracle or human guide is injected back into the agent's prompt. No fine-tuning, no learned dialogue policy — pure zero-shot MLLM reasoning, on simulated benchmarks and on a physical Unitree Go2.

66.0%
SR on R2R-CE, above the best supervised SOTA — fully zero-shot
+17.2
absolute SR gain over the best open-loop zero-shot method (GTA)
84.6%
SR on VLNVerse with a vanilla NavGPT agent (up from 19.3%)
62.0%
real-world SR on a Unitree Go2 quadruped
Motivation

Open-loop execution is an architectural failure mode

When people get lost in an unfamiliar building, they don't keep wandering — they describe what they see and ask a guide. This instinctive error-recovery loop is exactly what standard VLN frameworks lack. Prior dialogue-driven VLN work treats conversation as supplementary planning context inside simulators that assume perfect communication and zero physical execution error. T2E instead treats dialogue as an operational safety net for real deployment.

Paradigm shift from single-turn to failure-aware dialogue-driven VLN
Paradigm shift. Left: single-turn agents execute one-off instructions open-loop, architecturally brittle to physical uncertainty. Right: T2E closes the loop with proactive, kinematic-aware help-seeking.
Method

Monitor → trigger → ask → hint → re-plan

Overview of the Talk2Escape framework
The Talk2Escape framework. A kinematic monitor evaluates the base agent's spatial state at every step; on failure it pauses the robot, grounds the current observation into an active query, and injects the guide's corrective hint into the MLLM prompt as a Navigation Assistance block.
Results

Zero-shot, above the supervised ceiling

On R2R-CE Val-Unseen (100-episode protocol of Open-Nav), T2E + GTA reaches 66.0% SR — past the strongest supervised method (Efficient-VLN, 64.2%) without a single in-domain training trajectory. Even a vanilla NavGPT base jumps to 60.0% SR.

MethodTypeNE ↓OSR ↑SR ↑SPL ↑
Efficient-VLNSupervised4.1873.764.255.9
NavFoMSupervised4.6172.161.755.3
GTAZero-shot4.9556.248.841.8
VLN-ZeroZero-shot5.9751.642.426.3
Talk2Escape + NavGPTZero-shot4.8460.064.046.4
Talk2Escape + GTAZero-shot4.8066.072.049.4

R2R-CE Val-Unseen, 100 sampled episodes following the Open-Nav protocol. Full RxR-CE and complete baselines in the paper.

Dialogue beats internal reasoning by a wide margin

On the high-fidelity VLNVerse benchmark, passive Chain-of-Thought lifts NavGPT by only 9.4 points; active dialogue lifts it by 65.3.

AgentNE ↓SR ↑OSR ↑SPL ↑nDTW ↑
NavGPT6.1019.3061.508.1352.60
NavGPT + CoT5.2528.7050.509.9144.10
NavGPT + T2E2.9284.6088.3035.2060.50
MapGPT5.6225.5355.327.5743.38
MapGPT + CoT4.5142.1972.4011.3245.20
MapGPT + T2E4.1467.0274.4730.4653.62

Fine-grained instruction task on VLNVerse.

Real-World Deployment

From simulator to a Unitree Go2

We deploy T2E on a Go2 quadruped with a servo-mounted depth camera sweeping orthogonal views, decoupling MLLM cognition from high-frequency locomotion. Under a strict human-in-the-loop protocol mirroring the deviation trigger, T2E reaches 62.0% SR against 40.0% for the strongest open-loop baseline.

Real-world execution of the Talk2Escape paradigm
Rescue in the wild. The agent misreads the instruction and heads for the wrong door; a grounded hint via dialogue corrects the route and the robot reaches the intended target.
MethodTypeSR ↑ (%)NE ↓ (m)
VLN-BERTSupervised16.05.36
RDPSupervised20.05.45
SmartWayZero-shot32.04.85
GTAZero-shot40.03.66
Talk2Escape + GTAZero-shot62.03.31
Key Findings

When to speak matters as much as what to say

strategic sparsity

The deviation trigger is the best autonomous strategy (SR 66.0, SPL 49.4). The hybrid trigger fires too often (23.1% of steps) and drops to SR 54.0 — excessive querying disrupts the agent's contextual memory and causes "intervention fatigue".

agent capability

Weaker agents degrade linearly as guidance gets sparser; a strong agent like GTA shows an inverted U-shape — too-frequent polling interferes with its internal planning, too-sparse polling deprives it of rescue.

modality mismatch

MapGPT's dense absolute 3D coordinates conflict with the oracle's egocentric hints; a clean, action-centric history aligns far better with natural-language corrections than over-engineered geometric prompts.

asymmetric context

Human guides receiving only static frames struggle to reconstruct the robot's true orientation — the "asymmetric context" bottleneck. Future work: spatio-temporal video buffers and multi-view memory in the dialogue state.

Citation

BibTeX

🔖BibTeX coming soon — arXiv preprint in preparation