Systematic Multi‑Agent Vision‑and‑Language Navigation

Formulation, Benchmark, and Method

Yunzhe Xu  ·  Zhe Liu

Shanghai Jiao Tong University

Abstract

Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent’s exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, navigation, and execution, highlighting the challenges of coordinating under MAVLN task constraints.

Example episode

Three agents prepare a room for evening reading. The mission has six subtasks connected by a dependency graph, a presence lock and a holding chain, and is completed in two scheduling turns.

Agent 1Agent 2Agent 3Top-down map

Panels show each agent’s egocentric view, followed by the top-down map. Trail colors on the map encode elapsed steps, not agent identity.

Subtask dependencies

Turn 1 of 2
Six-subtask dependency graph for evening reading Agent 1 holds the grill cover at task 1 until Agent 2 fastens it at task 2. Task 2 releases the lock and enables task 5, switch on the lamp, and task 6, organize the sofa. Agent 3 retrieves books at task 3 and carries them to task 4. presence lockS2 releases Agent 1samecarrier 1Hold the grill coverMaintain presence until S2 2Fasten the strapsRelease the presence lock 3Retrieve the booksCarry to the coffee table 4Place the booksComplete the holding chain 5Switch on the lampAfter S2 6Organize the sofaAfter S2 Solid: prerequisite · Dashed: holding chain
Agent 1: cover, then sofa Agent 2: straps, then lamp Agent 3: books, then table

Turn 1: hold, fasten, retrieve

Agent 1 holds the cover. Agent 2 fastens it after the hold is established, releasing Agent 1. Agent 3 retrieves the books.

Turn 1 firing order: subtask 1, then 2, then 3.
Firing order S1 → S2 → S3. S2 releases the presence lock created by S1.

The dependency graph admits several valid orders; this is the order chosen in this episode.

Instruction regimes

The same mission is rendered in three styles: one instruction per agent; one instruction for the whole team; and one team instruction without agent attribution.

Show the instruction

Task

Each mission is a directed acyclic graph over subtasks. A subtask names a goal object instance, the agents permitted to execute it, its prerequisites, and resource attributes. Two resource primitives restrict what an agent may do while they are active. Under a holding chain, the agent carries an object and may complete nothing else until it delivers it. Under a presence lock, the agent must remain at its goal until a teammate releases it. Dependencies govern when subtasks may complete; resource constraints govern what an agent may do in the meantime. A mission succeeds only if every subtask succeeds under these semantics, so a team can reach every destination and still fail.

Solving the task therefore requires three capabilities beyond single-agent instruction following:

Dynamic subtask scheduling
Identify dependency and resource constraints from language and dispatch subtasks over successive rounds.
Simultaneous experience sharing
Reuse what teammates have already explored to plan in space an agent has not seen itself.
Inter-agent conflict resolution
Turn simultaneous navigation intentions into collision-free routes.
Comparison of embodied navigation and task planning with systematic multi-agent VLN, including scheduling, experience sharing, and conflict resolution.
Figure 1. Systematic multi-agent VLN combines constraint-aware scheduling, experience sharing, and conflict resolution in unseen environments.

The MAVLN benchmark

Episodes
11,724
Scenes (HM3D)
145
Avg. path length
35.9 m
Agents per team
up to 4

MAVLN is built on Habitat 3 with HM3D scenes. A four-stage pipeline grounds scene descriptions at the viewpoint and zone level, synthesizes missions from household profiles, grounds each subtask to a destination viewpoint and optimizes the makespan, and finally verifies and renders the instructions. Rendered instructions pass a closed-loop filter that reconstructs the arrival order from language and checks it against the ground-truth constraints.

The 3,908 unique missions are each rendered in three instruction regimes. Episodes carry five subtasks on average and up to eight. Holding chains appear in 92.9% of evaluation missions and presence locks in 28.3%, with 26.3% exercising both.

Scene-disjoint splits.
SplitMissionsScenes
Train3,079128
Val-unseen3867
Test-unseen44310
Four-stage MAVLN crafting pipeline from scene understanding to verified instruction rendering.
Figure 2. The crafting pipeline: (a) grounded scene understanding, (b) mission synthesis, (c) mission refinement with makespan optimization, and (d) verified instruction rendering in three styles.

TRISS

TRISS is a coordination-ready navigation system with three modules.

LLM-based subtask scheduler. An extractor turns the task instructions into a shared pool of atomic navigation instructions. A scheduler then assigns at most one instruction per agent in each round, reasoning about the constraints stated in language. Agents may be dispatched ahead of an unmet dependency and wait for an explicit firing signal, which overlaps travel with waiting.

Shared topological memory. All agents build one topological map on the fly. New viewpoint proposals are localized against the shared node set, so duplicate proposals from different teammates merge rather than fork. A DUET-style cross-modal planner reasons over this map with each agent’s own instruction.

Conflict-aware execution. Target viewpoints are allocated by bipartite matching so that no two agents choose the same destination. Conflict-based search then plans collision-free routes over the shared graph. Failure-verified exclusion and traversal-verified connection keep failed or drifting executions from corrupting the map that every teammate relies on.

TRISS architecture: atomic instruction extraction and scheduling, a navigation planner over shared topological memory, and conflict-aware execution.
Figure 3. (a) Atomic instructions are extracted and allocated by the scheduler. (b) The planner predicts viewpoint-level actions over the shared topology for each agent. (c) Conflict-free target allocation and path finding, with topology maintenance during execution.

Results

With oracle scheduling, TRISS raises mission success over an ETPNav navigator trained on MAVLN from 2.8 to 9.1 SR on val-unseen and from 5.6 to 9.3 on test-unseen. With LLM schedulers, mission success stays in the single digits in every instruction regime, and the remaining headroom is spread across scheduling, navigation and execution rather than concentrated in one module.

Test-unseen results. LLM results are means over three runs.
SchedulerNavigatorSR ↑SPL ↑CSR ↑CSPL ↑TC ↑MAC ↓
Ground-truth atomic instructions
OracleETPNav5.64.228.722.4100.00.85
OracleTRISS9.35.233.320.0100.01.42
RandomTRISS0.20.16.74.1100.01.12
SequentialTRISS2.51.221.812.6100.01.28
LLM scheduler, by instruction regime
Gemma4 · decentralizedTRISS8.34.433.219.595.81.22
Qwen3.6 · decentralizedTRISS7.94.132.118.797.41.26
Gemma4 · centralizedTRISS6.83.931.618.695.01.13
Qwen3.6 · centralizedTRISS8.44.232.419.196.51.03
Gemma4 · centralized-implicitTRISS7.13.729.617.092.80.97
Qwen3.6 · centralized-implicitTRISS6.83.528.816.596.51.03
Full results on val-unseen and test-unseen
SchedulerNavigatorVal-unseenTest-unseen
SR ↑SPL ↑CSR ↑CSPL ↑ISPL ↑TC ↑TS ↓MAC ↓SR ↑SPL ↑CSR ↑CSPL ↑ISPL ↑TC ↑TS ↓MAC ↓
Oracle (SA)Oracle67.163.174.968.591.6100.04120.0069.164.478.070.890.4100.04190.00
Oracle (LE)Oracle100.099.7100.099.592.6100.02611.16100.099.6100.099.492.0100.02481.39
OracleOracle100.099.8100.099.692.6100.02151.75100.099.6100.099.492.0100.02062.10
OracleETPNav2.81.927.221.630.5100.03940.875.64.228.722.430.6100.04030.85
Oracle (SA)TRISS5.43.326.015.325.2100.020560.005.93.825.217.427.7100.016650.00
OracleTRISS9.14.036.419.726.9100.015880.829.35.233.320.027.9100.014941.42
RandomTRISS0.30.17.94.323.6100.016210.910.20.16.74.124.1100.015471.12
SequentialTRISS3.41.623.612.824.9100.015930.902.51.221.812.626.5100.014431.28
Decentralized
Gemma4TRISS7.0 ± 0.92.9 ± 0.433.6 ± 0.918.2 ± 0.624.0 ± 0.494.6 ± 0.51570 ± 370.95 ± 0.18.3 ± 0.14.4 ± 0.033.2 ± 0.319.5 ± 0.326.5 ± 0.295.8 ± 0.21447 ± 131.22 ± 0.0
Qwen3.6TRISS7.3 ± 0.72.8 ± 0.334.7 ± 0.418.2 ± 0.124.9 ± 0.397.0 ± 0.31613 ± 590.87 ± 0.17.9 ± 0.84.1 ± 0.432.1 ± 0.318.7 ± 0.225.9 ± 0.497.4 ± 0.21521 ± 251.26 ± 0.1
Centralized
Gemma4TRISS7.5 ± 1.23.1 ± 0.534.3 ± 1.018.1 ± 0.523.9 ± 0.594.4 ± 0.51570 ± 400.83 ± 0.06.8 ± 0.13.9 ± 0.131.6 ± 0.118.6 ± 0.125.4 ± 0.195.0 ± 0.21475 ± 101.13 ± 0.0
Qwen3.6TRISS7.6 ± 0.13.1 ± 0.234.4 ± 0.318.2 ± 0.224.0 ± 0.495.7 ± 0.11619 ± 290.85 ± 0.08.4 ± 0.94.2 ± 0.532.4 ± 0.819.1 ± 0.726.3 ± 0.596.5 ± 0.41470 ± 71.03 ± 0.0
Centralized-implicit
Gemma4TRISS5.5 ± 0.12.3 ± 0.132.1 ± 0.816.5 ± 0.222.5 ± 0.292.5 ± 0.41611 ± 440.76 ± 0.07.1 ± 0.33.7 ± 0.229.6 ± 0.617.0 ± 0.523.3 ± 0.392.8 ± 0.41512 ± 370.97 ± 0.1
Qwen3.6TRISS6.2 ± 0.62.7 ± 0.331.5 ± 0.216.6 ± 0.023.2 ± 0.296.8 ± 0.21662 ± 160.78 ± 0.06.8 ± 0.73.5 ± 0.428.8 ± 0.616.5 ± 0.524.1 ± 0.596.5 ± 0.51562 ± 261.03 ± 0.1

Bold marks the best LLM scheduler per column. SA: single agent. LE: legacy scheduling without pre-allocation. Random and sequential schedulers dispatch ground-truth atomic instructions to free agents.

Metric definitions
SR
Fraction of missions in which every subtask succeeds under the task’s constraint semantics.
SPL
SR weighted by the ratio of the optimal path length to the total path traversed by all agents.
CSR
Fraction of subtasks whose arrival lies within the success radius while honoring all dependency, locking and holding constraints.
ISPL / CSPL
Per-subtask path efficiency, gated by success without and with task constraints, respectively.
TC
Fraction of subtasks the team consumes, irrespective of whether the arrival is correct.
TS
Synchronized simulation steps until termination: the realized makespan, including congestion, waiting and deadlock.
MAC
Time steps in which any two agents lie within twice the agent radius of each other, divided by TS times team size.

SR, SPL, CSR, CSPL, ISPL and TC are percentages. Arrivals are credited within 3 m of any navigable anchor of the target instance.

Qualitative example

A second episode, preparing for a movie night, spans five subtasks and two scheduling rounds. In the first round, agents A and B reach the refrigerator and the coffee machine and enter holding states for the snacks and drinks, while agent C reaches the couch. In the second round, A and B are assigned the side table and the dining table and complete the mission, keeping their assignments across the holding chains.

Three agents completing five subtasks across two scheduling rounds; A and B carry snacks and drinks to their assigned tables.
Figure 4. Scheduling and navigation process of TRISS. Flags mark subtask goals in each agent’s color.

Citation

    @article{xu2026mavln,
        title={Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method},
        author={Xu, Yunzhe and Liu, Zhe},
        journal={arXiv preprint arXiv:2609.35965},
        year={2026}}

Contact: xyz9911@sjtu.edu.cn.