Systematic Multi‑Agent Vision‑and‑Language Navigation
Formulation, Benchmark, and Method
Shanghai Jiao Tong University
Abstract
Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent’s exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, navigation, and execution, highlighting the challenges of coordinating under MAVLN task constraints.
Example episode
Three agents prepare a room for evening reading. The mission has six subtasks connected by a dependency graph, a presence lock and a holding chain, and is completed in two scheduling turns.
Panels show each agent’s egocentric view, followed by the top-down map. Trail colors on the map encode elapsed steps, not agent identity.
Subtask dependencies
Turn 1 of 2Turn 1: hold, fasten, retrieve
Agent 1 holds the cover. Agent 2 fastens it after the hold is established, releasing Agent 1. Agent 3 retrieves the books.
The dependency graph admits several valid orders; this is the order chosen in this episode.
Instruction regimes
The same mission is rendered in three styles: one instruction per agent; one instruction for the whole team; and one team instruction without agent attribution.
Show the instruction
Task
Each mission is a directed acyclic graph over subtasks. A subtask names a goal object instance, the agents permitted to execute it, its prerequisites, and resource attributes. Two resource primitives restrict what an agent may do while they are active. Under a holding chain, the agent carries an object and may complete nothing else until it delivers it. Under a presence lock, the agent must remain at its goal until a teammate releases it. Dependencies govern when subtasks may complete; resource constraints govern what an agent may do in the meantime. A mission succeeds only if every subtask succeeds under these semantics, so a team can reach every destination and still fail.
Solving the task therefore requires three capabilities beyond single-agent instruction following:
- Dynamic subtask scheduling
- Identify dependency and resource constraints from language and dispatch subtasks over successive rounds.
- Simultaneous experience sharing
- Reuse what teammates have already explored to plan in space an agent has not seen itself.
- Inter-agent conflict resolution
- Turn simultaneous navigation intentions into collision-free routes.
The MAVLN benchmark
- Episodes
- 11,724
- Scenes (HM3D)
- 145
- Avg. path length
- 35.9 m
- Agents per team
- up to 4
MAVLN is built on Habitat 3 with HM3D scenes. A four-stage pipeline grounds scene descriptions at the viewpoint and zone level, synthesizes missions from household profiles, grounds each subtask to a destination viewpoint and optimizes the makespan, and finally verifies and renders the instructions. Rendered instructions pass a closed-loop filter that reconstructs the arrival order from language and checks it against the ground-truth constraints.
The 3,908 unique missions are each rendered in three instruction regimes. Episodes carry five subtasks on average and up to eight. Holding chains appear in 92.9% of evaluation missions and presence locks in 28.3%, with 26.3% exercising both.
| Split | Missions | Scenes |
|---|---|---|
| Train | 3,079 | 128 |
| Val-unseen | 386 | 7 |
| Test-unseen | 443 | 10 |
TRISS
TRISS is a coordination-ready navigation system with three modules.
LLM-based subtask scheduler. An extractor turns the task instructions into a shared pool of atomic navigation instructions. A scheduler then assigns at most one instruction per agent in each round, reasoning about the constraints stated in language. Agents may be dispatched ahead of an unmet dependency and wait for an explicit firing signal, which overlaps travel with waiting.
Shared topological memory. All agents build one topological map on the fly. New viewpoint proposals are localized against the shared node set, so duplicate proposals from different teammates merge rather than fork. A DUET-style cross-modal planner reasons over this map with each agent’s own instruction.
Conflict-aware execution. Target viewpoints are allocated by bipartite matching so that no two agents choose the same destination. Conflict-based search then plans collision-free routes over the shared graph. Failure-verified exclusion and traversal-verified connection keep failed or drifting executions from corrupting the map that every teammate relies on.
Results
With oracle scheduling, TRISS raises mission success over an ETPNav navigator trained on MAVLN from 2.8 to 9.1 SR on val-unseen and from 5.6 to 9.3 on test-unseen. With LLM schedulers, mission success stays in the single digits in every instruction regime, and the remaining headroom is spread across scheduling, navigation and execution rather than concentrated in one module.
| Scheduler | Navigator | SR ↑ | SPL ↑ | CSR ↑ | CSPL ↑ | TC ↑ | MAC ↓ |
|---|---|---|---|---|---|---|---|
| Ground-truth atomic instructions | |||||||
| Oracle | ETPNav | 5.6 | 4.2 | 28.7 | 22.4 | 100.0 | 0.85 |
| Oracle | TRISS | 9.3 | 5.2 | 33.3 | 20.0 | 100.0 | 1.42 |
| Random | TRISS | 0.2 | 0.1 | 6.7 | 4.1 | 100.0 | 1.12 |
| Sequential | TRISS | 2.5 | 1.2 | 21.8 | 12.6 | 100.0 | 1.28 |
| LLM scheduler, by instruction regime | |||||||
| Gemma4 · decentralized | TRISS | 8.3 | 4.4 | 33.2 | 19.5 | 95.8 | 1.22 |
| Qwen3.6 · decentralized | TRISS | 7.9 | 4.1 | 32.1 | 18.7 | 97.4 | 1.26 |
| Gemma4 · centralized | TRISS | 6.8 | 3.9 | 31.6 | 18.6 | 95.0 | 1.13 |
| Qwen3.6 · centralized | TRISS | 8.4 | 4.2 | 32.4 | 19.1 | 96.5 | 1.03 |
| Gemma4 · centralized-implicit | TRISS | 7.1 | 3.7 | 29.6 | 17.0 | 92.8 | 0.97 |
| Qwen3.6 · centralized-implicit | TRISS | 6.8 | 3.5 | 28.8 | 16.5 | 96.5 | 1.03 |
Full results on val-unseen and test-unseen
| Scheduler | Navigator | Val-unseen | Test-unseen | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SR ↑ | SPL ↑ | CSR ↑ | CSPL ↑ | ISPL ↑ | TC ↑ | TS ↓ | MAC ↓ | SR ↑ | SPL ↑ | CSR ↑ | CSPL ↑ | ISPL ↑ | TC ↑ | TS ↓ | MAC ↓ | ||
| Oracle (SA) | Oracle | 67.1 | 63.1 | 74.9 | 68.5 | 91.6 | 100.0 | 412 | 0.00 | 69.1 | 64.4 | 78.0 | 70.8 | 90.4 | 100.0 | 419 | 0.00 |
| Oracle (LE) | Oracle | 100.0 | 99.7 | 100.0 | 99.5 | 92.6 | 100.0 | 261 | 1.16 | 100.0 | 99.6 | 100.0 | 99.4 | 92.0 | 100.0 | 248 | 1.39 |
| Oracle | Oracle | 100.0 | 99.8 | 100.0 | 99.6 | 92.6 | 100.0 | 215 | 1.75 | 100.0 | 99.6 | 100.0 | 99.4 | 92.0 | 100.0 | 206 | 2.10 |
| Oracle | ETPNav | 2.8 | 1.9 | 27.2 | 21.6 | 30.5 | 100.0 | 394 | 0.87 | 5.6 | 4.2 | 28.7 | 22.4 | 30.6 | 100.0 | 403 | 0.85 |
| Oracle (SA) | TRISS | 5.4 | 3.3 | 26.0 | 15.3 | 25.2 | 100.0 | 2056 | 0.00 | 5.9 | 3.8 | 25.2 | 17.4 | 27.7 | 100.0 | 1665 | 0.00 |
| Oracle | TRISS | 9.1 | 4.0 | 36.4 | 19.7 | 26.9 | 100.0 | 1588 | 0.82 | 9.3 | 5.2 | 33.3 | 20.0 | 27.9 | 100.0 | 1494 | 1.42 |
| Random | TRISS | 0.3 | 0.1 | 7.9 | 4.3 | 23.6 | 100.0 | 1621 | 0.91 | 0.2 | 0.1 | 6.7 | 4.1 | 24.1 | 100.0 | 1547 | 1.12 |
| Sequential | TRISS | 3.4 | 1.6 | 23.6 | 12.8 | 24.9 | 100.0 | 1593 | 0.90 | 2.5 | 1.2 | 21.8 | 12.6 | 26.5 | 100.0 | 1443 | 1.28 |
| Decentralized | |||||||||||||||||
| Gemma4 | TRISS | 7.0 ± 0.9 | 2.9 ± 0.4 | 33.6 ± 0.9 | 18.2 ± 0.6 | 24.0 ± 0.4 | 94.6 ± 0.5 | 1570 ± 37 | 0.95 ± 0.1 | 8.3 ± 0.1 | 4.4 ± 0.0 | 33.2 ± 0.3 | 19.5 ± 0.3 | 26.5 ± 0.2 | 95.8 ± 0.2 | 1447 ± 13 | 1.22 ± 0.0 |
| Qwen3.6 | TRISS | 7.3 ± 0.7 | 2.8 ± 0.3 | 34.7 ± 0.4 | 18.2 ± 0.1 | 24.9 ± 0.3 | 97.0 ± 0.3 | 1613 ± 59 | 0.87 ± 0.1 | 7.9 ± 0.8 | 4.1 ± 0.4 | 32.1 ± 0.3 | 18.7 ± 0.2 | 25.9 ± 0.4 | 97.4 ± 0.2 | 1521 ± 25 | 1.26 ± 0.1 |
| Centralized | |||||||||||||||||
| Gemma4 | TRISS | 7.5 ± 1.2 | 3.1 ± 0.5 | 34.3 ± 1.0 | 18.1 ± 0.5 | 23.9 ± 0.5 | 94.4 ± 0.5 | 1570 ± 40 | 0.83 ± 0.0 | 6.8 ± 0.1 | 3.9 ± 0.1 | 31.6 ± 0.1 | 18.6 ± 0.1 | 25.4 ± 0.1 | 95.0 ± 0.2 | 1475 ± 10 | 1.13 ± 0.0 |
| Qwen3.6 | TRISS | 7.6 ± 0.1 | 3.1 ± 0.2 | 34.4 ± 0.3 | 18.2 ± 0.2 | 24.0 ± 0.4 | 95.7 ± 0.1 | 1619 ± 29 | 0.85 ± 0.0 | 8.4 ± 0.9 | 4.2 ± 0.5 | 32.4 ± 0.8 | 19.1 ± 0.7 | 26.3 ± 0.5 | 96.5 ± 0.4 | 1470 ± 7 | 1.03 ± 0.0 |
| Centralized-implicit | |||||||||||||||||
| Gemma4 | TRISS | 5.5 ± 0.1 | 2.3 ± 0.1 | 32.1 ± 0.8 | 16.5 ± 0.2 | 22.5 ± 0.2 | 92.5 ± 0.4 | 1611 ± 44 | 0.76 ± 0.0 | 7.1 ± 0.3 | 3.7 ± 0.2 | 29.6 ± 0.6 | 17.0 ± 0.5 | 23.3 ± 0.3 | 92.8 ± 0.4 | 1512 ± 37 | 0.97 ± 0.1 |
| Qwen3.6 | TRISS | 6.2 ± 0.6 | 2.7 ± 0.3 | 31.5 ± 0.2 | 16.6 ± 0.0 | 23.2 ± 0.2 | 96.8 ± 0.2 | 1662 ± 16 | 0.78 ± 0.0 | 6.8 ± 0.7 | 3.5 ± 0.4 | 28.8 ± 0.6 | 16.5 ± 0.5 | 24.1 ± 0.5 | 96.5 ± 0.5 | 1562 ± 26 | 1.03 ± 0.1 |
Bold marks the best LLM scheduler per column. SA: single agent. LE: legacy scheduling without pre-allocation. Random and sequential schedulers dispatch ground-truth atomic instructions to free agents.
Metric definitions
- SR
- Fraction of missions in which every subtask succeeds under the task’s constraint semantics.
- SPL
- SR weighted by the ratio of the optimal path length to the total path traversed by all agents.
- CSR
- Fraction of subtasks whose arrival lies within the success radius while honoring all dependency, locking and holding constraints.
- ISPL / CSPL
- Per-subtask path efficiency, gated by success without and with task constraints, respectively.
- TC
- Fraction of subtasks the team consumes, irrespective of whether the arrival is correct.
- TS
- Synchronized simulation steps until termination: the realized makespan, including congestion, waiting and deadlock.
- MAC
- Time steps in which any two agents lie within twice the agent radius of each other, divided by TS times team size.
SR, SPL, CSR, CSPL, ISPL and TC are percentages. Arrivals are credited within 3 m of any navigable anchor of the target instance.
Qualitative example
A second episode, preparing for a movie night, spans five subtasks and two scheduling rounds. In the first round, agents A and B reach the refrigerator and the coffee machine and enter holding states for the snacks and drinks, while agent C reaches the couch. In the second round, A and B are assigned the side table and the dining table and complete the mission, keeping their assignments across the holding chains.
Citation
@article{xu2026mavln,
title={Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method},
author={Xu, Yunzhe and Liu, Zhe},
journal={arXiv preprint arXiv:2609.35965},
year={2026}}
Contact: xyz9911@sjtu.edu.cn.