DuoMind

Hierarchical Multi-Robot Coordination Framework

RoboPoly

Multi-Robot Benchmark

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Enabling multi-robot collaboration greatly unlocks the capability of robots in solving challenging tasks. We propose DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination.

DuoMind overview showing two robot orchestrators exchanging semantic messages, their action models, seven RoboPoly tasks, and average benchmark performance
(a) An illustration of DuoMind, a distributed framework combining an orchestrator and action model for multi-robot coordination. (b) RoboPoly is a benchmark we develop for long-horizon coordination under distributed robot control. (c) DuoMind consistently outperforms the baselines across both benchmarks.

Distributed hierarchy

Framework Overview

To achieve multi-robot coordination under distributed setting, we adopt a hierarchical framework. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots.

DuoMind architecture with one VLM orchestrator and one VLA action model per robot, low-frequency shared information, and high-frequency actions

Long-horizon benchmark

RoboPoly

We introduce RoboPoly, a benchmark of seven challenging multi-robot tasks that require long-horizon coordination under distributed control, with each robot acting independently from its available observations.

The videos above are from DuoMind evaluation results.

Quantitative evaluation

Experiment Result

We conduct extensive experiments to evaluate DuoMind across challenging cooperative manipulation tasks on both RoboPoly and RoboTwin. Our results demonstrate the effectiveness of the proposed hierarchical design in multi-robot coordination and further quantify the contributions of integrating orchestrator and inter-agent communication.

Table of success rates for DuoMind and action-model-only baselines on seven RoboPoly tasks
Table of success rates for DuoMind and action-model-only baselines on eight RoboTwin tasks

Ablations

Inter-agent Communication Ablation

To further study the effect of the shared messages, we conduct an ablation using DuoMind with π0.5 but without inter-agent communication.

The results show that the orchestrated π0.5 system consistently outperforms the non-communicating π0.5 system across all seven RoboPoly tasks. Without inter-agent communication, we observe more frequent conflicts and asynchronous behaviors between the two robots.

Bar chart comparing DuoMind with and without communication across seven RoboPoly tasks

Action Model Compatibility

To analyze the influence of the action model, we conduct an additional experiment using π0 in place of π0.5. DuoMind with π0 still substantially outperforms the π0-only baseline. By introducing high-level reasoning and subtask decomposition through the orchestrator, DuoMind provides more appropriate low-level instructions for the evolving multi-agent state and improves coordination over the action-model-only setting.

Action model compatibility chart comparing DuoMind with pi zero point five, DuoMind with pi zero, and pi zero only across three RoboPoly tasks

Citation

@misc{zhou_duomind,
  title = {{DuoMind}: Enabling Distributed Multi-Robot Coordination with Semantic Communication},
  author = {Zhou, Hanchu and Gao, Dechen and Wang, Hang and Lynch, Brendan and Zhao, Boqi and Ma, Qiyao and Goyal, Raman and Zhang, Junshan},
  url = {https://arxiv.org/abs/2610.02161}
}