Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers

Pushing Trapezium Along Cross

Pushing Star Along Spiral

Pushing Hexagon Along Square

Above videos show MATBC-FT-DEC being rolled out in Real World and Simulation. The real world video is sped up by 4x.

Abstract

Main method figure

Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due to high action-space redundancy, inter-robot cooperation, and dynamic object-robot interactions. This paper introduces a framework based on spatially conditioned Multi-Agent Transformers (MATs) to efficiently learn robust control policies for a DDM system grounded in an array of 64 soft delta robots arranged in an 8 × 8 grid. Our three core contributions are: (i) an MAT with adaptive layer norm for compute efficiency, (ii) spatial contrastive embeddings to ground transformer embeddings in the spatial configuration of the robots, and (iii) an MAT-based behavior cloning method fine-tuned using Soft Actor Critic. We also propose an action selection formulation to analyze the trade-off between task performance and the number of robots utilized. Our experiments show that MATs iteratively refine their actions through the stacked attention blocks. This further informs the benefit of spatial conditioning in transformers to learn DDM policies. We demonstrate long-horizon planar manipulation tasks with objects of various geometries in simulation and real-world. Finally, we show how action selection mitigates robot maintenance by reducing wear and tear due to inter-robot collisions while maintaining the ability to manipulate objects along various trajectories in the real-world, achieving an average error of ∼1.5 cm, while using ∼65% fewer robots.

Results

The overall GraphEQA method.

Trajectory tracking performance of each algorithm for each trajecotry, over all objects in simulation.

The overall GraphEQA method.

Trajectory tracking performance of each algorithm for each trajecotry, over all objects in real world.

The overall GraphEQA method.

Number of active robots used by each algorithm for each object over all trajectories in simulation.

The overall GraphEQA method.

Average tracking distance error by each algorithm for each object over all trajectories in simulation.

The overall GraphEQA method.

Number of attempts for completing a 20-step trajectory by each algorithm for each object over all trajectories in simulation (min 20, max 60).

The overall GraphEQA method.

Results for Hexagon, Star, and Trapezium from the real world experiments. (Left) Shows the number of active robots, (Middle) shows the average distance error, (Right) shows the number of attempts for each object.

BibTeX

@article{ddmmat2025,
    author    = {Anonymous Author},
    title     = {Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers},
    journal   = {arXiv},
    year      = {20XX},
  }