Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due to high action-space redundancy, inter-robot cooperation, and dynamic object-robot interactions. This paper introduces a framework based on spatially conditioned Multi-Agent Transformers (MATs) to efficiently learn robust control policies for a DDM system grounded in an array of 64 soft delta robots arranged in an 8 × 8 grid. Our three core contributions are: (i) an MAT with adaptive layer norm for compute efficiency, (ii) spatial contrastive embeddings to ground transformer embeddings in the spatial configuration of the robots, and (iii) an MAT-based behavior cloning method fine-tuned using Soft Actor Critic. We also propose an action selection formulation to analyze the trade-off between task performance and the number of robots utilized. Our experiments show that MATs iteratively refine their actions through the stacked attention blocks. This further informs the benefit of spatial conditioning in transformers to learn DDM policies. We demonstrate long-horizon planar manipulation tasks with objects of various geometries in simulation and real-world. Finally, we show how action selection mitigates robot maintenance by reducing wear and tear due to inter-robot collisions while maintaining the ability to manipulate objects along various trajectories in the real-world, achieving an average error of ∼1.5 cm, while using ∼65% fewer robots.
Trajectory tracking performance of each algorithm for each trajecotry, over all objects in simulation.
Trajectory tracking performance of each algorithm for each trajecotry, over all objects in real world.
Number of active robots used by each algorithm for each object over all trajectories in simulation.
Average tracking distance error by each algorithm for each object over all trajectories in simulation.
Number of attempts for completing a 20-step trajectory by each algorithm for each object over all trajectories in simulation (min 20, max 60).
Results for Hexagon, Star, and Trapezium from the real world experiments. (Left) Shows the number of active robots, (Middle) shows the average distance error, (Right) shows the number of attempts for each object.
@article{ddmmat2025,
author = {Anonymous Author},
title = {Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers},
journal = {arXiv},
year = {20XX},
}