ManiSkillFormer: Demonstration-Free Compositional Manipulation via Geometric Contracts and Agentic Skill Graph
Carnegie Mellon University
Abstract
Adapting robotic manipulation to new objects and tasks often requires additional demonstrations, policy fine-tuning, or manual engineering. Reusable manipulation skills can reduce this effort, but connecting their execution requirements to scene-specific geometry remains challenging. We present ManiSkillFormer, a framework for demonstration-free and compositional manipulation that connects perception and action through explicit geometric contracts. Building on reusable skill schemas, LLM agents generate contracts specifying the geometry primitives required by each skill, together with corresponding motion templates for semantic objects and task contexts. These contracts guide a perception module to ground task-relevant 3D geometry from observations, which is then used to instantiate reusable motion templates in a skill library. We evaluate ManiSkillFormer on a dual-arm robot across demonstration-free pick-and-place with 30 instances from 8 object categories, functional manipulation including unscrewing, pouring, pressing, and folding, and three long-horizon tasks. ManiSkillFormer achieves an average success rate of 88.97% for pick-and-place, 75.00% for functional manipulation, and completion rates of 50–80% across the long-horizon tasks, outperforming the evaluated baselines and two ablated pipelines. These results demonstrate the potential of explicit geometric contracts to support skill reuse and composition across objects and tasks without per-object policy fine-tuning or additional robot demonstrations.
Method
Given a language instruction, the planner retrieves a sequence of skills and their geometric contracts from AgenSkillGraph. Promptable 3D perception grounds the requested semantic primitives in calibrated multi-view observations. The grounded geometry then parameterizes reusable motion templates to generate Cartesian waypoints.
Agentic skill generation
Human experts define skill structures and contract schemas. During adaptation, a contract agent and a skill agent tailor these schemas to semantic skill–object pairs and their task context. At runtime, perception grounds the requested keypoints and surface normals to instantiate the motion templates.
Demonstrations
Real-robot experiments on the Galaxea R1-Lite. Playback speeds are indicated in the footage.
Functional manipulation
Multi-object sorting
Cone stacking
Long-horizon manipulation
Results
We evaluate object generalization, functional manipulation, and long-horizon composition. A1 removes explicit geometric contracts; A2 generates each atomic skill’s contract and motion template independently, without conditioning on subsequent skills.
Object generalization
Pick-and-place across 30 instances from 8 object groups, with 17 trials per group (136 trials per method). The π0 baseline is fine-tuned with 50 demonstrations per object category.
| Object group | Ours | MOKA | A1 | A2 | π0 |
|---|---|---|---|---|---|
| Bottle / Cup | 94.12 | 64.71 | 76.47 | 88.24 | 41.18 |
| Pen | 88.24 | 41.18 | 70.59 | 88.24 | 35.29 |
| Cables / Towels | 94.12 | 29.41 | 94.12 | 76.47 | 29.41 |
| Bowl | 100.00 | 64.71 | 88.24 | 100.00 | 82.35 |
| Toys | 88.24 | 58.82 | 76.47 | 88.24 | 64.71 |
| Tools | 82.35 | 47.06 | 94.12 | 76.47 | 35.29 |
| Cone | 94.12 | 64.71 | 64.71 | 70.59 | 41.18 |
| Lego structures | 70.59 | 47.06 | 58.82 | 58.82 | 35.29 |
| Average | 88.97 | 52.21 | 77.94 | 80.88 | 45.59 |
Functional manipulation
| Task | Ours | MOKA | A1 | A2 |
|---|---|---|---|---|
| Unscrew button | 76.47 | 23.53 | 58.82 | 52.94 |
| Pour | 82.35 | 47.06 | 70.59 | 58.82 |
| Fold cloth edge to center | 64.71 | 58.82 | 35.29 | 52.94 |
| Press button | 76.47 | 76.47 | 70.59 | 76.47 |
Long-horizon composition
| Task | Ours | MOKA | A1 | A2 |
|---|---|---|---|---|
| Pick and place 6 objects | 6/10 | 2/10 | 4/10 | 6/10 |
| Stack cones | 8/10 | 3/10 | 5/10 | 8/10 |
| Task C (11 steps) | 5/10 | 1/10 | 2/10 | 1/10 |
In Task C, the robot places yellow and blue Lego pieces into a cup, places a red Lego piece into a bowl, pours the cup into a box, and places the bowl into the box. Completion requires maintaining compatible intermediate states throughout the sequence.
BibTeX
@misc{yu2026maniskillformerdemonstrationfreecompositionalmanipulation,
title = {ManiSkillFormer: Demonstration-Free Compositional
Manipulation via Task-Conditioned Geometric Contracts},
author = {Peiqi Yu and Mosam Dabhi and Shangtao Li and
Bowei Li and Laszlo Jeni and Changliu Liu},
year = {2026},
eprint = {2609.16331},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.16331},
}