RPG builds simulation practice tasks from an offline dataset, develops reusable robot skills, and deploys them on a physical robot without model fine-tuning.
RPG uses an offline dataset to decide what to practice in simulation. Execution feedback guides new skills, improvements to existing skills, and prompt revisions. The resulting shared skill library and system prompt transfer to a physical robot after calibration and hardware adaptation.
We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks.

Reconstruct: build practice tasks. Practice: develop skills and validate revisions across tasks. Go Real: adapt to the physical robot.
The Runtime Agent is a multimodal LLM that calls reusable code functions, or skills, for perception and control. RPG revises these skills and the agent’s system prompt between practice rounds.
Identify capabilities from dataset videos and descriptions, then build simulation tasks to practice them. “Organize sunglasses,” for example, becomes closing a loaded case. Humans review the success checkers; tasks and checkers then stay fixed.
Attempt the tasks, diagnose failures, and revise shared skills and the prompt. Test each candidate and merged revision across tasks before retention.
Deploy after the same calibration and hardware adaptation used for CaP-Agent0. No task-specific tuning, simulator state, or offline failure analysis during physical execution.

The Privileged Agent uses the same model and skills with access to simulator state. Comparing its executions with the Runtime Agent’s helps diagnose failures. The Video Analyzer also uses related dataset videos when available; the Implementor proposes revisions, and the Merger selects and combines eligible changes.
Shared skills can help one task but hurt another. Test candidates and merged systems on five fixed initializations per task. Passing candidates are eligible for merging, not automatically accepted.
After practice, we test every saved round-end system on 22 tasks × 10 new starting configurations, held out from revision selection. Runtime: Gemini 3.8 Flash. Offline analysis and revision: Fable 5.1.
Click a heatmap row to show that task’s results on the curve.
Round 3 → 4: no skill edits; prompt and observation-handling revisions, +30.5 percentage points. Round 4 → 15: library changes with a fixed prompt, +21.4 points.
Each cell counts successful trials out of ten held-out initializations.
D: development. V: validation. H: feedback-held-out. Only D tasks generate task-specific revisions; all tasks enter regression checks.
Same initial system. Five practice rounds, then 220 held-out episodes per variant.
| Video Analyzer | Privileged Agent | Success |
|---|---|---|
| Yes | Yes | 78.2% |
| No | Yes | 51.8% |
| Yes | No | 50.9% |
| No | No | 50.9% |
Removing either component reduces success from 78.2% to about 51%.
Only the library changes. Model, prompt, perception, and budgets stay fixed. Ten paired trials per task.
| Task | Before | After |
|---|---|---|
| Two-arm lift | 6/10 | 10/10 |
| Place dish on rack | 5/10 | 9/10 |
| Nut assembly | 5/10 | 8/10 |
| Place bottles in bin | 10/10 | 10/10 |
| Success rate | 65.0% | 92.5% +27.5 pp |
A separate library-snapshot comparison: three tasks improve, and bottle-placement success is preserved.
| Task | CaP-Agent0 Gemini 3.8 Flash | CaP-Agent0 GPT-6 Astra Pro | CaP-Agent0 Opus 5 | CaP-Agent0 Fable 5.1 | RATs Gemini 3.8 Flash | ASPIRE Gemini 3.8 Flash | RPG Gemini 3.8 Flash |
|---|---|---|---|---|---|---|---|
| CaP-X tasks | |||||||
| Lift cube | 9/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| Nut assembly | 4/10 | 6/10 | 4/10 | 5/10 | 0/10 | 3/10 | 9/10 |
| Restack cubes | 6/10 | 9/10 | 8/10 | 7/10 | 5/10 | 10/10 | 10/10 |
| Spill wipe | 2/10 | 10/10 | 6/10 | 9/10 | 5/10 | 8/10 | 10/10 |
| Stack cubes | 8/10 | 10/10 | 9/10 | 7/10 | 10/10 | 10/10 | 10/10 |
| Two-arm handover | 5/10 | 3/10 | 3/10 | 6/10 | 2/10 | 6/10 | 10/10 |
| Two-arm lift | 2/10 | 3/10 | 3/10 | 2/10 | 6/10 | 8/10 | 10/10 |
| Group mean | 51.4% | 72.9% | 61.4% | 65.7% | 54.3% | 78.6% | 98.6% |
| ABC-derived and analogous tasks | |||||||
| Extract from fixture | 8/10 | 10/10 | 2/10 | 2/10 | 2/10 | 7/10 | 10/10 |
| Insert into fixture | 2/10 | 1/10 | 0/10 | 0/10 | 4/10 | 2/10 | 10/10 |
| Place dish on rack | 1/10 | 9/10 | 2/10 | 5/10 | 4/10 | 7/10 | 10/10 |
| Serve onto plate | 9/10 | 10/10 | 9/10 | 8/10 | 2/10 | 10/10 | 9/10 |
| Fold towel | 1/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 3/10 |
| Transfer object | 0/10 | 0/10 | 0/10 | 5/10 | 1/10 | 9/10 | 9/10 |
| Lift rod | 8/10 | 10/10 | 9/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| Pick up cup | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| Place bottles in bin | 0/10 | 1/10 | 0/10 | 0/10 | 0/10 | 3/10 | 10/10 |
| Close sunglasses case | 9/10 | 6/10 | 6/10 | 6/10 | 9/10 | 10/10 | 10/10 |
| Close drawer | 6/10 | 5/10 | 6/10 | 3/10 | 9/10 | 10/10 | 10/10 |
| Sort cubes (standard) | 3/10 | 2/10 | 4/10 | 5/10 | 1/10 | 10/10 | 9/10 |
| Sort cubes (extended) | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 8/10 | 10/10 |
| Stack blocks | 4/10 | 7/10 | 0/10 | 1/10 | 2/10 | 5/10 | 10/10 |
| Transport cup | 10/10 | 10/10 | 9/10 | 7/10 | 0/10 | 10/10 | 10/10 |
| Group mean | 47.3% | 54.0% | 38.0% | 41.3% | 36.0% | 74.0% | 93.3% |
| Mean success | 48.6% | 60.0% | 45.5% | 49.1% | 41.8% | 75.5% | 95.0% |
| Total | 107/220 | 132/220 | 100/220 | 108/220 | 92/220 | 166/220 | 209/220 |
Ten held-out initializations per task. CaP-Agent0 and ASPIRE are reproduced from released code. RATs uses its native any-time success criterion; the other methods are scored on the final outcome.
One episode per task, run by the round-15 RPG system in MuJoCo.
Three tasks, ten trials each, with the same calibration, hardware adaptation, and budgets for all methods. No task-specific tuning or system revisions. Autonomous retries are allowed within a trial; human assistance and manual resets are not.
Large counts: task completion. Smaller counts: ball placed in drawer or bowl lifted to center, before closure or handover. Across these two tasks, GPT-6 CaP-Agent0 finishes 5/20 trials; RPG finishes 20/20.
Store Ball in Drawer. Success requires fully closing the drawer with the ball inside. The videos show all ten trials at 15× speed, with the final tally at the end.
Fold Towel. Success requires a final footprint of 40–60% of the initial area and approximately parallel sides. This physical criterion differs from the simulation criterion.
Transfer Bowl. Success requires the right hand alone to hold the bowl off the table for at least three seconds after the left hand releases it. Intermediate support from the table is allowed.
The first five tasks are zero-shot: instruction only, with no task-specific real-world demonstrations, tuning, or retry-based adaptation. Sort Utensils is another qualitative example. Videos show execution at 15× speed, not measured success rates.
The revised skills verify grasp retention, object alignment, and placement, rather than relying on gripper motion alone. Results below use five development trials, separate from the held-out evaluation.

Bottle placement. A collision leads to full-path checks, release inside the bin, and placement verification. Development success improves from 1/5 to 5/5.
@misc{wang2026rpg,
title={Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents},
author={Yen-Jen Wang and Haozhe Jiang and Shuying Deng and Haoru Xue and Weirui Ye and Rocky Duan and Nika Haghtalab and S. Shankar Sastry and Pieter Abbeel and Haozhi Qi},
year={2026},
eprint={2610.02204},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2610.02204},
}
Acknowledgments. We thank Zi Wang and Bike Zhang for helpful discussions and feedback. The website template is adapted from vision-language-kinematics.github.io.