Reconstruct,
Practice,
Go Real.

Guided Self-Improvement for Embodied Agents

RPG builds simulation practice tasks from an offline dataset, develops reusable robot skills, and deploys them on a physical robot without model fine-tuning.

1UC Berkeley  ·  2Amazon FAR  ·  3MIT  ·  4University of Chicago
*Equal contribution  ·  †Equal advising  ·  ‡Work done while interning at Amazon FAR
01 / 06
Unscrew Bottle Capphysical deployment · 15× · 00:00 real time
0%
success on held-out initializations of 22 simulation tasks
0/30
successful physical trials across three tasks
0 rounds
of autonomous practice
0 entries
in the shared tool and skill library, up from 15
Overview

Develop robot skills through simulation practice.

RPG uses an offline dataset to decide what to practice in simulation. Execution feedback guides new skills, improvements to existing skills, and prompt revisions. The resulting shared skill library and system prompt transfer to a physical robot after calibration and hardware adaptation.

Read the abstract

We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks.

RPG overview: Reconstruct builds simulation practice tasks from a real-world dataset; Practice revises shared skills; Go Real deploys the frozen system on a physical robot.

Reconstruct: build practice tasks. Practice: develop skills and validate revisions across tasks. Go Real: adapt to the physical robot.

Method

Improve the skills and prompt, not the model weights.

The Runtime Agent is a multimodal LLM that calls reusable code functions, or skills, for perception and control. RPG revises these skills and the agent’s system prompt between practice rounds.

Stage 01

Reconstruct

Identify capabilities from dataset videos and descriptions, then build simulation tasks to practice them. “Organize sunglasses,” for example, becomes closing a loaded case. Humans review the success checkers; tasks and checkers then stay fixed.

Stage 02

Practice

Attempt the tasks, diagnose failures, and revise shared skills and the prompt. Test each candidate and merged revision across tasks before retention.

Stage 03

Go Real

Deploy after the same calibration and hardware adaptation used for CaP-Agent0. No task-specific tuning, simulator state, or offline failure analysis during physical execution.

RPG architecture: Runtime and Privileged agents share a skill library; a Video Analyzer diagnoses failures; an Implementor proposes candidates; a Merger integrates them.

The Privileged Agent uses the same model and skills with access to simulator state. Comparing its executions with the Runtime Agent’s helps diagnose failures. The Video Analyzer also uses related dataset videos when available; the Implementor proposes revisions, and the Merger selects and combines eligible changes.

One practice round

  1. Run five Runtime and five Privileged Agent episodes per task
  2. Diagnose failures on the 17 development tasks
  3. Propose at most one revision per failing development task
  4. Test each candidate across all 22 tasks
  5. Select eligible revisions, merge them, and test again
  6. Retain only if validation passes and the Merger accepts

Cross-task validation

pass if mean success ↑ and each task loses ≤ 1/5

Shared skills can help one task but hurt another. Test candidates and merged systems on five fixed initializations per task. Passing candidates are eligible for merging, not automatically accepted.

Simulation results

28.6% success after round 1. 95.0% after round 15.

After practice, we test every saved round-end system on 22 tasks × 10 new starting configurations, held out from revision selection. Runtime: Gemini 3.8 Flash. Offline analysis and revision: Fable 5.1.

Task success across practice rounds

Click a heatmap row to show that task’s results on the curve.

RPGBaseline results

Round 3 → 4: no skill edits; prompt and observation-handling revisions, +30.5 percentage points. Round 4 → 15: library changes with a fixed prompt, +21.4 points.

Per task, per round

Each cell counts successful trials out of ten held-out initializations.

010
swipe to see all rounds →

D: development. V: validation. H: feedback-held-out. Only D tasks generate task-specific revisions; all tasks enter regression checks.

How do diagnostic inputs affect improvement?

Same initial system. Five practice rounds, then 220 held-out episodes per variant.

Video AnalyzerPrivileged AgentSuccess
YesYes78.2%
NoYes51.8%
YesNo50.9%
NoNo50.9%

Removing either component reduces success from 78.2% to about 51%.

Do library revisions improve task success?

Only the library changes. Model, prompt, perception, and budgets stay fixed. Ten paired trials per task.

TaskBeforeAfter
Two-arm lift6/1010/10
Place dish on rack5/109/10
Nut assembly5/108/10
Place bottles in bin10/1010/10
Success rate65.0%92.5% +27.5 pp

A separate library-snapshot comparison: three tasks improve, and bottle-placement success is preserved.

Full per-task comparison with baselines
TaskCaP-Agent0
Gemini 3.8 Flash
CaP-Agent0
GPT-6 Astra Pro
CaP-Agent0
Opus 5
CaP-Agent0
Fable 5.1
RATs
Gemini 3.8 Flash
ASPIRE
Gemini 3.8 Flash
RPG
Gemini 3.8 Flash
CaP-X tasks
Lift cube9/1010/1010/1010/1010/1010/1010/10
Nut assembly4/106/104/105/100/103/109/10
Restack cubes6/109/108/107/105/1010/1010/10
Spill wipe2/1010/106/109/105/108/1010/10
Stack cubes8/1010/109/107/1010/1010/1010/10
Two-arm handover5/103/103/106/102/106/1010/10
Two-arm lift2/103/103/102/106/108/1010/10
Group mean51.4%72.9%61.4%65.7%54.3%78.6%98.6%
ABC-derived and analogous tasks
Extract from fixture8/1010/102/102/102/107/1010/10
Insert into fixture2/101/100/100/104/102/1010/10
Place dish on rack1/109/102/105/104/107/1010/10
Serve onto plate9/1010/109/108/102/1010/109/10
Fold towel1/100/100/100/100/100/103/10
Transfer object0/100/100/105/101/109/109/10
Lift rod8/1010/109/1010/1010/1010/1010/10
Pick up cup10/1010/1010/1010/1010/1010/1010/10
Place bottles in bin0/101/100/100/100/103/1010/10
Close sunglasses case9/106/106/106/109/1010/1010/10
Close drawer6/105/106/103/109/1010/1010/10
Sort cubes (standard)3/102/104/105/101/1010/109/10
Sort cubes (extended)0/100/100/100/100/108/1010/10
Stack blocks4/107/100/101/102/105/1010/10
Transport cup10/1010/109/107/100/1010/1010/10
Group mean47.3%54.0%38.0%41.3%36.0%74.0%93.3%
Mean success48.6%60.0%45.5%49.1%41.8%75.5%95.0%
Total107/220132/220100/220108/22092/220166/220209/220

Ten held-out initializations per task. CaP-Agent0 and ASPIRE are reproduced from released code. RATs uses its native any-time success criterion; the other methods are scored on the final outcome.

Watch the 22 simulation tasks

One episode per task, run by the round-15 RPG system in MuJoCo.

Lift cubeCaP-X
Nut assemblyCaP-X
Restack cubesCaP-X
Spill wipeCaP-X
Stack cubesCaP-X
Two-arm handoverCaP-X
Two-arm liftCaP-X
Extract from fixtureABC
Insert into fixtureABC
Place dish on rackABC
Serve onto plateABC
Fold towelABC
Transfer objectABC
Lift rodABC
Pick up cupABC
Place bottles in binABC
Close sunglasses caseABC
Close drawerABC
Sort cubes (standard)ABC
Sort cubes (extended)ABC
Stack blocksABC
Transport cupABC
Go Real

RPG succeeds in all 30 physical trials.

Three tasks, ten trials each, with the same calibration, hardware adaptation, and budgets for all methods. No task-specific tuning or system revisions. Autonomous retries are allowed within a trial; human assistance and manual resets are not.

RPGGemini 3.8 Flash
30/30
intermediate stage reached20/20
BaselineCaP-Agent0 · GPT-6 Astra Pro
14/30
intermediate stage reached14/20
BaselineCaP-Agent0 · Gemini 3.8 Flash
3/30
intermediate stage reached13/20

Large counts: task completion. Smaller counts: ball placed in drawer or bowl lifted to center, before closure or handover. Across these two tasks, GPT-6 CaP-Agent0 finishes 5/20 trials; RPG finishes 20/20.

RPG
00:00real time
Result · 10 trials10 / 10ball placed 10 / 10 · drawer closed 10 / 10
RPG (ours)
Gemini 3.8 Flash
15×
VS
Baseline
00:00real time
Result · 10 trials4 / 10ball placed 6 / 10 · drawer closed 4 / 10
CaP-Agent0
GPT-6 Astra Pro
15×
RPG
ball placed in drawer 10/10
→
drawer closed with ball inside 10/10
CaP-Agent0 · GPT-6
ball placed in drawer 6/10
→
drawer closed with ball inside 4/10

Store Ball in Drawer. Success requires fully closing the drawer with the ball inside. The videos show all ten trials at 15× speed, with the final tally at the end.

Additional physical tasks

Use the shared skills on additional physical tasks.

The first five tasks are zero-shot: instruction only, with no task-specific real-world demonstrations, tuning, or retry-based adaptation. Sort Utensils is another qualitative example. Videos show execution at 15× speed, not measured success rates.

15× · 00:00 real time
01
Open Scissors
align with handles→insert gripper→spread handles
15× · 00:00 real time
02
Uncap Marker
pick up marker→grasp cap→pull off cap
15× · 00:00 real time
03
Pull out Tissue
approach→grasp tissue→pull upward
15× · 00:00 real time
04
Erase Whiteboard
pick up eraser→hold board→wipe
15× · 00:00 real time
05
Unscrew Bottle Cap
hold bottle→grasp cap→twist cap
15× · 00:00 real time
06
Sort Utensils
spoon → blue plate→fork → yellow plate→rest
Examples of skill refinement

Check what the object does, not just where the gripper moves.

The revised skills verify grasp retention, object alignment, and placement, rather than relying on gripper motion alone. Results below use five development trials, separate from the held-out evaluation.

Case study: a failed bottle placement is diagnosed, the shared pick-and-place skill is revised, and development success rises from 1 of 5 to 5 of 5.

Bottle placement. A collision leads to full-path checks, release inside the bin, and placement verification. Development success improves from 1/5 to 5/5.

Align the nut’s hole with the peg

NUT ASSEMBLY · ROUND 4 → 5 · 1/5 → 5/5 · prompt unchanged
before
grasp→align gripper to peg→push down→release
after
grasp handle→test lift→verify retention→estimate hole offset→align hole to peg→incremental insertion→verify seating→release

Verify both grasps before lifting

TWO-ARM LIFT · ROUND 11 → LATER · 3/5 → 5/5
before
approach (each arm)→grasp→lift
after
localize handles→check both paths→verify both grasps→paired test lift→synchronized lift→verify object rise
Citation

BibTeX

@misc{wang2026rpg,
      title={Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents}, 
      author={Yen-Jen Wang and Haozhe Jiang and Shuying Deng and Haoru Xue and Weirui Ye and Rocky Duan and Nika Haghtalab and S. Shankar Sastry and Pieter Abbeel and Haozhi Qi},
      year={2026},
      eprint={2610.02204},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2610.02204}, 
}

Acknowledgments. We thank Zi Wang and Bike Zhang for helpful discussions and feedback. The website template is adapted from vision-language-kinematics.github.io.