FINDFind Something You Can’t Do

Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

A robot that learns what to practice next. FIND connects scene understanding, weakness-aware task selection, self-evaluation, and policy improvement in a persistent real-world workspace.

Yuan Fang1,*Zechu Li1,2,*Haolei Tong3Puze Liu4,5Georgia Chalvatzaki1,2

1 TU Darmstadt  ·  2 Hessian.AI  ·  3 University of Augsburg
4 Tongji University  ·  5 Shanghai Research Institute for Intelligent Autonomous Systems
* Equal contribution

FIND system overview: scene-conditioned task selection, paired-image evaluation, frozen VLA with residual policy, and asynchronous RL
FIND closes the loop between what the robot sees, what it practices, how it evaluates success, and how its policy improves.
55% → 71.9%
Human-assessed
success rate
8
Real-world
manipulation tasks
456
Episodes in a
6-hour run
426
Episodes without
human intervention
30
Scene-recovery
interventions

01Abstract

Vision–language–action (VLA) models provide strong priors for robotic manipulation, but they are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning.

We introduce FIND, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. Instead of restoring a predefined scene after each rollout, a vision–language agent identifies feasible tasks from a fixed library, prioritizes those with lower recent success rates, and evaluates outcomes using paired before and after observations.

Using a frozen π0.5 VLA with residual off-policy RL, FIND improves independent human-assessed success across eight real-world tasks from 55% to 71.9%. A representative 6-hour run completes 456 episodes with 30 scene-recovery interventions and no human-provided reward labels during online learning.

02FIND in the real world.

See the autonomous practice loop and robot learning setup in the accompanying video.

03HOW FIND works

We introduce FIND, an agentic real-world RL framework that closes the autonomous loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace.

FIND's autonomous loop connects scene-conditioned task selection, VLM outcome evaluation, a frozen VLA with a residual policy, and asynchronous reinforcement learning
01

Scene-conditioned task selection

Scene-conditioned task selection allows FIND to continue practicing without repeatedly restoring the manipulated objects to a predefined configuration.

02

Performance-aware curriculum

FIND instead tracks recent task performance and gives higher priority to tasks on which the current policy performs poorly.

03

Paired-image self-evaluation

Paired-image self-evaluation provides autonomous task-level supervision without requiring human outcome labels.

04

Policy Improvement via Residual Reinforcement Learning

The pretrained VLA remains frozen, while a lightweight residual policy learns bounded corrections from offline demonstrations and autonomously collected experience.

Scene-conditioned task selection · Performance-aware curriculum
Animated diagram of FIND updating task sampling across iterations as the scene and recent success rates change

04Eight tasks. One persistent workspace.

Reversible or mutually compatible tasks allow one rollout’s final scene to become the next rollout’s starting point.

01 Put a cube into the bowl
02 Take the cube out of the bowl
03 Stack one cube on the other
04 Take the top cube off
05 Open the drawer
06 Close the drawer
07 Hang the mug on the mug tree
08 Take the mug off the mug tree
REAL-WORLD ROLLOUTS

Eight tasks in action

TASK 01

Put a cube into the bowl

TASK 02

Take the cube out of the bowl

TASK 03

Stack one cube on the other

TASK 04

Take the top cube off

TASK 05

Open the drawer

TASK 06

Close the drawer

TASK 07

Hang the mug on the mug tree

TASK 08

Take the mug off the mug tree

05Improvement through interaction.

The full FIND system improves independent human-assessed task success across eight real-world tasks.

Human-assessed success rate
55% → 71.9%

Mean across five runs at the final checkpoint. Human-assessed evaluation rollouts were kept separate from training.

6 hinteraction
30scene-recovery interventions
48 minlongest uninterrupted run
Paper plots showing learning curves, task sampling distribution, and per-task success and sampling trends
Learning curves and adaptive task sampling from Fig. 3 of the paper.

In the representative 6-hour run, 426 of 456 episodes proceeded without human intervention. The 30 interventions were manual scene recoveries; online learning required no human-provided reward labels.

06Citation

@misc{fangfind,
  title={Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models},
  author={Fang, Yuan and Li, Zechu and Tong, Haolei and Liu, Puze and Chalvatzaki, Georgia},
  note={Manuscript under review}
}

Citation details are provisional; please check the paper for the latest version.