RAIN: Robotic Task Generalization
with Region-Aware Interaction Networks

1LG AI Research2Huawei Technologies Ltd.3Seoul National University4KAIST

*Equal contribution. Work done during internships at LG AI Research.†Corresponding authors.

When the goal changes, existing policies can fall back on familiar demonstration trajectories instead of following the semantics of the current observation and instruction. We introduce RAIN, Region-Aware Interaction Networks, to generalize robotic tasks through reusable, target-region-conditioned interactions.

The benchmark

LIBERO-Analogy

Train on LIBERO, then evaluate on LIBERO-Analogy: 60 tasks that are not directly seen during training, but can be solved by analogy to learned interactions. The benchmark tests robotic task generalization through Adapt, Compose, and Decompose.

Explore all 60 tasks

Adapt

Changed layouts and goals

TRAINING SETRelated tasks from LIBERO

Put the white mug on the left plate and put the yellow and white mug on the right plate
Put the white mug on the plate and put the chocolate pudding to the right of the plate

EVALUATIONLIBERO-Analogy · Adapt

Put the white mug on the right plate

Adapt learned interactions to changes in scene layout, object arrangement, and task goals.

Compose

Combine skills into longer tasks

TRAINING SETRelated tasks from LIBERO

Put the bowl on the stove
Turn on the stove
Open the top drawer and put the bowl inside

EVALUATIONLIBERO-Analogy · Compose

Put the cream cheese on the stove, then turn on the stove, then open the top drawer of the wooden cabinet.

Combine skills learned from LIBERO into new long-horizon tasks.

Decompose

Execute only the requested part

TRAINING SETRelated tasks from LIBERO

put the black bowl in the bottom drawer of the cabinet and close it

EVALUATIONLIBERO-Analogy · Decompose

Close the bottom drawer of the cabinet

Perform the requested part of a learned long-horizon task, leaving the rest of the sequence out.

From instruction to interaction

Robotic Task as Sequence of Interactions

A VLM plans the subtasks. SAM3 segments and tracks each target. RAIN acts on the region and decides when to move on.

  1. 01 Plan the interactions
  2. 02 Find the target
  3. 03 Act & transition
Task description

Pick up the potato and place it on the rightmost plate.

Planning Prompt

Split the task into ordered interactions. For each, specify the action type and target object or region.

VLM
SAM3
Initial observation
Robot observation

RAIN

Region encoder
Action
decoder
Transition
head
ActionTransition: No
Observe

Start with the robot’s current observation.

0.0 / 48.0 s

RAIN receives images, target masks, and the current action type. It predicts actions and a transition decision; a Yes activates the next subtask.

Recorded execution with SAM3 tracking and visual target initialization. The animation illustrates the control flow.

The method

Region-Aware Interaction Networks

Learn how to interact with a target region.

RAIN architecture. TCE conditions multi-stage, multi-view features on target masks with TarLN. PlanDiT predicts waypoints and action chunks. The Transition Head predicts subtask completion.

The Target-adaptive Cross-view Encoder (TCE) uses Target-adaptive Layer Normalization (TarLN) to condition each view’s multi-stage visual features on its target mask, then exchanges information across views. PlanDiT jointly predicts target-directed waypoints and action chunks. The Transition Head pools target-region features and combines the views through gated fusion to determine when to advance to the next subtask.

Reference-based Retargeting Strategy

RRS generates smooth approaches to alternative target regions by joining a new trajectory to a recorded reference, preserving its contact-rich interaction. This expands training coverage without collecting additional data.

Experiments

Results

Success rates on LIBERO and LIBERO-Analogy, followed by the VLM comparison.

Table 1Comparison on LIBERO and LIBERO-Analogy

Scroll horizontally to see all columns →

MethodTrainable parametersLIBERO-AnalogyLIBERO
Pre-trainPost-trainAdaptComposeDecomposeAvg.SpatialObjectGoalLongAvg.
Regular VLA models
SmolVLA (arXiv'25)0.1B0.1B5.60.05.43.790.096.092.071.087.3
OpenVLA-OFT (RSS'25)7B0.3B18.46.06.410.397.698.497.994.597.1
GR00T N1.6† (arXiv'25)3B3B27.510.69.816.096.6100.095.893.096.3
π₀ (RSS'25)3B3B13.61.115.510.196.898.895.885.294.2
π₀.₅ (CoRL'25)3B3B57.110.77.625.198.898.298.092.496.9
Evo-1 (CVPR'26)0.1B0.8B24.00.39.511.392.797.796.392.394.8
World-action models
Cosmos Policy (ICLR'26)-2B29.15.07.113.798.1100.098.297.698.5
VLA-JEPA (ECCV'26)2.4B2.4B42.15.87.818.696.299.697.295.897.2
Flex-π (arXiv'26)6B6B36.711.311.119.799.299.899.098.699.2
VLA models with spatial reasoning
InspireVLA (ICRA'26)1B1B29.20.066.431.990.794.388.373.386.7
Action-Sketcher (CVPR'26)3B3B35.82.06.614.897.299.694.896.096.9
MolmoAct2 (CoRL'26)4.4B5B53.929.522.435.397.8100.097.893.297.2
RAINQwen3.5-4B + SAM3-0.2B53.331.077.553.992.496.891.282.490.7
RAIN (GT masks)-0.2B62.437.182.760.793.699.095.293.695.4

Success rate (%). † Our reproduction with a single model jointly fine-tuned on the four LIBERO suites.

Table 2VLM comparison: task success and goal-mask tracking

Scroll horizontally to see all columns →

VLMLIBEROLIBERO-Analogy
SpatialObjectGoalLongOverallAdaptComposeDecomposeOverall
SRJ&FSRJ&FSRJ&FSRJ&FSRJ&FSRJ&FSRJ&FSRJ&FSRJ&F
Qwen3.5-4B92.480.096.876.191.255.982.457.890.765.453.360.931.051.977.549.453.953.7
Qwen3.5-9B91.680.795.875.490.055.381.457.689.765.253.259.530.052.877.449.853.554.0
Cosmos-Reason2-8B87.683.390.272.287.455.783.059.887.166.360.663.429.956.476.251.855.657.3

All VLMs use SAM3 for segmentation and tracking. SR (%) and tracking J&F (%) are measured on the same rollouts. J&F pools visible-target observations across both views and tasks.