StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

Ziyi Yin1†*Sangmin Woo2†Kang Zhou2Sungyeon Kim2Aosong Feng2Haibo Ding2Luke Huan2
1The Pennsylvania State University2Amazon AWS AI
†Co-first authors   *Work done during an internship at Amazon

TL;DR: StructRL rewards verifiable subtask progress, gated by the task's prerequisite structure and paced by demonstration timing, for online RL of long-horizon VLA policies.

Long-horizon task example and success rates
(a) An example long-horizon VLA task from RoboCasa365. (b) Success rates (%) of GR00T-N1.5 on three representative composite tasks. StructRL consistently outperforms both the SFT policy and the RL baseline SimpleVLA-RL.

Abstract

Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and π0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training.

Rollouts

Each video shows GR00T-N1.5 after SFT (top) and after online RL with StructRL (bottom), starting from the same RoboCasa365 scene. The demos use the decomposition with about five subtasks per task (N̄ = 5), the best setting of the density study in the paper (Fig. 4). The panel scores both rollouts with that StructRL reward, computed after the fact from the simulator's subtask checks (GR00T-N1.5 itself is trained by imitation, without any reward): a subtask pays once it is completed after its prerequisites, sooner completions pay more, and completing the task adds 2.0.

StoreLeftoversInBowl (1,700-step horizon) · click the video to pause

Method

StructRL overview
Left: An LLM decomposes the task command into candidate subtasks and assigns a dependency structure to the retained subtasks, represented by prerequisite sets X(vi). The verifiability check removes candidates without a reliable binary completion criterion (e.g., Approach Box), while the progress check removes candidates that do not by themselves indicate progress toward task completion (e.g., Open Gripper). Right: Structure-aware reward gating uses this dependency structure to determine whether a detected completion is eligible for reward, while dynamic reward pacing determines the magnitude of that reward. The resulting chunk-level rewards are used to optimize the VLA with PPO.

Results

Success rate (%). RoboCasa365: 16 Composite-Seen tasks, 100 episodes each. LIBERO-Long: 10 tasks, 50 episodes each.

Table 1 of the paper (overall success rate). Best per backbone in bold.
BackboneMethodRoboCasa365LIBERO-Long
GR00T-N1.5SFT38.690.6
Sparse-RL (PPO)40.291.2
SimpleVLA-RL (GRPO)41.592.4
StructRL (PPO)49.196.6
π0.5SFT39.390.6
Sparse-RL (PPO)41.192.6
SimpleVLA-RL (GRPO)41.994.0
StructRL (PPO)45.896.2
Reward-component ablation
Reward-component ablation with PPO and GR00T-N1.5, adding one component at a time to the terminal binary reward: fixed subtask rewards, dynamic reward pacing, and structure-aware gating, which together form StructRL. Panels report SR (%) by task-horizon bucket and overall on RoboCasa365 (a-d) and LIBERO-Long (e-h). Values are means over three evaluation runs.

BibTeX

@article{yin2026structrl,
  title   = {StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks},
  author  = {Yin, Ziyi and Woo, Sangmin and Zhou, Kang and Kim, Sungyeon and Feng, Aosong and Ding, Haibo and Huan, Luke},
  journal = {arXiv preprint arXiv:2609.36352},
  year    = {2026}
}