3 months ago

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan

Abstract

Vision-Language-Action (VLA) models enable robots to understand and perform complex tasks from multimodal input. Although recent work explores using reinforcement learning (RL) to automate the laborious data collection process in scaling supervised fine-tuning (SFT), applying large-scale RL to flow-based VLAs (e.g., π0, π0.5) remains challenging due to intractable action log-likelihoods from iterative denoising. We address this challenge with πRL, an open-source framework for training flow-based VLAs in parallel simulation. πRL implements two RL algorithms: (1) FlowNoise models the denoising process as a discrete-time MDP with a learnable noise network for exact log-likelihood computation. (2) Flow-SDE integrates denoising with agent-environment interaction, formulating a two-layer MDP that employs ODE-to-SDE conversion for efficient RL exploration. We evaluate πRL on LIBERO and ManiSkill benchmarks. On LIBERO, πRL boosts few-shot SFT models π0 and π0.5 from 57.6% to 97.6% and from 77.1% to 98.3%, respectively. In ManiSkill, we train πRL in 320 parallel environments, improving π0 from 41.6% to 85.7% and π0.5 from 40.0% to 84.8% across 4352 pick-and-place tasks, demonstrating scalable multitask RL under heterogeneous simulation. Overall, πRL achieves significant performance gains and stronger generalization over SFT-models, validating the effectiveness of online RL for flow-based VLAs.

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

HyperAI

3 months ago

Reinforcement Learning

Supervised Fine-Tuning

Multi-Task Learning

Method/Architecture

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan

Abstract

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

HyperAI

3 months ago

Reinforcement Learning

Supervised Fine-Tuning

Multi-Task Learning

Method/Architecture

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan

Abstract

Source PDF

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding

Ready-to-use GPUs

Best Pricing

Get Started View Pricing

HyperAI Newsletters

Subscribe to our latest updates

We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning

Command Palette

π𝚁𝙻: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan3 more

Abstract

Build AI with AI

HyperAI Newsletters

Command Palette

π𝚁𝙻: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan3 more

Abstract

Build AI with AI

HyperAI Newsletters

Command Palette

π𝚁𝙻: Online RL Fine-tuning for Flow-based Vision-Language-Action Models

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan3 more

Abstract

Build AI with AI

HyperAI Newsletters

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan

Kang Chen Zhihao Liu Tonghe Zhang Zhen Guo Si Xu Hao Lin Hongzhi Zang Quanlu Zhang Zhaofei Yu Guoliang Fan