RoboBrain-X0

A cross-embodiment vision-language-action model trained across heterogeneous robot platforms.

Overview

RoboBrain-X0 is a unified vision-language-action model for cross-embodiment action generation. It supports heterogeneous robot platforms through a shared action-token representation.

I contributed to the data pipeline rather than leading the project. My work focused on processing AgiBot-World and DROID data, including action normalization, tokenization, and structured conversion for large-scale VLA pretraining.

View the project on GitHub.