Vela sailboat logoVela

Scaling Vision-Language-Action Models
with Adaptive Action Curve Parametrization

Yifan Li1,2,*Jiaxu Wang3,*Dongming Wu3Yicheng Jiang3Ryan Ji4Xiangyu Yue3,†Yanwei Fu1,2,†
1Fudan University2Shanghai Innovation Institute3The Chinese University of Hong Kong4NovaXBot

* Equal contribution † Corresponding authors

Cooking egg cake

1:29

Tool use and coordinated bimanual manipulation across four stages.

Potato shredding

1:13

Sustained contact and precise coordination between the potato and grater.

The idea

From actions to trajectories.

Vela predicts a compact curve and its temporal span.
The same control-point budget covers long, smooth motions or shorter, finely resolved movements.

Vela takes visual observations, an instruction and proprioception, and jointly predicts control points and a horizon. A fixed spline budget represents long smooth motion or short precise motion, and the curve can be sampled at different rates.
Fixed representation. Flexible temporal support.

01 Compact geometry

A compact set of spline control points defines a continuous action trajectory.

02 Adaptive horizon

A motion-dependent horizon adapts the temporal span to smooth or precise movements.

03 Scalable pretraining

A shared action interface supports learning across heterogeneous robot embodiments.

Pretrained in trajectory space on approximately 40,000 hours of public and private robot data.

Evaluation

Across tasks. Across control demands.

From real-world bimanual manipulation to diverse simulation benchmarks.

Real-world manipulation

Mean subtask success rate · four stages · 10 trials per stage

Cooking egg cake

67.5%
Egg-cake success rates: Vela 100, 80, 60 and 30 percent across the four stages, compared with pi 0.5 and OpenWAM.

Potato shredding

87.5%
Potato-shredding success rates: Vela 90, 80, 90 and 90 percent across the four stages, compared with pi 0.5 and OpenWAM.

Task progressionCooking
egg cake

Four rows of keyframes showing Vela completing the egg-cake task.

Task progressionPotato
shredding

Four rows of keyframes showing Vela completing the potato-shredding task.

Simulation benchmarks

LIBERO-X and EBench

Vela retains the same architecture, trainable parameter count, and input modalities as π0.5, changing the output modeling space to action curves. Benchmark post-training follows the official π0.5 recipes.

LIBERO-X

Average success rate

+6.0 pp
45.3%
Vela
45.3%
π0.5
39.3%
03060%

EBench

Overall success rate

+8.3 pp
49.7%
Vela
49.7%
π0.5
41.4%
03060%

Gains are absolute percentage-point improvements over π0.5. Results follow the manuscript’s evaluation protocols.

Benchmark details

LIBERO-X

Success rate (%)
LIBERO-X · Success rate (%)
ModelL1L2L3L4L5Avg.
OpenVLA-OFT29.017.68.86.44.213.2
π029.421.911.07.65.115.0
X-VLA30.122.610.36.04.114.6
OpenWAM-α35.928.221.613.010.021.7
GR00T N1.543.332.918.713.39.723.6
τ0-VLA46.534.220.513.010.524.9
Xiaomi-R071.449.633.722.117.238.8
π0.565.253.236.024.118.039.3
π0.5-Vela66.153.540.727.322.542.0
Vela67.255.944.931.527.145.3

EBench

Success rate (SR) and task-progress Score (%)
EBench · Success rate (SR) and task-progress Score (%)
ModelTabletopPick & placeLong horizonOverall
SRScoreSRScoreSRScoreSRScore
X-VLA8.62450.0546.22523.736
InternVLA-A14.31143.04717.94623.936
FastWAM2.91349.55316.93825.637
Cosmos3-Edge18.63241.54424.14929.342
GigaBrain-0.737.95942.04520.03733.346
π032.94941.54525.74933.747
InternVLA-A1.512.12250.05333.75934.246
Qwen-RobotManip50.07056.56029.95545.660
π0.529.34556.56134.15541.454
π0.5-Vela34.35459.06328.55542.058
Vela44.36463.06739.16549.766

π0.5-Vela is initialized from the π0.5 checkpoint and post-trained in trajectory space; Vela is pretrained directly in trajectory space.

Citation

@misc{li2026vela,
  title = {Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization},
  author = {Yifan Li and Jiaxu Wang and Dongming Wu and Yicheng Jiang and Ryan Ji and Xiangyu Yue and Yanwei Fu},
  year = {2026}
}
Download BibTeX
Vela Method Figure