Rhoda’s new study examines how web-video pre-training affects Direct Video-Action policies for robot manipulation. The study varies model size and pre-training compute, then evaluates policies on an industrial bearing unpacking and sorting task with more than 200 hours of real-robot testing.
Scaling model size matters most in the reported task
Rhoda reports success rates rising from 4% for its smallest policy to 65%, 75% and 85% for larger XS, S, M and L configurations in the evaluated task. Each policy was tested over roughly 100 to 260 trials, according to the study page.
More pre-training helps, with diminishing returns
When model size is held in the study’s scaling comparison, increasing pre-training compute from 0.08× to 1× is associated with a reported rise from 58% to 75% task performance, with the largest differences appearing when demonstrations are limited.
A useful correlation, not a causal guarantee
The paper reports that a DINO feature-distance metric on held-out web video tracks robot performance in this setting. The result is evidence for a measurement relationship in one architecture family and task, not proof that web-video quality causes better real-world policies or that the result transfers to other robots.
