Long visuomotor context is one of THE most important missing parts for robotic foundation models, there is no way around it if we want generic policies that remember more than a couple of camera frames. However, adding visual context naively significantly increases inference time
We scaled robot policies to 8K timesteps of visuomotor context, orders of magnitude beyond current SoTAs, at constant inference latency.
Introducing RoboTTT 🤖
With minutes of experience in context, our robots:
🎥 one-shot imitate human video demos
📈 improve themselves during




