March 29, 2023
We present the largest and most comprehensive empirical study of pre-trained visual representations (PVRs) or visual ‘foundation models’ for Embodied AI. First, we curate CortexBench, consisting of 17 different tasks spanning locomotion, navigation, dexterous, and mobile manipulation. Next, we systematically evaluate existing PVRs and find that none is universally dominant. To study the effect of pre-training data scale and diversity, we combine over 4,000 hours of egocentric videos from 7 different sources (over 5.6M images) and ImageNet to train different-sized vision transformers using Masked Auto-Encoding (MAE) on slices of this data. Contrary to inferences from prior work, we find that scaling dataset size and diversity does not improve performance universally (but does so on average). Our largest model, named VC-1, outperforms all prior PVRs on average but does not universally dominate either. Finally, we show that task- or domain-specific adaptation of VC-1 leads to substantial gains, with VC-1 (adapted) achieving competitive or superior performance than the best known results on all of the benchmarks in CortexBench. These models required over 10,000 GPU-hours to train and can be found on our website for the benefit of the research community.
Written by
Franziska Meier
Aravind Rajeswaran
Jitendra Malik
Karmesh Yadav
Oleksandr Maksymets
Sergio Arnaud
Sneha Silwal
Vincent-Pierre Berges
Aryan Jain
Claire Chen
Jason Ma
Yixin Lin
Publisher
Arxiv
Research Topics
Robotics
June 11, 2025
Aaron Foss, Ammar Rizvi, Chloe Evans, Justine T. Kao, Koustuv Sinha, Sasha Mitts
June 11, 2025
June 11, 2025
Mojtaba Komeili, Sarath Chandar, Abha Gejji, Ada Martin, Adrien Bardes, Ammar Rizvi, Artem Zholus, Claire Roberts, Daniel Dugas, David Fan, Francisco Massa, Francois Robert Hogan, Franziska Meier, Kapil Krishnakumar, Koustuv Sinha, Marc Szafraniec, Matthew Muckley, Mido Assran, Michael Rabbat, Nicolas Ballas, Patrick Labatut, Piotr Bojanowski, Quentin Garrido, Russell Howes, Sergio Arnaud, Vasil Khalidov, Xiaodong Ma, Yann LeCun, Yong Li
June 11, 2025
April 17, 2025
Ruslan Partsey, Ayush Jain, Ang Cao, Ishita Prasad, Aravind Rajeswaran, Abha Gejji, Ada Martin, Arjun Majumdar, Daniel Dugas, Franziska Meier, Krishna Murthy Jatavallabhula, Mido Assran, Mikael Henaff, Mike Rabbat, Mrinal Kalakrishnan, Nicolas Ballas, Oleksandr Maksymets, Paul McVay, Phillip Thomas, Alexander Sax, Sergio Arnaud, Vincent-Pierre Berges
April 17, 2025
October 31, 2024
Nolan Black, Romeo Mercado, Norb Tydingco, Gregg Kammerer, Ricardo Chavira, Eric Sanchez, Yitian Ding, Roberto Calandra, Mike Lambeta, Alexander Sohn, Ali Sengül, Byron Taylor, Dave Stroud, Haozhi Qi, Jake Khatha, Jitendra Malik, Kevin Sawyer, Kurt Jenkins, Kyle Most, Neal Stein, Thomas Craven-Bartle, Tingfan Wu, Victoria Rose Most
October 31, 2024

Our approach
Latest news
Foundational models