Products

AI Research

Resources

About

Graphics

Computer Vision

Ego4D: Around the World in 3,000 Hours of Egocentric Video

October 14, 2021

Abstract

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,025 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 855 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception.
Project Page:https://ego4d-data.org/

Download the Paper

AUTHORS

Written by

Kristen Grauman

Andrew Westbury

Eugene Byrne

Zachary Chavis

Antonino Furnari

Rohit Girdhar

Jackson Hamburger

Hao Jiang

Miao Liu

Xingyu Liu

Miguel Martin

Tushar Nagarajan

Ilija Radosavovic

Santhosh Ramakrishnan

Fiona Ryan

Jayant Sharma

Michael Wray

Mengmeng Xu

Eric Zhongcong Xu

Chen Zhao

Siddhant Bansal

Vincent Cartillier

Sean Crane

Tien Do

Akshay Erapall

Christoph Feichtenhofer

Adriano Fragomeni

Qichen Fu

Christian Fuegen

Abrham Gebreselasie

Cristina Gonzalez

James Hillis

Xuhua Huang

Yifei Huang

Wenqi Jia

Weslie Khoo

Jachym Kolar

Anurag Kumar

Federico Landini

Chao Li

Zhenqiang Li

Karttikeya Mangalam

Raghava Modhugu

Jonathan Munro

Tullie Murrell

Takumi Nishiyasu

Will Price

Paola Ruiz Puentes

Merey Ramazanova

Leda Sari

Kiran Somasundaram

Audrey Southerland

Yusuke Sugano

Ruijie Tao

Minh Vo

Yuchen Wang

Xindi Wu

Takuma Yagi

Yunyi Zhu

Pablo Arbelaez

David Crandall

Dima Damen

Giovanni Maria Farinella

Bernard Ghanem

Vamsi Krishna Ithapu

C. V. Jawahar

Hanbyul Joo

Kris Kitani

Haizhou Li

Richard Newcombe

Aude Oliva

Hyun Soo Park

James M. Rehg

Yoichi Sato

Jianbo Shi

Mike Zheng Shou

Antonio Torralba

Lorenzo Torresani

Mingfei Yan

Jitendra Malik

Publisher

arXiv

Research Topics

Computer Vision

Graphics

Related Publications

July 13, 2026

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez, Devendra Singh Sachan, Barlas Oguz, Seungwhan Moon, Shang-Wen Li, Gargi Ghosh, Xin Dong, Wen-Tau Yih

July 13, 2026

July 03, 2026

Human & Machine Intelligence

Robotics

Interpreting Physics in Video World Models

Sonia Joseph, Quentin Garrido, Randall Balestriero, Matthew Kowal, Thomas Fel, Shahab Bakhtiari, Blake Richards, Mike Rabbat

July 03, 2026

May 26, 2026

Human & Machine Intelligence

Theory

Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images

Josephine Raugel, Max Seitzer, Marc Szafraniec, Huy V. Vo, Jérémy Rapin, Patrick Labatut, Piotr Bojanowski, Valentin Wyart, Jean Remi King

May 26, 2026

May 19, 2026

Human & Machine Intelligence

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

Dongyan Lin, Phillip Rust, Angel Villar Corrales, Alvin W. M. Tan, Mahi Luthra, Charles-Eric Saint-James, Rashel Moritz, Sheila Krogh-Jespersen, Vanessa Stark, Surya Parimi, Jiayi Shen, Youssef Benchekroun, Yosuke Higuchi, Martin Gleize, Tom Fizycki, Nicolas Hamilakis, Manel Khentout, Sho Tsuji, Balázs Kégl, Juan Pino, Michael C. Frank, Emmanuel Dupoux

May 19, 2026

June 11, 2019

Computer Vision

ELF OpenGo: An Analysis and Open Reimplementation of AlphaZero | Facebook AI Research

Yuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta, Zhuoyuan Chen, James Pinkerton, Larry Zitnick

June 11, 2019

April 30, 2018

NLP

Computer Vision

Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent | Facebook AI Research

Zhilin Yang, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston

April 30, 2018

October 10, 2016

Speech & Audio

Computer Vision

Polysemous Codes | Facebook AI Research

Matthijs Douze, Hervé Jégou, Florent Perronnin

October 10, 2016

June 18, 2018

Speech & Audio

Computer Vision

Low-shot learning with large-scale diffusion | Facebook AI Research

Matthijs Douze, Arthur Szlam, Bharath Hariharan, Hervé Jégou

June 18, 2018

Help Us Pioneer The Future of AI

We share our open source frameworks, tools, libraries, and models for everything from research exploration to large-scale production deployment.

About AI at Meta

Media Generation

Foundational models

Our approach

Our approach About AI at Meta People Careers

Research

Research Infrastructure Resources Demos

Meta AI

Meta AI Assistant Media Generation Vibes AI Studio

Latest news

Latest news Blog Newsletter

Foundational models

Meta © 2026