Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
November 24, 2025International Journal of Computer VisionOpen Access

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

View Full Paper
Ask AI
Bookmark
Share

Authors

KGKristen GraumanThe University of Texas at AustinAWAndrew WestburyMeta (United States)LTLorenzo TorresaniUniversidad del Noreste

Discussion

Loading...

Member takes

Implication

Multimodal benchmark study demonstrates synchronized first- and third-person captures across skilled human activities, highlighting new targets for fine-grained computer vision.

Key Points

  • Introduce a large-scale, multimodal, multiview dataset and benchmark suite to advance machine perception of complex, skilled human activities from both first-person and third-person viewpoints.
  • Recorded 740 participants across 13 cities and 123 natural scene contexts performing skilled activities, including sports, music, dance, and bike repair, in sessions lasting 1 to 42 minutes.
  • Synchronized egocentric and exocentric video streams alongside multichannel audio, eye gaze tracking, 3D point clouds, camera poses, inertial measurement unit (IMU) data, and expert coach commentaries.
  • Constructed evaluation benchmarks for fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand and body pose estimation.
  • Assembled a publicly accessible multimodal dataset providing 1,286 total hours of synchronized first- and third-person video across diverse real-world environments.
  • Established foundational benchmark tasks and annotations to support multi-angle activity recognition, skill assessment, and spatial human pose tracking.

Cite This Study

Grauman et al. (2025) studied this question.

synapsesocial.com/papers/6a7d24ec2f71449e10424e3chttps://doi.org/10.1007/s11263-025-02557-6
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Ego4D: Around the World in 3,000 Hours of Egocentric Video2022 · 600 citations
  2. 2SUN3D: A Database of Big Spaces Reconstructed Using SfM and Object Labels2013 · 856 citations
  3. 3Procedure-Aware Pretraining for Instructional Video Understanding2023 · 33 citations