First-person data,
at the scale robots need.
One of the largest egocentric datasets focused on real-world manual labor: roughly 1.9 million hours of first-person video, captured in live work environments and annotated with 3D pose and semantics for Embodied AI and world-model training.
Real labor, captured first-person
We've built one of the largest egocentric vision datasets focused on real-world manual labor: roughly 1.9M hours of first-person video, collected with custom head-mounted devices in live work environments. Participants wear the device through normal workflows (typically 5–20 hours each), so behavior stays natural across long sequences with minimal inductive bias. These are economically useful physical tasks, captured without scripting or artificial constraints.
It spans light-manufacturing assembly, warehouse logistics, construction and skilled trades, commercial cleaning, food service, and agriculture. The interaction data is dense and high-signal, and hard to replicate synthetically. Footage is 1080p with a wide ~180° field of view, paired with IMU and audio, and hands are visible 80–90% of the time, which suits manipulation-heavy modeling.
On top of the raw data we provide structured annotations and derived signals: 3D pose (with hands), point tracking, motion features, depth where available, and semantic labels covering tasks, environments, actions, objects/tools, and step structure. Everything is filtered with vision-language models for consistency and quality.
Where egocentric data sits in the pyramid
Web video is abundant but shallow, with no first-person, contact-rich signal. Real-robot teleoperation captures that signal but can't scale past lab hardware. Egocentric human data keeps the scale of human video and the manipulation fidelity robots need to learn.
Every clip, measured
Uniform capture parameters and metadata across the collection, so recordings are directly comparable and trainable out of the box.
What's in the data
Category mix across a representative sample of the collection.
All 26 categories in the collection are represented above; no single category exceeds 14.5% of the modeled sample.
Task and difficulty spread
Task and difficulty distribution across a representative sample of the collection.
Manufacturing/factory categories are broken out using the same 7 sub-categories and shares as the assembly-manufacturing 50K distribution, so the two line up directly.
Every clip runs this pipeline
Stages are cost-ordered — cheap checks run first, and a clip only reaches expensive model inference once it survives everything before it.
Task & environment labeling
structured taxonomy classification
An admission-independent pass assigns sector, task, and difficulty labels from a fixed taxonomy. Labeling never affects whether a clip was already accepted.
Production-scale 3D pose & semantics
Our post-processing turns raw first-person video into structured, learning-ready labels. Three annotation types, at production scale.
Filtered with vision-language models for consistency and quality; frame-accurate segmentation and language descriptions align visual observation, physical motion, and task semantics.
Delivered files
Every clip ships as a set of files sharing one clip id.
What you get, and how to get more
- Curated sample clips, playable inline
- Synchronized IMU stream per clip
- 3D hand-pose overlay and reconstructed pose video
- Per-clip task, environment, and rig metadata
Samples are for evaluation, not licensed for training use.
- ~1.9M hours of first-person video across the full category mix
- 3D hand + full-body pose, point tracking, motion features at production scale
- Frame-accurate semantic labels: task, environment, action, object/tool
- Custom collection scoped to your task/environment coverage needs
Questions about the data?
Modalities, licensing, task coverage, annotation formats, or a custom collection. Send a question and we'll reply by email.