---
title: "Egocentric Video Dataset"
description: "This egocentric dataset contains first-person videos of daily activities in home environments, designed for Physical AI, robotic systems, manipulation tasks, and egocentric vision research. The…"
url: "https://unidata.pro/datasets/egocentric-video/"
date_modified: "2026-06-16T15:50:32+03:00"
language: "en-US"
---
This egocentric dataset contains first-person videos of daily activities in home environments, designed for Physical AI, robotic systems, manipulation tasks, and egocentric vision research. The videos were captured using 2 hardware configurations: Pico + Motion Trackers (2,321 hours) and Zed + Pico + Motion Trackers (1,729 hours). Quaternion-based orientation supports 3D pose estimations and egocentric tracking from first-person.

## List of Parameters

- **Title:** Tasks — **Description:** Hand Activity Recognition, Egocentric Action Recognition, Hand-Object Interaction
- **Title:** Hours of recordings — **Description:** 4,050
- **Title:** Video source — **Description:** Pico 4 Ultra VR headset (egocentric) + 4 Zed stereo cameras (spatial depth)
- **Title:** Sensor data — **Description:** IMU signals — accelerometer, gyroscope, magnetometer
- **Title:** Activities — **Description:** Sorting, transferring, folding, assembly/disassembly, tool use, two-handed manipulation

## Dataset Structure

### The Numbers Section

**Numbered list:**

- **Number:** 4,050 — **Text:** Hours
- **Number:** 13 — **Text:** Scenarios

### Tooltips Section

**Tooltip items:**

- **Name:** Computer Vision
- **Name:** Human Activity Recognition
- **Name:** Robot Learning
- **Name:** Motion Analysis
- **Name:** Egocentric Vision

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | Egocentric video recordings of daily activities in home environments |
| Data types | Video |
| Tasks | Hand Activity Recognition, Egocentric Action Recognition, Hand-Object Interaction |
| Hours of recordings | 4,050 |
| Hardware setups | 2 |
| Setup 1 (Pico + Motion Trackers) | 2,321 hours (57.3%) — natural speed, slow-motion, and real-speed object transferring |
| Setup 2 (Zed + Pico + Motion Trackers) | 1,729 hours (42.7%) — scripted object transfer tasks with spatial depth + egocentric view |
| Scenarios | 13 (sorting unsorted items, arranging products by category, collecting items into a container, transferring from drawer to table, wardrobe & table & bag, transport box & display table, folding fabric items, lids & cookware & drawers, transferring with a spoon, transferring with tongs, packing into containers, two-handed sorting, assembly & disassembly) |
| Environments | Kitchen, bathroom, living room, and other home settings |
| Activities | Sorting, transferring, folding, assembly/disassembly, tool use, two-handed manipulation |

**Media Slider:** - **Video on Slayder:** [https://unidata.pro/wp-content/uploads/2026/05/egocentric\_new.mp4](https://unidata.pro/wp-content/uploads/2026/05/egocentric_new.mp4)

**Link to the sample:** [Download sample](https://drive.google.com/drive/folders/1XKuVs7T7cAbc7jT4yiyGHUzEPrYRDM8c)

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Video source | Pico 4 Ultra VR headset (egocentric) + 4 Zed stereo cameras (spatial depth) |
| Recording speed | Real-time / Slow-motion |
| Hand visibility | As needed / Both hands always in frame (slow-motion recordings) |
| Extension of labeling file | .txt per recording |
| Orientation format | Quaternions (from onboard sensor-fusion) |
| Sensor data | IMU signals — accelerometer, gyroscope, magnetometer |

**Source and data collection methodology:** Source and collection methodology: Data was captured using two different hardware configurations (Pico + motion trackers and Zed + Pico + motion trackers) while participants performed daily household activities in home environments.

### Statistics - Charts

**Charts with Titles:** - **Shortcode:** [ays_chart id='27'] — **caption above the graph:** Hardware configuration distribution

### LLM Languages

**Section Title:** Scenarios

**List of Statistics:**

- **Filter by:** Scenarios — **GIF image:** Top 20 Tags — **Table with data:**

| Variable | Hours | % |
| --- | --- | --- |
| Sorting unsorted items | 800 | 19.8% |
| Arranging products by category | 800 | 19.8% |
| Collecting items into a container | 400 | 9.9% |
| Transferring from drawer to table | 400 | 9.9% |
| Wardrobe, table & bag | 400 | 9.9% |
| Transport box & display table | 400 | 9.9% |
| Folding fabric items | 400 | 9.9% |
| Lids, cookware & drawers | 200 | 4.9% |
| Transferring with a spoon | 50 | 1.2% |
| Transferring with tongs | 50 | 1.2% |
| Packing into containers | 50 | 1.2% |
| Two-handed sorting | 50 | 1.2% |
| Assembly & disassembly | 50 | 1.2% |
| Total | 4,050 | 100% |

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Robotics & Physical AI — **Title:** Teaching Robotic Systems Real-World Manipulation — **Text:** First-person video paired with quaternion orientation data gives robotic systems the spatial context they need for manipulation tasks. The slow-motion subset captures detailed hand kinematics, making it practical training data for robotic arms learning object transfers. Thirteen scripted scenarios across varied conditions cover the range of daily household actions that real-world robotics actually encounters.
- **Industry:** Computer Vision & Egocentric Vision Research — **Title:** Benchmarking First-Person Action Recognition — **Text:** Researchers building egocentric action recognition models get six subsets recorded across varied speeds and hand visibility conditions — something most existing datasets don’t offer. Multimodal data covering hand-object interactions and 3D pose estimations support more accurate activity recognition. The large-scale egocentric footage spans dynamic scenes across kitchens, bathrooms, and living rooms.
- **Industry:** Healthcare & Rehabilitation — **Title:** Analyzing Human Motion for Clinical Applications — **Text:** Detailed hand motion data across scripted and naturalistic scenarios supports performance analysis in occupational therapy and motor rehabilitation. Pose estimations and egocentric tracking of hand-object interactions provide measurable visual data for assessing patient progress. The slow-motion subset makes fine-grained human motion visible in ways standard recording speeds miss.
- **Industry:** Augmented & Virtual Reality — **Title:** Improving Hand Tracking in Immersive Environments — **Text:** VR and AR developers can use this egocentric data to train hand tracking algorithms that hold up in real home environments. IMU signals and motion capture data from Pico VR headsets reflect how hands actually move during daily activity, covering the first-person perspectives immersive virtual applications need to feel responsive and accurate.

### Fact

**FAQs Heading:** FAQ

**List of Questions:**

- **Question:** How was the data collected? — **Answer:** Data was captured using two hardware configurations: Pico VR headset with motion trackers, and Zed + Pico + motion trackers, while participants performed daily household activities. Recording conditions vary across subsets, covering real-time speed, slow-motion kinematics, and scripted object transfer scenarios.
- **Question:** What makes this egocentric dataset suitable for robot learning? — **Answer:** Unlike generic action recognition datasets, this collection focuses on daily household manipulation tasks from a first-person perspective using wearable sensors. The combination of egocentric video, IMU signals, and realistic object interactions provides high-quality training data for Physical AI, embodied AI, and robot learning.
- **Question:** Is this dataset unique? — **Answer:** Yes. The combination of slow-motion kinematic recording, dual-camera setup (Zed + Pico), and scripted object transfer scenarios across six structured subsets is not replicated in existing datasets. The multimodal data, such as egocentric video, IMU signals, and quaternion orientation, make it distinct for robotic and Physical AI research.
- **Question:** Is this real-world or synthetic data? — **Answer:** This is entirely real-world data. Participants performed genuine daily household actions, including object transfers across surfaces, giving the dataset the natural human motion variability that synthetic datasets typically lack.
- **Question:** How are Unidata datasets licensed? — **Answer:** Unidata datasets follow a dual-licensing model: free samples are provided for trial and testing, while full datasets are available exclusively through purchase. Licensing terms are confirmed during the request and documentation process.
- **Question:** Do Unidata datasets comply with GDPR? — **Answer:** Yes. All datasets are curated in compliance with GDPR and applicable data protection laws. Data is collected from legally permissible sources to ensure ethical and lawful usage across research and commercial applications.
- **Question:** How are Unidata datasets stored? — **Answer:** All datasets are stored securely on AWS cloud infrastructure, aligned with ISO 27001 and ISO 27701 standards for information security and privacy management. This ensures high availability, scalability, and a privacy-focused environment for handling sensitive visual data.
- **Question:** How long does delivery take? — **Answer:** Once you submit a request, Unidata will contact you to review the details and complete the necessary documents. After signing and payment, the dataset is delivered within 3–10 days.

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
