A robot policy trained on the wrong kind of video will pass every offline evaluation and still fail on the workbench. The usual cause isn't the model — it's footage shot from a viewpoint the robot's own camera will never see, with no depth, no hand pose, and no way to tell a completed grasp from a failed one.
That's the specific gap egocentric data providers exist to close, and in the last year the market for buying this data — rather than shooting it yourself — has gone from a handful of research labs to a real, if still small, commercial category.

This guide compares the providers actually operating in that category today: full-stack capture-and-enrichment vendors, contributor-network marketplaces, annotation platforms that work with egocentric footage you already have, and the open academic datasets every buyer ends up benchmarking against before they spend a dollar.
Scope note: this covers egocentric data for robotics, embodied AI, and physical AI specifically — not the broader academic field of egocentric vision (life-logging, AR/VR research), which has a different provider landscape and different buying criteria.
What Are Egocentric Data Providers and Why Do They Matter for Robotics?
Egocentric data providers capture, enrich, and license first-person video and sensor data — footage recorded from the viewpoint of whoever (or whatever) is performing a task, rather than from a fixed external camera. In robotics specifically, this data trains and fine-tunes manipulation policies and vision-language-action (VLA) models, because the viewpoint, hand-object geometry, and attention signal it captures sit far closer to what a robot's onboard camera sees at deployment than third-person footage does.
What Do Egocentric Data Providers Actually Do?
Three functions, in order: capture (wearable or head-mounted rigs, sometimes robot-mounted cameras, recording real task performance), enrichment (depth, 3D pose, segmentation, and action annotation layered onto the raw video), and delivery (licensed, QA-verified datasets — either off a shelf or built to a custom scenario spec).
Why Can't Teams Just Use Ego4D or EPIC-Kitchens?
Usually not as the sole source. Academic egocentric datasets were built for action recognition and perception research, not for training a manipulation policy that has to predict what to do next. Ego4D and EPIC-Kitchens are head-mounted-only, use a daily-life action vocabulary, don't ship per-joint hand pose as a primary deliverable, and Ego4D's license restricts commercial use [10].
They're strong for pretraining perception components — weak as the primary training signal for a production policy. This distinction runs through the whole guide and gets its own section below.
How to Choose the Right Egocentric Data Provider
Nine criteria separate a provider that will actually move your policy's performance from one that just ships first-person video. Build a scorecard around these before you take a vendor call.
- Camera/sensor hardware and FOV matching. Does the rig match common robot sensor geometry — wrist-mounted, head-mounted, fisheye — or only one fixed configuration? A field-of-view mismatch between training footage and the robot's deployed camera is one of the most common causes of a policy that works in evaluation and fails on hardware.
- Multimodal synchronization. Is depth, IMU, and skeleton data captured and time-synced to the RGB stream, or is it RGB-only? Flat video without depth or camera pose is usable for pretraining, not for final fine-tuning.
- Annotation depth. Does the provider deliver sub-action temporal segmentation and per-joint hand pose, or just segment-level action labels? The gap between these two is the gap between trainable data and footage that only looks right.
- Commercial licensing and rights clearance. Verify explicit commercial-use rights on every participant and environment, not just "research use." This is the single biggest practical difference between a commercial provider and an academic dataset.
- Scenario diversity and scripting. Structured coverage of edge cases and failure/recovery sequences beats unstructured "natural activity" footage, which clusters around common actions and produces almost nothing of the hard cases a policy will fail on.
- Data scale and delivery cadence. Total hours or scenarios available or collectible, and how fast new batches ship once a project starts.
- Format compatibility. RLDS, LeRobot, raw MP4 plus metadata, or a visualization format like Rerun's .rrd — check it plugs into your training pipeline without a conversion step.
- Privacy and consent compliance. GDPR-equivalent consent protocols for first-person biometric video, consent from anyone who might appear in the background, and a stated anonymization approach.
- Sample access before commitment. Can you inspect real recordings — not marketing stills — before signing an NDA or contract? Ungated sample access is rarer than it should be in this market and is worth checking for directly.
Best Egocentric Data Providers for Robotics & Embodied AI
Provider Overview
| Provider | Best For | Capture Model | Commercial Licensing |
|---|---|---|---|
| Unidata | Production multimodal capture with the deepest verified sync | In-house rigs (Pico 4 Ultra + motion trackers + ZED) | Explicit commercial release per participant |
| Claru | Enriched, ready-to-license libraries at scale | Contributor network + custom collection | Commercial license stated per dataset |
| Objectways | Managed end-to-end egocentric pipelines | In-house/managed capture teams | Project-based, verify per contract |
| Labellerr | Capture plus annotation in one workflow | Managed capture + AI-assisted labeling | Project-based |
| Luel | Fast, rights-cleared marketplace access | Large global contributor network | Rights-cleared, consent-infrastructure model [4] |
| Awign | High-volume, budget-oriented collection | Gig workforce, India-based | Verify per contract |
| Build AI | Industrial/factory-floor scenarios | Open-source-adjacent large dataset + collection | Verify per contract |
| Lightwheel | High-volume collection across varied settings | Managed capture teams | Verify per contract |
| Encord, Scale AI, Appen, Labelbox | Annotating egocentric footage you already have | Annotation/infrastructure platform — not capture-first | Standard enterprise terms |
Technical Depth Comparison
| Provider | Multimodal Sync (depth/IMU/skeleton) | Annotation Schema Depth | Format Support | Notable Scale |
|---|---|---|---|---|
| Unidata | Yes — 6 synchronized streams (stereo, depth, hand tracking, 6DoF pose, full-body skeleton) | Auto-generated 3D pose + scene/action/object/hand-pose labels | .mp4, .rrd, .json, pose/hand-keypoint files; Parquet/WebDataset on request | 38,457 scenarios, 10,255 hours captured to date; 15,000+ scenario licensable base library |
| Claru | Depth, segmentation, pose as annotation layers | Multi-layer: depth, segmentation, pose, action | Not fully published — verify at contract stage | 500K+ clips claimed |
| Objectways | Varies by project; documented best-practice guidance on camera placement and capture completeness | Verify per project | Verify per project | Not publicly quantified |
| Labellerr | Capture + AI-assisted annotation, up to 99% claimed accuracy | Full-stack capture and annotation | Verify per project | Not publicly quantified |
| Luel | IMU sensor data paired with first-person video | Varies — marketplace model, not a fixed schema | Varies by dataset in catalog | 500K contributors across 96 countries claimed |
| Ego4D (open) | 3D scans, audio, some multimodal — not robot-camera-matched | Benchmark-level (episodic memory, hand-object interaction) | Academic release format | 3,670 hours, 923 participants, 74 locations |
Egocentric Data Providers, Ranked and Reviewed
Tier 1 — Full-Stack Capture and Enrichment Providers
Unidata — best for production-grade multimodal capture with verified specs.

Unidata runs two production capture rigs: Pico 4 Ultra plus motion trackers, and the same setup extended with a ZED multi-camera array (ZED X Mini plus three ZED X One units) for depth and wider coverage — eight device types and five rig configurations in active use, with up to 15 synchronized devices in a single session [19]. Every session ships six synchronized streams: stereo passthrough, hand tracking, 6DoF pose, and depth from the headset, plus full-body skeleton reconstructed from wrist, ankle, and pelvis trackers [19]. Running totals published on the service page put captured volume at 38,457 scenarios and 10,255 hours to date, with a licensable base library of 15,000+ structured scenarios across 7 room types and 90+ functional zones — coffee-making, cooking, object handling, everyday chores [19]. The packaged Egocentric Video Dataset alone ships 4,050 hours across 13 household scenarios, captured under two hardware configurations and delivered with quaternion-based 3D pose generated automatically from onboard sensor fusion, so no manual pose-annotation pass is required [20].
The detail worth checking directly: Unidata publishes live Rerun-viewer links to real scenario recordings, so a buyer can inspect actual multimodal data — RGB, depth, pose, hand tracking — before signing an NDA or a contract [19]. None of the other providers in this comparison offer an equivalent ungated sample. Custom collection projects typically start delivering within two weeks of scoping, and every participant signs a commercial-use release before recording, covering faces, voices, and home interiors for downstream commercial AI products [19]. Two supporting cases show the same rig in client work: a humanoid robot developer's project added instrumented tactile gloves to the Pico 4 Ultra and motion-tracker setup specifically because heterogeneous object mass and surface texture aren't visible in video alone [24], and a separate simulation-data engagement scanned the same physical spaces used for egocentric capture with LiDAR and photogrammetry, so a client's simulation environment and real footage point to the same room geometry [25].
Unidata's own production write-up on this work is unusually candid about what isn't solved yet — tactile/instrumented-glove integration is explicitly on the roadmap, not yet shipped as a standard deliverable, and force/gripper-pressure data is described as the next frontier the whole industry is missing [26]. The one dataset in Unidata's public lineup that isn't first-party captured is the Robotic Household Activities Dataset (1,000 hours, head-mounted video plus seven 9-axis IMU units), sourced through a verified Unidata partner rather than shot in-house [21] — worth knowing if first-party-only capture matters to your procurement process.
Claru — best for large, pre-enriched licensable libraries.

Claru positions itself explicitly as a commercial alternative to Ego4D, claiming 500K+ licensed first-person clips with depth, segmentation, and pose as standard annotation layers, collected across kitchen, workshop, warehouse, and retail categories via three capture pipelines using wearable cameras and smartphones [2]. The company's own comparison piece frames its edge as enrichment depth and delivery speed against providers that ship raw video only [1]. Verify current clip counts and licensing terms directly before a purchase decision — the AI training data market moves fast enough that vendor-published figures age quickly.
Objectways — best for buyers who want a fully managed pipeline.

Objectways runs egocentric collection as a structured, end-to-end service and has published detailed operational guidance on its own blog — correct camera placement, avoiding motion blur, capturing complete interaction sequences rather than partial task flows [7] [8]. This kind of public methodology content is a reasonable proxy for process maturity, though the company doesn't publish scale figures as concretely as Unidata or Claru do.
Labellerr — best for teams that want capture and annotation from one vendor.

Labellerr combines egocentric capture with AI-assisted annotation in a single workflow, claiming up to 99% annotation accuracy through combined human and model-assisted labeling [3]. Positioned as a full robotics-data platform rather than an egocentric specialist first.
Tier 2 — Marketplace and Contributor-Network Models
- Luel — best for fast, budget-conscious access via a large contributor base. Luel is a Y Combinator-backed marketplace (W2026) that launched in early 2026 and reported a $31.2M seed round with a claimed 500,000 contributors across 96 countries, built around a "rights-cleared" consent infrastructure for multimodal data including first-person video paired with IMU sensor data [4] [5] [6]. The marketplace model trades deep enrichment for speed and contributor-network scale — useful for teams that need volume fast and can do their own annotation, less suited to buyers who need contact-point or per-joint hand pose out of the box. As with any recently funded, fast-scaling marketplace, verify current contributor and dataset-quality figures directly rather than relying on launch-period press.
- Awign — best for high-volume collection on a tighter budget. India-based, gig-workforce model, positioned around scale and cost rather than enrichment depth [3].
- Build AI — best for industrial and factory-floor scenarios. Positioned around a large, open-source-adjacent dataset for industrial robotics use cases [3].
- Lightwheel — best for high-volume collection across varied environments. Positioned on output volume and environment diversity rather than annotation schema depth [3].
Tier 3 — Annotation and Infrastructure Platforms (Not Egocentric-Capture Specialists)
Encord, Scale AI, Appen, and Labelbox are established annotation and data-infrastructure platforms. Encord in particular markets itself as a fully multimodal annotation and curation tool — image, video, DICOM, document, text, and audio in one interface — and can annotate egocentric footage a client already has, but it is a curation and labeling platform, not a company that captures first-person video [9]. Scale AI and Appen offer some managed capture services as part of much broader data-collection catalogs, but neither is egocentric-capture-first the way the Tier 1 and Tier 2 names are.
This distinction matters: several existing "best egocentric provider" roundups list these platforms alongside true capture specialists without making clear that "can annotate egocentric video" and "collects egocentric video" are different capabilities — check which one you're actually buying before you sign.
The Open Academic Datasets Buyers Benchmark Against
These aren't vendors — you can't license them the way you'd license from a commercial provider — but every serious buyer compares commercial offers against them first, so they belong in this comparison with that distinction stated plainly.
- Ego4D (Meta AI, 2022) — 3,670 hours from 923 participants across 74 locations in 9 countries, supporting benchmarks for episodic memory, hand-object interaction, and forecasting. The reference dataset in the field. License restricts commercial use [10].
- Ego-Exo4D (Meta FAIR plus 15 university partners, 2023–24) — 1,286–1,422 hours of synchronized egocentric-and-exocentric video across skilled human activities (sports, music, cooking, bike repair), captured with Project Aria devices from over 740 participants in 13 cities [11] [12].
- EPIC-Kitchens — roughly 100 hours of kitchen activity with strong action-segment annotation; narrow domain, doesn't generalize past cooking tasks.
- Project Aria (Meta) — not a dataset but the reference hardware platform (glasses form factor, eye tracking, on-device SLAM) that several commercial providers' rigs get benchmarked against [14].
- EgoDex (Apple Research, 2025) — 829 hours of egocentric video with paired 3D hand and finger joint tracking across 194 tabletop tasks, captured with Apple Vision Pro and on-device SLAM — the closest academic dataset to production-grade manipulation annotation depth currently public [15].
Should You Use Open Datasets Instead of a Commercial Provider?
Usually not as your only source. The distribution gap between how humans capture data and how a robot's onboard camera observes the world is a hardware problem as much as a licensing one: a head-mounted camera and a wrist-mounted end-effector camera see fundamentally different geometry, and a policy trained on one perspective won't transfer cleanly to deployment on the other [26].
Ego4D-class datasets are genuinely useful for pretraining perception components — hand-object detectors, action forecasters — but weren't built to be the primary training signal for a manipulation policy, and commercial-use restrictions rule several of them out for production work outright.

The practical pattern that works: use open datasets for early pretraining and benchmarking, and commercial or custom-collected data — matched to your robot's actual camera geometry — for the fine-tuning stage that determines whether the policy works on hardware.
Egocentric Data Provider Pricing in 2026
None of the providers compared here publish flat per-hour rates for egocentric data — pricing is quoted per project, driven by four factors: capture complexity (RGB-only versus a synchronized multimodal rig), annotation depth requested, whether you need off-the-shelf licensing or custom scenario collection, and exclusivity terms. Treat any public "per hour" figure you see for this category with caution; it's rarely the number you'll actually be quoted.
For context on the wider cost envelope this sits inside: teleoperation and demonstration-collection hardware for robot data generally runs $6,000–$80,000 depending on setup complexity, and a production-grade robot training dataset of 500+ demonstrations has been estimated at $50,000–$200,000 all-in, covering hardware, operator labor, and post-processing [18].
One useful reframe from a recent industry analysis: cost per hour is the wrong metric to optimize for egocentric and robot data generally, because value differs by orders of magnitude across data types — a repetitive demonstration and a genuinely novel failure-recovery sequence cost the same to capture but are not remotely equal in training value [17]. Budget around diversity and verification, not raw hours purchased.
Egocentric Data Providers by Use Case
- Best for VLA/foundation-model pretraining at scale: providers with large existing licensed libraries — Claru's claimed 500K+ clip catalog, or open Ego4D/EgoDex for non-commercial pretraining work.
- Best for production manipulation-policy fine-tuning: providers offering robot-matched FOV, per-joint hand pose, and multimodal sync out of the box — Unidata's documented rig and annotation depth is the most concretely verifiable option here; Objectways is a reasonable second look for managed custom collection.
- Best for rapid, budget-constrained pilots: marketplace and contributor-network models — Luel for speed and rights-cleared access, Awign where cost matters more than enrichment depth.
- Best for teams that already have footage and need annotation only: Encord, Scale AI, or Labelbox — not capture providers, but capable of turning existing raw video into a labeled training set.
Egocentric data trends to watch in 2026 and beyond
- Scaling laws are now empirically grounded, not theoretical. Research on human egocentric video pretraining has found a log-linear relationship between data volume and downstream robot task performance, turning egocentric data acquisition into a budgeted infrastructure line item rather than an experimental spend [16].
- Force and tactile data are becoming the next enrichment layer providers will need to offer. Video captures what a hand looks like during a grasp; it can't capture how hard the grasp is or whether it's about to slip. Several providers, including Unidata, list tactile integration as an active roadmap item rather than a shipped feature today [26] — expect this gap to close over the next 12–18 months.
- Robot-matched camera geometry is becoming an explicit selling point. Wrist-mounted and fisheye rigs that mirror actual robot sensor placement are increasingly advertised directly, rather than left for the buyer to ask about.
- Consolidation risk is real in the marketplace tier. Several Tier 2 entrants (Luel among them) are very recently funded — seed or early Series A as of 2026 [6]. That's not a reason to avoid them, but it is a vendor-diligence question worth asking directly: what happens to a dataset license or an in-progress custom collection if the vendor is acquired or winds down.
📑 Article Disclaimer
The information contained in this article is provided for general informational and editorial purposes only. The content reflects the opinions, research, and editorial judgment of the author(s) at the time of publication and does not constitute professional, legal, financial, or business advice of any kind.
No Endorsement or Warranty
The mention, ranking, or listing of any company, product, or service within this article does not constitute an endorsement, recommendation, or guarantee of quality, reliability, or fitness for any particular purpose. The publisher makes no representations or warranties, express or implied, regarding the accuracy, completeness, timeliness, or suitability of the information provided.
Independence of Judgment
Readers are strongly encouraged to conduct their own independent research and due diligence before engaging with, contracting, or entering into any business relationship with any of the companies referenced herein. The inclusion or exclusion of any company does not imply a definitive assessment of its capabilities, compliance, or business conduct.
No Liability
To the fullest extent permitted by applicable law, the publisher, editors, authors, and any affiliated parties expressly disclaim all liability for any direct, indirect, incidental, consequential, or punitive damages arising from reliance on the information contained in this article, including but not limited to decisions made on the basis of company rankings or descriptions.
Third-Party Information
Certain information presented in this article may be sourced from third parties, publicly available data, or company self-disclosures. The publisher does not independently verify all such information and assumes no responsibility for errors, omissions, or changes occurring after the date of publication.
No Legal or Regulatory Advice
Nothing in this article should be construed as legal, regulatory, or compliance guidance. Data collection practices are subject to varying laws and regulations across jurisdictions. Readers should consult qualified legal counsel regarding their specific circumstances and applicable law.
Subject to Change
The data collection industry is dynamic. Company rankings, capabilities, and reputations are subject to change. This article represents a snapshot in time and may not reflect current market conditions.
By accessing and reading this article, you acknowledge and agree to the terms of this disclaimer.
- AI Training
- Robotics
- Data Annotation
- Python
- AWS
Sources & References
- [1] Claru — "7 Best Egocentric Video Data Providers for Robotics (2026)" — 2026 — https://claru.ai/blog/best-egocentric-data-providers
- [2] Claru — "Egocentric Video Data Collection at Scale" — 2026 — https://claru.ai/solutions/egocentric-video-data
- [3] Labellerr — "7 Top Egocentric Data Service Providers for Robotics 2026" — 2026 — https://www.labellerr.com/blog/top-egocentric-data-providers-robotics/
- [4] Luel — "Launching Luel: a rights-cleared marketplace and collection engine for multimodal data" — 2026 — https://www.luel.ai/resources/blog/launching-luel-yc-w26
- [5] Luel — "Inside the gig economy training the next generation of AI" — 2026 — https://www.luel.ai/resources/blog/scaling-egocentric-video-data-pipeline
- [6] StartupHub.ai — "Claude's Corner: Luel, The Web Is Scraped" — 2026 — https://www.startuphub.ai/ai-news/claudes-corner/2026/claudes-corner-luel-yc-w2026
- [7] Objectways — "Egocentric Data Collection Is Driving the Future of Robotics" — 2026 — https://objectways.com/blog/egocentric-data-collection-is-driving-the-future-of-robotics/
- [8] Objectways — "Training Robots in Real Life with Egocentric Data" — 2026 — https://objectways.com/blog/training-robots-in-real-life-with-egocentric-data/
- [9] Encord — "Multimodal Data Annotation Tool & Curation Platform" — 2026 — https://encord.com/multimodal/
- [10] Grauman, K., et al. "Ego4D: Around the World in 3,000 Hours of Egocentric Video" — CVPR — 2022 — https://ego4d-data.org/
- [11] Meta AI — "Introducing Ego-Exo4D: A foundational dataset for research on video learning and multimodal perception" — 2023 — https://ai.meta.com/blog/ego-exo4d-video-learning-perception/
- [12] Ego-Exo4D dataset — https://ego-exo4d-data.org/
- [13] Damen, D., et al. "Scaling Egocentric Vision: The EPIC-Kitchens Dataset" — ECCV — 2018 — https://epic-kitchens.github.io/
- [14] Meta — "Project Aria" — https://www.projectaria.com/
- [15] Xu, R., et al. "EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video" — Apple Research — 2025 — https://arxiv.org/abs/2505.11709
- [16] NVIDIA Research — "EgoScale: Scaling Egocentric Human Video for Robot Dexterity" — 2026 — https://arxiv.org/abs/2502.04349
- [17] Avala — "What robot data costs, and where budgets leak" — 2026 — https://avala.ai/news/robotics-data-pyramid-buyers-guide
- [18] Robotics Center of Silicon Valley — "How Much Does Robot Data Collection Cost in 2026?" — 2026 — https://www.roboticscenter.ai/en/blog/robot-data-collection-cost
- [19] Unidata — "Egocentric Video Data Collection for AI Training" — 2026 — https://unidata.pro/data-collection/egocentric-video-data/
- [20] Unidata — "Egocentric Video Dataset" — 2026 — https://unidata.pro/datasets/egocentric-video/
- [21] Unidata — "Robotic Household Activities Dataset" — 2026 — https://unidata.pro/datasets/robotic-household-activities/
- [22] Unidata — "Lerobot SO-101 Manipulations Dataset" — 2026 — https://unidata.pro/datasets/lerobot-so-101-manipulations/
- [23] Unidata — "Robotics Training Data & Manipulation Datasets" — 2026 — https://unidata.pro/robotics-training-data/
- [24] Unidata — "Egocentric Data Collection for Humanoid Robot Training" (case study) — 2026 — https://unidata.pro/cases/egocentric-data-collection-for-humanoid-robot-training/
- [25] Unidata — "Data for Simulations: 3D Scanning for Robot Training" (case study) — 2026 — https://unidata.pro/cases/data-for-simulations-3d-scanning-for-robot-training/
- [26] Meshyk, K. — "Egocentric Data Collection for Robot Training: What Actually Works in Production" — Unidata — 2026 — https://unidata.pro/blog/egocentric-data-collection-for-robot-training/
Frequently Asked Questions (FAQ)
Egocentric data is video and sensor data captured from the first-person viewpoint of whoever or whatever is performing a task, rather than from a fixed external camera. In robotics, it’s used to train and fine-tune manipulation policies and VLA models because it captures the hand-object geometry and viewpoint closer to what a robot’s own camera sees at deployment.
Third-person video doesn’t capture the fine hand-object geometry, gaze, and grip detail a manipulation policy needs, and it doesn’t match the viewpoint a robot’s onboard camera will see once deployed. Egocentric data closes that gap, which is a large part of why robots trained on mismatched-viewpoint data often fail in real-world deployment despite passing offline evaluation.
Collection is capturing the raw first-person video and sensor streams — camera, depth, IMU, motion tracking. Annotation is labeling that footage afterward with pose, action segments, object states, or contact points. Some providers (Unidata, Claru, Objectways) do both; others (Encord, Scale AI, Labelbox) primarily annotate footage a client already has.
Check the license before you plan around it — Ego4D’s license restricts commercial use, and academic datasets generally aren’t cleared for training a model that ships in a commercial product [10]. They’re well suited for research pretraining and benchmarking, not as the sole data source for a production policy.
Providers price by project rather than publishing flat per-hour rates, driven by capture complexity, annotation depth, and scenario scripting requirements. For context, teleoperation hardware for robot data collection generally runs $6,000–$80,000, and a production-grade demonstration dataset has been estimated at $50,000–$200,000 all-in [18] — egocentric-specific quotes vary from that baseline depending on rig complexity and enrichment requested.
Common setups include head-mounted VR/passthrough headsets (Pico 4 Ultra is widely used), stereo depth cameras (ZED series), body-worn motion trackers for skeleton reconstruction, and increasingly wrist-mounted cameras to match robot end-effector geometry. Some providers layer in tactile gloves for grip and pressure data, though this remains an emerging capability industry-wide rather than a standard deliverable.
A usable schema goes beyond segment-level action labels to include sub-action temporal segmentation (reach, pre-grasp, contact, lift, release), per-joint hand pose at key frames, object state changes, and failure flags for incomplete attempts. Some providers auto-generate 3D pose from onboard sensor fusion, reducing the manual annotation burden for that layer specifically.
At minimum: does the capture rig match your robot’s camera geometry, is depth/IMU/skeleton data synchronized to the video, what’s the actual annotation schema (not just “annotated”), what commercial rights does the license cover, and can you inspect real sample data before signing anything. A vendor unwilling to show real footage before a contract is worth treating as a caution flag.