What Is the Open X-Embodiment Dataset?

14 minutes read
What Is the Open X-Embodiment Dataset?

Robotics has a data problem, and it's not the one you'd guess. It's not that there isn't enough data out there. It’s that the data is fragmented and collected in different formats, under different conditions, and rarely designed to work together. A university lab spends months collecting demonstrations on one arm, a company does the same on a completely different one, and neither dataset is much use to the other. Every robot has been learning alone.

The Open X-Embodiment dataset is the attempt to fix that. Instead of one robot, one dataset, one narrow skill set, it brings together more than one million robot trajectories from dozens of labs and 22 robot embodiments into a common framework. The idea behind it is simple: if a model can learn from many different robots at once, it may become less dependent on the quirks of any one platform and learn skills that transfer across hardware it has not seen before.

This guide walks through what's actually inside the dataset, why it matters, how to get your hands on it, and where it falls short.

Overview and origins of the Open X-Embodiment Dataset

Picture 21+ institutions all showing up to the same robotics project, contributing 60 datasets from 34+ labs, with a published author list running into the hundreds. That's not a hypothetical. That's the Open X-Embodiment collaboration, led by Google DeepMind, and it's exactly why the Open X-Embodiment dataset carries so much weight in robotics circles today. [1][13] Worth keeping the numbers straight here: institutions, contributing labs, and individual researchers are three different counts, not interchangeable ones. This isn't one lab's pet project either way; it's closer to a group effort where everyone agreed to stop hoarding their robot data and pool it instead.

The idea driving all of it isn't new — it's borrowed. NLP and computer vision both got a lot smarter once researchers realized bigger, messier, more varied datasets beat small, tidy ones. Open X-Embodiment takes that same playbook and points it at robots. [1] The flagship result of all that pooled data is RT-X, a model family introduced to investigate whether training on data from many different robots could produce positive transfer across platforms.

What's inside the dataset: data composition

Here's the number that tends to make people sit up: the Open X-Embodiment (OXE) dataset packs in more than 1 million robot trajectories, pulled from 60 existing datasets contributed by 34+ labs. [1] Each trajectory is basically a mini-story: what the robot saw, what it was asked to do, and what it actually did about it.

The dataset is a registry of many separately collected datasets rather than one static file, so contributors track exactly which datasets are included through a public Google Sheet, linked right off the project's GitHub page. [2] Unglamorous, but it's the closest thing to a live inventory anyone can check.

Want to see how raw robot footage turns into something a model can actually learn from? It's basically a four-step handoff: the 60 source datasets from those 34+ labs and 22 robot embodiments are brought into a common representation containing more than one million trajectories. The data is represented using the RLDS episode format, while the RT-1-X model uses a seven-variable end-effector action representation, and can then be used for training and benchmarking models such as RT-1-X and RT-2-X. Here's what that looks like laid out:

What's inside the dataset: data composition

Robot embodiments included

Most datasets stick to one robot. OXE didn't get that memo. There are 22 different robot embodiments in the mix, including single-arm and bimanual setups and other robotic platforms. The names you’ll see most are Google Robot, WidowX, Franka Emika Panda, xArm, and UR5.

Why does mixing robots matter so much? Think of it like teaching someone to drive using only one car versus letting them practice in a sedan, a pickup, and a stick-shift hatchback. The second person adapts faster to a rental car they've never touched. Same logic here: a model that's only ever seen a WidowX arm gets thrown off by an unfamiliar camera angle or arm length. A model trained across Franka, xArm, UR5, and the Google Robot may be less dependent on the quirks of any one platform.

Tasks and skills covered

The task list is extensive: 527 distinct skills spread across 160,266 tasks [1] — pick-and-place, sure, but also wiping counters, assembling parts, yanking open drawers, and poking at appliance buttons. It's less "robot doing one trick really well" and more "robot getting shoved into the chaos of an actual kitchen."

Key features and technical specifications

Under the hood, OXE datasets are represented using the RLDS episode format and can be accessed through TensorFlow Datasets (TFDS). This provides a common structure for streaming and processing large collections of episodic robot data rather than requiring every source dataset to use its own interface. Many OXE datasets contain RGB observations, but the available sensor modalities vary across component datasets. Some include depth, proprioceptive state, or additional observations, while others provide a more limited set of signals. Before using a particular component dataset, check its documentation rather than assuming that every modality is available everywhere.

The gnarliest problem the team had to solve: robots don't all move the same way. A UR5 and a WidowX don't speak the same control "language." RT-1-X uses a seven-variable end-effector action representation: x, y, z, roll, pitch, yaw, and gripper opening. [1] That standardization doesn't mean every robot now shares one perfectly aligned coordinate system, though. Coordinate frames still differ across the source datasets, and depending on which dataset an episode came from, that same 7-D vector might represent an absolute position, a relative movement, or a velocity. It's not elegant. It's a compromise that works well enough, and in engineering, that's usually the whole game.

RT-X models: why this dataset matters

All that data collecting only matters if it makes robots better, and this is where RT-X shows up to make the case. Google DeepMind built two Transformer-based models on top of OXE: RT-1-X, the leaner generalist, and RT-2-X, which adds web-scale pretraining the way large language models soak up internet text before ever touching a robot arm. [1][4]

The headline result is specific, not a blanket claim: on an emergent-skills evaluation, where a Google Robot was tested on tasks that existed in other robots' training data but never in its own, RT-2-X scored roughly 3x higher than RT-2 (the version trained without the broader cross-robot dataset). [1][13] Researchers call this positive transfer. Ablation tests reinforced the pattern: a 55-billion-parameter model achieved a significantly higher emergent-skills success rate than the 5-billion-parameter model, web-scale pretraining was critical to performance, and including a short history of images improved generalization. [1]

One especially convincing test: researchers removed the Bridge dataset from RT-2-X training and found that performance on the corresponding hold-out tasks dropped substantially. Because those tasks were associated with Bridge’s WidowX data, the result suggests that experience from the WidowX platform contributed to the additional skills demonstrated by RT-2-X on the Google Robot. [1]

The two models don't share an architecture, so it's worth describing them separately. RT-1-X is built on the original RT-1 design: EfficientNet reads the camera image, the Universal Sentence Encoder turns the spoken instruction into something math-friendly, FiLM layers let that instruction actively reshape how the model looks at the image (so "pick up the red block" nudges attention toward, well, the red block), and a Transformer turns the fused result into a tokenized action. [4]

RT-2-X is a different animal. It's built on RT-2's vision-language-action approach: a large vision-language model, using a ViT-based vision encoder and a UL2 language model, co-fine-tuned on web-scale image-text data alongside robot trajectories, with robot actions represented as tokenized output rather than routed through a separate small Transformer head. [4][12] That web-scale pretraining is part of the RT-2-X setup, while the emergent-skills result specifically comes from the combination of RT-2's VLM capabilities and broader cross-embodiment robotics training.

Laid out side by side, here's how the two architectures actually compare:

RT-X Models: Why This Dataset Matters

People sometimes compare Open X-Embodiment's role in robotics to what ImageNet did for computer vision. That's an analogy, not a finding from the research itself, but it's a useful one: a shared dataset big enough to give an entire field the same starting line. [5]

RT-X was just the opening move. Several influential generalist robot policies, including OpenVLA and Octo, subsequently used OXE data or closely related cross-embodiment pretraining strategies. The dataset therefore became an important part of the broader shift toward training robot policies on diverse, multi-robot datasets. Here is how that lineage actually unfolded:

RT-X Models

Real-world applications and use cases

Since its release, OXE has been used as a pretraining source for several generalist robot policies. A few ways teams actually put it to work:

Benchmarking: a fair, shared test for comparing new robot policies.

Fine-tuning on your own robot: got a Franka Panda and only a couple hundred demos? Start from an OXE-trained checkpoint instead of collecting a mountain of data from scratch.

Academic research: university labs without a warehouse full of robots get to play in the same sandbox as the big players.

Startup prototyping: a “pretrain on broad robot data, then fine-tune on the target task” workflow is a practical strategy when collecting a large target-robot dataset is expensive.

That recipe, or a close variant of it, is how OpenVLA [6] and Octo [11] were built. pi0 is a bit different: OXE-derived data is part of its pretraining mixture, but pi0 also trains on other proprietary and third-party data, so it isn't accurate to call it simply "an OXE-trained model." [7]

If you're advising a team without deep pockets for data collection, pretrain-then-fine-tune is a workflow worth considering, and DROID-style, single-robot data is a common choice once you need a consistent fine-tuning set after the generalist checkpoint is already in hand. [8]

How to access and use the dataset

Two doors get you in: the project's GitHub repository, which points you to documentation and dataset links [2], or loading things directly through Google's tensorflow_datasets (TFDS), which already speaks RLDS fluently. [3]

If you're also juggling robot sensor logs in other formats, keep the Foxglove SDK on your radar. It's a common sidekick for visualizing and debugging that data, which is exactly what the next section is about. [9]

Data format and developer tools

The straightforward path is RLDS through tensorflow_datasets (TFDS). [3] A no-drama first setup looks roughly like this:

  1. Spin up a clean Python environment and install tensorflow_datasets plus tensorflow.
  2. Point TFDS at whichever OXE component datasets you actually want; the public tracking sheet lists the names. [2]
  3. Pull a handful of episodes first and eyeball them. Do the camera frames, instructions, and actions actually line up? Check before you commit.
  4. Once that looks right, stream the full thing rather than trying to download it, since size alone rules out a full download.

More teams these days skip TFDS entirely and go straight for PyTorch. HuggingFace Hub mirrors several OXE datasets, and LeRobot, also from HuggingFace, gives you PyTorch-native loaders built specifically for data like this. [10] Worth noting these are community-maintained mirrors, not the official OXE distribution channel, so cross-check against the GitHub registry if you need to confirm exactly what's included.

If you're a Python developer already comfortable in PyTorch, this route usually causes fewer headaches than bending TFDS pipelines to your will.

Quick tip: always eyeball one sample episode before you go big. Plot a few frames next to the recorded actions. It's five minutes of checking that saves you from a wasted training run.

Licensing, storage, and compute considerations

Two things will bite you if you skip them.

First: there is no single blanket license covering everything in OXE. Licensing is another detail worth checking carefully. The OXE repository itself uses Apache-2.0 for software and CC-BY-4.0 for other materials, but the individual datasets contributed to OXE may have their own licenses and usage conditions. Anyone planning commercial use should therefore check the license of every component dataset they intend to use rather than assuming that the repository-level license applies to the entire collection.

Second: this is a large dataset. Downloading the whole thing to a laptop isn't realistic, which is why streaming through TFDS, HuggingFace, or LeRobot is the normal move. [3][10] Exact compute requirements for pretraining versus fine-tuning depend heavily on model size, batch size, and which subset of OXE you're using. The project's own papers report the hardware used for their specific experiments, but there's no single GPU count that applies to every use case, so budget based on your own model and data volume rather than a rule of thumb.

Benefits and limitations

The good stuff is genuinely good: OXE is one of the foundational large-scale open datasets for cross-embodiment robot learning, with substantial embodiment diversity and demonstrated use in models such as RT-X, OpenVLA, and Octo. [1][6][11] Calling it an openly released dataset infrastructure is more accurate than calling it fully open access, though: the registry and code are open, but what you can actually do with any given component dataset depends on that dataset's own license.

But it's not flawless, and pretending otherwise wastes your time later:

  • Quality is uneven. Dozens of independent labs mean dozens of different camera setups, lighting choices, and demonstration habits.
  • Action spaces still need babysitting. The shared 7-dimensional format (built on RLDS) helps a ton, but it's a compromise, and some robots still need extra normalization work. [1][3]
  • Contribution volume isn't even across all 22 embodiments. Some platforms are far more heavily represented than others, and a model trained without correcting for that can pick up biases toward whichever robots dominate the mix.
  • New robots aren't automatically covered. Something outside the 22 embodiments already in the registry is a hopeful transfer target, not a guarantee.

Here's a hypothetical but realistic-sounding headache: reconciling a depth stream from a sensor like the ZED Mini against other contributing datasets that only ever recorded RGB. This specific scenario isn't documented in the OXE papers themselves, but it illustrates a real class of problem: a model trained assuming depth is always there can quietly stumble on RGB-only robots unless that mismatch gets caught first. It's a good reminder that OXE's diversity is both its superpower and its cleanup job. For more on data-quality challenges in robot learning generally, the DROID dataset paper is a solid companion read. [8]

Comparison with other robotics datasets

OXE isn't the only big robot dataset out there. Earlier efforts like RoboNet and BridgeData paved some of this ground, but the comparison everyone actually cares about is with DROID.

DROID comes from Stanford, UC Berkeley, and TRI (Toyota Research Institute). It takes a different approach: a standardized Franka-based hardware setup and a consistent collection protocol, while still emphasizing diversity across real-world scenes and tasks rather than diversity across robot hardware. It contains about 76,000 trajectories, or roughly 350 hours of interaction data. [8]

FeatureOpen X-Embodiment (OXE)DROID
Scale1M+ trajectories [1]~76,000 trajectories [8]
Embodiments
22 distinct robots [1]
1 (single embodiment) [8]
Data formatRLDS / tfrecord [3]HDF5 [8]
Primary strengthCross-robot generalization, pretrainingConsistency, reproducible fine-tuning
Backing institutionsGoogle DeepMind + 21 partner institutions [1]Stanford, UC Berkeley, TRI [8]
Best used forFoundation-model pretraining (e.g., OpenVLA, Octo, pi0) [6][7][11]Task-specific fine-tuning after pretraining

A workflow worth considering, based on how these tools tend to get used in practice: pretrain broadly on multi-robot data such as OXE, as in OpenVLA and Octo, then use more consistent target-robot data when task-specific adaptation is required. They are better understood as complementary datasets with different strengths: OXE emphasizes diversity across robot embodiments, while DROID emphasizes consistency in hardware and data collection. A reasonable workflow is to use broad multi-robot data for pretraining and then use more consistent target-robot data when task-specific adaptation is required. This is a practical strategy rather than a universal industry standard, and the best choice depends on the model, robot, and target task.

The future of cross-embodiment learning

Robotics is now testing a script that NLP and computer vision already ran: the idea that scaling up diverse, shared data drives real progress, more so than architecture tweaks alone. Open X-Embodiment’s positive-transfer results are a genuine, specific demonstration that cross-embodiment training can improve performance in the reported experiments, but they are not proof that the same scaling curve will hold everywhere in robotics the way it did for NLP and computer vision. [1] Still, it's a strong enough signal that more OXE-style data pools and more generalist policies built on top of them seem like a reasonable bet for where the field goes next, alongside robots that pick up new bodies or new tasks with less task-specific training data than today's standard.

Frequently Asked Questions (FAQ)

What is the Open X-Embodiment dataset?

It is a large collection of robot-learning datasets containing more than one million trajectories from 34+ labs. The project was created to standardize and make diverse robot data usable for training generalist robot policies across multiple embodiments, with RT-X being its flagship demonstration. [1]

What robot embodiments are included in the Open X-Embodiment dataset?

There are 22 different robot embodiments in the mix, including single arms, two-armed setups, and four-legged machines. [1] The names you’ll see most are Google Robot, WidowX, Franka Emika Panda, xArm, and UR5, a deliberately wide mix meant to stop models from overfitting to just one robot.

Is the Open X-Embodiment dataset free to use?

The OXE registry and code are openly released, but that’s separate from the individual contributed datasets inside it. OXE bundles 60 separate datasets, each with its own license [2], so anyone planning commercial use needs to check the specific licenses attached to whichever component datasets they’re actually pulling from.

How large is the Open X-Embodiment dataset?

Over 1 million robot trajectories, 527 distinct skills, and more than 160,000 tasks, all pulled from 60 source datasets. [1] It’s large enough that a full local download isn’t practical for most teams, which is why streaming access through TFDS, HuggingFace, or LeRobot is the usual approach.

How is the Open X-Embodiment dataset related to RT-X models?

RT-X, specifically RT-1-X and RT-2-X, is the model family trained on top of this dataset, though the two use genuinely different architectures. [1] On an emergent-skills evaluation, RT-2-X scored roughly 3x higher than RT-2 on tasks absent from its own robot’s training data but present in data from other robots, a result researchers call positive transfer. [1][13]

Where can the Open X-Embodiment dataset be downloaded?

Start with the project’s GitHub repository, or load it directly through Google’s tensorflow_datasets (TFDS) library, which natively understands its RLDS format. [2][3] HuggingFace Hub and the LeRobot library also mirror parts of it for anyone working in PyTorch. [10]

Is the Open X-Embodiment dataset relevant for robots not included in it?

OXE can be relevant to robots outside its original 22 embodiments because cross-embodiment training encourages transferable representations and skills. However, transfer to a genuinely new robot is not guaranteed. Differences in kinematics, sensors, action spaces, and coordinate conventions can still require normalization, adaptation, or fine-tuning.

What licensing terms apply to the Open X-Embodiment dataset?

The OXE repository itself is released under Apache-2.0 (software) and CC-BY-4.0 (other materials), but that’s separate from the individual contributed datasets inside it, each of which carries its own terms: CC-BY-4.0, Apache-2.0, MIT, and more. [2] As covered in the licensing section above, always read the fine print on each component dataset before any commercial use.

Insights into the Digital World

Robotics

What Is the Open X-Embodiment Dataset?

Datasets

What is Data Cleaning in Machine Learning?

Robotics

NVIDIA Isaac Sim for Robot Training: A Complete Guide

Datasets

What Is Crowdsourcing? Complete Guide

AI Training, Data Labeling

What Is 3D LiDAR SLAM Technology? How It Works and Where It’s Used 

Datasets

Structured vs. Unstructured Data: A Data Expert’s Guide to Maximizing Business Value

Datasets

Qualitative Data Collection Methods: A Practical Guide for ML and Product Teams

Datasets

How to Collect Data for Machine Learning: A Practical Guide for ML Teams

Ready to get started?

Tell us what you need — we’ll reply within 24h with a free estimate

    Andrew
    Head of Client Success

    — I'll guide you through every step, from your first
    message to full project delivery

    Thank you for your
    message

    It has been successfully sent!

    We use cookies to enhance your experience, personalize content, ads, and analyze traffic. By clicking 'Accept All', you agree to our Cookie Policy.