Data Collection

Data for Simulations: 3D Scanning for Robot Training

Image

How to close the relevant data gap for physical AI through photogrammetry and lidar scanning, and connect real scenes to simulation environments.

Image

The Problem

Humanoid robot developers keep running into the same wall: egocentric video data is scarce, expensive to collect, and slow to accumulate — nowhere near the volume needed to meaningfully advance training. Simulation solves the scale problem: you can spin up hundreds of parallel environments and generate millions of iterations in a short time. But those simulation environments need to be populated with realistic spaces and objects the robot will actually interact with.

Building environments by hand is costly and slow — it requires designers, developers, and virtual environment specialists. The goal was to collect simulation data directly from the real world, without a full production team.

The work split into two parallel streams:

Space scanning

Photorealistic 3D models of apartments and rooms, ready to load into simulation environments like IsaacSim — where a robot can be placed and its object interactions recorded.

Object scanning

3D models of everyday items — mugs, boxes, tools — to populate simulation scenes with geometrically accurate, properly textured objects.

Solution

Space Scanning

Building environments manually without a large team is not realistic. Instead, we scan real rooms and load them directly into the simulator. The setup uses a 360-degree camera with an integrated lidar. Lidar provides metric accuracy; the camera provides photorealistic textures.

Coverage is monitored in real time:

  • the operator walks through the space with the scanner
  • the software flags zones with insufficient coverage
  • data is uploaded to volumetric reconstruction software

Environments are static — drawers and cabinet doors do not open. For most object manipulation scenarios on surfaces, this is not a meaningful limitation. More importantly, scans are tied to the same spaces where egocentric footage is recorded. That connection is the direct integration point between the two data types.

Object Scanning

Lidar is not suitable for individual objects: resolution is insufficient for fine details and textures. The method of choice is photogrammetry with reconstruction via 3D Gaussian Splatting.

The pipeline works as follows:

  • approximately 150 shots from different angles under controlled lighting
  • processing in ColMap: computing the position of each frame relative to the object
  • reconstruction in 3DGS: a precise model with realistic textures as output

The hardest part is lighting. Three parameters are in constant tension, and there is no universal solution.

Here is what happens when each one is off:

  • High ISO reduces the need for light but introduces grain — ColMap stops correctly matching points between frames
  • Wide aperture lets in more light but narrows depth of field — part of the object goes out of focus and reconstruction in those zones degrades
  • Long exposure requires the camera to be completely still — any movement interferes with frame matching

After testing phone cameras, we switched to DSLRs: the larger sensor produces acceptable results in low light without a critical increase in ISO. Shooting parameters were calibrated separately for each object class.

Integration with Egocentric Data

Egocentric recordings and scans of the same spaces form a unified dataset where real data and simulation point to the same environment.

This gives the client capabilities that are unavailable when purchasing the two separately:

  • reproduce a scenario from egocentric video in simulation — with the same geometry and the same objects
  • adapt data to the physics of a specific robot: different grip, different height, different degrees of freedom
  • collect additional data in simulation without another field visit — adjust lighting, object placement, trajectories

A single egocentric data collection session becomes a scalable source of simulation data.

PhaseInputScope of WorkQuality Control
Preparation & CalibrationClient requirements, list of target spaces and objectsLidar scanning and photogrammetry setup, shooting parameters for lighting conditionsCoverage accuracy, ISO / aperture / shutter balance
Pilot ScanningTest room, set of objectsTrial scans of spaces and objects, identifying problem areas: dark corners, reflective surfaces, fine detailsTexture quality, geometry completeness, absence of reconstruction artifacts
Space Scanning360-degree camera with integrated lidarWalk-through with real-time coverage monitoring, upload to volumetric reconstruction softwareMetric dimensional accuracy, photorealistic textures, no uncovered zones
Object ScanningDSLR camera, interaction objects~150 frames from different angles, ColMap processing, 3DGS pipeline reconstructionFull object in focus, correct frame matching in ColMap
Reconstruction & ProcessingRaw scanning and photogrammetry dataBuilding final 3D models of spaces and objects, geometry and scale verificationGPU / RAM resources, no degradation in underlit zones
Egocentric IntegrationRoom scans, egocentric video from the same spacesAligning scans with recordings, test scene loading in IsaacSim, format compatibility checkGeometry match between scan and real space from video
Final DeliveryValidated 3D models and linked datasetPackaging with format documentation, handoff with instructions for simulator importIsaacSim compatibility, metadata completeness, pipeline reproducibility
Week 1
Preparation & Calibration
Week 2
Pilot Scanning
Weeks 3–5
Main Collection
Week 6
Reconstruction & Validation

The Results

  • Photorealistic 3D scenes of real spaces, ready to load into IsaacSim
  • A library of 3D objects with accurate geometry and textures for populating simulation environments
  • A linked dataset: egocentric video tied to scans of the same spaces
  • A reproducible scanning pipeline that does not require a large team of virtual environment specialists
The key decision was straightforward: record egocentric footage and scan the same space. The client does not receive two separate products, they receive one environment where real data and simulation point to the same place.
Martinian Letunovsky
Martinian Letunovsky
Head of IT Operations

Similar Cases

Banking Call Categorization for NLP Automation

  • Banking
  • 363,000 audio files
  • 50 days
Learn more

Urban Image Annotation for Waste Detection

  • Housing and Utilities
  • 8000 images
  • 2 week
Learn more

Pose Estimation for Proctoring

  • Education
  • 6000 images
  • 7 weeks
Learn more

Image Annotation for Ore Detection

  • Mining and Oil & Gas Industry
  • 300 annotated ore images
  • 2 weeks
Learn more

Audio Data Collection for Emotion-Sensitive Voice Systems

  • Development of child response systems for laughter and crying
  • 750 unique audio files featuring children's voices
  • 1 month
Learn more

Ready to get started?

Tell us what you need — we’ll reply within 24h with a free estimate

    Andrew
    Head of Client Success

    — I'll guide you through every step, from your first
    message to full project delivery

    Thank you for your
    message

    It has been successfully sent!

    We use cookies to enhance your experience, personalize content, ads, and analyze traffic. By clicking 'Accept All', you agree to our Cookie Policy.