---
title: "Real-World Data for a Synthetic-Trained PAD Model"
description: "How a global biometric vendor broke a two-year certification deadlock by replacing synthetic training data with iBeta-ready real-world recordings"
url: "https://unidata.pro/cases/real-world-data-for-a-synthetic-trained-pad-model/"
date_modified: "2026-09-02T06:42:41+03:00"
language: "en-US"
---
**The Problem**
---------------

For over two years, a global biometric vendor trained its PAD system entirely on synthetic datasets. They were easy to generate and scale — and they never reached certification.

- Generalization: the model struggled with real-world lighting, texture, and motion artifacts.
- Pre-check failures: advanced attacks, especially latex and high-quality 3D masks, consistently bypassed detection.
- Stalled R&D: two years of iteration brought the system no closer to iBeta Level 2, delaying market entry.

**Solution**
------------

The vendor replaced its synthetic pipeline with an iBeta-ready Level 2 dataset built to mirror certification conditions:

- 32,000+ high-resolution videos covering advanced attack types, including latex and composite masks.
- Target-device recordings, so the data matched the exact hardware used in the lab.
- Diverse actors across multiple ethnic backgrounds, filmed in real environments.
- Iterative integration into the client's pipeline, with feedback rounds to close gaps found during training.

| Phase | Input | Scope of Work | Quality Control |
|---|---|---|---|
| Model & Requirements Intake | Client's PAD model, target Level 2 attack scenarios | Access setup, mapping the synthetic-to-real gap across latex, composite, and 3D-mask attacks | Test scope matches the client's model and Level 2 attack matrix |
| Dataset Delivery | iBeta-ready Level 2 set (32,000+ videos) | Delivering advanced-attack footage recorded on the target certification device | Attack types and device coverage confirmed against iBeta conditions |
| Integration & Retraining | Delivered dataset + client pipeline | Folding real-world data into training, replacing synthetic-only samples | Real-world lighting, texture, and motion represented in the training set |
| Error Analysis | Model inference results | Accuracy breakdown by attack vector; isolating residual latex / 3D-mask misses | Systematic weak spots identified, not an aggregate score |
| Feedback Rounds | Error findings | Additional recordings targeting the vectors the model still missed | New data maps directly to a documented gap |
| Pre-Check & Handoff | Retrained model, final dataset | Level 2 pre-check; packaging data and findings for the client | Client sign-off; model cleared for lab submission |

## Main title

Real-World Data for a Synthetic-Trained PAD Model

## Description

A global biometric vendor spent two years training its PAD system on synthetic data. The model cleared every internal check yet failed iBeta pre-checks each time. Real-world recordings closed the gap in two months.

## Hero

**Project duration:** 7 weeks

## Progress - Results - Quote

### Progress - Steps

**List of Steps:**

- **Number of days:** Week 1 — **Step Description:** Model & Requirements Intake
- **Number of days:** Weeks 2–4 — **Step Description:** Dataset Integration & Retraining
- **Number of days:** Weeks 4 –6 — **Step Description:** Error Analysis & Feedback Rounds
- **Number of days:** Week 7 — **Step Description:** iBeta Pre-Check & Handoff

### Results

**List of Results:**

- Passed iBeta Level 2 on the first attempt.
- Reliable detection across the most complex spoofing scenarios, including latex and composite masks.
- 2.6% accuracy gain on critical attack vectors, beyond what synthetic data reached.
- Two years of stagnation resolved in two months of dataset use and retraining.

### Quote

**Quote:** The team spent two years tuning the model, but the real constraint was the data. We added real latex and composite masks to the pipeline and accuracy improved within days. A lot of teams get caught out by device specificity, iBeta runs its tests on one exact device, so a model trained on generic footage tends to underperform when it gets there.

**Author:** Elizabeth Karnaukhova

**Position:** Datamarket Project Manager

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
