The Problem
For over two years, a global biometric vendor trained its PAD system entirely on synthetic datasets. They were easy to generate and scale — and they never reached certification.
- Generalization: the model struggled with real-world lighting, texture, and motion artifacts.
- Pre-check failures: advanced attacks, especially latex and high-quality 3D masks, consistently bypassed detection.
- Stalled R&D: two years of iteration brought the system no closer to iBeta Level 2, delaying market entry.
Solution
The vendor replaced its synthetic pipeline with an iBeta-ready Level 2 dataset built to mirror certification conditions:
- 32,000+ high-resolution videos covering advanced attack types, including latex and composite masks.
- Target-device recordings, so the data matched the exact hardware used in the lab.
- Diverse actors across multiple ethnic backgrounds, filmed in real environments.
- Iterative integration into the client's pipeline, with feedback rounds to close gaps found during training.
| Phase | Input | Scope of Work | Quality Control |
|---|---|---|---|
| Model & Requirements Intake | Client's PAD model, target Level 2 attack scenarios | Access setup, mapping the synthetic-to-real gap across latex, composite, and 3D-mask attacks | Test scope matches the client's model and Level 2 attack matrix |
| Dataset Delivery | iBeta-ready Level 2 set (32,000+ videos) | Delivering advanced-attack footage recorded on the target certification device | Attack types and device coverage confirmed against iBeta conditions |
| Integration & Retraining | Delivered dataset + client pipeline | Folding real-world data into training, replacing synthetic-only samples | Real-world lighting, texture, and motion represented in the training set |
| Error Analysis | Model inference results | Accuracy breakdown by attack vector; isolating residual latex / 3D-mask misses | Systematic weak spots identified, not an aggregate score |
| Feedback Rounds | Error findings | Additional recordings targeting the vectors the model still missed | New data maps directly to a documented gap |
| Pre-Check & Handoff | Retrained model, final dataset | Level 2 pre-check; packaging data and findings for the client | Client sign-off; model cleared for lab submission |
The Results
- Passed iBeta Level 2 on the first attempt.
- Reliable detection across the most complex spoofing scenarios, including latex and composite masks.
- 2.6% accuracy gain on critical attack vectors, beyond what synthetic data reached.
- Two years of stagnation resolved in two months of dataset use and retraining.
The team spent two years tuning the model, but the real constraint was the data. We added real latex and composite masks to the pipeline and accuracy improved within days. A lot of teams get caught out by device specificity, iBeta runs its tests on one exact device, so a model trained on generic footage tends to underperform when it gets there.
- Elizabeth Karnaukhova
- Datamarket Project Manager