The Task
In identity verification, a single undetected spoof can admit a fraudulent user, which makes any gap in attack coverage a real exposure. For this project, Unissey wanted to broaden the presentation attacks its models had been trained against while avoiding additional spend on scenarios they already handled reliably.
The project had two clear objectives:
- extend coverage across the presentation-attack scenarios that existing data did not adequately address.
- include consistent, structured metadata across relevant demographic groups.
The Solution
Dataset Selection
The team compared several providers before selecting one. Unidata offered the coverage Unissey needed, and the breadth of attack types in its catalogue proved useful throughout the work.
Portfolio Coverage
Unidata maintains a broad biometric portfolio of 29 datasets across face and liveness. Its liveness collections are built around iBeta Level 1 and Level 2 PAD scenarios and cover the presentation attacks seen in production. These include low-effort spoofs such as printed and cut-out photos and paper or cardboard constructions, screen replays on standard consumer devices, and higher-fidelity physical artefacts such as hyper-realistic wearable masks designed to defeat systems already robust against simpler attacks. Several of these collections are among Unidata's flagship offerings. The datasets are built to the ISO/IEC 30107-3 standard for presentation attack detection and have supported projects that reached iBeta Level 1 and Level 2 completion.
Working with the Data
Unissey used the data to train its face anti-spoofing models, using image and video datasets. Integration was straightforward; the team downloaded the datasets and reformatted them to its own conventions. The consistent metadata supported tracking coverage across demographic groups, and the variety of capture conditions and devices matched the real-world range the models must handle.
| Phase | Input | Scope of Work | Quality Control |
|---|---|---|---|
| Requirements Review | Unissey’s existing liveness models | Identifying underrepresented presentation attacks and defining coverage gaps | Existing attack coverage identified |
| Dataset Selection | Unidata’s face liveness datasets | Selecting data covering the presentation attacks missing from existing training coverage | Attack types and available metadata reviewed |
| Dataset Delivery | Selected image and video datasets | Downloading and reformatting datasets to Unissey’s conventions | Consistent metadata maintained across relevant demographic groups |
| Model Training | Integrated liveness datasets | Using the data to train Unissey’s face anti-spoofing models | Previously underrepresented attack scenarios added to training data |
| Coverage Expansion | Retrained models and expanded training data | Broadening the range of presentation attacks represented in training | Coverage expanded without duplicating adequately covered scenarios |
| Final Handoff | Reformatted datasets | Delivering the targeted data for continued model development | Data aligned with the agreed technical and compliance requirements |
The Results
- A broader set of presentation-attack scenarios now represented in the training data.
- Consistent metadata across relevant demographic groups, allowing coverage to be assessed by subgroups.
- A focused purchase addressing the identified gaps rather than a broad re-buy of existing data.
Unidata has been a responsive and collaborative partner, providing high-quality data that met our technical needs. We also appreciated their willingness to engage constructively with our compliance and audit requirements and to provide the contractual framework needed to support them.
- Faouzy Soilihi
- Chief Product & Strategy Officer