---
title: "Training Data for Face Anti-Spoofing Certification"
description: "What did Biometric.Vision get the most value from: broader attack coverage or more usable data within existing categories? After receiving 35,800 videos, the answer was…"
url: "https://unidata.pro/cases/training-data-for-face-anti-spoofing-certification/"
date_modified: "2026-10-04T12:04:06+03:00"
language: "en-US"
---
[Biometric.Vision](https://biometric.vision/) needed broader coverage for its face anti-spoofing models ahead of iBeta certification. The result was 35,800 attack videos across 8 attack types, with the biggest value turning out to be somewhere less obvious: the sheer volume of usable recordings.

Biometric.Vision was preparing for iBeta certification and needed more training data for its face anti-spoofing models. The request to Unidata was about coverage: more attack types, more capture devices, enough range to meet the lab's requirements.

The coverage came through as ordered. Asked afterwards which part of the delivery had actually done work for them, though, the team went straight past the attack types and named the volume: thousands of usable recordings, most of them falling into categories the models were already training on, which is a different kind of gain from covering new ground and, going by their answer, the one that was actually short.

The Problem
-----------

Biometric.Vision develops face anti-spoofing: presentation attack detection and liveness detection. With certification coming up, the training set needed more range across attack types and across the devices those attacks get captured and presented on.

In a supplier, that meant breadth in both directions at once, and a willingness to build the set around the certification requirements instead of shipping a catalogue item as it stands. Those were the terms the choice was made on.

Solution
--------

### What Was Delivered

The iBeta Level 1 dataset: 35,800 videos, 50 actors, 8 attack types.

Subject variation in the set:

- Ethnicity: European (60%), Asian (20%), African (20%)
- Age: 18–29 (44%), 30–49 (50%), 50+ (6%)
- Appearance features: bald, beard and moustache, makeup, scar, piercing, none

On coverage the set did what it was bought for, with enough spread across attack types and capture devices to meet the certification requirements. What the client rates highest is the volume. The new recordings were usable ones and they went into attack types the team already worked with, so the training set got thicker in places it was already covering. The team also rated the quality of the spoof attacks themselves as good.

### The Metadata

Each video carried a fair amount besides its attack label: capture device, presentation device, camera angle and position, background, video resolution, and the demographics of the person on screen. With that in place, a set this size can be pulled apart and looked at by slice instead of used whole.

The metadata turned out to be very complete. Beyond the attack type itself, there was the capture device and the presentation device, the camera angle and position, the background, the video resolution, and the subject's demographics. That gave us the flexibility to filter and analyse the data by different slices instead of just using it as it came.

### Integration

Integration went smoothly. The data is in preprocessing at the moment, being brought into the same structure as Biometric.Vision's internal data.

Process
-------

| Phase | Input | Scope of Work | Quality Control |
|---|---|---|---|
| Requirements Intake | Client's certification targets and training-set gaps | Defining the attack types, capture devices, and subject coverage needed | Requested composition maps to iBeta Level 1 requirements |
| Composition Alignment | Coverage requirements | Adapting dataset composition and parameters to the client's certification requirements | Delivered composition matches the requirements the client stated |
| Dataset Assembly | iBeta Level 1 dataset (35,800 videos, 50 actors, 8 attack types) | Assembling videos across attack types, capture and presentation devices, and subject groups | Coverage confirmed across all eight attack categories |
| Metadata Preparation | Assembled videos | Labelling attack type, capture and presentation device, camera angle and position, background, resolution, demographics | Metadata complete and consistent across the full set |
| Delivery | Final dataset and metadata | Packaging and handoff | Delivered data matches the agreed requirements and the agreed timeline |
| Handoff | Delivered dataset | Handing the set over in a form the client can normalise alongside its own data | Client reports integration went smoothly; the set enters preprocessing on their side |

## Main Title

Training Data for Face Anti-Spoofing Certification

## Description

What did Biometric.Vision get the most value from: broader attack coverage or more usable data within existing categories? After receiving 35,800 videos, the answer was the latter.

## Hero

**Data:** 35,800 videos (50 actors, 8 attack types) **Project Duration:** 2.5 weeks

## Progress - Results - Quote

### Progress - Steps

**List of Steps:**

- **Number of days:** Requirements Intake — **Step Description:** 2 days
- **Number of days:** Composition Alignment — **Step Description:** 1.5 days
- **Number of days:** Dataset Assembly & Metadata — **Step Description:** 7 days
- **Number of days:** Delivery & QC — **Step Description:** 1 day

### Results

**List of Results:**

- 35,800 videos, 50 actors, 8 attack types, delivered and now in preprocessing
- The gain the client names first is volume: more usable material inside the attack types already in use
- Spread across attack types and capture devices sufficient for the certification requirements
- Labelling that supports slicing by device, angle, background, resolution or subject demographics

### Quote

**Quote:** We'd recommend Unidata to colleagues. The team responded to requests quickly and was willing to adjust the composition and parameters of the dataset to our specific requirements, and what we received matched what we expected to receive. No complaints on timelines, on communication, or on how well the data matched the stated requirements.

**Author:** Nurmuhammed

**Position:** ML Lead, Biometric.Vision

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
