Datasets

Training Data for Face Anti-Spoofing Certification

Image

What did Biometric.Vision get the most value from: broader attack coverage or more usable data within existing categories? After receiving 35,800 videos, the answer was the latter.

Image

Biometric.Vision needed broader coverage for its face anti-spoofing models ahead of iBeta certification. The result was 35,800 attack videos across 8 attack types, with the biggest value turning out to be somewhere less obvious: the sheer volume of usable recordings. 

Biometric.Vision was preparing for iBeta certification and needed more training data for its face anti-spoofing models. The request to Unidata was about coverage: more attack types, more capture devices, enough range to meet the lab's requirements.

The coverage came through as ordered. Asked afterwards which part of the delivery had actually done work for them, though, the team went straight past the attack types and named the volume: thousands of usable recordings, most of them falling into categories the models were already training on, which is a different kind of gain from covering new ground and, going by their answer, the one that was actually short. 

The Problem

Biometric.Vision develops face anti-spoofing: presentation attack detection and liveness detection. With certification coming up, the training set needed more range across attack types and across the devices those attacks get captured and presented on.

In a supplier, that meant breadth in both directions at once, and a willingness to build the set around the certification requirements instead of shipping a catalogue item as it stands. Those were the terms the choice was made on.

Solution

What Was Delivered

The iBeta Level 1 dataset: 35,800 videos, 50 actors, 8 attack types.

Subject variation in the set:

  • Ethnicity: European (60%), Asian (20%), African (20%)
  • Age: 18–29 (44%), 30–49 (50%), 50+ (6%)
  • Appearance features: bald, beard and moustache, makeup, scar, piercing, none

On coverage the set did what it was bought for, with enough spread across attack types and capture devices to meet the certification requirements. What the client rates highest is the volume. The new recordings were usable ones and they went into attack types the team already worked with, so the training set got thicker in places it was already covering. The team also rated the quality of the spoof attacks themselves as good.

The Metadata

Each video carried a fair amount besides its attack label: capture device, presentation device, camera angle and position, background, video resolution, and the demographics of the person on screen. With that in place, a set this size can be pulled apart and looked at by slice instead of used whole.

The metadata turned out to be very complete. Beyond the attack type itself, there was the capture device and the presentation device, the camera angle and position, the background, the video resolution, and the subject's demographics. That gave us the flexibility to filter and analyse the data by different slices instead of just using it as it came.

Integration

Integration went smoothly. The data is in preprocessing at the moment, being brought into the same structure as Biometric.Vision's internal data.

Process

PhaseInputScope of WorkQuality Control
Requirements IntakeClient's certification targets and training-set gapsDefining the attack types, capture devices, and subject coverage neededRequested composition maps to iBeta Level 1 requirements
Composition AlignmentCoverage requirementsAdapting dataset composition and parameters to the client's certification requirementsDelivered composition matches the requirements the client stated
Dataset AssemblyiBeta Level 1 dataset (35,800 videos, 50 actors, 8 attack types)Assembling videos across attack types, capture and presentation devices, and subject groupsCoverage confirmed across all eight attack categories
Metadata PreparationAssembled videosLabelling attack type, capture and presentation device, camera angle and position, background, resolution, demographicsMetadata complete and consistent across the full set
DeliveryFinal dataset and metadataPackaging and handoffDelivered data matches the agreed requirements and the agreed timeline
HandoffDelivered datasetHanding the set over in a form the client can normalise alongside its own dataClient reports integration went smoothly; the set enters preprocessing on their side

Requirements Intake
2 days
Composition Alignment
1.5 days
Dataset Assembly & Metadata
7 days
Delivery & QC
1 day

The Results

  • 35,800 videos, 50 actors, 8 attack types, delivered and now in preprocessing
  • The gain the client names first is volume: more usable material inside the attack types already in use
  • Spread across attack types and capture devices sufficient for the certification requirements
  • Labelling that supports slicing by device, angle, background, resolution or subject demographics
We'd recommend Unidata to colleagues. The team responded to requests quickly and was willing to adjust the composition and parameters of the dataset to our specific requirements, and what we received matched what we expected to receive. No complaints on timelines, on communication, or on how well the data matched the stated requirements.
Nurmuhammed
Nurmuhammed
ML Lead, Biometric.Vision

Similar Cases

License Plate Annotation for Vehicle Recognition System

  • 100,000 images with detailed license plate markup (bounding boxes, digits, regional symbols)
  • 3 weeks
Learn more

Hindi Speech Transcription Dataset for ASR Evaluation

  • 1.5 weeks
Learn more

Image Segmentation for Retail Applications

  • E-commerce and Retail
  • 1,000 high-resolution annotated images
    30+ object classes per image
  • 6 weeks
Learn more

Banking Call Categorization for NLP Automation

  • Banking
  • 363,000 audio files
  • 50 days
Learn more

Targeted Edge-Case Coverage for Face Presentation Attack Detection

  • 35.500 videos
Learn more

Ready to get started?

Tell us what you need — we’ll reply within 24h with a free estimate

    Andrew
    Head of Client Success

    — I'll guide you through every step, from your first
    message to full project delivery

    Thank you for your
    message

    It has been successfully sent!

    We use cookies to enhance your experience, personalize content, ads, and analyze traffic. By clicking 'Accept All', you agree to our Cookie Policy.