AI Model Testing Services

Your model scores 98% on your own data. We test it on ours. Independent evaluation on datasets your model has never seen shows what internal benchmarks hide.

Connect with our data expert
Image
Every dataset built & owned in-house
31 attack classes & 3 bona fide sets
Error rates per segment, condition & attack class
One module under NDA is enough to run the audit

AI Model Testing In Numbers

2 805 238
Labeled files
24
Labeled datasets
10+
Clients tested in biometrics & KYC

AI Testing Is QA for AI

Your QA team tests the system around the model: contracts, latency, fallbacks, whether the threshold config does what it says. We test the model behind it, with error rates per attack class and per user segment, measured on labeled data it has never seen.

Classic QA

  • Expected vs actual outcome
  • Acceptance criteria per feature
  • Regression suite
  • Security testing

AI model testing at unidata

  • Labeled ground truth at scale, supplied by us
  • Target metrics per module & per segment
  • Version-over-version testing on a fixed…
  • Presentation attack & deepfake resistance

Why Do You Need AI Model Testing in 2026?

Deployment without an audit

Your metrics come from your own data Every number from data your team collected
One number for the whole run The average hides the segments that fail
Paying customers get rejected A spoof score never shows rejected users
Certification booked from $25,000 Pass or fail, never the cause
Data bought before any diagnosis Full sets chosen on guesswork

With AI Model Audit from Unidata

Data from outside your pipeline Built & owned in-house, not from public sources
A number for every segment Per segment, per condition, per attack class
Real customers tested as hard as attackers False rejection reported for every segment
You book certification once The weak class, named before iBeta
Data bought after the report Only the 5–30% you actually need

Our AI Model Testing Services

One testing infrastructure. Most clients start with AI Model Audit and add the rest as the picture becomes clear

AI Model Audit

An independent baseline evaluation on Unidata test datasets.

Get a free consultation Start here if you have never tested outside your own data

Stress Testing

Hard, rare & degraded conditions: extreme angles, poor lighting, motion blur.

Get a free consultation Start here if the model works in the lab & fails with real users

Regression Testing

Each new version against the same fixed dataset, so you see what quietly got worse.

Get a free consultation Start here if you retrain often & release to production

Benchmarking

Two or more models on one dataset under identical conditions.

Get a free consultation Start here if you are choosing between vendors

Which Data Your Model Is Tested On

Pick a task to see which sets your model would be measured against, and how large each one is.

How Robust Is Your Training Data?

Lab-tested vs. real-world — not always the same.

Show me what's missing

What Your Model Is Tested On

Every sample carries metadata, so error rates are reported per attribute and per combination.

Ethnicity

Ethnicity

Capture devices

Capture devices

Gender

Gender

Capture conditions

Capture conditions

Age

Age

Biometric attacks

Biometric attacks

Deepfake Synthetic Overlay Silicone Mask Latex Mask 3D Mask Voluminous Screen Replay Phone Screen Replay Tablet Screen Replay Laptop Replay Second Device

Industries We Work With

Identity & Biometrics

Identity & Biometrics

Smart City

Smart City

Documents & OCR

Documents & OCR

Before You Book iBeta Certification

A lab evaluation is a paid, scheduled, pass-or-fail exam. Run it with us first.

Audit first with unidata, then the lab

01

Run the audit with Unidata

02

See what fails, and for whom

03

Buy only the 5–30% you need

04

Book the
 lab once

How the AI Model Audit Process Works at Unidata

Eight steps, two phases, one approval gate between them. Every step is marked with who owns it.

01 Briefing & Task Setup
Together
We discuss your model, define key metrics such as FAR, FRR and Accuracy, and identify focus areas.
02 NDA
Together
Signed before anything leaves your environment. Your IP and the terms of transfer are covered first.
03 Pilot & Estimation
Together
You provide access to the module. We return initial results with a timeline and a cost estimate.
04 Test Run Days 1–5
Unidata
Day 1 covers model & dataset intake. Days 2–5 run your module on our proprietary labeled datasets across every relevant attack class.
05 Error Analysis Day 6
Unidata
Results analysed across 20+ conditions: demographics, attack types, lighting, devices.
06 Validation Review Days 6–7
You
You review the performance breakdown by segment and condition before anything is final.
07 Report Delivery Day 7
Unidata
Visualisations, failure modes and recommendations for the model & for data selection.
08 Targeted Data Purchase Optional
You
You buy only the 5–30% of segments where the model struggles. Diagnosis comes first, data second.

Have questions about the process? Every project starts with a free consultation — no commitment required.

Book a call about your case
Image

Which Software We Use

Docker for the run, open-source frameworks for the evaluation, Streamlit and Grafana for the report. The checks behind every number are public, documented methods your own ML team can inspect.

The People Who Run Your Audit

Diagnosis Before Data

We do not recommend a dataset before the model has been run. First we find out where it fails, then we look at which data closes that gap. Buying a set because it sounds relevant is how a team spends a quarter and moves the metric by nothing.

Image
Martinian Letunovsky
Head of IT Department

The Other Half of Liveness

A liveness model can score well on standard attack datasets and still break on everyday variation: different lighting, glasses, an unusual angle. None of these are attacks, and a real customer is the one who gets rejected. We see this regularly, so we test the genuine-user side as hard as the attack side, with false rejection reported per segment. Turning those customers away is a cost, and it never shows up in a spoof-detection score.

Image
Elizabeth Karnaukhova
Project manager, AI model audit
Image
Example

How Your Dashboard Will Look

You open it yourself and read it without us in the room. Every row is a class or a segment, and every number is a rate, not a verdict.

Liveness Detection · V4.2
iBeta 31 attack classes 279,555 files sample data
Detection rate per attack class Acceptance rate per genuine-user segment
Printed photo attack 99.6%
Cut photo, eyes cut 98.9%
Screen replay - phone 97.2%
2D mask + glasses on top 91.4%
3D mask, voluminous 88.0%
Screen replay - tablet 75.0%
Latex mask 62.5%
Silicone mask 0.0%
Daylight · male · 30–40 99.8%
Age 60+ 91.2%
Children 7–12 88.6%
Glasses · low light 84.8%
Accuracy 96.4%
Far 3.1%
FRR 4.8%
Worst class 0.0%

Recommended Data To Close The Gaps

silicone mask sets glasses in low light age 60+ children 7-12

Roughly 12% of our biometric catalogue, selected from what the report above marks red

Get this report for my model

How We Protect Your Model & IP

NDA before transfer

A previous stable version is enough

Isolated VM per engagement

Your code never read or modified

No dataset leaves our infrastructure

Case Study: Liveness Model Audit for Biometric Security

  • Biometrics & Face Recognition
  • 1000+real-user videos with diverse spoof attacks
  • 2 months
Learn more
Liveness Verified

FAQ

What is an AI model audit?
An independent evaluation of a model on curated datasets with full metadata, run by a third party rather than by the team that built it. You provide a Docker container and a Python inference script, we run the model inside our infrastructure on data your team has never used, and you receive performance broken down by demographics, lighting, attack type and operational context, plus identified vulnerabilities and targeted data recommendations.
What do I have to hand over?
One module rather than your whole product. For example Liveness Detection, delivered as a Docker container with a Python inference script and, where needed, additional model files. If your current production build cannot leave your environment, the last stable version is enough: the failure modes are almost always the same. We run the container, record inputs and outputs, and measure them against labeled ground truth. Reverse engineering is outside the scope of the engagement and outside the terms of the NDA.
Why a Docker container?
The container carries your libraries and dependencies with it, so the model runs inside our infrastructure in the same environment as on your side. No code adaptation, no reimplementation on our end.
Why a Python inference script?
The script is the entry point. It takes an image or video, runs the model and returns a result, for example as JSON. During testing it runs automatically across the full dataset, and the outputs are what we analyse. Your model code is never modified.
Can you test our model through an API?
We run your model on our side rather than calling an endpoint on yours. An API test would mean sending our test datasets into your infrastructure, and those datasets are our intellectual property. Running locally also keeps conditions identical across every model we test, and it is usually lighter on your team: a container and an entry point, with nothing of your product exposed as a service.
Which metrics do you report?
FAR, FRR, accuracy and error rates, reported both in aggregate and per segment, per condition and per attack class. The exact set is agreed at briefing, because a liveness module and a document OCR pipeline need different definitions of quality.
How long does an audit take?
A standard audit of one module runs seven days from intake to report: one day to receive the model and prepare datasets, four days for the full test run, one day for error analysis, one day for delivery. The clock starts after the NDA and pilot, so scoping time sits outside it. Wider scopes run longer: the case study above covers 1000+ real-user videos across two months. You receive a firm timeline together with the pilot estimate.

Ready to get started?

Tell us what you need — we’ll reply within 24h with a free estimate

    Andrew
    Head of Client Success

    — I'll guide you through every step, from your first
    message to full project delivery

    Thank you for your
    message

    It has been successfully sent!

    We use cookies to enhance your experience, personalize content, ads, and analyze traffic. By clicking 'Accept All', you agree to our Cookie Policy.