---
title: "Selfie with ID Dataset"
description: "This document dataset contains selfie images paired with identity documents for training and evaluating AI models for facial recognition and identity verification tasks in KYC…"
url: "https://unidata.pro/datasets/selfie-with-id/"
date_modified: "2026-05-15T11:24:11+03:00"
language: "en-US"
---
This document dataset contains selfie images paired with identity documents for training and evaluating AI models for facial recognition and identity verification tasks in KYC applications. It includes technical metadata and demographic attributes such as age, gender, and country.

## Dataset Structure

### The Numbers Section

**Numbered list:**

- **Number:** 69,435 — **Text:** Files
- **Number:** 4,629 — **Text:** People
- **Number:** 85 — **Text:** Countries

### Tooltips Section

**Tooltip items:**

- **Name:** Re-identification
- **Name:** Facial Recognition
- **Name:** Computer Vision
- **Name:** Security
- **Name:** iBeta

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | Photos of individuals and their identification documents for facial recognition tasks. |
| Data types | Image |
| Tasks | Face recognition, Computer Vision, Biometric Verification |
| Total number of images | 69,435 |
| Total number of people | 4,629 |
| Number of files in a set | 15 (13 selfies and 2 photos of document) |
| Labeling | Only technical characteristics and metadata (age, gender, country) |
| Gender | Male, Female |
| Number of countries | 85 |
| Type of document | Passports, international passports, driver licenses, student cards, health certificates, membership/bank/transport cards, certificates, etc |

**Media Slider:**

- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/id_2.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/id_1.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_1.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_2.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_3.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_4.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_5.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_6.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_7.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_8.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_9.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_10.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_11.webp)
- **Image in the slider:** ![selfie with ID dataset](https://unidata.pro/wp-content/uploads/2024/10/selfie_12.webp)

**Link to the sample:** [Download sample](https://www.google.com/url?q=https://drive.google.com/drive/folders/1_r5aNHl-syhgd4DPb1NgTTxZYvp10Sb8&sa=D&source=editors&ust=1730370482220906&usg=AOvVaw1pWxK45jJ1an7d9TlpFJyX)

### LLM Languages

**Section Title:** Statistics

**List of Statistics:**

- **Filter by:** Age distribution — **GIF image:** Age Distribution — **Table with data:**

| Age | Count |
| --- | --- |
| Under 18 | 451 |
| 19 - 25 | 1411 |
| 26-32 | 1301 |
| 33 - 39 | 889 |
| 40 - 46 | 377 |
| 47 - 53 | 128 |
| 54 - 59 | 45 |
| 60 - 66 | 23 |
| 67+ | 4 |
- **Filter by:** Country Distribution — **GIF image:** Country — **Table with data:**

| Country | Count | Percent |
| --- | --- | --- |
| Russia | 3084 | 51.13% |
| Brazil | 214 | 3.55% |
| Nigeria | 155 | 2.57% |
| Ukraine | 155 | 2.57% |
| Venezuela | 78 | 1.29% |
| Colombia | 64 | 1.06% |
| India | 49 | 0.81% |
| Peru | 45 | 0.75% |
| Turkey | 42 | 0.70% |
| Spain | 36 | 0.60% |
| Kenya | 35 | 0.58% |
| Pakistan | 32 | 0.53% |
| Dominican Republic | 30 | 0.50% |
| Argentina | 29 | 0.48% |
| Ecuador | 27 | 0.45% |
| Moldova | 27 | 0.45% |
| Philippines | 24 | 0.40% |
| Kazakhstan | 20 | 0.33% |
| Mexico | 18 | 0.30% |
| Poland | 14 | 0.23% |
| Nicaragua | 13 | 0.22% |
| South Africa | 13 | 0.22% |
| Egypt | 12 | 0.20% |
| Ethiopia | 11 | 0.18% |
| Italy | 11 | 0.18% |
| Costa Rica | 11 | 0.18% |
| Panama | 11 | 0.18% |
| France | 10 | 0.17% |
| Bolivia | 9 | 0.15% |
| Germany | 9 | 0.15% |
| Guatemala | 9 | 0.15% |
| Uruguay | 9 | 0.15% |
| Chile | 8 | 0.13% |
| Honduras | 8 | 0.13% |
| Paraguay | 8 | 0.13% |
| United States | 8 | 0.13% |
| Uganda | 7 | 0.12% |
| Madagascar | 6 | 0.10% |
| Uzbekistan | 6 | 0.10% |
| Greece | 5 | 0.08% |
| Nepal | 5 | 0.08% |
| El Salvador | 5 | 0.08% |
| Bangladesh | 4 | 0.07% |
| Czech Republic | 4 | 0.07% |
| Ghana | 4 | 0.07% |
| Portugal | 4 | 0.07% |
| Romania | 4 | 0.07% |
| United Kingdom | 4 | 0.07% |
| Algeria | 3 | 0.05% |
| Azerbaijan | 3 | 0.05% |
| Other | 65 | 1.08% |

### Statistics - Charts

**Charts with Titles:**

- **Shortcode:** [ays_chart id='26'] — **caption above the graph:** Gender Distribution
- **Shortcode:** [ays_chart id='25'] — **caption above the graph:** Continent Distribution

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Image Extensions | Jpg, jpeg, heic |
| Devices | Xiaomi redmi note 10s, Infinix smart 8, Samsung, iPhone 11, Xiaomi Redmi 14C, iPhone X, Redmi and etc. |

**Source and data collection methodology:** Source and collection methodology: Data was collected via crowdsourcing platforms.

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Financial Services — **Title:** KYC and Remote Onboarding — **Text:** Selfie with ID Dataset supports banks and fintech platforms in improving identity verification during digital onboarding. The dataset contains paired selfie photos and ID documents, allowing recognition systems to confirm customer identities. This reduces fraud, strengthens verification systems, and ensures compliance with KYC requirements for secure financial transactions.
- **Industry:** Telecommunications — **Title:** Subscriber Verification and SIM Registration — **Text:** Telecom companies use the Selfies & ID Images Dataset to validate ID cards during new SIM activation and subscription management. By matching facial images from selfie photos with official ID photos, providers enhance document verification, prevent fraudulent accounts, and protect personal information while meeting regulatory compliance standards.
- **Industry:** E-Government Services — **Title:** Digital Identity Authentication — **Text:** The Selfies and Documents Dataset enables public agencies to strengthen identity verification in e-government platforms. Since the dataset consists of aligned facial images and ID documents, it helps train systems that authenticate citizens securely. This safeguards access to public services, reduces impersonation risks, and ensures safe handling of sensitive personal information.
- **Industry:** Technology and Biometrics Industry — **Title:** Training and Benchmarking Models — **Text:** For technology companies and research groups, the Selfie and ID Dataset provides high-quality biometric data for recognition technology development. It serves as training data and testing data for recognition models that compare facial features in selfie images with ID photos. This supports learning models, improves selfie apps, and advances global verification systems.

### Fact

**FAQs Heading:** FAQs

**List of Questions:**

- **Question:** What are the sources of data for Unidata datasets? — **Answer:** Selfie and ID Dataset is collected from legally permissible sources, including crowdsourcing platforms and in-house data capture teams. All selfies and documents dataset entries are validated to ensure high-quality facial images and ID documents for reliable recognition technology testing.
- **Question:** What types of annotations are provided? — **Answer:** The dataset includes annotations for facial landmarks, ID document metadata, and identity verification attributes. These annotations support facial recognition, verification systems, and biometric data analysis in benchmark datasets and recognition technology applications.
- **Question:** Can this dataset be used for biometric verification and KYC applications? — **Answer:** Yes. The dataset was specifically designed for biometric verification, face recognition, and Know Your Customer (KYC) workflows by pairing selfie images with identity documents from the same individual. It is suitable for training AI models used in digital onboarding, customer verification, and fraud prevention.
- **Question:** What technical formats and devices are used in the dataset? — **Answer:** Images are provided in JPG, JPEG, and HEIC formats and were captured using a wide range of smartphones, including iPhone, Samsung, Xiaomi, Redmi, and Infinix devices. This variety reflects real-world image quality and device characteristics commonly encountered in KYC and remote identity verification workflows.
- **Question:** Can I request a sample of the dataset before purchasing? — **Answer:** Yes. Unidata provides a sample of the selfie with ID dataset so you can evaluate the dataset quality, selfie photos, and ID documents coverage. This allows you to test its applicability for document verification and facial recognition before making a purchase.
- **Question:** Do Unidata datasets follow GDPR or other data privacy regulations? — **Answer:** Yes. Every dataset is curated in accordance with GDPR and applicable privacy laws. Data is obtained only from legal and legitimate sources to ensure proper usage.
- **Question:** How are Unidata datasets licensed? — **Answer:** Our datasets are offered under a dual licensing framework. Free samples can be used for evaluation, but the full datasets require purchase.
- **Question:** How are Unidata datasets stored? — **Answer:** All Unidata datasets are securely stored on AWS, benefiting from its resilient cloud infrastructure. We follow ISO 27001 and ISO 27701 frameworks to maintain compliance with international security and privacy requirements. This ensures users can rely on a safe, and privacy-first environment for their data.

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
