---
title: "Synthetic Passports Dataset"
description: "The passport dataset comprises synthetic document images from multiple countries with metadata and is designed for training AI models in face recognition, identity verification, and…"
url: "https://unidata.pro/datasets/synthetic-passports/"
date_modified: "2025-10-01T08:27:51+03:00"
language: "en-US"
---
The passport dataset comprises synthetic document images from multiple countries with metadata and is designed for training AI models in face recognition, identity verification, and document analysis to detect fake passports and prevent identity fraud

## Dataset Structure

### The Numbers Section

**Numbered list:**

- **Number:** 100 000 — **Text:** Images
- **Number:** 100+ — **Text:** Countries

### Tooltips Section

**Tooltip items:**

- **Name:** Know your customer
- **Name:** Data generation
- **Name:** Computer Vision
- **Name:** Security
- **Name:** Anti-spoofing

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | Generated passports for training a neural network to identify a document |
| Data types | Image |
| Tasks | Face recognition, Computer Vision |
| Total number of files | 100 000 |
| Number of countries | 100+ |
| Labeling | Metadata (background) |
| Gender | Male, Female |
| Countries | Albania, Algeria, Andorra, Argentina, Armenia, Australia, Austria, Azerbaijan, Bahrain, Bangladesh, Belarus, Belgium, Bosnia and Herzegovina, Brazil, Brunei, Bulgaria, Cambodia, Canada, China, Croatia, Cyprus, Czech Republic, Denmark, Egypt, Estonia, Finland, France, Georgia, Germany, Greece, Hong Kong (HKSAR), Hungary, India, Indonesia, Iraq, Ireland, Israel, Italy, Japan, Jordan, Kazakhstan, Kuwait, Kyrgyzstan, Laos, Latvia, Lebanon, Liechtenstein, Lithuania, Luxembourg, Malaysia, Malta, Mexico, Moldova, Mongolia, Montenegro, Morocco, Myanmar, Nepal, Netherlands, New Zealand, North Macedonia, Norway, Oman, Pakistan, Palestine, Philippines, Poland, Portugal, Qatar, Romania, Russia, Russia Standart, Saudi Arabia, Serbia, Singapore, Slovakia, Slovenia, South Africa, South Korea, Spain, Sri Lanka, Sweden, Switzerland, Taiwan, Tajikistan, Thailand, Tunisia, Turkey, Turkmenistan, UAE, UK, Ukraine, USA, Uzbekistan, Vatican, Vietnam and others |

**Media Slider:**

- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2024/10/11.webp)
- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2024/10/12.webp)

**Link to the sample:** [Download sample](https://drive.google.com/drive/folders/1T1nmPTG0i57dd2W6sGUWjSTVjQKRIwAT?usp=sharing)

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Image Extensions | png |
| Data Type | generated |

**Source and data collection methodology:** Source and collection methodology. Data was AI-generated.

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Financial Services — **Title:** Training Identity Verification Systems — **Text:** Banks and fintech companies use Synthetic Passport Dataset to strengthen fraud detection and onboarding workflows. Since the dataset includes passport images from different countries, models learn to identify ID documents accurately without exposing personal data. This enables safer, compliant, and scalable identity verification solutions built on synthetic generation technology.
- **Industry:** Border Control & Security — **Title:** Document Recognition for Immigration Systems — **Text:** Security agencies apply passport datasets to improve automated checks at airports and border crossings. The dataset contains images generated to represent multiple document types, ensuring systems recognize variations in ID cards and identity documents. By using synthetic data, authorities train verification systems without relying on sensitive personal information.
- **Industry:** Technology & AI Development — **Title:** Building Machine Learning Models — **Text:** This dataset provides a rich training base for learning models focused on document analysis and text recognition. Since the dataset consists of thousands of high-quality synthetic ID images, researchers can experiment with different types of layouts, fonts, and structures. This supports advancements in computer vision and intelligent recognition systems.
- **Industry:** Education & Research — **Title:** Training Without Sensitive Data Exposure — **Text:** Universities and research labs rely on the synthetic passport dataset as an alternative to public datasets containing real identity documents. Because the images generated are free from personal data, students and professionals can explore document analysis methods safely. This promotes innovation in synthetic ID research while maintaining compliance with privacy standards.

### Fact

**FAQs Heading:** FAQs

**List of Questions:**

- **Question:** What is Synthetic Passports Dataset used for? — **Answer:** It is designed for training learning models in computer vision, fraud detection, and identity verification. It helps develop systems for detecting fake documents, preventing identity fraud, and improving verification systems that process passport photos and other identification documents.
- **Question:** Does the dataset contain real personal information? — **Answer:** No. The dataset is fully AI-generated and does not contain real passport holders or genuine personal identity information. This makes it suitable for AI development where privacy, compliance, and data security are important considerations.
- **Question:** Is it possible to request a custom dataset? — **Answer:** Yes, Unidata provides the option to create custom datasets tailored to your project needs. You can specify the data sources, annotation types, or formats, and the dataset will be collected and labeled according to your requirements.
- **Question:** What should I consider before buying this dataset? — **Answer:** When purchasing it, consider the document types, document layout, and country coverage to ensure it meets your project’s needs. Since the dataset consists of synthetic datasets rather than real identity documents, it is best suited for training anti-spoofing, verification solutions, and computer vision models without handling actual personal information.
- **Question:** Is this a real-world dataset or synthetic data? — **Answer:** This passport dataset is synthetic data, not a real-world collection. All passport images are AI-generated and include metadata such as gender and background.
- **Question:** Do Unidata datasets follow GDPR or other data privacy regulations? — **Answer:** Yes. All our datasets follow GDPR standards and meet privacy legislation requirements. Data collection is performed solely through lawful methods.
- **Question:** How is the data stored? — **Answer:** All Unidata datasets are safeguarded on AWS’s reliable cloud infrastructure, designed for availability and growth. We operate under ISO 27001 and ISO 27701 standards, internationally recognized for strong security and privacy measures. This provides clients with a dependable and compliant data environment.
- **Question:** How long does it take to receive the dataset? — **Answer:** Once you submit a request, our team will contact you to confirm details and complete the necessary documents. After the agreement is signed and payment is processed, the dataset is typically delivered within 3–10 days.

## List of Parameters

- **Title:** Tasks — **Description:** Face recognition, Computer Vision
- **Title:** Labeling — **Description:** Metadata (background)
- **Title:** Total number of files — **Description:** 100 000
- **Title:** Number of countries — **Description:** 100+
- **Title:** Data type — **Description:** Image (PNG)

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
