---
title: "Synthetic Printed Mexican Passports Dataset"
description: "Includes a diverse collection of AI-generated Mexican passport images replicating authentic passport layouts, fonts, and visual features. Designed as a synthetic ID dataset for machine…"
url: "https://unidata.pro/datasets/synthetic-printed-mexican-passports/"
date_modified: "2026-02-16T22:40:42+03:00"
language: "en-US"
---
Includes a diverse collection of AI-generated Mexican passport images replicating authentic passport layouts, fonts, and visual features. Designed as a synthetic ID dataset for machine learning and OCR training, it supports document verification, data extraction, and identity recognition across varied lighting, backgrounds, and camera perspectives.

## Dataset Structure

### The Numbers Section

**Numbered list:** - **Number:** 5 000 — **Text:** Images

### Tooltips Section

**Tooltip items:**

- **Name:** PII
- **Name:** Data generation
- **Name:** Security
- **Name:** Anti-spoofing
- **Name:** Computer Vision

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | Printed synthetic passport images for training ML models in PII extraction |
| Data types | Image |
| Tasks | OCR, Computer Vision |
| Total number of files | 5 000 |
| Number of files in a set | 96 (Angles - 3, Lighting - 4, Backgrounds - 4, Distances - 2) |
| Angles | 0°, 25°, 45° |
| Lighting | Natural-daylight, Office-LED, Warm-indoor, Dim-light |
| Backgrounds | Neutral wall, Textured desk, Outdoor pavement, Docs-on-docs |
| Distance | Close (80-90 % frame), Medium (50-60 %) |
| Labeling | Metadata (Passport ID, Sample ID, Class, Country, Gender, Age Group, Angle, Distance, Category, Resolution, Camera, Light Condition, Background, Timestamp) |
| Gender | Male, Female |

**Media Slider:**

- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2025/10/mexican-passport-dataset2-scaled.webp)
- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2025/10/mexican-passport-dataset3-scaled.webp)

**Link to the sample:** [Download sample](https://drive.google.com/drive/folders/1WePUG8AYmZr2-XGA6OVe8GF6rmSpelUZ)

### Statistics - Charts

**Charts with Titles:** - **Shortcode:** [ays_chart id='65'] — **caption above the graph:** Distribution by gender

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Image Extensions | HEIC, JPG |
| Data Type | generated |

**Source and data collection methodology:** Source and collection methodology: Data was AI-generated.

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Government and Border Control — **Title:** Automated Passport Verification Systems — **Text:** Synthetic Printed Mexican Passports Dataset is ideal for training AI systems in passport recognition and border verification. It helps models identify visual and textual details from Mexican passports, improving accuracy in document validation, visa processing, and citizenship verification for secure border and immigration systems.
- **Industry:** Financial Services and Banking — **Title:** Identity Verification and KYC Automation — **Text:** Banks and fintech platforms can use this synthetic passport dataset to train models that detect and verify Mexican passports during digital onboarding. It strengthens fraud prevention, supports document authenticity checks, and helps improve KYC compliance by enabling automated extraction of identity data from high-resolution passport images.
- **Industry:** Travel and Airline Industry — **Title:** Streamlined Passenger Check-In Systems — **Text:** This dataset supports machine learning models used in airport kiosks and self-check-in systems. By providing images with varied lighting, angles, and backgrounds, it helps improve recognition performance during passport scanning, visa validation, and passenger data collection for smoother travel experiences.
- **Industry:** Research and AI Development — **Title:** Synthetic Data for Document Recognition Research — **Text:** AI researchers can use Mexican Passport Dataset to test OCR, data extraction, and document segmentation algorithms. Its AI-generated samples allow experimentation with security feature detection, data accuracy, and synthetic identity modeling, offering a reliable training resource without involving real personal data.

### Fact

**FAQs Heading:** FAQs

**List of Questions:**

- **Question:** What should I consider before purchasing this dataset? — **Answer:** Before purchasing the Synthetic Printed Mexican Passports Dataset, review your project’s requirements for OCR, computer vision, and PII extraction tasks. Check the dataset’s structure, image formats (HEIC, JPG), and metadata fields to ensure it aligns with your model training and document recognition objectives.
- **Question:** Can I request a sample of the dataset before purchasing? — **Answer:** Yes. A free sample is available so you can evaluate the image quality, metadata labeling, and synthetic generation accuracy. This helps ensure the dataset fits your technical and research needs before making a full purchase.
- **Question:** Is it possible to request a custom dataset? — **Answer:** Yes. Custom synthetic passport datasets can be created upon request to match specific parameters such as country, image angle, background, lighting, or data volume. This option allows developers and researchers to generate AI-ready data suited to unique project goals.
- **Question:** How are Unidata datasets licensed? — **Answer:** Unidata datasets follow a dual-licensing model - free samples are offered for testing and evaluation, while complete datasets are available through purchase. This ensures accessible data for both research and commercial use.
- **Question:** Do Unidata datasets follow GDPR or other privacy regulations? — **Answer:** Yes. All Unidata datasets are developed in compliance with GDPR and international data protection laws. Since this is a synthetic dataset, no personal or government records are involved, ensuring legal, safe, and privacy-compliant data usage.
- **Question:** How long does it take to receive the dataset? — **Answer:** Once you submit your request, our team will contact you to confirm details and finalize documentation. After signing and payment, the dataset is securely delivered within 3–10 business days.
- **Question:** Is this a real-world dataset or synthetic data? — **Answer:** This is a fully synthetic dataset, created using AI-based synthetic generation techniques. It simulates Mexican passports with realistic textures, lighting, and positioning, ideal for training document recognition systems without involving any real individuals or government data.
- **Question:** Why is synthetic data a good choice for document AI development? — **Answer:** Synthetic data provides realistic training samples while eliminating privacy risks associated with real identity documents. It also allows developers to create balanced datasets with consistent annotations, improving the quality and reproducibility of machine learning experiments.

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
