---
title: "DeepFake Videos Dataset"
description: "The deepfake dataset contains real and AI-generated deepfake videos, featuring diverse subjects with detailed metadata on age, gender, and ethnicity to help train powerful deepfake…"
url: "https://unidata.pro/datasets/deepfake-videos-dataset/"
date_modified: "2026-02-16T22:37:41+03:00"
language: "en-US"
---
The deepfake dataset contains real and AI-generated deepfake videos, featuring diverse subjects with detailed metadata on age, gender, and ethnicity to help train powerful deepfake detectors

## Dataset Structure

### The Numbers Section

**Numbered list:**

- **Number:** 5,000 — **Text:** files
- **Number:** 5,000 — **Text:** people

### Tooltips Section

**Tooltip items:**

- **Name:** Facial Recognition
- **Name:** Computer Vision
- **Name:** Machine learning
- **Name:** Data generation
- **Name:** Security

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | Real video of people with AI-generated faces, where individuals turn their heads in different directions |
| Data types | Video |
| Tasks | Facial recognition, Computer Vision |
| Total number of files | 5,000 |
| Number of people | 5,000 |
| Video generation sites | aisaver.io, faceswapvideo.ai, magichour.ai |
| Labeling | Metadata (age, gender, ethnicity) |
| Gender | Male, Female |
| Ethnicity | Asian (30%), African (70%) |
| Age | Min = 18, max = 80, mean = 45 |

**Media Slider:**

- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2025/03/imagefordeep.webp) — **Video on Slayder:** <https://unidata.pro/wp-content/uploads/2025/03/deepfakevideo.webm>
- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2025/03/imagefordeep2.webp)

**Link to the sample:** [Download sample](https://drive.google.com/drive/folders/1FMoDhxNNLFeq33OIT3qQ12jMyw718yBg?usp=sharing)

### Statistics - Charts

**Charts with Titles:**

- **Shortcode:** [ays_chart id='29'] — **caption above the graph:** Distribution by age
- **Shortcode:** [ays_chart id='30'] — **caption above the graph:** Duration of the video duration
- **Shortcode:** [ays_chart id='31'] — **caption above the graph:** Distribution by gender
- **Shortcode:** [ays_chart id="32"] — **caption above the graph:** Distribution by ethnicity

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Video extension | mp4, MOV |
| Video Resolutions | 1920 x 1080p, 480 x 360p, 1280 x 720p, 720 x 480p, 640 x 480p, 1920 x 920p |
| Video duration | Mean = 9, median = 9, min = 2, max = 34 |
| Frames per second | Mean = 26.6 |
| Devices | iPhone 13 (30%), Google Pixel (70%) |

**Source and data collection methodology:** Source and collection methodology. Data was collected by overlaying generated faces onto real videos using the following websites: aisaver.io, faceswapvideo.ai, and magichour.ai.

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Cybersecurity & Digital Forensics — **Title:** Detecting Deepfake Content and Fake Videos — **Text:** Deepfake Videos Dataset provides critical training data for developing deepfake detectors and detection algorithms. Containing both real videos and fake videos generated using advanced deepfake technology, it allows analysts to train detection systems that identify synthetic media and protect against identity fraud, misinformation, and other digital threats.
- **Industry:** AI & Machine Learning Research — **Title:** Training Models for Deepfake Detection — **Text:** This deep fake detection dataset is widely used in machine learning and deep learning projects. The dataset consists of thousands of video clips and face images, including manually labelled examples. Models trained on this data achieve better accuracy in spotting AI-generated videos and distinguishing between real and synthetic video datasets.
- **Industry:** Media & Journalism — **Title:** Verifying Video Content Authenticity — **Text:** News organizations use such datasets to enhance video detection tools that verify YouTube videos, interviews, and shared clips. By training recognition systems on datasets containing both source videos and generated faces, journalists can validate footage, identify manipulated content, and strengthen trust in digital reporting.
- **Industry:** Technology & App Development — **Title:** Building Safer Recognition and Verification Systems — **Text:** Tech companies rely on the synthetic video dataset to test facial recognition and object detection systems against deepfake content. The dataset comprising high-quality video frames, fake images, and synthetic data helps in creating more reliable generative AI detection methods. This improves authentication solutions and delivers better results in protecting digital platforms.

### Fact

**FAQs Heading:** FAQs

**List of Questions:**

- **Question:** How large is DeepFake Videos Dataset compared to other available datasets? — **Answer:** With 5,000 video clips and diverse demographic coverage, this collection is one of the largest datasets of its kind. Its scale allows models trained on it to achieve higher accuracy in deepfake content detection and facial recognition tasks.
- **Question:** What devices and resolutions are represented in the dataset? — **Answer:** Videos were recorded on iPhone 13 devices (30%) and Google Pixel devices (70%), then processed with deepfake technology. The dataset covers multiple resolutions, including 1080p, 720p, and 480p, supporting a wide range of video detection methods.
- **Question:** How was the data collected? — **Answer:** The dataset was built by generating fake faces using AI models and overlaying them on real videos (using the following tools: aisaver.io, faceswapvideo.ai, and magichour.ai).
- **Question:** What metadata is included with the dataset? — **Answer:** Each video includes technical metadata such as age, gender, and ethnicity, supporting demographic analysis and balanced model training. The dataset also preserves video characteristics including resolution, duration, frame rate, and recording device for benchmarking deep learning models.
- **Question:** Is it possible to request a custom deepfake dataset? — **Answer:** Custom datasets can be created on request, allowing you to specify generation methods, annotation formats, or target demographics. This flexibility ensures better results for applications such as face recognition, synthetic video detection, or generative AI model training.
- **Question:** How can this dataset improve deepfake detection models? — **Answer:** The dataset provides paired real-world video characteristics with AI-generated facial manipulations, making it valuable for training deepfake detection and facial recognition models. Its diverse demographics, multiple video resolutions, and realistic head movements help improve model robustness across different attack scenarios.
- **Question:** How are Unidata datasets licensed? — **Answer:** Unidata datasets follow a dual-licensing model. Free samples are available for testing and evaluation, while full datasets, including the Deepfake Videos Dataset, can be accessed exclusively through purchase.
- **Question:** How are Unidata datasets stored? — **Answer:** Unidata securely stores datasets on AWS cloud infrastructure, ensuring high availability and scalability. Our storage practices comply with ISO 27001 and ISO 27701 standards, guaranteeing internationally recognized information security and privacy management for sensitive data.
- **Question:** How long does it take to receive the dataset? — **Answer:** Once you submit a request, our team will contact you to confirm details and finalize documentation. After signing and payment, the dataset will be delivered within 3–10 business days.
- **Question:** Is this a real-world dataset or synthetic data? — **Answer:** The Deepfake Videos Dataset is a synthetic video dataset created by combining real-world source videos with AI-generated faces. This hybrid approach provides realistic training material for developing deepfake detection methods, recognition models, and generative AI research.

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
