---
title: "SWE-Bench Coding Tasks Dataset"
description: "Extended programming languages dataset based on SWE-Bench, featuring broader language coverage, golden and test patches, and real-world coding tasks such as bug fixing, code completion,…"
url: "https://unidata.pro/datasets/swe-bench-coding-tasks/"
date_modified: "2025-11-25T15:34:58+03:00"
language: "en-US"
---
Extended programming languages dataset based on SWE-Bench, featuring broader language coverage, golden and test patches, and real-world coding tasks such as bug fixing, code completion, and automated code review. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets.

## Dataset Structure

### The Numbers Section

**Numbered list:**

- **Number:** 8,712 — **Text:** files
- **Number:** 6 — **Text:** programming languages

### Tooltips Section

**Tooltip items:**

- **Name:** Programming languages
- **Name:** Machine Learning
- **Name:** Automated Code Review
- **Name:** Bug Fixing

### Dataset Information

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Description | An extended benchmark of real-world software engineering tasks with enhanced artifacts and broader language coverage |
| Data types | Text |
| Tasks | Bug fixing, code completion, pull request generation, automated code review |
| Total number of files | 8,712 |
| Total number of people | 30 |
| Labeling | Annotated with golden patches, test patches, post-patch reference states, and metadata stored in parquet files (e.g., repository name, issue/PR identifier, diffs, test results) |
| Programming languages | C#, Go, PHP, Rust, Kotlin, Ruby |

**Media Slider:**

- **Image in the slider:** ![](https://unidata.pro/wp-content/uploads/2025/09/fermatix-swe-bench-slider.webp)
- **Image in the slider:** ![Fermatix SWE-Bench dataset](https://unidata.pro/wp-content/uploads/2025/09/fermatix-swe-bench-slider-2.webp)

**Link to the sample:** [Dataset sample](https://drive.google.com/drive/folders/15xIWOrDGL4j5A7goF-V7M-SqAPE8pdcR?usp=sharing)

### Technical Specifications

**Table with data:**

| Characteristic | Data |
| --- | --- |
| Files Extensions | parquet (metadata), .patch (golden/test patches), .txt/.xml (reference outputs), .yml (docker-compose), Dockerfile, Makefile, .env |
| Models | Compatible with original Multi-SWE-Bench execution tools and models designed for code understanding and generation |
| File Size | 8.85 GB |

**Source and data collection methodology:** Source and collection methodology. Data was collected from permissively-licensed, non-utilitarian GitHub repositories to ensure diversity and reduce bias.

### Dataset Use Cases - Slider

**Industry Cards:**

- **Industry:** Software Development — **Title:** Improving Automated Code Generation — **Text:** This dataset supports developers working on automated coding tools by providing verified issue–resolution pairs from real Python repositories. It helps models understand bug patterns, code structure, and patch creation across large codebases. With consistent annotations and real-world examples, it strengthens model training for reliable code generation and automated software repair tasks.
- **Industry:** Machine Learning & AI — **Title:** Training and Evaluating Coding Agents — **Text:** This SWE-bench verified dataset helps researchers train large language models on software engineering challenges. With tasks requiring Python repositories and large codebases, the Multi-SWE-Bench framework delivers benchmark scores and evaluation results, making it a reliable resource for building better coding models and testing new benchmarks in real software environments.
- **Industry:** Software Engineering Research — **Title:** Benchmarking Real-World Engineering Tasks — **Text:** SWE-Bench Dataset introduces a robust SWE benchmark for analyzing engineering tasks across large codebases. By including GitHub repository issues and patches, it creates realistic conditions for testing developer tools, assessing coding tasks, and validating language models against existing benchmarks, enhancing reliability in software engineering research and development practices.
- **Industry:** Developer Tools & Testing — **Title:** Enhancing Reliability in Software Projects — **Text:** With nearly 9,000 files and curated annotations, this dataset helps improve developer tools for bug fixing and pull request generation. It supports testing coding agents in python projects, refining evaluation results, and addressing real-world coding challenges, strengthening recognition of pass rates across software development and testing pipelines.

### Fact

**FAQs Heading:** FAQs

**List of Questions:**

- **Question:** What makes SWE-Bench Coding Tasks Dataset different from existing benchmarks? — **Answer:** Unlike existing benchmarks limited to one language, this one expands to multiple languages and provides enhanced metadata. This makes it a more robust option for advanced coding evaluation, engineering tasks, and software development research.
- **Question:** Which programming languages are supported? — **Answer:** The dataset covers multiple programming languages including Python, C#, Go, PHP, Rust, Kotlin, and Ruby. This broader scope makes it suitable for multilingual coding benchmarks and large codebases.
- **Question:** How large is the dataset and in what format is it available? — **Answer:** The dataset size is 8.85 GB, with files in .parquet, .patch, .yml, .txt, .xml, along with Dockerfiles and Makefiles. This structure ensures compatibility with Multi-SWE-Bench execution tools and reproducible workflows.
- **Question:** What types of annotations are provided? — **Answer:** Annotations include golden patches, test patches, diffs, test results, and repository metadata. This ensures models are evaluated against verified coding benchmarks with clear pass/fail criteria, supporting transparent evaluation results.
- **Question:** Do Unidata datasets follow GDPR and other privacy regulations? — **Answer:** Yes. All Unidata datasets comply with GDPR and applicable international privacy laws. Data is collected from legally permissible sources and curated to ensure ethical and responsible usage.
- **Question:** How are Unidata datasets stored? — **Answer:** All datasets are stored securely on AWS cloud infrastructure, following ISO 27001 and ISO 27701 standards. This ensures high security, availability, and reliable long-term dataset management.
- **Question:** Is this dataset real-world or synthetic? — **Answer:** This is a real-world dataset, sourced from genuine GitHub issues, pull requests, and repository histories. Each task reflects actual engineering challenges encountered in live software projects.
- **Question:** How are Unidata datasets licensed? — **Answer:** Unidata datasets follow a dual-licensing model. Free samples are available for testing and trial use, while full datasets are provided exclusively through purchase for commercial or large-scale research applications.
- **Question:** Why is SWE-Bench useful for evaluating coding LLMs? — **Answer:** Unlike synthetic programming benchmarks, SWE-Bench measures whether large language models can solve realistic software engineering problems using existing project codebases, dependencies, and repository context.

[Full list of this site's AI-readable pages](https://unidata.pro/llms.txt)
