SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Alexander Golubev · Feb 27, 2026 · Citations: 0

Abstract

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). However, RL training is constrained by the scarcity of large-scale task collections with reproducible execution environments and reliable test suites. Although a growing number of benchmarks have emerged, datasets suitable for training remain limited in scale and diversity or often target a limited set of high-resource language ecosystems. We introduce SWE-rebench V2, a language-agnostic automated pipeline for harvesting executable real-world SWE tasks and constructing RL training environments at scale. The pipeline synthesizes repository-specific installation and test procedures via an interactive setup agent, and filters unsound instances using an ensemble of LLM judges, validated against human-verified SWE-bench annotations. Using this pipeline, we construct a dataset of 32,000+ tasks spanning 20 languages and 3,600+ repositories, with pre-built images for reproducible execution. To further scale training data, we additionally release 120,000+ tasks with installation instructions, fail-to-pass tests and rich metadata, where the problem statement is generated based on the original pull request description. We validate the collected instances through a diagnostic study that covers a subset of tasks in five programming languages across seven popular models, and provide instance-level metadata that flags common confounders such as overly restrictive tests and underspecified descriptions. We release the datasets, the collection and execution code, and associated artifacts to enable large-scale training of SWE agents across diverse languages and repositories.

HFEPX Relevance Assessment

This paper appears adjacent to HFEPX scope (human-feedback/eval), but does not show strong direct protocol evidence in metadata/abstract.

Eval-Fit Score

0/100 • Low

Treat as adjacent context, not a core eval-method reference.

Human Feedback Signal

Not explicit in abstract metadata

Evaluation Signal

Weak / implicit signal

HFEPX Fit

Adjacent candidate

If you are doing eval pipeline work, start here:

Human Eval Hub LLM-as-Judge Hub Pairwise Preference Hub Tool-Use Eval Hub

Human Data Lens

Uses human feedback: No
Feedback types: None
Rater population: Unknown
Unit of annotation: Unknown
Expertise required: Coding
Extraction source: Runtime deterministic fallback

Evaluation Lens

Evaluation modes:
Agentic eval: None
Quality controls: Not reported
Confidence: 0.25
Flags: low_signal, possible_false_positive, runtime_fallback_extraction

Protocol And Measurement Signals

Benchmarks / Datasets

SWE-benchswe-rebench

Reported Metrics

No metric terms were extracted from the available abstract.

Research Brief

Deterministic synthesis

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). HFEPX protocol signal is limited in abstract-level metadata, so treat it as adjacent context. Updated from current HFEPX corpus.

Generated Mar 3, 2026, 10:58 PM · Grounded in abstract + metadata only

Key Takeaways

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL).
Although a growing number of benchmarks have emerged, datasets suitable for training remain limited in scale and diversity or often target a limited set of high-resource language…

Researcher Actions

Treat this as method context, then pivot to protocol-specific HFEPX hubs.
Cross-check benchmark overlap: SWE-bench, swe-rebench.
Verify metric definitions before comparing against your eval pipeline.

Caveats

Generated from title, abstract, and extracted metadata only; full-paper implementation details are not parsed.
Low-signal flag detected: protocol relevance may be indirect.

Recommended Queries

human-eval protocol design pairwise preference data quality inter-rater agreement adjudication

Research Summary

Contribution Summary

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL).
Although a growing number of benchmarks have emerged, datasets suitable for training remain limited in scale and diversity or often target a limited set of high-resource language ecosystems.
We introduce SWE-rebench V2, a language-agnostic automated pipeline for harvesting executable real-world SWE tasks and constructing RL training environments at scale.

Why It Matters For Eval

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL).
Although a growing number of benchmarks have emerged, datasets suitable for training remain limited in scale and diversity or often target a limited set of high-resource language ecosystems.

Researcher Checklist

Gap: Human feedback protocol is explicit

No explicit human feedback protocol detected.
Gap: Evaluation mode is explicit

No clear evaluation mode extracted.
Gap: Quality control reporting appears

No calibration/adjudication/IAA control explicitly detected.
Pass: Benchmark or dataset anchors are present

Detected: SWE-bench, swe-rebench
Gap: Metric reporting is present

No metric terms extracted.

Category-Adjacent Papers (Broader Context)

These papers are nearby in arXiv category and useful for broader context, but not necessarily protocol-matched to this paper.

SkillCraft: Can LLM Agents Learn to Use Tools Skillfully? Category Neighbor

Citations: 0 Relevance: 3.20
- Shared arXiv category (cs.SE, cs.CL)
InnerQ: Hardware-aware Tuning-free Quantization of KV Cache for Large Language Models Category Neighbor

Citations: 0 Relevance: 2.05
- Shared arXiv category (cs.CL)
- Shared terminology (scale)
Natural Language Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift Category Neighbor

Citations: 0 Relevance: 2.05
- Shared arXiv category (cs.CL)
- Shared terminology (scale)

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote