PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 20250 citationsOpen Access

Automated Benchmark Generation for Repository-Level Coding Tasks

View Full Paper
KVKonstantinos VergopoulosMMMark Niklas MüllerMVMartin Vechev

Key Points

  • SetUpAgent generates datasets that help code agents better handle repository-level coding tasks, finding 40% lower success rates.
  • The new datasets, SWEE-Bench and SWA-Bench, enhance diversity over SWE-Bench, using hundreds of repositories for evaluation.
  • Significant distributional differences were observed in issue descriptions and fix complexities, affecting agent performance.
  • Historically accurate dependency setups and test execution are now automated, reducing manual efforts in benchmark generation.

Abstract

Code Agent development is an extremely active research area, where a reliable performance metric is critical for tracking progress and guiding new developments. This demand is underscored by the meteoric rise in popularity of SWE-Bench. This benchmark challenges code agents to generate patches addressing GitHub issues given the full repository as context. The correctness of generated patches is then evaluated by executing a human-written test suite extracted from the repository after the issue's resolution. However, constructing benchmarks like SWE-Bench requires substantial manual effort to set up historically accurate execution environments for testing. Crucially, this severely limits the number of considered repositories, e.g., just 12 for SWE-Bench. Considering so few repositories, selected for their popularity runs the risk of leading to a distributional mismatch, i.e., the measured performance may not be representative of real-world scenarios potentially misguiding development efforts. In this work, we address this challenge and introduce SetUpAgent, a fully automated system capable of historically accurate dependency setup, test execution, and result parsing. Using SetUpAgent, we generate two new datasets: (i) SWEE-Bench an extended version of SWE-Bench encompassing hundreds of repositories, and (ii) SWA-Bench a benchmark focusing on applications rather than libraries. Comparing these datasets to SWE-Bench with respect to their characteristics and code agent performance, we find significant distributional differences, including lower issue description quality and detail level, higher fix complexity, and most importantly up to 40% lower agent success rates.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vergopoulos et al. (2025) studied this question.

synapsesocial.com/papers/68d90a0f41e1c178a14f69d2https://doi.org/10.48550/arxiv.2503.07701
Ask AI
Helpful
Bookmark
Share
View Full Paper