PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 17, 20243 citationsOpen Access

REPOEXEC: Evaluate Code Generation with a Repository-Level Executable Benchmark

View Full Paper
NHNam Le HaiDNDung Manh NguyenNBNghi D. Q. Bui

Key Points

Key points are not available for this paper at this time.

Abstract

The ability of CodeLLMs to generate executable and functionally correct code at the repository-level scale remains largely unexplored. We introduce, a novel benchmark for evaluating code generation at the repository-level scale, emphasizing executability and correctness. provides an automated system that verifies requirements and incorporates a mechanism for dynamically generating high-coverage test cases to assess the functionality of generated code. Our work explores a controlled scenario where developers specify necessary code dependencies, challenging the model to integrate these accurately. Experiments show that while pretrained LLMs outperform instruction-tuning models in correctness, the latter excel in utilizing provided dependencies and demonstrating debugging capabilities. aims to provide a comprehensive evaluation of code functionality and alignment with developer intent, paving the way for more reliable and applicable CodeLLMs in real-world scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hai et al. (2024) studied this question.

synapsesocial.com/papers/68e64686b6db6435875d82dahttps://doi.org/10.48550/arxiv.2406.11927
Ask AI
Helpful
Bookmark
Share
View Full Paper