Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 17, 2025Open Access

NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition

View Full Paper
Ask AI
Bookmark
Share

Authors

LDLe DengZJZhonghao JiangJCJialun Cao

Discussion

Loading...

Member takes

Overview

Benchmark evaluates task success rates of large language models in no-code development, highlighting challenges.

Key Points

  • The best large language models achieved a task success rate of only 28.07%, indicating significant limitations.
  • NoCode-bench includes 634 tasks based on software documentation updates and 114k associated code changes.
  • Despite high token usage, challenges persist in codebase understanding, cross-file editing, and tool invocation.
  • The benchmark establishes a foundation for future advancements in natural language-driven no-code development.

Cite This Study

Deng et al. (2025) studied this question.

synapsesocial.com/papers/68f19f20de32064e504ddf2ehttps://doi.org/10.48550/arxiv.2507.18130
View Full Paper
Ask AI
Bookmark
Share