PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 27, 2022618 citations

Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models

View Full Paper
PVPriyan VaithilingamTZTianyi ZhangEGElena L. Glassman

Key Points

  • This research aims to evaluate the usability of LLM-based code generation tools in real-world programming tasks.
  • Conducted a within-subjects user study with 24 participants
  • Analyzed participants' perceptions and experiences using Copilot in programming tasks
  • Identified usability challenges and suggestions for improvement based on user feedback.
  • No significant improvement in task completion time or success rate with Copilot
  • Participants preferred using Copilot for its ability to provide useful starting points
  • Users faced challenges in understanding, editing, and debugging code snippets, affecting task-solving effectiveness.

Abstract

Recent advances in Large Language Models (LLM) have made automatic code generation possible for real-world programming tasks in general-purpose programming languages such as Python. However, there are few human studies on the usability of these tools and how they fit the programming workflow. In this work, we conducted a within-subjects user study with 24 participants to understand how programmers use and perceive Copilot, a LLM-based code generation tool. We found that, while Copilot did not necessarily improve the task completion time or success rate, most participants preferred to use Copilot in daily programming tasks, since Copilot often provided a useful starting point and saved the effort of searching online. However, participants did face difficulties in understanding, editing, and debugging code snippets generated by Copilot, which significantly hindered their task-solving effectiveness. Finally, we highlighted several promising directions for improving the design of Copilot based on our observations and participants’ feedback.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vaithilingam et al. (2022) studied this question.

synapsesocial.com/papers/69d9d9d0a1d151c65f6854cehttps://doi.org/10.1145/3491101.3519665
Ask AI
Helpful
Bookmark
Share
View Full Paper