PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 2026Science China Information Sciences12 citationsOpen Access

A survey on large language models for software engineering

QZQuanjun ZhangNanjing University of Science and TechnologyCFChunrong FangNanjing University
Yi Xie
Yi XieHebei University of Engineering

Key Points

  • This research aims to systematically survey the current applications and challenges of large language models in software engineering.
  • Conducted a survey of 62 large language models used in software engineering tasks.
  • Categorized 15 pre-training objectives and 16 downstream tasks relevant to software engineering.
  • Reviewed 926 studies related to code tasks and their integration into software engineering processes.
  • Identified key model architectures and their applications in software engineering tasks.
  • Highlighted critical aspects such as empirical evaluation and security challenges in LLM integration.
  • Presented potential opportunities for future research, including improved domain tuning and dataset evaluation.

Abstract

Abstract Software engineering (SE) is the systematic design, development, maintenance, and management of software applications underpinning the digital infrastructure of our modern world. Very recently, the SE community has seen a rapidly increasing number of techniques employing large language models (LLMs) to automate a broad range of SE tasks. Nevertheless, existing information on the applications, effects, and possible limitations of LLMs within SE is still not well-studied. In this paper, we provide a systematic survey to summarize the current state-of-the-art research in the LLM-based SE community. We summarize 62 representative LLMs of Code across three model architectures, 15 pre-training objectives across four categories, and 16 downstream tasks across five categories. We then present a detailed summarization of the recent SE studies for which LLMs are commonly utilized, including 926 studies for 112 specific code-related tasks across five crucial phases within the SE workflow. We also discuss several critical aspects during the integration of LLMs into SE, such as empirical evaluation, benchmarking, security and reliability, domain tuning, compressing, and distillation. Finally, we highlight several challenges and potential opportunities in applying LLMs for future SE studies, such as exploring domain LLMs and constructing clean evaluation datasets. Overall, our work can help researchers gain a comprehensive understanding about the achievements of the existing LLM-based SE studies and promote the practical application of these techniques. Our artifacts are publicly available and will be continuously updated at the living repository https://github.com/iSEngLab/AwesomeLLM4SE .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/69ba42cf4e9516ffd37a3750https://doi.org/10.1007/s11432-025-4670-0
Ask AI
Helpful
Bookmark
Share
View Full Paper