With the increasingly common use of Large Language Models (LLMs) in software development comes new concerns about the security of AI-generated code. These tools are commonly used to speed up development without explicit security prompts, and are often used by the developer to achieve production speeds. This paper is an empirical audit of the default security posture of code produced by two prominent LLMs: ChatGPT, by OpenAI, and Gemini, by Google. These models are evaluated by comparing their ability to deal with risky operations in two programming languages: C++ (unmanaged memory) and Java (managed memory). A programming prompt dataset was created and analyzed with five highly targeted programming prompts, and five code samples were collected from the dataset and analyzed with Static Application Security Testing (SAST) tools, namely Flawfinder and SpotBugs. The findings have shown that both models can still create serious vulnerabilities in C++, such as OS Command Injection (CWE-78), although managed languages such as Java can help minimize this problem. Moreover, there are medium-severity design flaws in both models: Absolute paths in Java. This study sheds light on the ongoing threat of relying on AI-generated code without proper human-in-the-loop security checks.
No takes yet. Share an insight, caveat, or question.
Jadhav Prabuddha (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: