SE Perspective on LLMs: Biases in Code Generation, Code Interpretability, and Code Security Risks

5Citations
Citations of this article
50Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Large Language Models (LLMs) are transforming the world with their ability to generate diverse content, including code, but embedded biases raise significant concerns. In this perspective piece, we critique the wide-spreading view of LLMs as infallible tools by examining how biases in training data can lead to discriminatory code generation, opaque code interpretation, and heightened security risks, ultimately impacting the trustworthiness of LLM-generated software. Through a reflective analysis grounded in existing literature, including case studies and theoretical frameworks from software engineering and AI ethics, we examine the specific manifestations of bias in code generation, focusing on how training data contribute to these issues. We investigate the challenges associated with interpreting LLM-generated code, highlighting the lack of transparency and the potential for hidden biases, and explore the security risks introduced by biased LLMs, namely vulnerabilities that may be exploited by malicious actors. We provide several recommendations for mitigating these challenges, emphasizing the need to refine training data and involve humans-in-the-loop.

Cite

CITATION STYLE

APA

Krasniqi, R., Xu, D., & Vieira, M. (2025). SE Perspective on LLMs: Biases in Code Generation, Code Interpretability, and Code Security Risks. ACM Computing Surveys, 58(5). https://doi.org/10.1145/3774324

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free