Abstract
Generative Artificial Intelligence (AI) is transforming industries by producing high-quality text, images, music, and code. Its applications extend to natural language processing, computer vision, and creative arts. However, assessing these systems' performance and impact remains challenging due to their complexity, subjectivity, and open-ended outputs. This paper comprehensively reviews evaluation methods for generative AI, beginning with its evolution and major applications, including advanced models like GPT, DALL·E, and AlphaCode. It categorizes evaluation approaches into quantitative metrics (such as BLEU and FID) and qualitative methods (human assessment and user-centered testing). Key challenges, such as subjectivity, bias, and scalability, are explored alongside emerging trends like automated evaluation tools, ethical impact assessments, and multimodal techniques. Through real-world case studies, this paper highlights practical evaluation strategies and their limitations. By integrating current best practices and identifying future research opportunities, this study aims to guide the development of reliable, fair, and comprehensive evaluation frameworks for generative AI systems.
Cite
CITATION STYLE
-, L. R. (2025). Evaluating Generative AI: Challenges, Methods, and Future Directions. International Journal For Multidisciplinary Research, 7(1). https://doi.org/10.36948/ijfmr.2025.v07i01.37182
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.