Abstract
The financial sector is undergoing a profound transformation as AI technologies, particularly large language models (LLMs), revolutionize financial analysis through efficient and accurate natural language processing (NLP). This study investigates the efficacy of LLMs in addressing financial question-answering tasks, focusing specifically on two state-of-the-art models: Llama2-7b by Meta and Gemma-7b by Google. Despite their established prowess in general NLP tasks, their suitability for domain-specific applications, such as financial question answering, necessitates further exploration. Employing a comprehensive evaluation approach encompassing zero-shot prompt engineering, few-shot prompt engineering, and supervised fine-tuning methodologies, this study assesses the performance of Llama2 and Gemma using key metrics, including ROUGE-L, cosine similarity, and human evaluation. The preliminary findings reveal significant distinctions between the two models. Llama2 demonstrates a higher frequency of correct answers, but it is prone to hallucinations, often producing incorrect or incomplete information. In contrast, Gemma's performance is notably inferior, struggling to respond accurately to most queries. These observations highlight the need for continued research to enhance LLMs' ability to answer financial questions, while this study offers key insights into their strengths and weaknesses in managing financial inquiries.
Author supplied keywords
Cite
CITATION STYLE
Mridul Provakar, M., & Hashi, E. K. (2025). Exploring the Effectiveness of Large Language Models in Financial Question Answering. In ICCA 2024 - 3rd International Conference on Computing Advancements, 2024 (pp. 498–505). Association for Computing Machinery, Inc. https://doi.org/10.1145/3723178.3723244
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.