Abstract
In this paper various LLMs are tested in a specific domain using a Retrieval-Augmented Generation (RAG) system. The study focuses on the performance and behavior of the models and was conducted in Spanish. A questionnaire based on The Bible, which consists of questions that vary in complexity of reasoning, was created in order to evaluate the reasoning capabilities of each model. The RAG system matches a question with the most similar passage from The Bible and feeds the pair to each LLM. The evaluation aims to determine whether each model can reason solely with the provided information or if it disregards the instructions given and makes use of its pretrained knowledge.
Author supplied keywords
Cite
CITATION STYLE
Cadena-Bautista, Á., López-Ponce, F. F., Ojeda-Trueba, S. L., Sierra, G., & Bel-Enguix, G. (2025). Exploring the Behavior and Performance of Large Language Models: Can LLMs Infer Answers to Questions Involving Restricted Information? Information (Switzerland), 16(2). https://doi.org/10.3390/info16020077
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.