“I'd Like to Have an Argument, Please”: Argumentative Reasoning in Large Language Models

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We evaluate two large language models (LLMs) ability to perform argumentative reasoning. We experiment with argument mining (AM) and argument pair extraction (APE), and evaluate the LLMs' ability to recognize arguments under progressively more abstract input and output (I/O) representations (e.g., arbitrary label sets, graphs, etc.). Unlike the well-known evaluation of prompt phrasings, abstraction evaluation retains the prompt's phrasing but tests reasoning capabilities. We find that scoring-wise the LLMs match or surpass the SOTA in AM and APE, and under certain I/O abstractions LLMs perform well, even beating chain-of-thought-we call this symbolic prompting. However, statistical analysis on the LLMs outputs when subject to small, yet still human-readable, alterations in the I/O representations (e.g., asking for BIO tags as opposed to line numbers) showed that the models are not performing reasoning. This suggests that LLM applications to some tasks, such as data labelling and paper reviewing, must be done with care.

Cite

CITATION STYLE

APA

de Wynter, A., & Yuan, T. (2024). “I’d Like to Have an Argument, Please”: Argumentative Reasoning in Large Language Models. In Frontiers in Artificial Intelligence and Applications (Vol. 388, pp. 73–84). IOS Press BV. https://doi.org/10.3233/FAIA240311

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free