Explore the Reasoning Capability of LLMs in the Chess Testbed

0Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Reasoning is a central capability of human intelligence. In recent years, with the advent of large-scale datasets, pretrained large language models have emerged with new capabilities, including reasoning. However, these models still struggle with long-term, complex reasoning tasks, such as playing chess. Based on the observation that expert chess players employ a dual approach combining long-term strategic play with short-term tactical play along with language explanation, we propose improving the reasoning capability of large language models in chess by integrating annotated strategy and tactic. Specifically, we collect a dataset named MATE1, which consists of 1 million chess positions with candidate moves annotated by chess experts for strategy and tactics. We finetune the LLaMA-3-8B model and compare it against state-of-the-art commercial language models in the task of selecting better chess moves. Our experiments show that our models perform better than GPT, Claude, and Gemini models. We find that language explanations can enhance the reasoning capability of large language models.

Cite

CITATION STYLE

APA

Wang, S., Ji, L., Wang, R., Zhao, W., Liu, H., Hou, Y., & Wu, Y. N. (2025). Explore the Reasoning Capability of LLMs in the Chess Testbed. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 2, pp. 611–622). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-short.52

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free