Comments on "Co-evolution in the successful learning of backgammon strategy"

14Citations
Citations of this article
57Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

The results obtained by Pollack and Blair substantially underperform my 1992 TD Learning results. This is shown by directly benchmarking the 1992 TD nets against Pubeval. A plausible hypothesis for this underperformance is that, unlike TD learning, the hillclimbing algorithm fails to capture nonlinear structure inherent in the problem, and despite the presence of hidden units, only obtains a linear approximation to the optimal policy for backgammon. Two lines of evidence supporting this hypothesis are discussed, the first coming from the structure of the Pubeval benchmark program, and the second coming from experiments replicating the Pollack and Blair results. © 1998 Kluwer Academic Publishers.

Cite

CITATION STYLE

APA

Tesauro, G. (1998). Comments on “Co-evolution in the successful learning of backgammon strategy.” Machine Learning. https://doi.org/10.1023/A:1007469231743

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free