Deep reinforcement learning for modeling chit-chat dialog with discrete attributes

Chinnadhurai Sankar; Sujith Ravi

Conference ProceedingsOPEN ACCESS

Deep reinforcement learning for modeling chit-chat dialog with discrete attributes

SIGDIAL 2019 - 20th Annual Meeting of the Special Interest Group Discourse Dialogue - Proceedings of the Conference (2019) 1-10

DOI: 10.18653/v1/w19-5901

17Citations

112Readers

Abstract

Open domain dialog systems face the challenge of being repetitive and producing generic responses. In this paper, we demonstrate that by conditioning the response generation on interpretable discrete dialog attributes and composed attributes, it helps improve the model perplexity and results in diverse and interesting non-redundant responses. We propose to formulate the dialog attribute prediction as a reinforcement learning (RL) problem and use policy gradients methods to optimize utterance generation using long-term rewards. Unlike existing RL approaches which formulate the token prediction as a policy, our method reduces the complexity of the policy optimization by limiting the action space to dialog attributes, thereby making the policy optimization more practical and sample efficient. We demonstrate this with experimental and human evaluations.

References Powered by Scopus

View more at Scopus

Cited by Powered by Scopus

View more at Scopus

Cite

CITATION STYLE

APA

Sankar, C., & Ravi, S. (2019). Deep reinforcement learning for modeling chit-chat dialog with discrete attributes. In SIGDIAL 2019 - 20th Annual Meeting of the Special Interest Group Discourse Dialogue - Proceedings of the Conference (pp. 1–10). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w19-5901

Readers over time

Readers' Seniority

PhD / Post grad / Masters / Doc 45

79%

Researcher 7

12%

Lecturer / Post doc 4

Professor / Associate Prof. 1

Readers' Discipline

Computer Science 53

84%

Linguistics 5

Engineering 3

Social Sciences 2

Deep reinforcement learning for modeling chit-chat dialog with discrete attributes

Abstract

References Powered by Scopus

Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning

SWITCHBOARD: Telephone speech corpus for research and development

A diversity-promoting objective function for neural conversation models

Cited by Powered by Scopus

Survey on reinforcement learning for language processing

A Large-Scale Dataset for Empathetic Response Generation

A modular data-driven architecture for empathetic conversational agents

Register to see more suggestions

Cite

Readers over time

Readers' Seniority

Readers' Discipline