Cascaded LSTMs based deep reinforcement learning for goal-driven dialogue

Yue Ma; Xiaojie Wang; Zhenjiang Dong; Hong Chen

Conference Proceedings

Cascaded LSTMs based deep reinforcement learning for goal-driven dialogue

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2018) 10619 LNAI 29-41

DOI: 10.1007/978-3-319-73618-1_3

1Citations

21Readers

Get full text

Abstract

This paper proposes a deep neural network model for jointly modeling Natural Language Understanding and Dialogue Management in goal-driven dialogue systems. There are three parts in this model. A Long Short-Term Memory (LSTM) at the bottom of the network encodes utterances in each dialogue turn into a turn embedding. Dialogue embeddings are learned by a LSTM at the middle of the network, and updated by the feeding of all turn embeddings. The top part is a forward Deep Neural Network which converts dialogue embeddings into the Q-values of different dialogue actions. The cascaded LSTMs based reinforcement learning network is jointly optimized by making use of the rewards received at each dialogue turn as the only supervision information. There is no explicit NLU and dialogue states in the network. Experimental results show that our model outperforms both traditional Markov Decision Process (MDP) model and single LSTM with Deep Q-Network on meeting room booking tasks. Visualization of dialogue embeddings illustrates that the model can learn the representation of dialogue states.

Author supplied keywords

Cite

CITATION STYLE

APA

Ma, Y., Wang, X., Dong, Z., & Chen, H. (2018). Cascaded LSTMs based deep reinforcement learning for goal-driven dialogue. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 10619 LNAI, pp. 29–41). Springer Verlag. https://doi.org/10.1007/978-3-319-73618-1_3

Cascaded LSTMs based deep reinforcement learning for goal-driven dialogue

Abstract

Author supplied keywords

Cite

Register to see more suggestions