Planning for Risk-Aversion and Expected Value in MDPs

Marc Rigter; Paul Duckworth; Bruno Lacerda; Nick Hawes

Conference ProceedingsOPEN ACCESS

Planning for Risk-Aversion and Expected Value in MDPs

Proceedings International Conference on Automated Planning and Scheduling, ICAPS (2022) 32 307-315

DOI: 10.1609/icaps.v32i1.19814

7Citations

11Readers

Abstract

Planning in Markov decision processes (MDPs) typically optimises the expected cost. However, optimising the expectation does not consider the risk that for any given run of the MDP, the total cost received may be unacceptably high. An alternative approach is to find a policy which optimises a risk-averse objective such as conditional value at risk (CVaR). However, optimising the CVaR alone may result in poor performance in expectation. In this work, we begin by showing that there can be multiple policies which obtain the optimal CVaR. This motivates us to propose a lexicographic approach which minimises the expected cost subject to the constraint that the CVaR of the total cost is optimal. We present an algorithm for this problem and evaluate our approach on four domains. Our results demonstrate that our lexicographic approach improves the expected cost compared to the state of the art algorithm, while achieving the optimal CVaR.

Cite

CITATION STYLE

APA

Rigter, M., Duckworth, P., Lacerda, B., & Hawes, N. (2022). Planning for Risk-Aversion and Expected Value in MDPs. In Proceedings International Conference on Automated Planning and Scheduling, ICAPS (Vol. 32, pp. 307–315). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/icaps.v32i1.19814

Planning for Risk-Aversion and Expected Value in MDPs

Abstract

Cite

Register to see more suggestions