Querying to find a safe policy under uncertain safety constraints in Markov decision processes

Shun Zhang; Edmund H. Durfee; Satinder Singh

Conference ProceedingsOPEN ACCESS

Querying to find a safe policy under uncertain safety constraints in Markov decision processes

AAAI 2020 - 34th AAAI Conference on Artificial Intelligence (2020) 2552-2559

DOI: 10.1609/aaai.v34i03.5638

5Citations

10Readers

Abstract

An autonomous agent acting on behalf of a human user has the potential of causing side-effects that surprise the user in unsafe ways. When the agent cannot formulate a policy with only side-effects it knows are safe, it needs to selectively query the user about whether other useful side-effects are safe. Our goal is an algorithm that queries about as few potential side-effects as possible to find a safe policy, or to prove that none exists. We extend prior work on irreducible infeasible sets to also handle our problem’s complication that a constraint to avoid a side-effect cannot be relaxed without user permission. By proving that our objectives are also adaptive submodular, we devise a querying algorithm that we empirically show finds nearly-optimal queries with much less computation than a guaranteed-optimal approach, and outperforms competing approximate approaches.

Cite

CITATION STYLE

APA

Zhang, S., Durfee, E. H., & Singh, S. (2020). Querying to find a safe policy under uncertain safety constraints in Markov decision processes. In AAAI 2020 - 34th AAAI Conference on Artificial Intelligence (pp. 2552–2559). AAAI press. https://doi.org/10.1609/aaai.v34i03.5638

Querying to find a safe policy under uncertain safety constraints in Markov decision processes

Abstract

Cite

Register to see more suggestions