Imitation learning is probably existentially safe

0Citations
Citations of this article
1Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Concerns about extinction risk from AI vary among experts in the field. However, AI encompasses a very broad category of algorithms. Perhaps some algorithms would pose an extinction risk, and others would not. Such an observation might be of great interest to both regulators and innovators. This paper argues that advanced imitation learners would likely not cause human extinction. We first present a simple argument to that effect, and then we rebut six different arguments that have been made to the contrary. A common theme of most of these arguments is a story for how a subroutine within an advanced imitation learner could hijack the imitation learner's behavior toward its own ends. However, we argue that each argument is flawed and each story implausible.

Cite

CITATION STYLE

APA

Cohen, M. K., & Hutter, M. (2025). Imitation learning is probably existentially safe. AI Magazine, 46(4). https://doi.org/10.1002/aaai.70040

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free