Abstract
Today's artificial intelligence (AI) systems rely heavily on Artificial Neural Networks (ANNs), yet their black box nature induces risk of catastrophic failure and harm. In order to promote verifiably safe AI, my research will determine constraints on incentives from a game-theoretic perspective, tie those constraints to moral knowledge as represented by a knowledge graph, and reveal how neural models meet those constraints with novel interpretability methods. Specifically, I will develop techniques for describing models' decision-making processes by predicting and isolating their goals, especially in relation to values derived from knowledge graphs. My research will allow critical AI systems to be audited in service of effective regulation.
Author supplied keywords
Cite
CITATION STYLE
Kierans, A. (2023). Benchmarked Ethics: A Roadmap to AI Alignment, Moral Knowledge, and Control. In AIES 2023 - Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (pp. 964–965). Association for Computing Machinery, Inc. https://doi.org/10.1145/3600211.3604764
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.