A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

본문 미리보기

arXiv:2607.00155v1 Announce Type: new Abstract: We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward function, while the AI privately knows the quality of the action it proposes. This is the kind of asymmetry that arises naturally when an autonomous robot or software agent has inspected a situation its human supervisor cannot directly assess. Building on Cooperative Inverse Reinforcement Learning (CIRL) and the Ov

A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry

본문 미리보기

관련 글

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

Bounded Morality: Defining the Space of Moral Computation

The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons

Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection