판 이력 — Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models | AIChainDay