콘텐츠 모더레이션 평가가 'human label과의 agreement'에 의존하는데 rule-governed 환경에서는 다중 결정이 정책에 모두 logically consistent — agreement metric이 valid 결정을 penalize하고 ambiguity를 error로 오해석하는 'Agreement Trap' 발생. 평가를 policy-grounded correctness로 재정의하고 Defensibility Index(DI)·Ambiguity Index(AI) 도입. audit token logprob 기반 PDS(Probabilistic Defensibility Signal)로 reasoning 안정성 추가 audit pass 없이 추정. Reddit 193,000+ 결정 검증 — agreement 기반과 policy-grounded 메트릭 사이 33~46.6pp 격차, false negative의 79.8~80.6%가 실제 정책 기반 결정.
- •Rule-governed 환경에서 'agreement = correctness' 가정의 실패 — Agreement Trap 형식화.
- •Policy-grounded correctness 재정의 + Defensibility Index·Ambiguity Index 도입.
- •Token logprob 기반 PDS로 audit pass 없이 reasoning 안정성 추정.
- •Reddit 193K 결정 검증 — agreement vs policy-grounded 33~46.6pp 격차.
- •False negative의 79.8~80.6%가 실제로는 policy 기반 valid 결정 — 기존 평가의 체계적 오류.
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI
- 1.AI 콘텐츠 중재 평가 한계
- 2.규칙 기반 AI의 특성 고려
- 3.방어 가능성 신호 제안
왜 중요한가?
이 연구는 규칙 기반 AI 시스템, 특히 콘텐츠 중재에서 기존의 평가 방식이 가지는 한계를 지적하고, AI의 의사결정 과정을 더 정확하게 평가할 수 있는 새로운 '방어 가능성 신호'를 제시하여 AI의 신뢰성을 높입니다.
본문 미리보기
arXiv:2604.20972v1 Announce Type: new Abstract: Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multiple decisions may be logically consistent with the governing policy, and agreement metrics penalize valid decisions while mischaracterizing ambiguity as error -- a failure mode we term the Agreement Trap. We formalize evaluation as policy-grounded correctness and introduce the Defensibility Index (D
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:29AI 초안

