LLM의 기만(거짓말) 탐지 프로브 성능에 영향을 주는 요인을 체계적으로 분석한 연구다. 표현 깊이, 프로브 표현력, 희소 특징 표현, 학습 데이터의 거짓말 유형 네 요인을 조작(fabrication)·누락(omission)·과장(exaggeration) 예시를 포함한 보강 데이터셋과 7가지 프로브 유형으로 실험했다. 최적 표현 깊이는 데이터셋에 크게 의존하고, 표현력 높은 프로브는 선형 베이스라인 대비 선택적 이득만 제공했으며, 희소 오토인코더(SAE) 특징도 밀집 은닉 상태와 비슷한 성능에 그쳤다. 결국 학습 데이터의 선택과 거짓말 유형이 탐지 가능성을 좌우하는 핵심 변수로, 기만 탐지가 표현 의존적 문제임을 보여준다—OOD 일반화가 안 되는 기존 프로브의 한계를 설명하는 결과다.
- •표현 깊이·프로브 표현력·희소 특징·거짓말 유형 4개 요인이 기만 탐지에 미치는 영향을 체계 분석
- •조작·누락·과장 등 다양한 기만 유형을 담은 보강 데이터셋으로 7가지 프로브 유형 실험
- •최적 표현 깊이는 데이터셋 의존적, 표현력 높은 프로브는 선형 대비 선택적 이득에 그침
- •SAE 희소 특징은 밀집 은닉 상태와 유사한 성능—희소화 자체는 해결책 아님
- •학습 데이터와 거짓말 유형 선택이 탐지 가능성을 실질적으로 좌우
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
본문 미리보기
arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. In this work, we conduct a systematic study on how various factors impact detection performance: representation depth, probe expressivity, sparse featur
전체 내용이 궁금하다면?
원문을 직접 읽어보세요