OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
- 1.LLM 멀티에이전트 시스템의 진화형 공격에 대응하는 공진화 지속 방어 OpenEvoShield 제안
- 2.공격측 빠른 학습률과 정상측 느린 학습률을 분리하는 비대칭 속도 제어기 등 4개 모듈 구성
- 3.노드·서브그래프·그래프 수준 증거를 융합하는 에너지 기반 탐지기로 신종 공격을 OOD로 분류
- 4.100라운드 배포·5개 벤치마크·4개 토폴로지에서 미확인 공격 대부분 탐지, 오탐율 낮게 유지
왜 중요한가?
기존 방어가 학습 범위를 벗어난 분포 변화에 급격히 무너지는 폐쇄 세계 가정을 깨고, 공격 진화와 정상 행동 드리프트가 동시에 일어나는 개방 환경을 처음부터 상정한 설계다. 에이전트 간 통신을 통한 악성 지시 전파가 현실 위협이 되는 만큼 시의성이 크다.
🏷️ 언급 프로젝트
본문 미리보기
arXiv:2607.19351v1 Announce Type: new Abstract: LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions through inter-agent communication to propagate harmful behaviors. Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed defenses while normal-agent behavior drifts with system expansion. Existing defenses treat deployment as a closed-world prob
전체 내용이 궁금하다면?
원문을 직접 읽어보세요