0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Visual Memory Attacks Can Persist Through The KV Cache
- 1.적대적 이미지가 컨텍스트에서 사라져도 KV캐시를 통해 공격 효과가 지속됨을 실증
- 2.P-VMI: 어텐션에서 마스킹된 후에도 백도어 행동을 유지하도록 이미지를 최적화
- 3.Qwen3-VL-8B-Instruct에서 최강 설정 시 목표 성공률 약 90% 달성
- 4.같은 모델이 생성한 요약으로 압축(compaction)해도 공격이 생존함을 확인
왜 중요한가?
위험한 입력을 컨텍스트에서 지우면 안전하다는 통상적 방어 가정이 KV캐시 수준에서는 성립하지 않음을 보여, 장기 대화·메모리 압축을 쓰는 멀티모달 에이전트의 보안 설계에 근본적 재검토가 필요함을 시사한다.
언급 프로젝트
본문 미리보기
arXiv:2610.09027v1 Announce Type: new Abstract: Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its source.We consider the Visual Memory Injection (VMI; Schlarmann
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 11:24AI 초안

