0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Safety Invariants for Agents Orchestrating Irreversible State Transitions: A Four-Dimensional Formalism Evaluated on Public Ledgers
- 1.온체인 자금 이동을 지갑·체인·주소·프로토콜 4차원 상태 전이로 형식화해 '실행 충실성' 보장을 증명
- 2.장애 모델 하에서 세션의 실제 원장 효과는 '무효과' 또는 '사용자에게 표시된 전이 정확히 1회'로 한정
- 3.N=60 적대적 테스트에서 쓰기 공격적 모델 기준 naive-ReAct 대비 통과율 약 74%p 향상
- 4.8개 체인, 108건의 프로덕션 쓰기 작업으로 실패 분류 체계를 검증, 실제 배포 중
왜 중요한가?
ReAct 등 기존 에이전트 프레임워크가 벤치마크 성공률에 집중하는 사이, 되돌릴 수 없는 온체인 쓰기 작업의 안전을 수학적으로 보장하려는 시도다. 모델별 효과 격차(74%p vs 3%p)를 근거로 단일 모델 평가로는 에이전트 안전 스택 검증이 사실상 불가능하다는 점도 보여줘 평가 방법론에 시사점이 크다.
본문 미리보기
arXiv:2608.00783v1 Announce Type: new Abstract: Autonomous agents are increasingly asked to produce irreversible effects on external systems - transferring funds, writing to durable storage, actuating hardware. Existing agent frameworks (ReAct, Reflexion, MCP) optimize task success on benchmarks and give little attention to the safety of irreversible side-effects. We formalize one such setting, movement of value across public ledgers, as state transitions in a four-dimensional space indexed by
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 11:35AI 초안



