AgentPatch는 여러 에이전틱 멀티모달 LLM을 하나의 범용 모델로 병합할 때 발생하는 비대칭 능력 손실과 핵심 행동 망각 문제를 해결하는 훈련 없는(training-free) 복구 프레임워크다. 안정적인 병합 백본을 선택한 뒤 Weak-Task Unique Residual Recovery로 약화된 작업별 신호를 복원하고, Agent-Guided Behavior-Critical Patch로 결정적 행동을 보호하며 되살린다. 라우팅이나 앙상블 없이 단일 정적 체크포인트만 생성하며, 6개의 에이전틱·멀티모달 벤치마크에서 다양한 병합 백본의 성능을 개선했다. 여러 전문 모델을 성능 저하 없이 하나로 통합해야 하는 실무 환경에 실용적 대안을 제시한다.
- •비대칭 능력 손실과 행동 망각이라는 두 가지 병합 실패 유형을 정의
- •Weak-Task Unique Residual Recovery로 약화된 작업 신호를 복원
- •Agent-Guided Behavior-Critical Patch로 핵심 행동을 보호하며 복구
- •훈련 없이 단일 정적 체크포인트로 병합 완료
- •6개 에이전틱·멀티모달 벤치마크에서 성능 개선 확인
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
본문 미리보기
arXiv:2608.06699v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dynamic environments. Yet current models are specialized for particular tools or environments, complicating consolidation into a single generalist. We formulate Agentic MLLM Merging and identify two challenges: asymmetric capability preservation, whereby capabilities with different interaction complexity are retai
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 10:11AI 초안

