LAMO는 경량 MLLM(Multimodal LLM)에 GUI 특화 지식과 멀티롤 오케스트레이션을 부여하는 프레임워크로, 온디바이스 GUI 에이전트의 비용·확장성 딜레마를 해결한다. 역할 지향 데이터 합성과 2단계 학습(SFT + Perplexity 가중 Cross-Entropy, 이후 RL)으로 구성되며, 결과로 개발된 LAMO-3B는 단독 실행과 MAS(Multi-Agent System) 오케스트레이션 모두 지원한다. 고급 플래너와 플러그앤플레이 결합 시 추가 성능 향상 가능.
- •경량 MLLM용 GUI 에이전트 프레임워크 LAMO 제안
- •역할 지향 데이터 합성 + 2단계 학습(SFT + RL)
- •LAMO-3B: 단독 실행 + MAS 오케스트레이션 모두 지원
- •고급 플래너와 플러그앤플레이 결합으로 성능 상한 확장
- •엔드유저 디바이스 GUI 자동화 비용 효율화
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Towards Scalable Lightweight GUI Agents via Multi-role Orchestration
본문 미리보기
arXiv:2604.13488v1 Announce Type: new Abstract: Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. While scaling both parameters and data has yielded substantial gains, advanced methods still suffer from prohibitive deployment costs on resource-constrained devices. When facing complex in-the-wild scenarios, lightweight GUI agents are bottlenecked by limited capacity and poor task scalability under e
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 15:12AI 초안

