AgentSearchBench — 약 10,000개 실제 AI 에이전트로 구성된 대규모 에이전트 검색 벤치마크. 기존 벤치마크가 specified functionalities·controlled candidate pools 가정을 깔던 것과 달리 in-the-wild 시나리오를 대상으로 함. 에이전트 검색을 retrieval + reranking 문제로 정형화하고 execution-grounded 성능 신호로 평가. 실험 결과 의미적 유사도와 실제 에이전트 성능 사이에 일관된 격차 — description 기반 검색의 한계 노출. lightweight behavioral signals(execution-aware probing)이 ranking 품질을 크게 개선해, 에이전트 디스커버리에 execution signals 통합이 필요함을 시사.
- •10,000개 실제 에이전트로 구성된 in-the-wild 에이전트 검색 벤치마크 AgentSearchBench 공개.
- •에이전트 검색을 retrieval + reranking 문제로 정형화 — execution-grounded 성능 신호로 평가.
- •의미적 유사도와 실제 에이전트 성능 간 일관된 격차 — description-only 검색은 한계.
- •execution-aware probing 같은 lightweight behavioral signals가 ranking 품질 크게 개선.
- •에이전트 마켓플레이스·플랫폼 설계에 execution signal 기반 디스커버리가 핵심 요건임을 입증.
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
- 1.실 세계 10,000개 에이전트 데이터 기반 에이전트 검색 벤치마크 AgentSearchBench 구축
- 2.실행 가능 작업 쿼리 및 고수준 작업 설명 양쪽에서 검색·재랜킹 문제로 정형화
- 3.의미론적 유사성과 실제 에이전트 성능 간의 일관된 격차 확인, 설명 기반 검색의 한계 노옶
- 4.실행 인식 프레빙 등 경량 행동 신호가 랜킹 품질을 크게 개선함을 입증
왜 중요한가?
에이전트 생태계 확장에 따라 적합한 에이전트를 찾는 검색 문제가 부상하고 있으며, 실행 기반 신호를 활용한 접근법이 필수적임을 시사한다.
언급 프로젝트
본문 미리보기
arXiv:2604.22436v1 Announce Type: new Abstract: The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task. Unlike traditional tools, agent capabilities are often compositional and execution-dependent, making them difficult to assess from textual descriptions alone. However, existing research and benchmarks typically assume well-specified functionalities, controlled candidate pool
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:45AI 초안

