Test-time compute를 정적 할당이나 고정 분포 샘플링이 아닌, 어디에 어떻게 쓸지 동적으로 적응시키는 프레임워크. warm-up 단계에서 쉬운 쿼리 식별 + 테스트셋 자체로부터 question-response 풀 구성, adaptive 단계에서 미해결 쿼리에 추가 컴퓨트 집중하면서 의미적으로 관련된 쿼리의 성공 응답을 in-context demonstration으로 제공해 generation 분포 reshape. 수학·코딩·추론 벤치마크에서 기존 baseline 일관 능가, inference compute는 훨씬 적게 소비.
- •Test-time compute 할당을 'where + how' 둘 다 동적 적응시키는 프레임워크.
- •Warm-up 단계에서 쉬운 쿼리 식별 + 테스트셋 자체에서 question-response 풀 구성.
- •Adaptive 단계에서 미해결 쿼리에 컴퓨트 집중 + 관련 성공 응답을 in-context demo로 제공.
- •수학·코딩·추론 벤치마크 일관 능가 + inference compute 절감.
- •정적 할당·고정 분포 샘플링의 한계를 넘어선 evolving demonstration 접근.
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Adaptive Test-Time Compute Allocation with Evolving In-Context Demonstrations
- 1.AI 모델 성능 개선
- 2.추론 시 컴퓨팅 자원 조절
- 3.동적 자원 할당 프레임워크
왜 중요한가?
이 연구는 AI 모델의 추론 성능을 최적화하기 위한 동적인 컴퓨팅 자원 할당 방식을 제안하여, 자원 효율성을 높이고 다양한 환경에서 모델의 유연성을 강화합니다.
본문 미리보기
arXiv:2604.21018v1 Announce Type: new Abstract: While scaling test-time compute can substantially improve model performance, existing approaches either rely on static compute allocation or sample from fixed generation distributions. In this work, we introduce a test-time compute allocation framework that jointly adapts where computation is spent and how generation is performed. Our method begins with a warm-up phase that identifies easy queries and assembles an initial pool of question-response
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:29AI 초안

