AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
- 1.어텐션 헤드별 학습형 회전 주파수·스케일링을 부여하는 AdaRoPE 제안
- 2.기능적 역할이 다른 헤드는 서로 다른 주파수 대역이 필요함을 이론·실험으로 입증
- 3.사전학습 시 partial RoPE·NoPE 등 기존 변형을 일관되게 능가
- 4.YaRN류 균일 스케일링이 컰텍스트 확장에 차선임을 보이며 단문맥 성능도 보존
왜 중요한가?
RoPE의 균일 주파수 스케줄이라는 표준 관행 자체를 헤드 단위로 재설계해, 긴 컨텍스트 확장에서 짧은 문맥 성능 희생 없이 개선을 얻었다. YaRN 등 널리 쓰이는 컨텍스트 확장 기법의 균일 스케일링 가정이 차선이라는 지적은 장문맥 모델 설계에 직접적 시사점을 준다.
본문 미리보기
arXiv:2607.19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effecti
전체 내용이 궁금하다면?
원문을 직접 읽어보세요