LLM 훈련 불안정성 해결을 위해 AdamW 위에서 작동하는 경계 자율 훈련 제어 계층 LBW-Guard를 소개한다. 옵티마이저 업데이트 규칙을 대체하지 않고 훈련 원격 측정을 관찰해 불안정 민감 체제를 해석하며, 훈련 목표를 유지하면서 경계 제어를 적용한다. Qwen2.5 중심 스트레스 테스트에서 7B 기준 퍼플렉시티를 18.7% 개선하고 처리 시간을 단축했으며, 강한 학습률(LR=3e-3)에서 AdamW가 퍼플렉시티 1,885로 붕괴할 때 LBW-Guard는 11.57을 유지했다.
- •AdamW 위에서 작동하는 경계 자율 훈련 제어 계층 LBW-Guard 도입, 옵티마이저를 대체하지 않음
- •훈련 원격 측정으로 불안정 체제를 감지하고 경계 제어를 적용해 안정적 훈련 유지
- •7B 기준 설정에서 퍼플렉시티 18.7% 개선, 처리 시간 1.10x 단축
- •LR=3e-3 강한 스트레스에서 AdamW는 퍼플렉시티 1,885로 붕괴, LBW-Guard는 11.57로 훈련 가능 상태 유지
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency
본문 미리보기
arXiv:2605.19008v1 Announce Type: new Abstract: Modern language-model training is increasingly exposed to instability, degraded runs, and wasted compute, especially under aggressive learning-rate, scale, and runtime-stress conditions. This paper introduces Learn-by-Wire Guard (LBW-Guard), a bounded autonomous training-control governance layer that operates above AdamW. Rather than replacing the optimizer update rule, LBW-Guard observes training telemetry, interprets instability-sensitive regime
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:10AI 초안

