DC-Leap은 디퓨전 LLM(dLLM)의 병렬 디코딩을 학습 없이 가속하는 프레임워크다. 기존 병렬 디코딩은 결합 확률 의존성 오류(JPDE) 탓에 보수적인 신뢰도 임계값을 쓸 수밖에 없어 불필요한 디노이징 반복이 발생했는데, DC-Leap은 엄격한 순서의 인과 제약을 통합한 동적 연속 검증(Dynamic Contiguous Verification)으로 토큰 의존성을 점진 검증해 JPDE를 무력화한다. 여기에 드래프트가 여러 토큰을 앞질러 문맥을 확장하는 드래프트 유도 디코딩으로 양방향 어텐션의 구조적 이점을 유지한다. 표준 벤치마크에서 MBPP 장문 생성 기준 최대 53.19배, KV 캐시 결합 시 최대 105.02배 속도 향상을 생성 품질 저하 없이 달성했다. dLLM의 실용 배포를 가로막던 추론 속도 문제를 크게 완화하는 결과이며 코드도 공개됐다.
- •dLLM 병렬 디코딩의 병목인 결합 확률 의존성 오류(JPDE)를 학습 없이 해결
- •인과 순서 제약을 통합한 동적 연속 검증으로 중간 신뢰도 구간에서도 안정적 가속
- •드래프트 유도 디코딩으로 룩어헤드 문맥 확보, 양방향 어텐션 이점 유지
- •MBPP 장문 생성 최대 53.19배, KV 캐시 결합 시 최대 105.02배 가속—품질 유지, 코드 공개
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
본문 미리보기
arXiv:2607.20467v1 Announce Type: new Abstract: While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by overly conservative confidence thresholds. These thresholds, necessitated by the Joint Probability Dependence Error (JPDE), result in redundant denoising iterations and suboptimal inference speeds. To overcome this, we propose DC-Leap, a training-free framework that enables reliable acceleration of dLLMs in the
전체 내용이 궁금하다면?
원문을 직접 읽어보세요