프런티어 LLM 11개·15개 분포 대상으로 native probabilistic sampling 능력을 통계적으로 감사한 첫 대규모 연구. 두 프로토콜(Batch Generation: 한 응답에 N=1000 샘플 / Independent Requests: stateless 1000 호출)에서 sharp 비대칭 — Batch는 7% 중간 통과율, Independent는 11개 모델 중 10개가 모든 분포에서 fail. 분포 복잡도와 N 증가에 따라 sampling 충실도 단조 감소. MCQ 답안 위치 균일성 위반·인구통계 텍스트-이미지 프롬프트 시스템 편향 등 다운스트림 task에 직접 영향. 통계 보장이 필요한 응용에는 외부 sampler 필수.
- •11개 프런티어 LLM·15개 분포에서 native probabilistic sampling 통계적 감사 — 첫 대규모 연구.
- •Batch Generation 7% 중간 통과 vs Independent Requests 11개 중 10개 전부 실패 — sharp 비대칭.
- •분포 복잡도·sampling horizon N 증가에 따라 sampling fidelity 단조 감소.
- •MCQ 답안 위치·인구통계 T2I 프롬프트 등 다운스트림 task에 systematic bias 전파.
- •통계 보장 필요한 응용에는 LLM 내부 sampler 대신 외부 도구 필수.
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
- 1.11개 LLM 대상 15개 확률분포 샘플링 능력 대규모 감사에서 심각한 취약점 확인
- 2.배치 생성 7% 패스율, 독립 요청은 11개 중 10개가 전체 분포에서 무펨 수준
- 3.분포 복잡도 증가 및 샘플링 범위 확대에 따라 성능이 단조 감소하는 패턴 확인
- 4.LLM 내부 샘플러 부재로 통계적 보장이 필요한 응용에는 외부 도구 필수
왜 중요한가?
LLM이 통계적 확률분포 샘플링을 신뢰성 있게 수행할 수 없다는 사실은 에이전트 파이프라인 설계 및 편향 리스크 관리에 중요한 시사점을 제공한다.
본문 미리보기
arXiv:2601.05414v3 Announce Type: cross Abstract: As large language models (LLMs) transition from chat interfaces to integral components of stochastic pipelines and systems approaching general intelligence, the ability to faithfully sample from specified probability distributions has become a functional requirement rather than a theoretical curiosity. We present the first large-scale, statistically powered audit of native probabilistic sampling in frontier LLMs, benchmarking 11 models across 15
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:45AI 초안

