기존 ATP(자동 정리 증명) 벤치마크는 최종 답이 형식 명제 안에 이미 들어있는 'Easy Mode' 설계로 모델 능력을 과대평가한다. 이 연구는 답을 직접 발견한 뒤 형식 증명을 구성하는 엄격한 'Hard Mode'를 도입. MiniF2F-Hard, FIMO-Hard 재주석 버전 공개 + DAP(Discover And Prove) 에이전틱 프레임워크 제안. CombiBench에서 SOTA를 7→10으로 끌어올렸고, PutnamBench에서 Hard Mode로 36개 정리를 최초 증명. 동시에 LLM은 답 정확도 80%+ vs 형식 증명자 10% 미만의 큰 격차 노출.
- •ATP 벤치마크 Hard Mode — 답 자체를 먼저 발견해야 하는 설정
- •MiniF2F-Hard, FIMO-Hard 재주석 벤치마크 공개
- •DAP 에이전틱 프레임워크 — LLM 자연어 추론 + 자기 반성
- •CombiBench SOTA 7→10, PutnamBench에서 Hard Mode 최초 36개 증명
- •LLM 답 정확도 80%+ vs 형식 증명자 10% 미만의 격차 노출
0단 자동
AI가 규칙대로 쓰고 그대로 게시했습니다. 사람이 따로 보지 않았습니다.
- 규칙 판
- 규칙 판 도입 이전 기사입니다.
- 남기는 것
- 규칙 판 · 모델 · 시각
- 판 기록
- 아직 없습니다.
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
- 1.수학 자동 증명의 Hard Mode 벤치마크 도입
- 2.답 탐색 + 형식 증명을 분리한 에이전트 접근
- 3.LLM 자연어 추론과 형식 증명자 간 큰 격차 드러냄
왜 중요한가?
AI 수학 증명은 DeepMind AlphaProof, OpenAI 등 최전선 경쟁 분야. Hard Mode 도입은 '진짜 발견' 능력을 측정해 과대평가된 벤치마크 생태계를 정화한다.
본문 미리보기
arXiv:2604.15839v1 Announce Type: new Abstract: Most ATP benchmarks embed the final answer within the formal statement -- a convention we call "Easy Mode" -- a design that simplifies the task relative to what human competitors face and may lead to optimistic estimates of model capability. We call the stricter, more realistic setting "Hard Mode": the system must independently discover the answer before constructing a formal proof. To enable Hard Mode research, we make two contributions. First, w
전체 내용이 궁금하다면?
원문을 직접 읽어보세요
이 글이 만들어진 과정
- 13:32AI 초안

