LLM 5종(ChatGPT-4o, Gemini 2.5 Flash, DeepSeek, Kimi, Llama 3.1 8B)이 다중 센서 물리적 위험 데이터를 어떻게 평가하는지 측정한 벤치마크 연구다. 3개 범주 60개 시나리오, 온도 0.0으로 1,800회 API 호출을 실험한 결과, 모든 모델이 단일 센서 임계값 위반은 거의 완벽하게 잡아냈지만(0.975~1.000), 여러 센서가 각각 한계치 이하로 동시에 상승한 복합 위험 시나리오에서는 예방적 경고를 전혀 내지 못했다(점수 0.000~0.592). 표 형식 입력이 일관된 이점을 주지 않았고 ChatGPT-4o는 오히려 산문 입력에서 유의하게 더 나았다(p=0.001). LLM을 물리 안전 모니터링에 도입하려는 실무자에게 다중 센서 복합 위험 판단의 맹점을 경고하는 결과다.
- •5개 LLM을 60개 시나리오·1,800회 API 호출(temperature 0.0)로 다중 센서 위험 평가 능력 측정
- •단일 센서 임계값 위반 탐지는 거의 완벽(0.975~1.000)
- •여러 센서가 개별 한계 이하로 동시 상승하는 복합 위험에서는 예방 경고 실패(0.000~0.592)
- •표 형식 입력이 산문보다 낫다는 증거 없음—ChatGPT-4o는 산문에서 유의하게 우수(p=0.001)
- •물리 안전 모니터링에 LLM 도입 시 복합 위험 판단 맹점을 보완해야 함
Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
본문 미리보기
arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three categories - multi-sensor joint assessment, response proportionality, and pattern disambiguation - with 1,800 API calls at temperature 0.0, we find that all tested models consistently produced no precautionary warning signal across the tested scenarios where multiple sensors are simultaneously eleva
전체 내용이 궁금하다면?
원문을 직접 읽어보세요