판 이력 — Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning | AIChainDay