판 이력 — Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning | AIChainDay