Segment-Aligned Policy Optimization for Multi-Modal Reasoning
A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.
lei-gao
A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.
System review of reward hacking.
A unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals .