Reinforcement Learning

Segment-Aligned Policy Optimization for Multi-Modal Reasoning featured image

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

lei-gao
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges featured image

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

System review of reward hacking.

xiaohua-wang
Omniquality-R: Advanced Reward Models through All-Encompassing Quality Assessment featured image

Omniquality-R: Advanced Reward Models through All-Encompassing Quality Assessment

A unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals .

yiting-lu