Segment-Aligned Policy Optimization for Multi-Modal Reasoning
A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.
lei-gao
A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.
A scalable and low-cost multi-modal pipeline to boost existing LMMs with domain-specific experts.
A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.
An all-in-one visual to text mapping tool.
A multi-modal large model focusing on geometric understanding and reasoning tasks with formalized visual-language pre-training.