Multimodal LLM

Segment-Aligned Policy Optimization for Multi-Modal Reasoning featured image

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

lei-gao
Chimera: Improving Generalist Model with Domain-Specific Experts featured image

Chimera: Improving Generalist Model with Domain-Specific Experts

A scalable and low-cost multi-modal pipeline to boost existing LMMs with domain-specific experts.

tianshuo-peng
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs featured image

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

jiakang-yuan
OmniCaptioner: One Captioner to Rule Them All featured image

OmniCaptioner: One Captioner to Rule Them All

An all-in-one visual to text mapping tool.

yiting-lu
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training featured image

GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

A multi-modal large model focusing on geometric understanding and reasoning tasks with formalized visual-language pre-training.

renqiu-xia