Paper-Conference

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models featured image

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models

Test-time reinforcement learning for vision-language-action models.

siyao-chen
Segment-Aligned Policy Optimization for Multi-Modal Reasoning featured image

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

lei-gao
FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow featured image

FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow

A multi-agent system built upon the dynamic structured knowledge flow.

yusong-hu
Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction featured image

Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction

A framework that allows users to dynamically regulate memory reliance.

muzhao-tian
Omniquality-R: Advanced Reward Models through All-Encompassing Quality Assessment featured image

Omniquality-R: Advanced Reward Models through All-Encompassing Quality Assessment

A unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals .

yiting-lu
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework featured image

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

State-of-the-art text-to-image generation model.

qi-qin
Chimera: Improving Generalist Model with Domain-Specific Experts featured image

Chimera: Improving Generalist Model with Domain-Specific Experts

A scalable and low-cost multi-modal pipeline to boost existing LMMs with domain-specific experts.

tianshuo-peng
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs featured image

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

jiakang-yuan
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing featured image

SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing

SurveyForge automatically generates and refines the content of surveys.

xiangchao-yan
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback featured image

Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback

A closed-loop auto-research framework that enables autonomous scientific research through thinking, practice, and feedback.

jiakang-yuan