Jiakang Yuan

Jiakang Yuan

Ph.D. Candidate

Fudan University

About Me

I am current a fourth-year Ph.D. student (Sep. 2022 - Jun. 2027, expected) in the School of Information Science and Technology, Fudan University, supervised by Prof. Tao Chen. Before this, I obtained my Bachelor’s degree in Electronic Engineering also from Fudan University (Sep. 2018 - Jun. 2022). I’m currently working in the fields of agentic model, multi-agent system, and multi-modal large language models. My research pursues to develop AI systems that can understand, reason, and interact with the real world.

Interests

Long-Horizon Agent Multi-Agent System Multimodal LLM Auto-Research

News

  • 2026.08: Qwen3.8-Max, a 2.4T-parameter model with 95B active parameters, is released with strong improvements in coding, research, and long-horizon autonomous tasks.
  • 2026.06: Agents-A1, an open-source 35B MoE model specialized on long-horizon tasks is out on Github and Huggingface, checkout the technical report.
  • 2026.06: One papers (T^2VLA) is accepted by ECCV 2026 which is about test-time learning of VLA models.
  • 2026.04: Three papers (MME-Reasoning, VisualScore, and, SAPO) are accepted by ICML 2026. Two of them are about multimodal reasoning, another is about visual quality assessment.
  • 2026.04: Bi3D++ is accepted by IEEE T-PAMI 2026.
  • 2026.04: Our survey about reward hacking is out on arxiv and github. Feel free to check it.
  • 2026.04: Two papers (FlowSearch and Controllable Memory Usage) are accepted by ACL 2026. One is about DeepResearch, the other is about agent memory.
  • 2026.03: Intern-S1-Pro technical report is released.
  • 2026.02: InternAgent-1.5 technical report is released.
  • 2025.12: SciEvalKit (An Open-source Evaluation Toolkit for Scientific General Intelligence) is released.
  • 2025.10: Codes of InternAgent 1.0 is released on github.
Show more ↓
  • 2025.07: SPOT is accepted by IEEE T-PAMI 2025.
  • 2025.06: Two papers (Lumina Image 2.0 and Chimera) are accepted by ICCV 2025.
  • 2025.05: Two papers (SurveyForge and Dolphin) are accepted by ACL 2025.
  • 2025.05: We release InternAgent (NovelSeek), a unified closed-loop multi-agent framework for Automatic Scientific Research.
  • 2025.02: One paper (CST-Stereo) is accepted by CVPR 2025.
  • 2024.12: One paper (GeoX) is accepted by ICLR 2025
  • 2024.12: One paper (AIOStereo) is accepted by AAAI 2025.
  • 2024.10: I receive the National Scholarship.
  • 2024.09: Two papers (AdaptiveDiffusion and 3DET-Mamba) are accepted by NeurIPS 2024.
  • 2024.07: One paper (Reg-TTA3D) is accepted by ECCV 2024.
  • 2024.01: One paper (ReSimAD) is accepted by ICLR 2024.
  • 2023.09: One paper (AD-PT) is accepted by NeurIPS 2023.
  • 2023.02: Two papers (Bi3D and Uni3D) are accepted by CVPR 2023.
  • 2022.07: One paper (HelixFormer) is accepted by ACM’MM 2022.

Selected Publications & Projects

Qwen3.8-Max: A New Bar for Coding and Cowork
Long-Horizon Agent Large Language Model
Blog

Qwen3.8-Max: A New Bar for Coding and Cowork

Qwen Team, Alibaba Group

Qwen 3.8-Max scales to 2.4 trillion parameters with comprehensive improvements across coding, work, research, and long-horizon tasks.

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Agentic AI Large Language Model
arXiv

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Agents-A1 Team, Shanghai AI Lab

Open-source Agentic models specialized on Long-horizon tasks.

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator
Agentic AI Distribution Shift
arXiv

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

Xiaohua Wang*, Jiakang Yuan*, Zisu Huang, Muzhao Tian, Changze Lv, Kaitao Song, Tao Chen, Xiaoqing Zheng

Graph-search-based end-to-end MLE task solver.

Intern-S1-Pro: Scientific multimodal foundation model at trillion scale
Large Language Model Agentic AI
arXiv

Intern-S1-Pro: Scientific multimodal foundation model at trillion scale

Intern-S1 Team, Shanghai AI Lab

Intern-S1-Pro: one-trillion-parameter scientific multimodal foundation model.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
Multimodal LLM Benchmark Reasoning
ICML 2026

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

Jiakang Yuan*, Tianshuo Peng*, Yilei Jiang, Yiting Lu, Renrui Zhang, Kaituo Feng, Chaoyou Fu, Tao Chen, Lei Bai, Bo Zhang, Xiangyu Yue

A comprehensive benchmark specifically designed to evaluate the reasoning capability of MLLMs.

Experience

Research Intern

Qwen Team, Alibaba Group

Long-horizon Agentic Task.

Research Intern

Hunyuan, Tencent

Agentic Image Captioning.

Research Intern

Shanghai AI Laboratory

Agentic Model, Multi-agent System.

Research Intern

Shanghai AI Laboratory

Autonomous Driving, 3D Perception.

Education

Ph.D. Candidate

Fudan University

Expected

B.Eng. in Electronic Engineering

Fudan University

Invited Talks

  • Efficient Pre-training of Autonomous Driving, 2023.09 — Watch Video
  • Towards 3D General Representation, Techbeat, 2023.07 — Watch Video
  • Transferability of Autonomous Driving, 2023.03 — Watch Video

Academic Services

Reviewer
CVPR ICCV ECCV NeurIPS ICML ICLR ACM'MM T-PAMI T-IP T-CSVT T-MM