Haodong Duan

Researcher · Multimodal Learning & Evaluation

ByteDance Seed Singapore

dhd.efz@gmail.com

欢迎围绕多模态学习、大模型评测与视频理解开展学术合作,也欢迎对开源评测工具和基准感兴趣的研究者与工程师联系我。

I am open to academic collaborations on multimodal learning, LLM/LMM evaluation, and video understanding. Please feel free to reach out by email.

Research Interests

My research focuses on multimodal learning, LLM/LMM evaluation, and video understanding. I build open-source evaluation infrastructure and benchmarks that make model capabilities easier to measure, compare, and reproduce.

Latest News

All News →
  • LEGO-Puzzles is published at ECCV 2026.

  • Seed2.1 is released, advancing multimodal reasoning and agentic productivity across the Seed model family.

  • Five papers are published at CVPR 2026, covering GUI-agent evaluation, spatial self-supervised RL, agentic reward modeling, visual reasoning, and poster intelligence.

  • Seed2.0 is officially launched, strengthening multimodal understanding, complex instruction execution, and real-world agent capabilities.

  • Three papers are published at ICLR 2026: MMSI-Bench, VisualPRM, and MM-HELIX.

  • Seed1.8 is officially released as a generalized agentic model with multimodal, search, coding, and GUI capabilities.

  • I joined ByteDance Seed in Singapore, after two years at Shanghai AI Laboratory working on OpenCompass and large-model evaluation.

  • RISEBench is accepted by NeurIPS 2025 Datasets & Benchmarks as an Oral presentation.

  • InternVL3.5 is released with stronger reasoning, efficiency, and deployment support.

Selected Works

All Publications →
First / Co-First Author Corresponding Author Project Lead
Figure from MMSI-Bench ICLR MMSI-Bench Spatial Reasoning · Multi-Image Reasoning · Multimodal Evaluation

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

MMSI-Bench evaluates whether MLLMs can integrate evidence across multiple images to solve grounded, real-world spatial reasoning problems.

Spatial ReasoningMulti-Image ReasoningMultimodal Evaluation

Sihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang, Mo Li, Jingli Lin, Chenming Zhu, Xiaochen Chen, Haodong Duan, Xiangyu Yue, Dahua Lin, Tai Wang, Jiangmiao Pang

ICLR, 2026

Figure from LEGO-Puzzles ECCV LEGO-Puzzles Spatial Reasoning · Visual Reasoning · Multimodal Evaluation

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

LEGO-Puzzles probes spatial understanding and multi-step planning through interpretable LEGO assembly tasks that remain difficult for frontier MLLMs.

Spatial ReasoningVisual ReasoningMultimodal Evaluation

Kexian Tang, Junyao Gao, Yanhong Zeng, Haodong Duan, Yanan Sun, Zhening Xing, Wenran Liu, Kai Chen, Kaifeng Lyu

ECCV, 2026

Figure from Visual-RFT ICCV Visual-RFT Reinforcement Learning · Reward Modeling · Model Adaptation

Visual-RFT: Visual Reinforcement Fine-Tuning

Visual-RFT brings reinforcement fine-tuning with verifiable rewards to visual perception tasks, improving data-efficient adaptation with limited examples.

Reinforcement LearningReward ModelingModel Adaptation

Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, Jiaqi Wang

IEEE/CVF International Conference on Computer Vision (ICCV), 2025

Figure from RISEBench NeurIPS RISEBench Multimodal Generation · Visual Reasoning · Multimodal Evaluation

Envisioning beyond the pixels: Benchmarking reasoning-informed visual editing

RISEBench measures whether image editors can follow reasoning-heavy temporal, causal, spatial, and logical instructions while preserving visual quality.

Multimodal GenerationVisual ReasoningMultimodal Evaluation

Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Xiaorong Zhu, Hao Li, Wenhao Chai, Zicheng Zhang, Renqiu Xia, Guangtao Zhai, Junchi Yan, Hua Yang, Xue Yang, Haodong Duan

NeurIPS D&B Oral, 2025

Figure from OmniAlign-V ACL OmniAlign-V Model Alignment · Reward Modeling · Data-Centric Learning

Omnialign-v: Towards enhanced alignment of mllms with human preference

OmniAlign-V improves multimodal assistants with diverse preference-alignment data and introduces MM-AlignBench for open-ended alignment evaluation.

Model AlignmentReward ModelingData-Centric Learning

Xiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosong Cao, Weiyun Wang, Jiaqi Wang, Xinyu Fang, Wenhai Wang, Guangtao Zhai, Haodong Duan, Hua Yang, Kai Chen

ACL, 2025

Figure from MMBench ECCV MMBench Multimodal Evaluation · Benchmark Design · Visual Reasoning

MMBench: Is Your Multi-modal Model an All-around Player?

MMBench provides a bilingual, ability-structured benchmark and CircularEval protocol for robustly diagnosing the capabilities of multimodal models.

Multimodal EvaluationBenchmark DesignVisual Reasoning

Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, Dahua Lin

European Conference on Computer Vision (ECCV), 2024

Figure from MMStar NeurIPS MMStar Benchmark Design · Multimodal Evaluation · Visual Reasoning

Are We on the Right Way for Evaluating Large Vision-Language Models?

MMStar curates genuinely vision-dependent questions to expose data leakage and inflated gains in large vision-language model evaluation.

Benchmark DesignMultimodal EvaluationVisual Reasoning

Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, Feng Zhao

Conference on Neural Information Processing Systems (NeurIPS), 2024

Figure from VLMEvalKit ACM MM VLMEvalKit Evaluation Infrastructure · Multimodal Evaluation

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

VLMEvalKit unifies model inference, benchmark access, and standardized scoring into a widely adopted toolkit for reproducible multimodal evaluation.

Evaluation InfrastructureMultimodal Evaluation

Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, Dahua Lin, Kai Chen

ACM International Conference on Multimedia (MM), 2024

Figure from ShareGPT4Video NeurIPS ShareGPT4Video Video Understanding · Data-Centric Learning · Multimodal Generation

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

ShareGPT4Video scales detailed video captioning to improve both video-language understanding and text-to-video generation.

Video UnderstandingData-Centric LearningMultimodal Generation

Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, Li Yuan, Yu Qiao, Dahua Lin, Feng Zhao, Jiaqi Wang

Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2024

Figure from ProSA EMNLP ProSA LLM Evaluation · Benchmark Design · Model Alignment

ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

ProSA measures prompt sensitivity at the instance level and analyzes why semantically equivalent instructions can produce unstable LLM performance.

LLM EvaluationBenchmark DesignModel Alignment

Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, Kai Chen

Empirical Methods in Natural Language Processing (EMNLP), Findings, 2024

Figure from MMBench-Video NeurIPS MMBench-Video Video Understanding · Long-Context Modeling · Multimodal Evaluation

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

MMBench-Video evaluates long-form, multi-shot video understanding with free-form questions and a fine-grained taxonomy of temporal capabilities.

Video UnderstandingLong-Context ModelingMultimodal Evaluation

Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, Kai Chen

Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2024

Figure from BotChat ACL BotChat LLM Evaluation · Agentic Systems · Model Alignment

BotChat: Evaluating LLMs' Capabilities of Having Multi-Turn Dialogues

BotChat evaluates multi-turn conversational ability through scalable bot-to-bot interaction and fine-grained dialogue-quality assessment.

LLM EvaluationAgentic SystemsModel Alignment

Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, Kai Chen

Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024

Figure from Prism NeurIPS Prism Visual Reasoning · Multimodal Evaluation · Model Efficiency

Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Prism decouples perception and reasoning to diagnose VLM capability bottlenecks and combine compact specialist components more effectively.

Visual ReasoningMultimodal EvaluationModel Efficiency

Yuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang, Lin Chen, Songyang Zhang, Jiaqi Wang, Dahua Lin, Kai Chen

Conference on Neural Information Processing Systems (NeurIPS), 2024

Figure from Ada-LEval ACL Ada-LEval Long-Context Modeling · LLM Evaluation · Benchmark Design

Ada-LEval: Evaluating Long-Context LLMs with Length-Adaptable Benchmarks

Ada-LEval adapts task length to model context windows, enabling controlled evaluation of retrieval, ordering, and full-text comprehension at scale.

Long-Context ModelingLLM EvaluationBenchmark Design

Chonghua Wang, Haodong Duan, Songyang Zhang, Dahua Lin, Kai Chen

Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2024

Figure from JourneyDB NeurIPS JourneyDB Image Understanding · Multimodal Generation · Multimodal Evaluation

JourneyDB: A Benchmark for Generative Image Understanding

JourneyDB pairs millions of generated images with prompts and annotations to benchmark captioning, prompt inversion, retrieval, and visual question answering.

Image UnderstandingMultimodal GenerationMultimodal Evaluation

Keqiang Sun, Junting Pan, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, Jifeng Dai, Yu Qiao, Limin Wang, Hongsheng Li

Conference on Neural Information Processing Systems (NeurIPS), Datasets & Benchmarks Track, 2023

Figure from TransRank CVPR TransRank Self-Supervised Learning · Video Understanding · Action Recognition

TransRank: Self-supervised Video Representation Learning via Ranking-based Transformation Recognition

TransRank reframes transformation recognition as relative ranking to learn stronger self-supervised video representations from temporal and spatial transformations.

Self-Supervised LearningVideo UnderstandingAction Recognition

Haodong Duan, Nanxuan Zhao, Kai Chen, Dahua Lin

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

Figure from OmniSource ECCV OmniSource Data-Centric Learning · Video Understanding · Action Recognition

Omni-sourced Webly-supervised Learning for Video Recognition

OmniSource unifies images, short clips, and untrimmed web videos to make large-scale web supervision more data-efficient for video recognition.

Data-Centric LearningVideo UnderstandingAction Recognition

Haodong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu, Dahua Lin

European Conference on Computer Vision (ECCV), 2020

Open-Source Projects

An all-in-one evaluation toolkit for large vision-language and multimodal models. It connects model inference, benchmark execution, scoring, and result comparison in one reproducible workflow used by both researchers and practitioners.

Multimodal EvaluationEvaluation Infrastructure

A comprehensive platform for evaluating large language and multimodal models. Its modular configs, task scheduling, evaluators, and reporting tools make broad and reproducible model comparisons easier to run and extend.

LLM EvaluationBenchmarking Platform

A unified toolbox for skeleton-based action recognition, with strong baselines spanning GCN- and CNN-based methods. It packages training recipes, pretrained models, and benchmark results so new methods can be compared on consistent footing.

Skeleton ActionVideo Recognition

OpenMMLab's modular video-understanding toolbox for action recognition, localization, skeleton-based recognition, and related tasks. It offers reusable components, extensive model coverage, and standardized training and inference pipelines.

Video UnderstandingOpenMMLab