Kuan Zhang(张宽)

About Me

I am a Ph.D. Candidate at the College of AI, Tsinghua University.

My research focuses on general game agents and general game world models. I am now working with Prof. Yiming Li at THUSI-Lab on building agents that can play and reason across diverse games, and I am currently a Research Intern at miHoYo (HoYoverse).

Previous: B.Eng. in Software Engineering at Beijing Institute of Technology (BIT); BIT-DataLab with Prof. Chengliang Chai — label noise learning and pretraining data selection.

Outside of research, I enjoy esports, AAA games, music, novels, and animation. Proud Master-tier League of Legends player. 🎮

General Game Agent General Game World Model

Education

College of AI, Tsinghua University

2026 – Present

Ph.D. Candidate  ·  Intern in THUSI-Lab

Beijing Institute of Technology

2022 – 2026

B.Eng. in Software Engineering  ·  Intern in BIT-DataLab

Chengdu No.7 High School

2019 – 2022

High School

Research Experience

miHoYo (HoYoverse)

Mar 2026 – Present

Research Intern

Publications

* denotes equal contribution, † denotes corresponding author

EMNLP 2026 Findings 2026

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Zirong Chen*, Fuda Ye*, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang†

SnapBench Paper Figure

The first paired benchmark for robust snap-and-ask multimodal retrieval, with 1,145 queries and 9,085 gallery items under 53 controlled corruption conditions. Image artifacts substantially degrade retrieval, while coarse user text can drag down joint retrieval; we propose MOOR for reliability-aware modality calibration.

ECCV 2026 2026

Towards Spatial Supersensing in the Wild

Tianjun Gu*, Tianyu Xin*, Kuan Zhang*, Bowen Yang, Kok-Chung Chua, Peize Li, Xinran Zhang, Yupeng Chen, Qiyue Zhao, Qinlei Xie, Jianhang Liu, Yucheng Lu, Yinan Han, Marco Pavone, Yiming Li†

VSI-Super-Wild Paper Figure

A spatial supersensing benchmark of 442 in-the-wild long-form videos and 6,980 QA pairs, showing that even strong multimodal models fail at coherent world-state tracking.

arXiv 2026 2026

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

Kuan Zhang*, Dongchen Liu*, Qiyue Zhao*, Tianyu Xin*, Yue Su*, Haisheng Wang, Han Yin, Hongbo Ma, Peize Li, Tianjun Gu, Xiangnan Wu, Xinran Zhang, Yongxuan Li, Zirong Chen, Yiming Li†

Game Multiverse Paper Figure

A systematic survey tracing the full lifecycle of generalist game players across four interdependent pillars — Dataset, Model, Harness, and Benchmark — and charting a five-level roadmap from single-game mastery toward the ultimate creator stage in the game multiverse.

ICML 2026 2026

GameVerse: Can Vision-Language Models Learn from Video-based Reflection?

Kuan Zhang*, Dongchen Liu*, Qiyue Zhao, Jinkun Hou, Xinran Zhang, Qinlei Xie, Miao Liu†, Yiming Li†

GameVerse Paper Figure

A comprehensive video game benchmark enabling a reflective visual interaction loop for VLMs, with cognitive hierarchical taxonomy spanning 15 globally popular games, dual action space, and milestone evaluation.

NeurIPS 2025 2025

Handling Label Noise via Instance-Level Difficulty Modeling and Dynamic Optimization

Kuan Zhang, Chengliang Chai, Jingzhe Xu, Chi Zhang, Han Han, Ye Yuan, Guoren Wang, Lei Cao

Label Noise Paper Figure

An efficient, hyperparameter-free, instance-level optimization framework for label noise in image classification.

ICLR 2025 Spotlight 2025

Harnessing Diversity for Important Data Selection in Pretraining Large Language Models

Chi Zhang*, Huaping Zhong*, Kuan Zhang, Chengliang Chai, Rui Wang, Xinlin Zhuang, Tianyi Bai, Qiu Jiantao, Lei Cao, Ju Fan, Ye Yuan, Guoren Wang, Conghui He

Harnessing Diversity Paper Figure

Combine multi-arm bandit with influence function to select data for pretraining process of large language models.