Wei Yao


Hi there, welcome! I am currently a final-year Ph.D. student at the Gaoling School of Artificial Intelligence, Renmin University of China, where I am honored to be advised by Prof. Yong Liu. Here is my CV. I am expected to graduate in 2027 and am open to research-oriented opportunities in academia and industry, including postdoctoral and research scientist positions.

From November 2025 to August 2026, I was a visiting Ph.D. student at the National University of Singapore, where I had the pleasure of working with Prof. Yunbei Xu. From October 2023 to March 2024, as a research intern at Shanghai AI Laboratory, I was fortunate to work under the guidance of Prof. Jing Shao. Prior to my Ph.D. studies, I earned my Bachelor of Engineering in Software Engineering from Huazhong University of Science and Technology in June 2022. I was fortunate to be advised by Prof. Kun He. During my undergraduate studies, I was honored to receive the National Scholarship (2019).

Research Interests


My long-term goal is to deepen our understanding of superhuman LLMs and help align them with human intentions. Toward this goal, I combine theoretical analysis and empirical studies to understand how LLMs learn and behave, and use these insights to improve their trustworthiness. My research interests include generalization in teacher-student learning, interpretability, robustness and fairness.

My recent work focuses on weak-to-strong generalization, a concrete setting for studying a central challenge in superalignment: how weak supervisors can align stronger models. Our work investigates how pre-training enables this phenomenon through representations and learning dynamics (arXiv:2605.05710), building on our earlier interpretability work on how representations relevant to trustworthiness evolve during LLM pre-training (ACL24). We also use a Bregman bias–variance analysis to characterize when models can generalize beyond their supervisors (ICML26), and examine how learning objectives shape generalization under weak supervision (ACL25). My earlier research explored robustness and fairness, including the theoretical principles underlying adversarial transferability (ICML25) and how learning objectives and network structure affect fairness (TMLR24, CVPR23).

Selected Publications


(* indicates equal contribution, # indicates corresponding authors)

Weak-to-Strong Generalization via Bregman Bias–Variance Decomposition

Gengze Xu*, Wei Yao*, Ziqiao Wang#, Yong Liu#
ICML 2026

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

Wei Yao*, Wenkai Yang*, Ziqiao Wang, Yankai Lin, Yong Liu#
ACL 2025 (Findings)

Understanding Model Ensemble in Transferable Adversarial Attack

Wei Yao*, Zeliang Zhang*, Huayi Tang, Yong Liu#
ICML 2025

Understanding Fairness Surrogate Functions in Algorithmic Fairness

Wei Yao*, Zhanke Zhou*, Zhicong Li, Bo Han, Yong Liu#
TMLR 2024 (Presented at ICLR 2025, Journal-to-Conference Track)

Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models

Chen Qian*, Jie Zhang*, Wei Yao*, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu#, Jing Shao#
ACL 2024 (Findings)

Fair Scratch Tickets: Finding Fair Sparse Networks without Weight Training

Pengwei Tang*, Wei Yao*, Zhicong Li, Yong Liu#
CVPR 2023

Service


Reviewer: ICML, NeurIPS, ICLR, ACL, CVPR.

Misc


Beyond academic research, I enjoy applying the same research mindset to everyday phenomena.