Research

Understanding model behavior and making reasoning verifiable.

As AI systems become more capable, using them safely and effectively increasingly depends on our ability to understand and verify their behavior. I study this challenge through two complementary directions: understanding how model behavior arises from internal mechanisms, and exploring verifiable reasoning methods to reduce the cost of obtaining reliable feedback.

01

Mechanistic Understanding for AI Safety

I study how weights, activations, and circuits give rise to model behavior, and how they change during training and adaptation. I draw on existing analysis tools and develop new methods when needed to characterize and compare these mechanisms.

SWD · sparse read/write components as intervention units.
02

Verifiable Reasoning for AI Safety

I study verifiable reasoning for AI safety, using formal languages and automated verifiers to reduce the cost of obtaining reliable feedback on model-generated solutions. I explore how this feedback can support scalable training and evaluation for tasks with explicit, checkable specifications.

Re:Form reinforcement-learning pipeline with specification-subset and verification rewards
Re:Form · verifier-backed reinforcement learning.

Publications & Preprints

2026

Sparse Weight Decomposition for Efficient Circuit Extraction

Chuanhao Yan*, Xuhan Huang*, Yawen Duan, Zhenfei Yin, Hang Zhao, Bryan Dai, Jie Fu

Preprint · 2026 Paper Code Blog
2026

CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling

Zhengyang Tang*, Zihan Ye*, Chenyu Huang*, Xuhan Huang, Chengpeng Li, Sihang Li, Guanhua Chen, Ming Yan, Zizhuo Wang, Hongyuan Zha, Dayiheng Liu, Benyou Wang

ICML · 2026 Paper
2026

Re:Form—Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

Chuanhao Yan*, Fengdi Che*, Xuhan Huang*, Xu Xu*, Xin Li*, Yizhi Li*, Xingwei Qu*, Jingzhe Shi, Chenghua Lin, Yaodong Yang, Binhang Yuan, Hang Zhao, Yu Qiao, Bowen Zhou, Jie Fu

TMLR · 2026 Paper Code
2026

Differentiable Evolutionary Reinforcement Learning

Sitao Cheng*, Tianle Li*, Xuhan Huang*, Xunjian Yin, Difan Zou

ICLR RSI · 2026 Poster Paper Code
2026

VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code

Lingfei Zeng*, Fengdi Che*, Xuhan Huang, Fei Ye, Xu Xu, Binhang Yuan, Jie Fu

ICLR · 2026 Paper Code
2026

Federated Linear Dueling Bandits

Xuhan Huang, Yan Hu, Zhiyan Li, Zhiyong Wang, Benyou Wang, Zhongxiang Dai

AAAI · 2026 Paper
2025

LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages

Xuhan Huang, Qingning Shen, Yan Hu, Anningzhe Gao, Benyou Wang

NAACL Findings · 2025 Paper Code

* Equal contribution.