Understanding model behavior and making reasoning verifiable.
As AI systems become more capable, using them safely and effectively increasingly depends on our ability to understand and verify their behavior. I study this challenge through two complementary directions: understanding how model behavior arises from internal mechanisms, and exploring verifiable reasoning methods to reduce the cost of obtaining reliable feedback.
01
Mechanistic Understanding for AI Safety
I study how weights, activations, and circuits give rise to model behavior, and how they change during training and adaptation. I draw on existing analysis tools and develop new methods when needed to characterize and compare these mechanisms.
SWD · sparse read/write components as intervention units.02
Verifiable Reasoning for AI Safety
I study verifiable reasoning for AI safety, using formal languages and automated verifiers to reduce the cost of obtaining reliable feedback on model-generated solutions. I explore how this feedback can support scalable training and evaluation for tasks with explicit, checkable specifications.
Re:Form · verifier-backed reinforcement learning.
Publications & Preprints
2026
Sparse Weight Decomposition for Efficient Circuit Extraction
Chuanhao Yan*, Xuhan Huang*, Yawen Duan, Zhenfei Yin, Hang Zhao, Bryan Dai, Jie Fu