Do Models Read What They Write? Causal Registers in Scratchpad Reasoning
B. Shih, J. Winnicki, and E. Darve
My research spans mechanistic interpretability and theoretical and scientific machine learning.
At Stanford, I am advised by Eric Darve in the DASH Lab, where I work on mechanistic interpretability: reverse-engineering the internal mechanisms of neural networks to understand how they represent information and arrive at their outputs.
Previously at Brown, I worked in the CRUNCH group with Zhongqiang Zhang and George Em Karniadakis on neural operators for differential equations, including transformer-based operator learning in finite-regularity settings.
B. Shih
Current research in the DASH Lab at Stanford, advised by Eric Darve. I work on mechanistic interpretability — reverse-engineering the internal computations of neural networks — across several directions in how models represent information and produce their behavior.
DASH Lab / Eric Darve
Previous work at Brown on neural operators for differential equations with the CRUNCH group, advised by Zhongqiang Zhang and George Em Karniadakis.
Earlier research on genome-wide association studies of neurodegenerative diseases with Dr. Li-San Wang at the University of Pennsylvania Wang Lab.