I'm a Member of Technical Staff at Anthropic, working on interpretability.
Previously, I was a research fellow at Anthropic, working on mechanisms of introspection in large language models.
Before that, I was a research scholar at MATS, working with Neel Nanda on reasoning model interpretability.
Before transitioning to safety and alignment research, I was a technology entrepreneur with an exit, having founded ventures in healthcare and education.
I also conducted research at Mila, where I focused on machine learning methods for neuroimaging.
I did my undergrad at Columbia University, studying computer science and applied mathematics.