About
Trying to reduce risks from the highly capable machine learning systems around the corner. I work on interpretability, evaluation, and (mis)generalization properties of machine learning systems. Previously: tensor networks. I wrote a guide to graphical tensor notation for mechanistic interpretability, did a machine learning/ neuroscience internship in 2020/2021, attended the Machine Learning for Alignment Bootcamp (MLAB) in Berkeley, 2022, and the ML Alignment & Theory Scholars (MATS) Program in 2024, supervised by Lee Sharkey and Dan Braun from Apollo Research. Recently I interned at the NTT Physics & Informatics (PHI) Laboratories under Jess Riedel, before another internship at Center for Human-Compatible Artificial Intelligence under Erik Jenner.
Links
Badges
