Verbalizing Subliminal Learning Effects Using Text Optimization
Nathan Hu, Sanmi Koyejo, Christopher Potts
Preprint, 2026
Transcoder Adapters for Reasoning-Model Diffing
Nathan Hu, Jake Ward, Thomas Icard, Christopher Potts
Preprint, 2026
LLMs Can Annotate Attribution Graphs
Ameen Patel*, Max Zhang*, Nathan Hu
ICML Workshop on Mechanistic Interpretability, 2026
Interpreting Language Model Parameters
Lucius Bushnaq*, Dan Braun*, Oliver Clive-Griffin*, Bart Bussmann, Nathan Hu, Michael Ivanitskiy, Linda Linsefors, Lee Sharkey
Goodfire, 2026
Measuring Sparse Autoencoder Feature Sensitivity
Claire Tian, Katherine Tian, Nathan Hu
NeurIPS Workshop on Mechanistic Interpretability (Spotlight), 2025
Training on Documents About Reward Hacking Induces Reward Hacking
Nathan Hu, Benjamin Wright, Carson Denison, Samuel Marks, Johannes Treutlein, Jonathan Uesato, Evan Hubinger
Anthropic Alignment Science Blog, 2025
Meta-Learning Online Adaptation of Language Models
Nathan Hu*, Eric Mitchell*, Christopher D. Manning, Chelsea Finn
EMNLP, 2023