youtube.nixfred.com nixfred.com
Creator

FAR.AI

AI safety research organization that runs the Alignment Workshop series and publishes its talks.

1video

← All videos

FAR.AI

Chris Olah - Looking Inside Neural Networks with Mechanistic Interpretability

Chris Olah lays out the case for treating a neural network as an object to be studied under a microscope rather than a black box to be probed from outside: features as the basic unit, circuits as the connections between them, and the awkward discovery that neurons are usually polysemantic. His answer is superposition, the hypothesis that models pack far more features than they have dimensions by using nearly orthogonal directions for sparse features, and that dictionary learning can pull those features back out. The safety payoff is auditing: finding out what a model is doing rather than asking it.

AIDeep LearningScienceSep 1, 2023