Finding an Induction Circuit in a Two-Layer Transformer
I am an ML practitioner, and I have always been drawn to the question of what happens under the hood inside a deep neural network. To explore that question hands-on, I started working through ARENA’s mechanistic interpretability exercises. I am writing this post to share my understanding as I develop it, beginning with one of the simplest interesting mechanisms found in transformers: induction circuits.