Lucas Sun

3 posts shown

An Associative Introduction to Deep Learning

One sum, values weighted by key-query brackets, accounts for almost every parameter in a modern network. Starting from the correspondence between a single neuron and a rank-one outer product, this post uses bra-ket notation to derive MLPs, softmax attention, effective rank, linear attention, the delta rule, Gated DeltaNet and PaTH attention as one associative memory wearing different coefficients.