Linear attention is arguably the most naive RNN, but still massively outperforms traditional RNNs by maintaining a matrix state. So... what if we use a (triadic) outer product of three vectors and maintain a three-dimensional state?
Introducing: Triadic Linear Attention ð§ĩ
PhD Student @MIT_CSAIL | Past: @togethercompute @LTIatCMU @MITIBMLab @UNCNLP, @SFResearch, @BaiduResearch | Machine Learning, NLP.






